Shengshu Technology Launches Vidu S2 With Real-Time Editing, Dynamic Reference Images and VR Headset Streaming
- Shengshu Technology released Vidu S2 on September 15 as a two-model series: S2-Avatar for continuous interactive digital characters and S2-Editing for real-time editing of input video streams.
- S2-Avatar raises real-time output from 540P to 720P and lets users add item, clothing or background reference images mid-stream, so the character can pick up a cup, change into specified clothing, or enter a new scene.
- S2-Editing covers four tasks (style transfer, virtual try-on, character replacement, background replacement) and can convert a normal monocular video into spatial video or edit existing spatial video, streaming synchronized left and right eye views to VR headsets.
- Training relies on Self-Replay Forcing (SRF), which samples and re-noises segments from longer self-generated trajectories to stop identity drift and action breaks; a VLM agent checks action completion, a single-step Refiner restores detail, and TurboDiffusion plus TurboServe cut compute cost.
- S2-Avatar led nine indicators in StreamAV-Bench, while S2-Editing scored 4.26 in the joint OpenVE and RefVIE evaluation and took best results in all indicators on Sparkle-Bench and the ViViD virtual try-on test.