ShengShu Technology's Vidu S2 splits into a real-time avatar model and a real-time video editing model, at 720p output
- ShengShu Technology's Vidu S2 ships as two models: S2-Avatar, a real-time interactive digital-character model, and S2-Editing, which edits a continuous incoming video stream from text instructions and optional reference images.
- Vidu S2-Avatar raises real-time output resolution from 540p to 720p over Vidu S1 and lets users drop in a new reference image at any point during generation, so a speaking character can pick up a pictured cup, change into a pictured outfit, or move into a pictured scene.
- S2-Avatar targets state preservation across chained instructions, such as "pick up the cup, then smile" or "take off the hat and put it back on", using visual feedback and prompt adjustments to keep actions reliable.
- S2-Editing covers four real-time tasks: style transfer, outfit change, subject replacement, and background replacement, each required to follow motion in the source stream such as a raised arm, a turn, or camera movement while keeping subject-background spatial relationships.
- The end-to-end work was led by Jintao Zhang, a PhD student advised by Professor Jun Zhu and head of streaming video generation and inference at ShengShu Technology; the team also explored real-time spatial video generation and editing for VR headsets.