ShengShu unveils Vidu S2, adding real-time avatar interaction and live stream editing
- ShengShu Technology unveiled Vidu S2 on September 15, 2026, split into two models: S2-Avatar for continuous real-time digital-character interaction and S2-Editing for editing incoming video streams.
- S2-Avatar raises real-time output resolution from 540p to 720p over Vidu S1 and lets users drop a new reference image at any point mid-stream, so a speaking character can pick up a pictured cup, change into a given outfit, or enter a new scene.
- S2-Editing supports four real-time tasks on live video: style transfer, outfit change, subject replacement and background replacement, tracking arm movements, turns and camera motion so the edited output follows the source footage.
- ShengShu connected S2-Avatar output to a spatial video pipeline for synchronized left- and right-eye views on VR headsets, while S2-Editing can either edit monocular video before conversion or edit existing spatial video.
- Jintao Zhang, a PhD student advised by Professor Jun Zhu and head of streaming video generation at ShengShu, led end-to-end development; training stabilized backgrounds on dance footage and annotated actions chronologically to teach state preservation across instructions.