ShengShu launches Vidu S1, a real-time interactive AI video model running on consumer GPUs
ShengShu launches Vidu S1, a real-time interactive AI video model running on consumer GPUs
AI video model releases
Sign in to create alerts.
ShengShu Technology unveiled Vidu S1 at the 2026 Global Digital Economy Conference in Singapore, a video model that generates continuous, real-time interactive video instead of fixed single clips.
Vidu S1 uses an autoregressive diffusion (AR + Diffusion) architecture to generate video frame by frame based on prior frames, live voice input, and conversational context, allowing unlimited-duration interaction rather than fixed-length clips.
The model runs at 540P (960x540) resolution at 25 FPS (up to 42 FPS) and lets users turn a single image (person, anime character, or pet) into a voice-controllable interactive avatar.
Voice input drives full avatar control (lip sync, facial expressions, eye movement, gestures, body posture) rather than just lip movement, by interpreting semantic meaning and emotional context of speech.
ShengShu achieved real-time performance on consumer-grade GPUs using acceleration techniques including TurboDiffusion, low-bit SageAttention, SLA, SpargeAttention, and the TurboServe inference serving engine.