JD.com Open-Sources JoyAI-EchoWM for Interactive Video and Audio World Generation
- JD.com has open-sourced JoyAI-EchoWM and the lighter EchoWM-Flash, audiovisual world models that generate synchronized video, ambient sound, music, and speech during real-time user navigation.
- A unified camera-intent interface converts keyboard or controller commands into continuous six-degree-of-freedom trajectories for first-person exploration and third-person follow shots.
- JoyAI-EchoWM reported an average 81.6 to 81.7 on WBench Navigation; the four-step causal streaming EchoWM-Flash scored near 81.0 while retaining high interaction scores.
- The models use a trailing audiovisual context window across exploration turns, but the team states they lack explicit persistent 3D memory, so output can drift during longer sessions.
- Training combined gameplay logs, human play footage, Unreal Engine pose-ground-truth simulations, and internet video, followed by audiovisual pretraining, action tuning, joint tuning, and autoregressive post-training.