JD.com Open Sources Echo-WM, a Navigable Video, Audio, Music and Speech World Model
AI video model releases
JD.com Open Sources Echo-WM, a Navigable Video, Audio, Music and Speech World Model
JD.com's research group released Echo-WM under the JoyAI-Echo repository, generating video, environmental sound, music, and speech together as users continuously navigate a scene.
The current model uses a modified Lightricks LTX-2.3 backbone; its roadmap plans a move to LTX-2.5 and lower-cost long rollouts through attention kernels, variable-length KV caching, and FP8 or TensorRT decoding.
A separate causal variant uses chunk-causal attention and a four-step Flash rollout to reduce latency for continuing scenes, but JD.com lists this real-time-oriented path as roadmap work rather than a shipped result.
The same repository contains Echo-LongVideo, a separate project intended to preserve multi-shot continuity across clips of roughly five minutes.
Echo-WM is restricted to academic research and non-commercial use under the LTX-2 Community License Agreement, and the repository provides no independent benchmarks, latency data, or comparisons with other world models.