JD.com Open Sources Echo-WM, a Navigable Video, Audio, Music and Speech World Model
- JD.com's research group released Echo-WM under the JoyAI-Echo repository, generating video, environmental sound, music, and speech together as users continuously navigate a scene.
- The current model uses a modified Lightricks LTX-2.3 backbone; its roadmap plans a move to LTX-2.5 and lower-cost long rollouts through attention kernels, variable-length KV caching, and FP8 or TensorRT decoding.
- A separate causal variant uses chunk-causal attention and a four-step Flash rollout to reduce latency for continuing scenes, but JD.com lists this real-time-oriented path as roadmap work rather than a shipped result.
- The same repository contains Echo-LongVideo, a separate project intended to preserve multi-shot continuity across clips of roughly five minutes.
- Echo-WM is restricted to academic research and non-commercial use under the LTX-2 Community License Agreement, and the repository provides no independent benchmarks, latency data, or comparisons with other world models.