JD.com open-sources EchoWM interactive audio-visual world model and shows JoyAI-Echo1.5 clearing 10 minutes of continuous generation
- JD Explore Research Institute unveiled two JoyAI-Echo developments at the JDD conference on September 9: the upgraded JoyAI-Echo1.5 long video model and the newly open-sourced interactive audio-visual world model EchoWM.
- Echo1.5 uses an audio-visual Memory Bank architecture to hold character appearance, voice, and story line across multi-camera, long-duration generation, sustaining continuous audio-visual output beyond 10 minutes.
- Long-video training raised inference speed up to 7.5 times, letting the model generate 24 fps at 480P on only 2 H200 GPUs (real-time 1:1), while a Director Agent handles story development, shot planning, generation review, and final assembly from natural language.
- EchoWM generates native 720P video plus ambient sound, music, and voice that shift with scene, camera, and subject motion, and a unified camera intention drives first-person navigation, third-person camera and subject collaboration, and multi-round exploration.
- EchoWM and its four-step causal streaming variant EchoWM-Flash took the top two places on the WBench Navigation evaluation, with EchoWM first overall across 158 multi-round navigation cases; JD also showed the JoyAI-Video foundation model, planned for official release in October.