Meituan Open-Sources LongCat-Video-Avatar 1.5, Claims Wins Over Kling, OmniHuman, HeyGen
AI video model releases
Meituan Open-Sources LongCat-Video-Avatar 1.5, Claims Wins Over Kling, OmniHuman, HeyGen
Meituan's LongCat large model team open-sourced LongCat-Video-Avatar 1.5, a commercial-grade digital human video generation model available on GitHub, HuggingFace, and Modelscope.
The model switches audio encoding from Wav2Vec2 to Whisper-large and uses GRPO with a first-frame hand detection mechanism to fix lip sync, hand distortion, and identity drift issues.
DMD (Distributed Matching Distillation) compresses generation from 50 steps to 8 steps, and a shared base model plus LoRA adapters cuts VRAM use, yielding roughly 15x faster inference (a 10-second video in about one minute).
On the EvalTalker benchmark (770 evaluators, 10 domain experts), the model recorded user preference win rates of 65.9% over Kling Avatar 2.0, 61.1% over OmniHuman-1.5, and 54.3% over HeyGen.
Video stability metrics show a 23.1% subject deformation rate, 9.4% background deformation rate, and a 0.8% frame skipping rate, which the team says is the lowest among compared models.