xAI updates Grok Imagine Video 1.5 with image and voice references, text-to-video, native 1080p
AI video model releases
xAI updates Grok Imagine Video 1.5 with image and voice references, text-to-video, native 1080p
xAI adds image and voice references, text-to-video, and native 1080p generation to Grok Imagine Video 1.5, up from the version launched the prior month.
Voice consistency lets users pass a character image plus a voice reference so the same face and voice persist across every scene in a generated video.
Multi-Reference supports up to seven reference images per generation, each locking in one element such as a face, product, or location, so users can swap scene, character, or action independently.
Text-to-video now pairs xAI's image generation with image-to-video, letting users generate video from a text prompt alone without a starting image.
Image references, text-to-video, and native 1080p are live now in the xAI API under model grok-imagine-video-1.5, with voice reference support available on request; consumer rollout starts in the US for SuperGrok Heavy and Plus tiers on grok.com/imagine and iOS, expanding to all tiers within days.