xAI adds character references and 1080p to Imagine Video 1.5

Imagine Video 1.5 adds up to seven image references, voice consistency, prompt-only generation, and native 1080p for Grok users.

· 2 min read
Grok

xAI has expanded Imagine Video 1.5 with image and voice references, prompt-only video generation, and native 1080p output. The update gives users several ways to shape a clip: describe a shot without a starting image, guide it with reference images, or combine a character image with a voice sample.

Image and voice references are starting in the United States for SuperGrok Heavy and SuperGrok Plus subscribers on Grok’s Imagine website and iOS app. xAI says the tools will roll out to all tiers over the next few days. Text-to-video and native 1080p generation are already generally available through Grok Imagine on the web, iOS, and Android, with 1080p supported for both text-to-video and image-to-video workflows.

Multi-Reference lets creators assign separate visual anchors to a generation. One image can hold a face in place, another can preserve a product, and another can define a location. Users can keep a character while changing the scene, retain a scene while swapping the character, or preserve both while altering only the action. The model accepts up to seven references in a single generation.

0:00
/0:11

Voice consistency adds another layer of continuity. When a character image and voice reference are supplied together, xAI says the model can maintain the same face and voice across scenes. That targets a common production need: generating multiple shots around a recurring character without rebuilding identity cues for every clip.

Developers can access image references, text-to-video, and native 1080p through the xAI API using the grok-imagine-video-1.5 model. Voice references require a request to xAI. The API example shows a prompt, reference image URL, duration, aspect ratio, and resolution passed into a video generation call.

The rollout extends the Imagine Video 1.5 release introduced last month, which xAI presented as an advance in motion, physics, and audio. The latest additions shift the focus toward tighter control over identity, scene, and product while expanding the model from image-led animation to videos generated directly from text.

Source