Grok Imagine Video 1.5 Workflow: Start Frames, References, Native Audio, and 1080p Limits
A creator workflow for Grok Imagine Video 1.5: pick start-frame or reference mode, stay inside official 15-second and 1080p limits, and know what to generate on ClipCanva instead.
Grok Imagine Video 1.5 Workflow: Start Frames, References, Native Audio, and 1080p Limits
A Grok Imagine Video 1.5 workflow is not “type a cinematic prompt and hope.” xAI’s 16 June 2026 launch put grok-imagine-video-1.5 in general availability on the Imagine API: start from an image, describe the motion, pick duration and resolution, and get sound effects, ambience, and dialogue in the same pass. The official generation docs (updated 18 August 2026) then split the job into modes. Text-to-video on 1.5 is a hidden first-frame render plus animation. Image-to-video locks that first frame. Reference-to-video can take up to seven stills and three preset voices, but it caps at 720p. If you mix those modes in one request, the API will not guess what you meant.
ClipCanva is not affiliated with xAI, Runway, Canva, ByteDance, or Kling. Grok Imagine Video 1.5 is not a live ClipCanva generator. Generate the still and the scene card here, then send the packet to xAI or to Runway’s grok_imagine_1_5 endpoint. For a hands-on clip on ClipCanva today, use Seedance 2 or the image-to-video tool.
Official facts for Grok 1.5 and nearby tools
| Source | Public claim | Workflow implication |
|---|---|---|
| xAI Grok Imagine Video 1.5 (16 Jun 2026) | Generally available on the Imagine API as grok-imagine-video-1.5. Video 1.5 Fast is on grok.com/imagine plus iOS and Android. Audio, speech, motion, and physics are generated in one pass. Fast: a 6-second 720p clip in about 25 seconds, down from 40+ seconds. |
Treat Fast as a draft lane. Do not quote 25 seconds as a studio SLA. |
| xAI video generation (updated 18 Aug 2026) | Duration 1–15 seconds. Aspect ratios: 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3. Resolution: 480p (default), 720p, 1080p. Text-to-video on 1.5 is text-to-image, then image-to-video. The intermediate still is not returned. Audio is on by default. |
If the first frame is wrong, do not keep regenerating the whole clip. Lock a still first. |
| Same docs | 1080p is supported on 1.5 for text-to-video and image-to-video. Reference-to-video is capped at 720p. Video editing keeps the input duration (max 8.7 seconds), aspect, and resolution, with resolution capped at 720p. |
1080p is not a global switch. It is a mode privilege. |
| xAI image-to-video | Animate a still from a public URL, a data URI, or a Files API file_id. Native 1080p is supported on 1.5. The output defaults to the still’s aspect ratio; setting aspect_ratio stretches the image. |
Crop the still before you generate. Do not stretch a 4:5 product crop into 16:9. |
| xAI reference-to-video (updated 5 Aug 2026) | Up to 7 reference images. Up to 3 preset voice_ids, tagged as <AUDIO_0>–<AUDIO_2>. Custom voice files are partner-only. Max 15 seconds, max 720p. Cannot combine with image-to-video or video editing. |
References guide identity. They do not lock frame one. That is a different mode. |
| xAI video extension | Input must be 2–15 seconds. The duration field is the extension only (2–10 seconds, default 6). Output matches the source aspect and resolution, capped at 720p. |
A 10-second keeper plus duration: 5 is a 15-second file, not a 5-second recut. |
| xAI Imagine pricing | grok-imagine-video-1.5: $0.08/sec at 480p, $0.14/sec at 720p, $0.25/sec at 1080p, plus $0.01 per input image. |
Price the resolution, not the adjective “cinematic.” |
| Runway API changelog (7 Aug 2026) | grok_imagine_1_5 on text-to-video and image-to-video. Optional native audio. Up to 7 image references and 3 audio references (3–15 seconds each). Audio references require at least one image. Requests with image references cap at 720p. Image-to-video uses a single first frame. Billing: 10 / 16 / 29 credits per second at 480p / 720p / 1080p, plus 1 credit per image or audio reference. |
Runway IDs and xAI IDs are different. Do not paste grok-imagine-video-1.5 into a Runway field. |
| Canva AI Video Generator | Create a Video Clip, powered by Google Veo-3: one 16:9 clip per prompt, up to eight seconds, with synchronized audio. | An 8-second layout clip is not a 15-second Grok reference pack. |
These are vendor-documented limits, not a promise that your bottle, face, or spoken line will hold. Access, region, safety filters, and credit rules still sit on the live xAI, grok.com, or Runway project.
Pick the mode before you attach files
Grok 1.5 fails when every brief is treated as one text box. Official docs are explicit: only one mode is active per request.
| Job | Official mode | What you must supply | What the model may invent | Fail condition |
|---|---|---|---|---|
| New world from a sentence | Text-to-video | One scene card. Know that 1.5 will invent a first frame you never see. | The still, the motion, and the audio bed | Unreadable label, extra SKU, new wardrobe |
| Animate an approved still | Image-to-video | One keeper still plus one verb | Motion and sound only | The product is redesigned |
| Keep a person or product without locking frame one | Reference-to-video | Named stills (<IMAGE_1>…), optional preset voice. Cap: 7 images, 3 voices, 720p |
Camera path and in-between motion | Mixing this with a start-frame upload |
| Continue a keeper | Video extension | A 2–15s clip; duration = extra seconds only |
The next beat | Asking for 1080p on the extension |
| Recut an existing clip | Video editing | Source video only. Duration stays, max 8.7s, 720p cap | Local changes, not a new story | A 15-second recut |
| Hands-on ClipCanva take | Seedance 2 or image-to-video | The same scene card, on a live model | Whatever that model documents | Pretending Grok is on ClipCanva |
| 8-second layout clip | Canva Create a Video Clip (Veo-3) | One 16:9 prompt | An 8-second shot with synced audio | Asking Canva for seven references |
If the brief is “this 30ml bottle must not change,” lock the still first. Approve the crop. Then run image-to-video. If the brief is “this person in this shirt, walking a new set,” use reference-to-video and tag the files. Do not upload a start frame and a seven-image zip in the same call.
Delivery card before you pay $0.25 a second
Write the card before you hit generate. Official xAI pricing jumps from $0.08/sec at 480p to $0.25/sec at 1080p.
Job: [animate approved still | new shot | identity refs | extend keeper]
Surface: [xAI API grok-imagine-video-1.5 | Runway grok_imagine_1_5 | grok.com Fast]
Mode: [text-to-video | image-to-video | reference-to-video | extend | edit]
Duration: [1–15s]. If extending, duration = extra seconds only (2–10).
Resolution: [480p draft | 720p review | 1080p only if T2V/I2V]
Aspect: follow the still. Do not stretch.
Audio: [SFX + room | quoted line | preset voice_id]
Identity lock: [SKU / face / wardrobe]. Fail if a second product appears.
Success: same object, one camera path, readable label, audio lands on the action.
Starter for a 8-second product keeper (original; not copied from xAI or Runway samples):
Mode: image-to-video. Do not send reference_images.
Source still: 4:5 crop, 30ml frosted bottle, gold pump, black label “NORTH HARBOR 01”.
Motion: slow push-in, one pump press, one bead of serum on one hand.
Camera: locked tabletop, north-window key, pale oak, no extra props.
Audio: pump click + quiet room. No voiceover. No music sting.
Duration: 8. Resolution: 720p for review. 1080p only after the SKU crop test passes.
Avoid: second bottle, slogan overlay, fake UI, extra language on glass.
If the still fails the label crop, stop. Do not animate a broken frame.
Build that still on ClipCanva with prompt ideas or a live image model, then move the file. Do not paste resolution: "1080p" into a ClipCanva prompt box. That field is not there.
A prompt formula that names the mode
Keep one skeleton. Change only the lock.
Mode: [image-to-video | reference-to-video | text-to-video]
Subject: [one object or one person]. No extras.
Action: [one verb].
Camera: [push-in / pan / locked-off]. One move.
Light: [window / overcast / practical].
Audio: [quoted line] or [SFX only].
Refs: still = frame one. <IMAGE_1> = identity. <AUDIO_0> = voice.
Duration / res: [Ns] / [480p|720p|1080p].
Fail if: [second SKU, new face, stretched crop, unreadable type].
For a talking 9:16 take on official reference-to-video:
Mode: reference-to-video. 9:16. 8 seconds. 720p. Do not attach a start frame.
<IMAGE_1> is the speaker. <IMAGE_2> is the navy shirt. <AUDIO_0> is eve.
The person from <IMAGE_1> faces camera in a quiet kitchen and says,
"Two pumps. Wait ten seconds. That is the whole routine."
One slow push-in. Hands stay below frame. No second person. No logo invent.
Draft the spoken line first on the AI script generator. Grok will not invent a legal-safe claim for you.
Creator checklist
- Decide the job: new world, locked still, identity refs, extend, or edit.
- Confirm the surface. xAI API, grok.com Fast, and Runway
grok_imagine_1_5do not share one ID. - Crop the still to the delivery ratio before image-to-video. Official docs will stretch if you override
aspect_ratio. - Count files. Seven images and three voices is the official reference ceiling. Audio files on Runway also need at least one image.
- Do not ask reference-to-video for 1080p. Official cap is 720p.
- Quote the line or write “SFX only.” Audio is on by default.
- Draft at 480p or 720p. Spend 1080p only on an approved T2V or I2V keeper.
- If you need a clip inside ClipCanva today, generate on Seedance 2 or reference-to-video. Compare the same brief later. The Grok vs Seedance page is a research page, not a dual generator.
- If the next beat is longer than 15 seconds, extend a keeper. Do not prompt for a minute.
- Legal review still sits with you. Official pages describe capabilities. They do not license a face, a song, or a competitor logo.
Limits to keep in the brief
- Launch is not the same as your login. 16 June 2026 is API general availability. Fast is a consumer lane. Runway added the model on 7 August 2026. Region and safety filters still apply.
- ClipCanva does not silently add Grok. The compare page is explicit. If that changes, re-read the live page.
- 1080p is mode-gated. Official: text-to-video and image-to-video only. Reference, edit, and extend sit at 720p.
- Text-to-video hides the first frame. If identity matters, generate the still yourself.
- One mode per request. Official: reference-to-video cannot combine with image-to-video or video editing.
- Custom voice files are not generally available. Preset
voice_ids are. Partner audio clones are request-only. - Video URLs expire. Official docs say generated files are temporary. Download keepers.
- Canva is a different job. Eight seconds, 16:9, one clip per prompt, Veo-3. Do not treat it as a Grok front end.
FAQ
What is Grok Imagine Video 1.5?
Grok Imagine Video 1.5 is xAI’s image-and-text video model, generally available on the Imagine API since 16 June 2026 as grok-imagine-video-1.5. Official materials describe same-pass audio, 1–15 second clips, 480p / 720p / 1080p on text-to-video and image-to-video, and a separate reference-to-video path capped at 720p. ClipCanva is a separate product and does not currently run this model.
Can I generate 1080p with references?
Not on the official reference-to-video path. xAI caps that mode at 720p. 1080p is documented for text-to-video and image-to-video only. Runway’s changelog matches the pattern: requests with image references are capped at 720p.
Is Grok Imagine Video 1.5 on ClipCanva?
No. Use the Grok vs Seedance comparison for documented differences, then generate on Seedance 2 or image-to-video if you need a clip here today.
How is Runway’s Grok different from xAI’s API?
Same family, different IDs and wrappers. xAI uses grok-imagine-video-1.5. Runway uses grok_imagine_1_5 and bills in credits (10 / 16 / 29 per second at 480p / 720p / 1080p). Runway image-to-video takes a single first frame. Confirm the live endpoint before you copy a JSON body across vendors.
Should I use Canva’s AI video clip instead?
Use Canva when you need one short 16:9 clip with synchronized audio inside a layout. Canva’s public feature page caps that clip at eight seconds and names Google Veo-3. Use Grok 1.5 when the brief is a 1–15 second shot with a locked still or named references, and you have xAI or Runway access.
Sources
- Grok Imagine Video 1.5, xAI, 16 June 2026
- Video Generation, xAI Docs, updated 18 August 2026
- Image-to-Video, xAI Docs
- Reference-to-Video, xAI Docs, updated 5 August 2026
- API Changelog, Runway, 7 August 2026 Grok entry
- Canva AI Video Generator, Canva