ClipCanva

MiniMax H3 Workflow: First and Last Frames, Native Audio, and 2K Clips

Use MiniMax H3 for 5–15 second 2K clips with native stereo audio. Pick text-to-video or first/last-frame mode, write a beat sheet, and generate on ClipCanva.

August 22, 2026ClipCanva Editorial

MiniMax H3 Workflow: First and Last Frames, Native Audio, and 2K Clips

A MiniMax H3 workflow is not “write a longer prompt.” MiniMax’s 31 July 2026 launch treats H3 as a general-purpose multimodal video model: it reads text, images, video, and audio as one context and returns video with native stereo sound, up to 15 seconds, at 768P or 2K. The official video-generation guide then splits the job into three modes: text-to-video, first/last-frame image-to-video, and reference generation. If you dump nine stills, a motion plate, and a voice take into one box without naming what each file controls, you will spend credits on a recast.

ClipCanva is not affiliated with MiniMax, Hailuo, Canva, Kling, OpenAI, or Google. On ClipCanva, MiniMax H3 is a live generator for text-to-video and image-to-video. Image-to-video accepts one opening still or opening and closing frames. The official API also documents a 12-file reference mode. That mode is not exposed in ClipCanva’s H3 interface yet. Write the packet once. Run the frames you can attach today.

Official facts for MiniMax H3 and nearby tools

Source Public claim Workflow implication
MiniMax H3 launch post (31 Jul 2026) General-purpose multimodal model. Unified context across text, images, video, and audio. Video with native stereo sound, up to 15 seconds at 2K. Consumer hands-on on Hailuo. Quote the spoken line. Silence is a setting. Do not assume a silent render.
MiniMax video-generation guide Three modes: text-to-video; first/last-frame image-to-video; reference generation. Output: 768P / 2K. Duration: 4–15 seconds, integers only. Pick one mode. Do not mix first/last frames with a 12-file reference pack in the same request.
Same API guide First/last-frame entry: 0, 1, or 2 images. Reference entry: ≤9 images, ≤3 videos, ≤3 audio clips, 12 files total. Video/audio clips 2–15s each, combined duration ≤15s. Prompt ≤7,000 characters. Count files. Label every attachment. Leave empty slots empty.
H3 open-source note (3 Aug 2026) Weights released. Same public output envelope: up to 2K, up to 15 seconds, native stereo, wide aspect-ratio set including 21:9 through 9:16. Self-hosting is a legal and hardware decision. Hosted generation still follows the same mode rules.
Canva AI Video Generator Create a Video Clip, powered by Google Veo-3: one 16:9 clip per prompt, up to eight seconds, with synchronized audio. An 8-second layout clip and a 10-second H3 first-to-last-frame shot are different jobs.
ClipCanva MiniMax H3 Live text-to-video and image-to-video. Image-to-video: one opening image or opening and closing frames. Current UI: 5- or 10-second duration, fixed 2K, native audio. Plan on official 4–15s. Generate today at 5s or 10s.

These are vendor-documented capabilities, not a guarantee that your SKU, face, or logo will hold for 15 seconds. Access, region, credits, and exact limits change. Confirm the live Hailuo, MiniMax API, or ClipCanva panel before you promise a client 15 seconds, 2K, or a 12-file reference pack.

Pick the job before you attach frames

H3 fails when you treat every brief as text-to-video with extra attachments. The official guide is explicit: with no image, the first/last-frame entry becomes text-to-video. With no image, video, or audio, reference generation also becomes text-to-video. The mode is inferred from what you attach.

Job Official mode What you must supply What the model is allowed to invent Fail condition
New 2K establishing shot Text-to-video One scene card, one camera move, quoted line or “SFX only” World, motion, and audio bed Extra products, new wardrobe, unreadable type
Animate an approved still First-frame image-to-video One keeper still plus one verb Motion and sound only The still is redesigned
Start here, end there First + last frame Opening still and closing still The interpolation and the audio A third room or a third SKU appears between them
Lock cast + motion + voice Reference generation (API / Hailuo) Named stills, optional motion clip, optional audio take. Cap: 9 / 3 / 3, 12 files total In-between motion Unlabeled zip; a third face appears
8-second layout clip Canva Create a Video Clip (Veo-3) One 16:9 prompt An 8-second shot with synced audio Asking Canva to hold a 15-second first-to-last-frame path

If the brief is “this bottle must not change,” that is first-frame image-to-video from a locked still, not a longer text prompt. If the brief is “open on the closed box, end on the label,” that is first + last frame. If you already have a motion plate and a voice take, that is official reference mode — on Hailuo or the API, not in ClipCanva’s current H3 form.

First-to-last card before you spend the 2K pass

Write the card before you attach the second still. Official first/last-frame copy is simple: the opening or ending frame is controlled; the model fills the path. That only works if both stills share identity.

Job: 10-second 16:9 product shot, first frame to last frame, native audio on.
Viewer takeaway: same bottle, one pump, same table.
Locked identity: approved pack shot. Label text is real, not generated.
Mode: first + last frame. Do not attach extra reference videos.
First frame: closed 30ml frosted bottle, three-quarter, label readable, no hand.
Last frame: same bottle, same table, one bead of product on the back of one hand.
Beats:
  0–3s  static tabletop, room tone, no cut.
  3–7s  one right hand enters, one pump press, click on the press.
  7–10s hand turns, bead sits on skin, end on the last-frame still.
Camera: one slow push-in. No orbit. No second angle.
Audio: pump click + quiet room. Quoted line max 8 words, or none.
Avoid: extra SKUs, extra hands, price overlay, second language on glass.
Review: would a buyer recognize the SKU with the logo cropped out?

If the last frame needs a new city or a new talent, stop. That is a new text-to-video generation, not a first-to-last interpolation.

A prompt formula that travels across t2v, first/last, and reference

Keep one skeleton. Change only the attachment block.

Model: MiniMax H3 (Hailuo 3.0). Duration: [5 | 10 | 15] seconds. Aspect: [16:9 | 9:16].
Mode: [text-to-video | first-frame | first+last | reference].
Audio: [quoted line / SFX only / silence]. Native stereo.

Subject: [SKU or person, materials, colors, distinctive marks].
Locked identity: [which still defines face / label / silhouette].
Story beats: 0–[s] [shot]. [s]–[s] [shot]. End state: [what is in frame].

Attachments:
  First frame: [role, or none].
  Last frame: [role, or none].
  If official reference mode: Images [n/9] @Image 1 = [role].
  Videos [n/3] @Video 1 = [motion or plate].
  Audio [n/3] @Audio 1 = [voice / bed / SFX].
  Total files: [n/12].

Camera: [one path]. Lighting: [source, direction]. Hold it.
Text on screen: [exact words that must stay, or “no generated type”].
Avoid: [extra products, extra limbs, fake UI, new logos, extra languages].

Starter for a 10-second first-to-last product clip (original; not copied from MiniMax demos):

Model: MiniMax H3. 10 seconds, 16:9, photoreal product, native audio on.
Mode: first + last frame.
First frame: approved 30ml frosted glass serum, gold pump, black sans-serif label “NORTH HARBOR 01”, pale oak table, north window, no hand.
Last frame: same bottle, same table, same window, one clear bead on the back of one right hand. Label still readable.
0–3s: 35mm tabletop wide, dust in the light, room tone only.
3–7s: one right hand enters from camera right, one pump press, a click on the press. No second bottle.
7–10s: hand turns palm-up, bead sits on skin, slow 10cm push-in, hold the last frame.
No slogan, no price, no second language on glass.

On ClipCanva, set duration to 10 seconds and attach those two stills. Do not paste a 12-file reference list into a form that only accepts frames.

For Canva’s 8-second Veo clip, cut the same card to one beat. Canva’s public feature page is one 16:9 clip, up to eight seconds, with dialogue and sound effects. It is a hook generator, not a first-to-last interpolation.

Creator / operator checklist

Run this before you pay for 2K.

  1. Name the end frame. If you cannot describe the last still, you are not ready for first+last mode.
  2. Lock identity offline. Export the approved pack shot. Do not ask H3 to invent the label, then “keep it consistent.”
  3. Choose one official mode. Text, first/last frames, or the 12-file reference pack. The API treats empty attachments as text-to-video. Mixing jobs in one prompt is how logos melt.
  4. Count the official budget. 9 images / 3 videos / 3 audio / 12 files. A folder of 40 screenshots will be truncated or ignored.
  5. Quote the spoken line. Launch copy is native stereo: dialogue, ambience, and effects in the same pass. Unspecified audio invents a jingle.
  6. One camera idea per pass. “Slow push-in” is a control. “Cinematic, dynamic, orbit, crane, handheld” is four fights.
  7. Draft length on purpose. Official range is 4–15 whole seconds. ClipCanva’s H3 generator currently exposes 5 and 10. A 15-second beat needs Hailuo, the API, or a later UI.
  8. Do not promise 4K. Official docs list 768P and 2K. Third-party pages that add 4K are not MiniMax’s spec sheet.
  9. Keep reference mode honest. If you need video or audio attachments, use Hailuo or the official API. ClipCanva’s live H3 page is text plus up to two frames.
  10. Compare like-for-like. Use model comparison when you are testing live models. Do not score a 10-second H3 first-to-last shot against an 8-second Canva clip and call one “better.”

How this maps to ClipCanva today

You need Live ClipCanva route H3-specific step
Script and beat sheet AI script generator 0–3 / 3–7 / 7–10 with one verb each
Shot language Prompt ideas Camera + lighting only; no extra SKUs
Approved still → motion Image to video Same still you will attach as the first frame
Talking-head or mouth lock Lip sync video Different tool. H3 native audio is not a dedicated lip-sync endpoint
Connected H3 generator MiniMax H3 Text, or one to two frames, 5s or 10s, fixed 2K
Other live video models AI video generator Use whatever model is actually live
Side-by-side choice Compare Same brief, different model, same review checklist

The packet should survive a model swap. If ClipCanva later exposes official reference uploads, you add labeled files. You do not rewrite the beat sheet.

Nearby models, without fake leaderboards

Need Better default Why
10s first-to-last product path with native audio MiniMax H3 Official first/last-frame mode plus stereo in one pass
Live H3 on ClipCanva MiniMax H3 Text or two frames, 5s/10s, fixed 2K
8s layout-ready clip with dialogue Canva Create a Video Clip (Veo-3) Public cap is eight seconds, 16:9
30s one-take with a large labeled reference pack Seedance 2.5 on official ByteDance surfaces Different duration and a 30/10/10 reference budget
Connected Seedance on ClipCanva Seedance 2.0 Live route; shorter clip; same identity packet
Dedicated mouth lock to an uploaded track Lip sync video H3 can generate speech. It is not a “drop WAV, lock lips” tool

Do not invent list prices, 4K output, or a 30-second H3 mode. Those numbers show up on aggregator pages. They are not in MiniMax’s 31 July launch post or the official video-generation guide.

Limits to keep in the brief

  • Launch is not the same as your account. Official 31 July access is Hailuo plus the MiniMax platform API. Region, login, safety filters, and pay-as-you-go rules still apply.
  • ClipCanva does not silently unlock the 12-file pack. The H3 model page is explicit: video and audio reference uploads are not in this interface. If that changes, re-read the page.
  • Duration is an integer. Official: 4–15 seconds. ClipCanva currently: 5 or 10. A 12.5-second idea is a 10-second shot plus an edit, not a fractional API call.
  • First/last frames set the shape. Official image-to-video follows the still. Do not fight a 9:16 first frame with a 16:9 prompt.
  • Reference files have size and length caps. Official: images ≤30 MB; videos ≤50 MB and 2–15s; audio ≤15 MB and 2–15s; combined video duration ≤15s; combined audio duration ≤15s.
  • Open weights are not a commercial free pass. The 3 August open-source note publishes weights. License, territory, and brand rights still sit with you.
  • Legal review still sits with you. Official pages describe capabilities. They do not license a face, a song, or a competitor logo.

FAQ

What is MiniMax H3?

MiniMax H3 is MiniMax’s general-purpose multimodal video model, launched on 31 July 2026. Official materials describe native stereo audio, 768P or 2K output, 4–15 second clips, and three modes: text-to-video, first/last-frame image-to-video, and multimodal reference generation. The consumer surface is Hailuo. Some pages also call it Hailuo 3.0. ClipCanva is a separate product.

How long can one H3 generation be?

The official API guide says 4–15 seconds, integer values only. ClipCanva’s live H3 generator currently offers 5 or 10 seconds. Multi-minute output is an edit, not one unbounded prompt.

Can I use first and last frames?

Yes, in official first/last-frame mode: one opening still, one closing still, or either alone. On ClipCanva, attach one opening image or both frames. Do not combine those frames with a separate 12-file reference pack in the same request.

Is MiniMax H3 on ClipCanva today?

Yes, as a live generator. Use MiniMax H3 for text-to-video or image-to-video. Confirm the on-page FAQ for the current KIE modes before you promise a client video or audio attachments.

Should I use Canva’s AI video clip instead?

Use Canva when you need one short 16:9 clip with synchronized audio inside a layout. Canva’s public feature page caps that clip at eight seconds and names Google Veo-3. Use MiniMax H3 when the brief is a 5–15 second shot with a locked first or last frame and native stereo. They are not drop-in replacements.

Sources