ClipCanva

Gemini Omni 1.1 Flash Workflow: First and Last Frames, 40-Second Scene Extension, and 4K Upscale

Gemini Omni 1.1 Flash adds first/last-frame interpolation, 10-second scene extension up to 40 seconds, 360p drafts, and 4K upscale. Operator workflow, limits, and ClipCanva handoff.

August 28, 2026ClipCanva Editorial

Gemini Omni 1.1 Flash Workflow: First and Last Frames, 40-Second Scene Extension, and 4K Upscale

A Gemini Omni 1.1 Flash workflow is not “chat until the clip feels cinematic.” Google’s 27 August 2026 developer post ships the controls the earlier Omni Flash API did not: first-and-last-frame interpolation, scene extension that reads up to 10 seconds of prior footage, 10-second extend steps up to a 40-second cumulative clip, 360p drafts, and 1080p/4K upscales. The Gemini API docs (updated the same day) pin the live model id as gemini-omni-1.1-flash. ClipCanva is not affiliated with Google, Canva, Adobe, Figma, or Runway, and Omni 1.1 is not a live ClipCanva generator. Lock the stills and the spoken line here, then send the packet to Google AI Studio or the Gemini API. For a clip you can run on ClipCanva today, use Veo 3.1 or image-to-video.

Official facts for Omni 1.1 Flash and nearby tools

Source Public claim Workflow implication
Gemini Omni 1.1 Flash (27 Aug 2026) Production-ready on the Gemini API in Google AI Studio. Scene extension analyzes up to 10 seconds of prior context (previous models used the last second). Extend in 10-second increments to a 40-second cumulative length. First and last frames for orbits, zooms, and loops. 360p drafts: up to 60% faster and about one third the 720p cost. Upscale to 1080p or 4K. Video references up to three seconds. Adobe Firefly, Figma Weave, GMI Cloud, and Runway are named as production surfaces. Treat 40 seconds as stacked 10-second tails, not one 40-second render. 4K is an upscale path, not the default generate.
Generate and edit videos with Gemini Omni Flash (updated 27 Aug 2026) Model id gemini-omni-1.1-flash. Default output 720p. Resolutions: 360p, 720p, 1080p (upscaled), 4k (upscaled). Aspect 16:9 or 9:16. First/last interpolation: two images in input plus a transition prompt. Extension: 3–10 second continuation, end-of-clip only. Uploaded input for edit/extend must be ≤10 seconds. Video refs: max 3 clips, ≤3 seconds each; audio on those refs is ignored. Uploaded-video edit/extend is not available in EEA, Switzerland, and the UK. You cannot add extra dialogue when extending an uploaded talking clip. Voice editing and uploaded audio refs are unsupported. All outputs carry SynthID. Name the job. A start still is not a 3-second video ref. Do not prepend.
Gemini API pricing (updated 27 Aug 2026) Paid-tier Omni Flash: input $1.50 / 1M tokens (text/image/video/audio); text output $9.00 / 1M; video output $17.50 / 1M. 720p video is billed at 5,792 tokens per second, about $0.10 per second. Free tier is not available for this model. Price the second and the resolution. Draft at 360p. Do not 4K every take.
Veo 3.1 Ingredients to Video (13 Jan 2026) Veo 3.1 stills-to-clip plus native 9:16 for Ingredients to Video. Separate Google product from Omni. A Veo Ingredients brief is not an Omni 40-second extend stack.
Canva AI Video Generator Create a Video Clip, powered by Google Veo-3: one 16:9 clip per prompt, up to eight seconds, with synchronized audio. An 8-second layout clip is not first/last interpolation and not a 40-second Omni tail.

These are vendor-documented limits, not a promise that your SKU, face, or quoted line will hold across four extend steps. Access, region, safety filters, and credits still sit on the live Google project.

Pick the job before you attach files

Omni 1.1 fails when every brief is treated as one chat box. The 27 August controls are four different desks.

Job Where it actually runs What you must supply What the model may invent Fail condition
Animate between two approved stills Gemini API gemini-omni-1.1-flash, two images + transition prompt First frame and last frame you have rights to, plus one camera path In-between motion and audio A third product, a new face, a jump cut you did not ask for
Continue a keeper Multi-turn previous_interaction_id, or Files API upload + “Continue the scene” A clip ≤10s if uploaded; one new beat The next 3–10 seconds at the tail Prepend, mid-clip insert, extra dialogue on an uploaded talking take
Stack a longer story Repeat 10-second extends Cumulative cap 40 seconds Continuity from up to 10s of prior context Treating 40s as one prompt
Cheap iteration resolution: "360p" One variable per take A draft, not a master Shipping 360p as the client file
Delivery master Generate 720p, then 1080p or 4K upscale The approved 720p take Sharper pixels, same shot 4K-upscaling a still whose label already failed
Hands-on ClipCanva take Veo 3.1, image-to-video, or Seedance 2.5 The same scene card on a live model Whatever that model documents Pasting gemini-omni-1.1-flash into a ClipCanva box
8-second layout clip Canva Create a Video Clip (Veo-3) One 16:9 prompt An 8-second shot with synced audio Asking Canva for a 40-second Omni extend

If the brief is “this 30ml bottle must start here and end there,” lock two stills first. Approve both crops. Then interpolate. If the brief is “keep talking after the keeper,” extend only at the tail, on a model-generated clip if you need new speech. Do not upload a 25-second talker and ask Omni to add a sentence — the docs block extra dialogue on uploaded talking footage.

Delivery card before you pay $0.10 a second

Write the card before you hit generate. Official 720p video output is about $0.10 per second. A 10-second 720p take is about $1.00 of video output before input tokens. 360p is documented as roughly a third of that 720p cost.

Job: [first-last interpolate | extend tail | 360p draft | 720p keeper | 1080p/4K upscale]
Surface: [AI Studio | Gemini API gemini-omni-1.1-flash | Firefly | Runway | Flow]
Model: gemini-omni-1.1-flash
Resolution path: 360p draft → 720p keeper → upscale only if the take is approved
Aspect: 16:9 or 9:16. Do not crop after the fact if you can set it now.
Duration: one 3–10s generate, or 10s extend steps, cumulative ≤40s.
Identity lock: [SKU / face / wardrobe / quoted line]
Stills: first.jpg = frame one. last.jpg = frame last. Refs are not frame one.
Extend: end of clip only. No prepend. No mid-roll insert.
Region: if the source is an upload, skip EEA/CH/UK for extend/edit.
Success: same object, readable type, camera path matches the card, audio lands on the action.

Starter for a first-to-last product move (original; not copied from Google samples):

Two stills. First frame: 30ml frosted bottle, gold pump, black label “NORTH HARBOR 01”, 3/4 on pale oak, north-window key.
Last frame: same bottle, same label, same table, camera has dollied to a tight label crop. No second bottle.
One continuous shot, no jump cuts. Slow push-in, one pump press in the middle of the move.
Audio: pump click + quiet room. No dialogue. No music sting.
Fail if the label becomes glyphs, a second language appears, or the last frame does not match the still.

Build those stills on ClipCanva with GPT Image 2 or collect lines on prompt ideas. Write the spoken beat, if you need one, on the AI script generator. Then move the files. Do not paste resolution: "4k" into a ClipCanva prompt. That field is not there.

First and last frames are a path, not two moodboards

Google’s interpolation path is two images plus a sentence about the move. It is not a pile of style refs. The API example is explicit: first image, last image, then a transition prompt. If you need identity without locking frame one, that is a different job — subject refs or ≤3-second video refs — and those refs are not the start still.

Keep one skeleton:

First frame: [file]. Last frame: [file]. Same SKU. Same wardrobe.
Move: [push-in / orbit / whip-pan / loop back to frame one].
Shot: one continuous take, no jump cuts.
Audio: [quoted line] or [SFX only].
Fail if: [new face, extra product, unreadable type, extra language].

A loop is just last frame = first frame, plus “seamless loop, no jump cuts.” If the stills disagree on the label, stop. Interpolating two broken packs will not invent a third correct pack.

For a talking 9:16 take, freeze the stills, quote the line, then interpolate or image-to-video. Do not ask Omni to invent a new spokesperson mid-orbit.

Scene extension is a tail, stacked to 40 seconds

The 27 August post is specific: Omni 1.1 reads up to 10 seconds of what already happened, then writes the next beat. Docs add the operator rules: continuation is 3–10 seconds, end-of-clip only, uploaded source ≤10 seconds, and uploaded extend/edit is blocked in EEA, Switzerland, and the UK. Model-generated multi-turn extend is the lane that still supports new speech.

Work it as a stack:

  1. Approve a 6–10 second keeper at 360p, then regenerate that same card at 720p.
  2. Extend once: one new action, one camera note, “Keep everything else the same.”
  3. Watch the join. If the bottle redesigns at second 11, do not stack step 4.
  4. Repeat only while the cumulative length stays ≤40 seconds.

Do not write a 40-second screenplay into the first prompt. Omni’s default generate still wants a short clip; length comes from tails. Do not ask it to insert a beat in the middle. Do not upload a talker from another tool and prompt extra dialogue — that path is documented as unsupported.

If you need a longer native take without stacking Omni tails, that is a different model. Seedance 2.5 is the ClipCanva page for longer reference-led clips. Veo 3.1 remains the live Google-family generator on ClipCanva. Pick one desk per shot.

Resolution path: 360p, 720p, then upscale

Google’s 1.1 post splits draft and master on purpose. 360p is for side-by-side tests — change one verb, one light, one camera path. 720p is the default generate. 1080p and 4K are upscaled outputs, not a different story model.

A practical spend card for one product shot:

Step Resolution Why
3–4 variants 360p About one third of 720p cost; up to 60% faster throughput
1 keeper 720p Documented default; ~$0.10/s video output
Delivery 1080p or 4K upscale Only after the 720p take survives the label crop

URI delivery is required once the file is larger than 4MB (docs: typical above 720p). If you need another conversational edit after the master, do not set store=false on the keeper — that flag drops the interaction you would extend from.

Creator / operator checklist

  • [ ] Model id is gemini-omni-1.1-flash, not the old gemini-omni-flash-preview card from June.
  • [ ] Job is one of: interpolate two stills, extend the tail, edit one thing, or upscale an approved take.
  • [ ] First and last stills match on SKU, face, and type. Cropped before upload.
  • [ ] Aspect is set (16:9 or 9:16) instead of cropping a landscape master into a Reel.
  • [ ] Drafts ran at 360p. Only the keeper paid 720p.
  • [ ] Extend prompts say “Continue the scene” plus one beat. End of clip only.
  • [ ] Cumulative length is counted. Stop at 40 seconds.
  • [ ] Uploaded talking clips are not asked for new dialogue.
  • [ ] EEA/CH/UK operators are not extending uploaded footage.
  • [ ] Audio is prompted (room, SFX, quoted line). Uploaded audio refs are not used.
  • [ ] Video refs, if any, are ≤3 clips, ≤3 seconds, likeness only. Their audio is ignored.
  • [ ] Spoken lines were drafted on the script generator and quoted exactly.
  • [ ] Hands-on ClipCanva fallback is queued: AI video generator or image-to-video.
  • [ ] Client deck does not call Omni a ClipCanva model.

FAQ

Does Gemini Omni 1.1 Flash replace Veo 3.1? No. They are separate Google products. Omni 1.1 is the Gemini API / AI Studio control surface with interpolation, 40-second stacked extends, and conversational edits. Veo 3.1 remains the Ingredients-to-video / Flow / Gemini-app video model, and it is the Google-family generator you can open on ClipCanva’s Veo 3.1 page.

Can I render a 40-second clip in one prompt? Not as a single generate. Official language is 10-second extend increments to a 40-second cumulative length, with each continuation in the 3–10 second band. Plan four tails, not one screenplay.

Is 4K native generation? The 27 August post and the API table describe 1080p and 4K as upscaled outputs. Default generate is 720p. Draft at 360p. Upscale after the take is approved.

Can I extend footage I shot on a phone? Only if the upload is ≤10 seconds, you are outside EEA/Switzerland/UK, you append at the tail, and you do not add new speech on a talking clip. Multi-turn extend on a model-generated keeper is the cleaner lane for new dialogue.

Is Omni 1.1 on ClipCanva? No. ClipCanva is independent. Use image-to-video or Veo 3.1 here, then take the stills and scene card to Google AI Studio if you have Omni 1.1 access.

Google will keep moving the model card. Re-read the live Omni docs before you quote duration, region, or price in a client estimate. The 27 August 1.1 controls are real; they are still not a 4K SLA on your label.