How to Write Cinematic AI Video Prompts: Shot Types, Lighting, and Scene Cards
A practical 2026 guide to cinematic AI video prompts: shot types, lighting, one camera move, and scene cards for Veo, Kling, Runway, Canva, and Luma workflows.
How to Write Cinematic AI Video Prompts: Shot Types, Lighting, and Scene Cards
A cinematic AI video prompt is not a mood paragraph. It is a shot brief: subject, action, setting, one camera move, lighting, style, and a review rule. Write that brief before you generate. Then choose text-to-video, image-to-video, or a storyboard-first path based on what the scene must preserve.
This matters more in 2026 than another model-vs-model argument. Official pages now treat prompting as production control. Luma’s August 13, 2026 prompt guide tells creators to front-load the first 20–30 words, use one camera movement per clip, and keep readable text or logos out of generation. Google DeepMind’s Veo page shows long, shot-level prompts with camera, atmosphere, and native audio. Canva’s AI video generator asks for style, framing, and lighting on clips of up to eight seconds. The common lesson is simple: the model executes a scene card. It does not invent a finished campaign.
ClipCanva is not affiliated with Luma, Google, Canva, Runway, Kling, or Kapwing. Use this guide to plan prompts, then draft the message in the AI Script Generator, save reusable language in Prompt Ideas, generate clips in the AI Video Generator, and lock product identity with Image to Video.
Quick facts for prompt-led video
| Source | Official signal | Prompt implication |
|---|---|---|
| Luma, Aug 13, 2026 | A six-part prompt: subject, action, setting, camera, lighting, style. Front-load the first 20–30 words. One camera move per clip. Do not ask the model to render readable text or logos. | Write camera notes, not adjectives. Split complex movement into separate shots. |
| Luma Scenes, Aug 11, 2026 | Approve storyboard keyframes before render. Ray 3.2 often returns several short clips; Seedance 2 often returns one longer clip. | Review the stills first. Do not pay for a full render until the sequence holds. See Introducing Luma Scenes. |
| Google Veo 3.1 | Official page emphasizes cinematic video, native audio, physics, realism, and prompt adherence. Example prompts include shot type, wardrobe, camera push-in, ambient sound, and spoken lines. | If sound carries the idea, write dialogue, room tone, and SFX in the prompt. Review audio as carefully as picture. |
| Canva AI Video | Create a Video Clip is powered by Google Veo-3. Official FAQ: up to eight seconds, 16:9, synchronized audio including dialogue and sound effects, one video per prompt. | Keep the prompt to one beat. Add captions, end cards, and brand text after generation. |
| Runway | Product page groups Gen-4.5, Multi-Shot, editing, and model routing in one toolkit. | Prompt for a shot you can edit. Leave headroom for text, product callouts, and cuts. |
| Kapwing | Official generator flow: enter a prompt or start frame, review a storyboard, then edit captions, voiceover, and export. | Treat the first prompt as a shot list, not a finished film. |
What a cinematic prompt actually decides
A weak prompt asks the model to “make it cinematic.” A usable prompt decides six things a director would decide on set.
| Prompt part | Decision you must make | Weak language | Stronger language |
|---|---|---|---|
| Subject | Who or what stays recognizable | “a product” | “matte black 350ml water bottle, white wordmark on the left third” |
| Action | One verb the viewer can read | “looking premium” | “condensation slides down the bottle as a hand lifts it” |
| Setting | Place, time, and clutter | “nice kitchen” | “sunlit apartment counter, pale oak, one plant, no extra products” |
| Camera | One move or a locked frame | “dynamic cinematic camera” | “slow push-in, 50mm, locked horizon” |
| Lighting | Source, direction, quality | “beautiful lighting” | “soft window key from camera-left, warm practical lamp in the back” |
| Style | Grain, palette, and what to avoid | “high-end commercial” | “35mm still-life look, muted earth tones, no lens flare, no fake badges” |
If any row is empty, the model fills it with a default look. That default is why so many AI clips feel interchangeable.
Scene cards beat one giant prompt
Do not put a 30-second ad into one paragraph. Write a scene card for each cut. Luma’s August 11 Scenes launch is useful here even if you never open that product: approve the stills, then generate. Kapwing’s official flow does the same thing with a storyboard before credits are spent.
A scene card should answer five questions:
- What should the viewer understand in two seconds?
- What input do I already have: script, still, product photo, or nothing?
- What is the one camera move?
- What must stay true: product shape, face, outfit, room, or claim?
- What fails the shot: extra logos, unreadable UI, distorted hands, invented metrics?
Example card for an ecommerce hook:
Scene 1 / 6 seconds / 9:16
Viewer takeaway: One still product photo can become a launch clip.
Input: approved bottle photo.
Camera: slow push-in only.
Preserve: bottle silhouette, label placement, color.
Audio: quiet fridge hum, then one line: “One still. Three launch angles.”
Avoid: extra bottles, discount badges, melting label, fake 4.9-star UI.
Review: would a shopper still recognize the SKU with sound off?
Turn that card into a generation prompt only after the script line is approved. Draft the hook, voiceover, and CTA first with the AI Script Generator. Save the working prompt pattern in Prompt Ideas so the next SKU does not start from a blank box.
A prompt formula that travels across models
Use the same skeleton for Veo-style, Kling-style, Runway-style, Canva, and Luma clips. Change the job of the scene, not the grammar.
Create a [duration] [aspect ratio] scene.
Subject: [specific object or person, materials, colors, distinctive marks].
Action: [one readable verb].
Setting: [place, time of day, what is not in frame].
Camera: [one shot type + one move or static].
Lighting: [source, direction, hardness].
Style: [film/still-life/doc look, palette, grain].
Audio: [room tone, one SFX, optional spoken line].
Preserve: [product, face, outfit, room geometry].
Avoid: [text, logos, extra products, fake UI, extra limbs].
Original starter for a product still-life:
Create an 8-second vertical scene.
Subject: matte black 350ml water bottle with a white wordmark on the lower third.
Action: a thin ribbon of condensation slides down the metal as a hand lifts the bottle.
Setting: pale oak apartment counter at morning; one plant; no other products.
Camera: medium still-life, slow push-in, 50mm, locked horizon.
Lighting: soft window key from camera-left, warm lamp in the background.
Style: 35mm commercial still, muted earth tones, shallow depth of field.
Audio: quiet room tone, light glass click, voiceover: “One still. Three launch angles.”
Preserve: bottle silhouette, wordmark placement, color.
Avoid: extra logos, discount badges, readable phone UI, distorted fingers.
That prompt is reusable. Swap the subject. Keep the camera, lighting, and avoid list. That is how a prompt library stays useful instead of becoming a pile of one-off adjectives.
Route the prompt by scene job
The same sentence should not be sent to every model. Official positioning is different enough to change the brief.
| Scene job | Better routing signal | Prompt emphasis |
|---|---|---|
| Cinematic realism and native audio | Veo 3.1 / Canva Veo-3 clip | Shot type, wardrobe, camera, spoken line, room tone |
| Continuity across a sequence | Storyboard-first path, then generate | Shared outfit, palette, room, and “preserve” list on every card |
| Product identity | Image-to-video from an approved still | Motion only; lighting inherited from the photo |
| Editable social ad | Runway-style or Kapwing-style finish path | Leave negative space for captions and CTA |
| Eight-second concept test | Canva-style short clip | One beat, one move, no on-screen text |
Use Compare when two models could do the same card. Do not compare “cinematic quality” in the abstract. Compare the same scene card, same duration, same review rule.
If the source is a webinar, tutorial, or long demo, do not prompt from memory. Pull the strongest beats with the AI Video Summarizer, then write new scene cards. If the asset is a still, start in Image to Video and describe motion only.
Shot types and lighting that models actually follow
Use one shot type. Then add lighting as a source, not a vibe.
| Shot type | Use it when | Pairing that usually holds |
|---|---|---|
| Extreme close-up | Texture, condensation, fabric, a button press | Soft side light; static or tiny push-in |
| Close-up | Face, product label, hands on the object | Window key from one side; avoid dual camera moves |
| Medium | Person plus product in context | Slow push-in or static; keep horizon locked |
| Wide | Place the viewer in the room or street | Static or gentle pull-back; do not also orbit |
| Tracking / follow | Subject walks or the bottle is carried | One follow path; no crane at the same time |
| Orbit | 360 product turntable energy | Slow orbit only; clean background |
Lighting recipes that survive generation:
- Soft window key, camera-left, warm fill from a practical lamp
- Overcast daylight, low contrast, no hard rim
- Night interior, one practical as key, dark negative space for captions
- Studio still-life, large soft source above-front, controlled reflection on metal
Do not stack three camera moves and four movie references. Luma’s official guide is blunt on this: one move per clip, and style anchors should be specific. If you need a pan after a push-in, generate two clips and cut them.
Never ask the model to write the price, URL, app UI, or logo lockup. Canva, Kapwing, and any timeline editor will do that more reliably after the clip exists.
Creator/operator checklist
- Write the viewer promise in one sentence before any prompt.
- Split the video into scene cards with one job each.
- Front-load subject, action, and camera in the first 30 words.
- Specify one shot type and one camera move, or lock the camera.
- Name the light source and direction.
- Put spoken lines, SFX, and caption needs in an Audio block.
- Use image-to-video when the product or face must stay recognizable.
- Keep readable text, logos, prices, and UI out of generation.
- Review with sound off: does the shot still communicate?
- Review with sound on: does the line match a claim you can stand behind?
- Save the winning skeleton in Prompt Ideas.
- Compare models on the same card, not on different ideas.
Recommended ClipCanva workflow
- Script the beat. Draft hook, voiceover, and CTA in the AI Script Generator.
- Write scene cards. One card per cut: input, camera, preserve, avoid, review rule.
- Reuse language. Store the working skeleton and variations in Prompt Ideas.
- Generate the right way. Concept shots go to the AI Video Generator. Approved stills go to Image to Video.
- Route if needed. Use Compare when the card needs a different balance of realism, audio, or continuity.
- Repurpose long footage. Summarize first with the AI Video Summarizer, then prompt only the missing scenes.
The goal is not a prettier first take. The goal is a prompt you can rerun on the next SKU, the next avatar, or the next 15-second cutdown without starting over.
FAQ
What is the best structure for an AI video prompt?
Use six parts: subject, action, setting, one camera move, lighting, and style. Add a separate audio block and a preserve/avoid list. Front-load the first 20–30 words with the subject and camera. That structure matches how current official guides describe usable prompts.
How long should a cinematic AI video prompt be?
Long enough to decide the shot, short enough to keep one idea. A dense scene card of 80–160 words is usually clearer than a 400-word mood essay. If you need a second camera move or a new location, start a new clip.
Should I mention a specific model in the prompt?
Only if you are actually generating in that system and the field asks for it. The transferable work is the scene card. Veo-style prompts can carry native audio lines. Image-to-video prompts should describe motion, not rebuild the product. Compare outputs on the same card.
Why do AI video prompts fail on text, hands, and logos?
Readable text and logos are a known weak spot across generators. Hands fail when they become the subject without a still that already shows them clearly. Keep type and logos for the edit. For hands, start from a photo or crop tighter so fingers are not the hero.
How do I keep prompts consistent across a campaign?
Lock the constants: camera grammar, lighting recipe, palette, and avoid list. Swap only the subject and the spoken line. Save those constants as a reusable pattern in Prompt Ideas and generate variants from the same approved still when identity matters.
Sources
- How to Write AI Video Prompts: The Complete Guide, Luma, August 13, 2026
- Introducing Luma Scenes, Luma, August 11, 2026
- Veo 3.1, Google DeepMind
- AI Video Generator, Canva
- Runway product
- Kapwing AI Video Generator