Seedance 2.0 on Runway: 2026 Prompt Guide
Use Seedance 2.0 inside Runway with text, image, video, and audio references. Learn input limits, prompt structure, settings, and review steps.

Seedance 2.0 is available inside Runway as a third-party video model. It accepts combinations of text, image, video, and audio inputs, so the most important choice happens before prompting: decide whether you need open-ended generation, a controlled start and end, or several references with clear roles.
Runway documents clips of 5–15 seconds and three creation modes: References, Start / End frames, and Text to Video. Exact availability, resolution, credits, and settings can depend on the current account and product surface. Check the live interface before planning a large production batch.
What Seedance 2.0 on Runway supports
Runway's official Seedance 2.0 guide describes support for text, image, video, and audio inputs. It lists up to nine images, three videos, and three audio files in the relevant reference workflow, with the combined number of files limited by the current product rules. It also documents 480p, 720p, and 1080p output options.
ByteDance presents the underlying model on the official Seedance 2.0 page. Use Runway's help article for the controls available inside Runway and ByteDance's page for model context. Do not assume that every capability shown by the model creator appears identically in a third-party interface.
The model can combine several media types, but more inputs do not automatically create more control. Every reference needs a specific job. Remove any file whose role cannot be stated in one sentence.
Choose the right creation mode
| Mode | Use it when | Prompt focus | Main risk |
|---|---|---|---|
| References | Several images, clips, or audio files must influence one result | Assign a role to each reference and explain their relationship | Conflicting style, identity, motion, or timing cues |
| Start / End frames | The opening and closing composition matter | Describe the transition, action, and camera path between frames | Incompatible frames causing drift in the middle |
| Text to Video | You need to explore a scene without prepared assets | Describe subject, action, camera, environment, timing, and sound | Too much visual freedom and weaker continuity |
Use Reference to Video when product identity, character appearance, visual style, motion, or sound needs an explicit source. Use a start and end pair when the transition itself is the idea. Use text only when exploration is more valuable than preserving an approved asset.
Prepare references before prompting
Clean inputs reduce contradictions. Crop images to the intended composition, remove accidental overlays, and use the highest practical source quality. A product reference should show the shape and details that must remain stable. A style reference should communicate light, color, texture, and framing without introducing a competing subject.
Trim video references to the movement you actually want. A long clip containing several actions creates an unclear instruction. Use audio references with a defined role such as speech rhythm, ambience, a sound effect, or music timing. Confirm that you have permission to use every uploaded asset.
Name the role of each file in your working notes:
- Image one establishes the main subject and product shape.
- Image two establishes the location and lighting.
- Video one supplies the camera movement.
- Audio one supplies timing and ambience.
If two files define different faces, products, camera directions, or lighting, resolve the conflict before generation.
A reusable prompt structure
Start with the job of the clip, then describe the visible sequence:
Purpose: a short product introduction for a vertical social post
Subject: preserve the product shape, material, color, and label position from image one
Action: a hand places the product on the table, then turns it slightly toward camera
Camera: slow push-in, stable horizon, no sudden zoom
Environment: warm morning kitchen light, uncluttered background
Reference roles: image one is the subject; image two is lighting only; video one is camera motion
Audio: use the rhythm of audio one; keep dialogue clear
End state: product centered with clean space above for a caption
Avoid: extra objects, altered packaging, invented text, badges, or logos
The prompt should explain change over time. Do not waste most of it repeating details already visible in the source images. State what must remain fixed, what should move, how the camera behaves, and how the shot ends.
Prompting References mode
References mode works best when the prompt assigns each asset a narrow responsibility. Separate subject, environment, motion, and audio roles. When one image should influence only lighting or style, say so. When a video supplies only camera movement, do not also ask it to replace the subject.
Begin with two or three essential files. Generate a test and inspect which characteristics transferred. Add another reference only when it solves a named failure. This is faster than uploading the maximum number of files and trying to diagnose a mixed result.
For a character scene, keep identity references visually consistent. For a product scene, use angles that agree on shape, color, and packaging. For a style scene, avoid references with prominent people or objects unless you want those elements to influence the result.
Prompting Start / End frames
The opening and closing frames should look like two moments from the same shot. Match aspect ratio, subject identity, scale, wardrobe, product details, lighting direction, and major background geometry. Large contradictions force the model to invent an explanation.
Describe the path between frames instead of describing each still. Name the subject action, camera movement, pace, and any important intermediate beat. Keep the sequence achievable within 5–15 seconds. A complex transformation with several unrelated actions is better split into separate clips.
Review the middle frames carefully. A convincing first and last image can hide identity changes, warped objects, or a physically impossible transition.
Prompting Text to Video
Text to Video gives the model the most freedom. Use it for mood, location, composition, and early motion exploration before spending time on exact references. A strong prompt contains a clear subject, one main action, one camera behavior, environment, light, pace, audio intention, and final composition.
Avoid long lists of adjectives and multiple scene changes. A short clip cannot reliably contain an establishing shot, dialogue exchange, transformation, product close-up, and final logo reveal. Turn that idea into a shot list and generate each shot separately.
When a text result establishes the right visual direction, save a frame and use it as a controlled input for the next iteration.
Settings and test strategy
Choose duration based on the action, not the maximum allowed length. A simple camera push or product turn may need only a short clip. A longer duration creates more frames in which identity and geometry can drift. Select resolution according to the testing stage and final delivery need; do not spend final-output credits while the prompt is still unstable.
Use the AI Video Generator workflow to keep the prompt, references, settings, and output together. Generate a small matrix:
- Baseline with the minimum essential references.
- One prompt revision that clarifies motion or continuity.
- One reference revision that removes a conflicting asset.
- A higher-resolution generation only after the direction passes review.
Change one meaningful variable at a time. If the prompt, reference set, duration, and resolution all change, you will not know what fixed or caused the result.
Review the result
Watch at normal speed for message clarity, then scrub frame by frame. Check subject identity, hands, faces, product geometry, labels, reflections, background continuity, and the relationship between motion and sound. Verify that the ending frame leaves usable space for captions or a call to action.
Do not rely on generated footage for exact prices, claims, legal statements, certifications, interface text, or packaging copy. Add factual text and brand elements in an editor. Confirm source rights, likeness permissions, audio rights, provider terms, and destination-channel policies before publishing.
Keep rejected outputs long enough to record why they failed. A short review note such as “product label changed during camera move” is more useful for the next prompt than “looks wrong.”
Common problems and fixes
| Problem | Likely cause | Next test |
|---|---|---|
| Subject identity changes | Conflicting subject references or too much motion | Use one primary identity reference and simplify the action |
| Product shape bends | Complex camera move or weak product reference | Use a clearer angle and a slower, narrower movement |
| Style is inconsistent | Several references compete for visual direction | Keep one style source and state that its role is lighting and color |
| Start and end look right but middle fails | Frames are too different or transition is overloaded | Align the frames and request one achievable transition |
| Audio feels unrelated | Audio role is not stated or timing conflicts with action | Name the audio role and simplify the visible sequence |
| Caption area disappears | Composition was not specified | Reserve negative space in the prompt and source frame |
FAQ
Is Seedance 2.0 really available on Runway?
Runway publishes a dedicated guide for creating with Seedance 2.0 inside its product. Availability and settings may still depend on plan, account, region, and the current interface.
How many references can I use?
Runway documents up to nine images, three videos, and three audio files in the relevant workflow, subject to its combined file limit and current product rules. Start with fewer files and give each one a clear role.
Which mode should I choose?
Use References for several controlled assets, Start / End frames for a planned transition, and Text to Video for open-ended scene exploration.
What duration does Seedance 2.0 support on Runway?
Runway documents 5–15 seconds. Choose the shortest duration that comfortably contains one clear action.
Should I generate at 1080p immediately?
Not during early prompting. Prove the composition, motion, identity, and reference roles first, then use the output setting appropriate for final delivery.
Can Seedance render exact text and product claims?
Do not trust generated video for factual or exact text. Add prices, claims, labels, captions, logos, and calls to action as editable layers after generation.