Seedance 2.5 Prompt Guide: Direct Better 30-Second Videos
Learn how to write Seedance 2.5 prompts for text-to-video, first and last frames, multimodal references, timed scenes, camera movement, and native audio.

A useful Seedance 2.5 prompt does more than describe a beautiful frame. It gives the model a job to complete over time: who or what appears, what changes, how the camera observes it, what the scene sounds like, and where the clip should finish.
The shortest reusable structure is:
Goal + subject + visible action + setting + camera + look + timing + audio + ending state
You will not need every field in every prompt. A simple eight-second product turn may fit in three sentences. A 30-second scene with several beats needs a timeline. Reference media changes the prompt again: instead of describing everything from scratch, you must explain what each uploaded image, video, or audio file should control.
That is the focus of this Seedance 2.5 prompt guide. It covers the three workflows currently available in Videoify's Seedance 2.5 workspace: text-to-video, image-to-video with an optional last frame, and reference-to-video with image, video, and audio inputs. The templates are original starting points, not promises of identical results. Treat each generation as a draft that must be watched, checked, and revised.
Choose the workflow before you write
The prompt cannot compensate for choosing the wrong kind of input. Start by deciding what the model must invent and what you can provide.
| Videoify workflow | Start here when | Your prompt should spend its detail on |
|---|---|---|
| Text to video | You have no required starting image | Subject, environment, action, shot plan, look, sound, and final composition |
| Image to video | You have an approved opening frame and perhaps a required last frame | Motion, camera behavior, the path between frames, and details that must remain stable |
| Reference to video | Identity, movement, setting, composition, or rhythm must come from existing media | The role of each file, the new scene to create, what to preserve, and what not to transfer |
Videoify lets you set a whole-second duration from 1 to 30 seconds and choose 480p, 720p, or 1080p. Text and reference workflows offer adaptive, 16:9, 4:3, 1:1, 3:4, 9:16, and 21:9 framing. Image-to-video follows the shape of its required first frame, so prepare that image in the intended aspect ratio before uploading it.
Set those output choices in the interface. Use the prompt to direct what happens inside the selected frame and duration.
Build the prompt in layers
Seedance 2.5 does not require a magic phrase or rigid syntax. The model's official examples use ordinary production language: observable action, camera direction, references, time ranges, sound, and an ending. A layered brief makes those decisions easier to inspect before generation.
Give the clip one job
Write a one-line objective before you add visual detail:
-
Reveal a travel bag, demonstrate how it opens, and finish on a clean product frame.
-
Follow a chef from the prep counter to the table in one continuous move.
-
Use an approved apartment entrance and window view to create a walkthrough between them.
-
Keep a presenter consistent while borrowing the pacing of a reference video.
The objective is not necessarily part of the final prompt. It is an editing filter. If a detail does not help the clip complete that job, it may be noise.
Describe action as a sequence, not a mood
“A mysterious woman at a station” is a still image idea. “A woman in a charcoal coat crosses the empty platform, notices an abandoned suitcase, slows down, and reaches for its handle” supplies events that can be staged.
Use verbs the viewer can see. Put them in playback order. When two actions must overlap, say so: “As the door opens, the camera moves past her shoulder.” When the clip must pause on a useful frame, name that hold rather than hoping the motion stops in the right place.
Fix the spatial relationships that matter
Longer prompts often describe several subjects without saying where they are. Give the model a simple blocking plan:
The courier stays in the foreground on the right. The market stall remains behind her on the left. The camera tracks parallel to the aisle and never crosses to the opposite side.
You do not need to map every prop. Lock only the relationships required for continuity, product accuracy, or a planned transition.
Direct the camera with a destination
“Dynamic camera” leaves both movement and framing unresolved. A stronger instruction names the shot, movement, speed, and destination:
Begin in a waist-high medium shot. Track backward at walking speed as the courier approaches, then settle into a close-up when she stops beneath the blue awning.
Use one dominant movement per shot. If the scene contains multiple shots, give each one its own camera behavior. Avoid combining instructions that fight each other, such as a locked tripod and handheld shake in the same beat, or a continuous take followed by several cuts.
Describe light through a source
Words such as “cinematic” and “premium” are too broad to define a look by themselves. Name a practical source, direction, palette, or material response:
Cool predawn window light from camera left, one warm desk lamp behind the subject, soft reflections on brushed metal, restrained contrast, natural skin color.
That gives the model physical relationships to maintain. Add style labels only after the scene is understandable.
Treat audio as part of the action
Seedance 2.5 is an audio-video joint generation model. Write sound cues where they occur instead of placing “cinematic audio” at the end.
The latch clicks as the case opens. Room tone stays quiet. A low musical pulse begins only when the watch face lights up and resolves before the final hold.
For dialogue, identify the speaker, exact line, delivery, and timing. Keep the wording plausible for the duration. Generated speech, pronunciation, timing, and lip movement still need review; an audio reference should be treated as guidance, not as a guaranteed dubbing track.
End on an edit-ready state
The last seconds should have a purpose. Ask for the camera to settle, a gesture to finish, or the important object to reach a readable angle. If the result will later carry typography, leave clean negative space and plan to add precise text in an editor rather than asking the video model to redraw a long slogan while the camera moves.
Use timecodes when the scene has real beats
Seedance 2.5 can generate up to 30 seconds in one pass, according to ByteDance's official model page and launch material. The extra duration is useful only when the prompt gives it somewhere to go. Do not stretch a single five-second action across half a minute just because the maximum is available.
For a scene with several events, divide the selected duration into a small number of beats. Each range should answer three questions: what changes, where the camera is, and what sound marks the moment.
0–5s: Establish the subject and space. Keep the camera simple.
5–13s: Begin the main action and move closer to the important detail.
13–22s: Show the consequence, reveal, or change in direction.
22–30s: Resolve the action and settle on the final composition.
This is a planning pattern, not a compulsory four-shot formula. A continuous tracking shot can still use time ranges to mark changes within one move. A quiet ten-second clip may not need timestamps at all.
Here is a complete 20-second example:
20 seconds, 16:9, one continuous take. A watchmaker in a dark green apron works alone at a walnut bench just before sunrise. Keep his round glasses, grey hair, apron, and the unfinished silver watch consistent throughout.
0–5s: Begin with a still medium-wide view from the end of the bench. Cool window light outlines the tools while one warm desk lamp illuminates his hands. He selects a tiny brass gear with tweezers.
5–12s: The camera slowly pushes along the bench into a close view as he places the gear inside the open watch movement. Keep the tool positions and watch geometry stable.
12–17s: He turns the crown once. The mechanism begins moving, and he pauses to listen. Rack focus from the turning gears to his eyes.
17–20s: He closes the case, places the finished watch in the center of a clean leather pad, and withdraws his hands. Settle into a locked overhead product frame.
Audio: quiet room tone, delicate metal contact, one precise mechanical click, and a faint watch tick at the end. No dialogue, no background music, no subtitles, no extra hands, no abrupt camera movement.
Copy-ready prompts for each Videoify workflow
Replace the bracketed fields and delete anything that does not affect the result. These templates are scaffolding, not syntax requirements.
Text-to-video template
Use this when Seedance 2.5 must invent the complete visual scene.
[Duration and aspect ratio]. [Single take or number of connected shots].
[Subject with stable identifying details] is in [specific setting].
[Subject] [performs visible actions in playback order].
Camera: [starting shot and position], then [one movement with speed and destination].
Look: [practical light source, palette, texture, and realism or animation style].
Timing: [optional time ranges for the important beats].
Audio: [ambience, synchronized effects, dialogue, and/or music behavior].
End with [stable final action or composition].
Preserve [critical identity, product, or environment details]. Avoid [specific unwanted behavior].
Example:
12 seconds, 9:16, one continuous handheld shot. A florist in a faded blue apron carries a bundle of yellow tulips through a narrow shop at opening time. She sets the flowers on the counter, unwraps the brown paper, and turns the stems toward the window. The camera follows from behind at shoulder height, moves around her left side, and settles into a close view of the flowers. Soft overcast window light, natural color, quiet documentary texture. Audio: paper rustle, bucket water, distant street traffic, no music. End with her hands leaving the frame while the tulips remain centered and still. Keep her apron, hairstyle, counter layout, and flower color consistent. No subtitles or extra customers.
Image-to-video template
The first image already establishes appearance and composition. Describe the change rather than narrating every visible detail back to the model.
Begin exactly from the uploaded first frame. [Subject or object] [starts with subtle motion], then [main action].
The camera [movement, speed, and destination].
Keep [identity, geometry, wardrobe, product, background, and light details] unchanged.
[If a last frame is supplied:] Move continuously toward the uploaded last frame by [observable intermediate actions]. Match its [pose, object position, camera angle, or composition] at the end.
Audio: [sound tied to the action]. End with [stable state].
Example with first and last frames:
Begin exactly from the uploaded first frame of the closed glass door and finish on the uploaded last frame looking out from the balcony. The door handle turns, the door opens inward, and the curtain lifts gently in the incoming air. The camera moves forward at walking speed through the doorway without cutting, rises slightly over the threshold, and settles on the balcony view. Preserve the room layout, floor pattern, door proportions, curtain color, daylight direction, and furniture positions. Match the horizon and final camera height of the last frame. Audio: handle click, door movement, light fabric, and distant city ambience. No people, no sudden zoom, no change in weather.
If the first and last frames disagree on perspective, lighting, room layout, or object identity, the prompt must solve several discontinuities at once. Choose more compatible endpoints or generate the transition as separate clips.
Reference-to-video template
Videoify's reference workflow accepts up to 30 images, 10 videos, and 10 audio files. Capacity is not a target. Start with the smallest coherent set that establishes the subject, environment, motion, or sound you actually need.
Use Image 1 for [character or product identity].
Use Image 2 for [wardrobe, location, composition, or material].
Use Video 1 only for [movement, performance timing, or camera path].
Use Audio 1 for [rhythm, ambience, or vocal quality].
Create a new [duration]-second [aspect ratio] video in which [subject] [action] in [setting].
Timing: [optional sequence of beats].
Camera: [shot plan]. Look: [lighting, palette, texture, and style].
Preserve [details] from [named reference]. Do not transfer [unwanted subject, setting, wardrobe, or sound] from [named reference].
Audio: [how the reference and generated sound should function].
End with [stable final state].
Example:
Use Image 1 for the cyclist's face and short dark hair. Use Image 2 for the mustard rain jacket and black commuter bicycle. Use Image 3 for the narrow brick underpass and wet pavement. Use Video 1 only for the low side-tracking camera and natural pedaling speed. Use Audio 1 for percussion rhythm, not for vocals.
Create a new 15-second 21:9 video. The cyclist enters the underpass at steady speed, passes through bands of reflected streetlight, and emerges into pale morning light. The camera tracks beside the bicycle at wheel height, then gradually pulls back to show the exit. Keep the cyclist's face, hair, jacket, bicycle frame, and direction of travel consistent. Do not copy the rider, location, or signage from Video 1. Align the passing light bands with the percussion accents. Add tire spray and tunnel ambience; no dialogue. End on a wide rear three-quarter view as the cyclist rides into the bright street.
Give every reference one responsibility
References help only when the prompt makes their hierarchy clear. A useful assignment looks like this:
| Reference | Good responsibility | What to clarify |
|---|---|---|
| Character image | Face, hair, age, wardrobe | Which traits must remain unchanged |
| Product image | Shape, material, color, packaging | Which angle or markings matter most |
| Location image | Layout, architecture, palette | Whether people and objects in the source should transfer |
| Video | Motion, acting, pacing, or camera path | What to copy and what subject or setting to ignore |
| Audio | Beat, ambience, timing, or vocal character | Whether it guides rhythm or should be heard in the new scene |
Conflicting references create an unresolved production decision. If five images show different jackets, hairstyles, and lighting, the model must choose among them. Reduce the set, name one primary identity source, and state which secondary references supply only pose, location, or style.
Use constraints to protect the brief
A closing constraint sentence is useful when it protects a real requirement. Keep it short and specific.
Weak:
No bad quality, no mistakes, no ugly motion, no artifacts, no weird things.
Useful:
Keep the same face, green coat, and bicycle in every shot. No subtitles, no background crowd, no camera roll, and no transfer of the rider from Video 1.
Positive direction should still carry most of the prompt. A long blacklist can compete with the scene itself. If the instructions contain contradictions, remove the conflict instead of adding more negative language.
Diagnose the first visible failure
Do not rewrite the entire prompt after one weak result. Watch the clip with sound and identify the first point where it leaves the brief.
| What happens | Likely prompt problem | Change for the next run |
|---|---|---|
| The opening is good but the ending drifts | The final state is missing or the last beat contains too many actions | Reserve time for the ending and describe the final composition explicitly |
| Actions feel rushed | Too many events are packed into the chosen duration | Remove a secondary event or lengthen the beat that carries the main action |
| The camera invents cuts | Shot structure is ambiguous | State “one continuous take” or specify the exact shot count and time ranges |
| A reference changes the wrong feature | Its role is not limited | Say what the file controls and what must not transfer from it |
| Character or product identity changes | References disagree or fixed traits are vague | Use a smaller consistent set and repeat the critical traits at major beats |
| Audio lands late | Sound is detached from the event | Tie the cue to an action or time range |
| The frame becomes crowded | Too many subjects and props have equal priority | Return to one clip objective and remove anything that does not serve it |
| Text is unstable | Detailed typography is moving or repeatedly redrawn | Reduce generated text and add precise titles or labels later in an editor |
Change one control layer at a time. A practical order is references and identity first, then action and timing, then camera, then look and audio. If you change everything at once, a better result will not tell you which instruction helped.
A final check before generating
Read the prompt once without looking at your source files. You should be able to answer:
-
What is the clip supposed to accomplish?
-
What visibly happens, and in what order?
-
Does the action fit the selected duration?
-
Where does the camera begin, move, and stop?
-
Which visual relationships must remain stable?
-
What should the scene sound like, and when do key sounds occur?
-
What does the final frame look like?
-
If references are attached, does each file have one named responsibility?
-
Are any instructions asking for opposite things?
If those answers are clear, the prompt is probably ready for a controlled first run. Open the Seedance 2.5 generator on Videoify, choose the workflow that matches your inputs, and keep a record of the prompt, references, duration, framing, and the one change you make between attempts.