NewSeedance 2.5 is live · create 4-to-30-second videos with native audio
Try Now

Best Seedance 2.5 Prompts: 14 Real Examples With Videos

A
Ananay Batra
24 min read
A glowing prompt card streaming light into AI-generated video frames of a concert stage, a serum bottle and a woodworking shop

TL;DR

Most Seedance 2.5 prompt lists are adjectives. Cinematic. Hyper-detailed. 8K. Award-winning lighting. None of them show you the clip that came out. That's the only evidence that matters.

Most Seedance 2.5 prompt lists are adjectives. Cinematic. Hyper-detailed. 8K. Award-winning lighting.

None of them show you the clip that came out. That's the only evidence that matters.

So this post does something different. Every prompt below was actually run on Seedance 2.5, either by ByteDance's own Seed team or by people who published the raw output and the exact prompt. The videos are embedded next to the prompts. You can watch what each sentence did.

We pulled from five sources: ByteDance's launch post (July 31, 2026), ByteDance's official prompt guide on BytePlus (published August 7, 2026), two hands-on test write-ups, a structured model benchmark, and a short-film production diary. Where a source tested something carefully, we tell you how. Where we're making a recommendation of our own, we label it.

Snapshot date: October 5, 2026. Seedance 2.5 is moving fast, so check the dates on anything you copy.

What changed in Seedance 2.5

Seedance 2.5 is ByteDance's video model, released July 31, 2026. It builds on the audio-video architecture of Seedance 2.0. Three changes matter for prompting.

Length. One generation now runs 4 to 30 seconds. Seedance 2.0 topped out at 15. You can also extend an existing clip in more rounds, so a single story can run for minutes.

References. You can attach up to 30 images, 10 video clips and 10 audio clips to one request. That's 50 assets. Reference video and audio clips run 2 to 30 seconds each, with a combined cap of 30 seconds per type.

Timestamps. ByteDance's guide says it plainly: Seedance 2.0 "does not respond to timestamps and only responds to shot numbers, while Seedance 2.5 supports integer-second timestamps." You can now write `0–4s:` and the model listens.

A few hard limits round it out. Output is 24 fps. Aspect ratios are 16:9, 4:3, 1:1, 3:4, 9:16, 21:9 and adaptive. Audio is generated in the same pass as the video. And the official docs say reference images and videos containing real human faces can't be uploaded, which matters a lot if you make creator-style ads. More on that later.

The 30-second number is the one people fixate on. The timestamp support is the one that changes how you write.

The official four-block prompt structure

A "six-part formula" went around in early August: subject, action, scene, style, camera, audio. It's fine. ByteDance published something different.

ByteDance's guide opens with one instruction: "Treat Seedance 2.5 as a visual content producer, and write structured prompts with a visual storytelling mindset." The Chinese edition says to write "with a director's thinking."

Then it gives four blocks.

1. Asset referencing. If you attached files, number them by upload order and say what each one is for. "Image 1 is the protagonist. Image 1 uses the timbre of Audio 1." "Strictly refer to the action and camera movement in Video 1."

2. One-sentence summary. Subject, location, event, genre or style, camera movement. One line that tells the model what the whole clip is.

3. Detailed plot. Either a shot list (`Shot 1:`, `Shot 2:`) or a timeline (`0–5s:`, `6–10s:`). You can mix them, as in `Shot 3 (6–10s):`. For each segment, describe the visuals, camera move, action, dialogue and sound.

4. Additional notes. Things that hold across the whole clip: camera position, environment, atmosphere, sound design, look.

Here's what that looks like filled in. This is our template, written to match the guide, for a 20-second product ad:

Image 1 is the product: a frosted glass serum bottle with a gold dropper.
Image 2 is the bathroom set.
A woman in her late 20s applies serum in a sunlit bathroom, warm natural look, handheld medium shots.
0–5s: Close-up of the bottle on the marble counter from Image 2. A hand enters frame and lifts it. Soft clink of glass on stone.
6–12s: Medium close-up. She squeezes the dropper, three drops land on her fingertips, she presses them into her cheek. She says: "Two weeks in and I stopped wearing foundation."
13–20s: Push in to her face as she smiles at the mirror. Hold on the bottle in the foreground, label facing camera.
Morning light from the window on the left throughout. Quiet bathroom room tone, water dripping, no music. No subtitles.

Three rules from the guide do most of the work.

Keep the timeline continuous and in whole seconds. Don't skip from 5s to 9s. Don't write 2.5s. And don't use timestamps for fast repetition. The guide's own example of what fails is "shake your head three times per second."

Size each segment honestly. Under-fill a time range and the model invents things to cover it. Over-fill it and you get choppy cutting or dropped beats.

Use negatives only for subtitles and audio. There's no negative prompt field. The guide sanctions exactly two kinds of negation: "No subtitles" and audio control like "No BGM; generate only environmental sounds and action sounds." Everything else, write as what you want to see.

One more rule people miss: don't put labels inside your reference images. The guide's example is writing "John" on a character's image and then writing "John is at school" in the prompt. You get duplicated or confused characters. Bind the asset to its role in the text.

Camera words the model already knows

The guide lists vocabulary you can use as-is. It's the most useful page in the whole document.

CategoryTerms
Shot sizeextreme wide shot, wide shot, medium shot, medium close-up, close-up
Movementpush in, pull out, pan, track, follow, orbit, dive, pull back, tilt up, handheld shake
Anglelow angle, overhead shot, first-person perspective
Techniqueone-shot / long take, Hitchcock zoom (dolly zoom), aerial perspective, FPV, bullet time, handheld shot, speed ramp

For anything outside that list, ByteDance says to write the term plus a plain description. Their example: "Rack focus: the focus of the frame shifts smoothly, the trees in the foreground that were sharp become blurred, and the figure in the background gradually comes into focus."

Same with transitions. Say when, and say how. "At 5s, a fast leftward lateral transition, wipe left plus a natural dissolve."

Now let's look at what real prompts produced.

1. The 30-second one-take

This is ByteDance's showcase for long-form text-to-video. No reference images. Just a prompt.

One-take handheld gimbal tracking shot. The camera slowly pushes in through a gap in a heavy red curtain and enters a warm-toned backstage dressing room. A young female singer, with her back to the camera, is adjusting her earpiece as a staff member reminds her it's time to go on. She turns toward the camera and starts singing citypop. The camera pulls back and tracks her as she passes through the curtain into a dim backstage corridor, interacting naturally with her dancers along the way; one staff member hands her a microphone. She and the dancers then step onto the stage, and the camera arcs around to the back, gradually revealing the red-and-black stage design, LED screens, spotlights, haze, and reflective floor. The camera finally pulls out to a wide shot of the arena, showing the packed audience, light boards, glow sticks, and cheering crowd, capturing the youthful, free-spirited climax of the concert.
Seedance 2.5 text-to-video, 30 seconds, one take. Dressing room, corridor, dancers, stage, arena. Video: ByteDance Seed launch post, July 31, 2026.

Watch it and you'll see the prompt played back in order. Mirror and dressing room. Corridor with dancers. Microphone handoff. Stage. Wide shot of the crowd.

The prompt is written as a route. Every clause is a place the camera goes next, joined with "then" and "finally." ByteDance's post says the point of 30 seconds is a story with "setup, development, turning points, and resolution," and this prompt has all four in one paragraph.

Notice what's missing. No "8K." No "masterpiece." The only style words are "warm-toned" and "red-and-black stage design," and both describe something you can point at.

2. Extending a clip into a minute

Seedance 2.5 can continue a video it already made. You pass the clip back in as a reference and describe what happens next.

Extend the video. Continue from the visuals and subjects in @Video 1 and generate another 30-second clip, keeping the character subjects, scene, visual style, and sound effects consistent. The little boy runs along the train carriage holding a soccer ball. When the subway stops, the side door opens and he immediately dashes out, with the male lead chasing after him. The two run across the platform and out onto the street, startling passersby and vehicles along the way. The male lead finally catches up and grabs him. The boy looks up, aggrieved. The male lead's anger slowly fades; he pats the boy's head and shows a helpless smile.
The original subway clip plus a 30-second extension: 65 seconds total. Same man, same boy, same soccer ball from the carriage to the street. Video: ByteDance Seed.

The extension prompt opens with the job ("Extend the video"), then the source ("@Video 1"), then the list of things to keep ("character subjects, scene, visual style, and sound effects"). Only after that does it describe the new action.

The ending is the clever part. "The male lead's anger slowly fades; he pats the boy's head and shows a helpless smile." That's an emotional resolution, and it gives the extension a place to land. Extensions without an ending tend to just keep running.

If you're cutting ads, this is how you'd get a 45 or 60-second spot without stitching three unrelated clips. Write the first 30 seconds, then extend with a clear end state.

3. Timestamps on a stage performance

This one uses four reference images and integer-second ranges. It's the clearest official example of timestamp control.

16:9 widescreen, cinematic texture, single continuous take, smooth camera movement, no cuts. Scene reference: @Image 4. 0–5s: Open with a close-up of the Overlord from @Image 2. The camera slowly circles his upper body and transitions into a medium shot. The Overlord spins and turns, his body and back flags sweeping quickly past the lens to form a natural occlusion, and the camera follows through to Consort Yu's side in @Image 1. 6–10s: The camera steadily circles Consort Yu in a medium shot from @Image 1, following her water sleeves through the arc. She raises her arm, flicks her wrist, unfurls the sleeves, and half-turns. She then draws the sleeves back, holds the pose, and looks sideways toward the Overlord. 11–20s: The male warrior from @Image 3 enters with an aerial flip. The Overlord takes center stage while the warrior advances and retreats on the opposite side in a combat exchange. Consort Yu stands slightly behind and to the side of the Overlord, weaving in water-sleeve movements to set softness against strength. The camera slowly pulls back from a medium-close shot of the warrior to a full stage view. At the end, all three face the audience and strike a synchronized Peking opera finale pose.
Reference-to-video with four images and three timestamp ranges, 20 seconds, one continuous take. Video: ByteDance Seed.

Two tricks are worth stealing.

First, the "natural occlusion." The Overlord's flags sweep past the lens, and the camera uses that moment to move to Consort Yu. That's how you change subjects inside a one-take without a visible cut. Real camera operators do the same thing with a passing body or a door frame.

Second, look at how the time is split: 5 seconds, 5 seconds, then 10 seconds for the busiest action. The fight with three characters gets the most room. That's the guide's sizing rule in practice.

Also look at the first line. Format, texture, take style and "no cuts" all come before any action. The model knows the rules before it reads the plot.

4. Eighteen reference images in one shot

Here's the reference system pushed hard. One venue, one pianist, one cello, one violin, a lead singer, five orchestra images, four choir images and four audience images.

A 30-second concert sequence in 16:9 landscape, with cinematic realism, authentic concert hall lighting and shadows, warm golden stage lighting, and the atmosphere of a formal classical concert. Use @Image 1 for the venue. Reference @Image 2 for the pianist. Reference @Image 3 for the cello. Reference @Image 4 for the violin. The lead vocalist must strictly follow @Image 5. Reference @Images 6 to 10 for the rest of the orchestra. Reference @Images 11 to 14 for the choir. Reference @Images 15 to 18 for the audience seating. The lead vocalist walks from center stage toward the front edge. The pianist is positioned by the piano. The orchestra is arranged on both sides and toward the rear. The choir stands at the back of the stage. Open with a high-angle wide shot of the full concert hall. The pianist strikes the keys, and the lead vocalist steps into the spotlight and begins singing. The camera naturally moves across the violin, cello, and orchestra as they perform together, with the violin feeling bright and the cello warm. In the latter part, the choir joins in. The lead vocalist briefly makes eye contact with front-row audience members, who respond with a smile and a slight nod. In the closing shot, the camera pulls back. The singing ends, and the audience joins in the applause.
Eighteen reference images bound to roles in one request. Video: ByteDance Seed.

Read it twice and you'll see three passes.

Pass one binds every image to a role. Pass two sets the blocking: who stands where. Pass three is the action, in order.

Look at the verbs too. Most assets get "reference." The singer gets "must strictly follow." That's how you rank your references. The lead is the one thing that can't drift, so it gets the strongest instruction.

For product work, this is the pattern. Product images get "must strictly follow." Set and props get "reference."

5. Green-screen replacement

You shoot a subject on green, then let Seedance 2.5 build the world around them.

Using @Video 1, render the green-screen background, obstacles, wardrobe, and supporting characters. 0–4s: outdoor training, replace the obstacles with rocks, bricks, tires, and wooden crates. 4–10s: locker room, friends offering encouragement. 10–15s: international match, replace the training poles with original defenders and a goalkeeper, and the protagonist scores. Overall photorealistic, cinematic quality.
The clip opens on ByteDance's prompt card, then shows the green-screen source turning into a training ground, a locker room and a match. Video: ByteDance Seed.

This prompt is short. That works because the source video already carries the motion. All the prompt has to do is say what replaces the green in each time window.

Note the word "original" in "original defenders and a goalkeeper." Asking for invented players keeps the model away from real teams and real athletes, which the content filters are strict about.

ByteDance's launch post says the model also handles how the subject reacts to the new environment: clothes, hair, gait and light. In practice that means you don't need to prompt "his shirt moves in the wind." Describe the environment and let it do the physics.

6. Re-shooting only the camera

This is the edit mode that will matter most for ads. Keep the action. Change the camera.

Edit @Video 1. Keep the characters, actions, and visual style unchanged. Adjust only the camera movement. A 15-second segmented camera plan: 0–4s, a micro-FPV move skims tightly past the pan, then follows the popping toast and whip-pans to the coffee; 4–7s, push in and track laterally along the rim of the pan, following the fried egg as it flips up and lands back in place; 7–11s, rapidly rise to a top-down view, then descend at a steady pace, sweeping across the plate and keys; 11–15s, use a handheld close-up to follow the hands with a fast lateral whip, then push in on the breakfast and pull back to a medium two-shot. Keep the entire sequence smooth, continuous, and stable.
Camera-only edit of a clay-style breakfast scene: FPV past the toaster, push along the pan, top-down sweep, handheld close-up. Video: ByteDance Seed.

The first three sentences do the locking. "Edit @Video 1." "Keep the characters, actions, and visual style unchanged." "Adjust only the camera movement."

That word "only" matters. ByteDance's guide separates tasks where your input video becomes part of the output timeline from tasks where it's just a reference. Use editing language and the model treats the request as an edit. Mix editing words into what you meant as a loose reference and you can get an error instead of a clip.

Every time window names a camera term from the official list: FPV, push in, track, top-down, handheld. And each one is tied to an object in frame: the pan, the toast, the egg, the plate, the hands. The camera always has something to follow.

7. Write the prompt as the dolly path

The test ran eight generations in matched pairs on the same seed, duration, resolution and aspect ratio, changing one clause at a time. The first finding: the order you list objects becomes the order the camera passes them.

A single continuous tracking shot moving left to right through a film equipment rental warehouse: past a wall of lens cases, then a row of tripods, then a stack of flight cases, ending on a technician wiping down a camera body with a cloth. No cuts. Quiet warehouse room tone, footsteps on concrete and the click of a lens cap, no music.
8 seconds, 720p, 16:9, seed 42. Scene detection found zero cuts. Credits: Segmind
Eight frames from the dolly-path test: lens cases, then tripods, then flight cases, then the technician

Frames one to three are the lens cases. Four and five are the tripods. Six and seven are the flight cases. Eight is the technician. The prompt, read aloud, in order.

Two things make it work. Writing the beats as a sequence with "then" and "ending on." And writing "No cuts," which the model honored in every single-take prompt in the test.

They also corrected a common myth. A thin prompt doesn't get chopped into random shots. Give Seedance 2.5 six vague words and it'll invent one competent continuous move and make every other decision for you. The take stays continuous. The decisions just stop being yours.

We'd steal this for every product hero shot. List the product's features in the order you want the camera to find them, and end on the thing you want the viewer to remember.

8. Frame relative to the subject

This pair is the one that'll save you the most wasted renders. Same carpenter, same seed, same six seconds. One clause changed.

Absolute framing:

A carpenter planing a long board at a bench in a sunlit workshop. Camera locked at chest height. No cuts. Quiet workshop room tone, the rasp of the plane and wood shavings falling to the floor, no music.
"Camera locked at chest height." The model put the camera at chest height. The head is out of frame for all six seconds. Credits: Segmind
Six frames from the absolute-framing clip, head cropped in every frame

Subject-relative framing:

A carpenter planing a long board at a bench in a sunlit workshop. Camera holds him from the chest to just above the head for the whole shot. No cuts. Quiet workshop room tone, the rasp of the plane and wood shavings falling to the floor, no music.
"Holds him from the chest to just above the head." Same scene, working medium shot, face in frame throughout. Credits: Segmind
Six frames from the subject-relative clip, head held in frame throughout

Seedance 2.5 executes camera instructions close to literally. "Chest height" is a height. The model obeyed it, and geometry doesn't know where the face is.

The fix is to anchor every framing instruction to the body. "Holds her from the waist up." "Frames him head to knees." "Stays tight on the hands."

This one matters a lot for talking-head ads. If your creator's face leaves the frame, the take is dead. Write "holds her from mid-chest to just above her head" and you've removed a whole category of failed renders.

9. Give every shot a time budget

Multi-shot prompts work. `Shot 1:`, `Shot 2:`, `Shot 3:` and the model cuts between them. What the test measured is how the model splits the time when you don't say.

Shot 1: a ceramicist centres a lump of wet clay on a spinning wheel, close on her hands. Shot 2: she pulls the walls of the pot upward, water running down her wrists. Shot 3: a wide of the finished pot alone on the wheel in the quiet studio. Quiet studio room tone, the wet slap of clay and the low hum of the wheel, no music.
Six frames from the unbudgeted clip: the finished pot only appears in the last frame

Scene detection put the cuts at 6.08s and 11.54s. That's a 6.08-second opener, a 5.46-second middle and a 3.53-second payoff. The shot of the finished pot, the whole point of the sequence, got the scraps.

Now add five words to each shot:

Shot 1, five seconds: a ceramicist centres a lump of wet clay on a spinning wheel, close on her hands. Shot 2, five seconds: she pulls the walls of the pot upward, water running down her wrists. Shot 3, five seconds: a wide of the finished pot alone on the wheel in the quiet studio. Quiet studio room tone, the wet slap of clay and the low hum of the wheel, no music.
Budgeted version, 15 seconds. Cuts landed at 4.88s, 5.16s and 5.03s, within a fifth of a second of the requested 5/5/5. Credits: Segmind
Six frames from the budgeted clip: the finished pot holds for two frames

The takeaway: the model front-loads. It spends time on the setup and squeezes whatever comes last. In an ad, the last shot is usually the product and the offer. That's the shot you can least afford to lose.

On Seedance 2.5 you can also write the same thing as timestamps, `0–5s:`, `5–10s:`, `10–15s:`. Either way, put the seconds in every shot.

10. End every audio line with "no music"

Seedance 2.5 writes its own soundtrack in the same pass as the picture. That's useful. It also has a trap.

The content-safety check runs on the finished audio, after the render. This prompt was fired on purpose to show it:

A violinist plays alone on a rooftop at dusk, the city wide behind her. Warm cinematic music swells throughout the clip.

It rendered for 145 seconds, then came back as an error: the generated audio was blocked, category "copyright." Nothing in the prompt named a song, an artist or a studio. It asked for music, the model wrote some, and the filter wouldn't release it.

The fix is in the wording. Describe the sounds the scene itself makes, and end the audio clause with "no music." Every other prompt in the run ended that way, and every one passed. You've seen the pattern in each prompt above: "quiet workshop room tone, the rasp of the plane and wood shavings falling to the floor, no music."

This lines up with ByteDance's own guide, which gives "No BGM; generate only environmental sounds and action sounds" as a sanctioned instruction.

Two more audio findings from the same test.

Mixes come back quiet. Integrated loudness across their clips ran from −26.7 to −41.9 LUFS. Streaming platforms target around −14. The warehouse take came back close to inaudible on a phone. Normalize every clip before you post it.

If you need music, add it in the edit. Generate the scene with room tone and dialogue, then lay a licensed track over it.

11. A 480p draft tests your words, and only your words

The common advice is to draft at 480p and re-render at 720p once you like the result. The same prompt and seed were run at both resolutions:

A barista pours a rosetta into a flat white on a marble counter, morning light from a window on the left. Slow push in. No cuts. Quiet cafe room tone, the hiss of the steam wand and cups set on saucers, no music.

The two clips were different takes. The 480p one played out on dark stone with the barista out of frame and hard backlight. The 720p one had pale marble, the barista's torso in shot and a softer room.

What carried over was everything the prompt named: the pour, the rosetta, the window light from the left, the slow push in, no cuts. What the prompt left open got re-rolled.

So use a 480p pass to check whether your wording produces the right action, light and props. Don't use it to approve composition. And if you want a specific look to survive, name it in the prompt. Anything you leave out is up for grabs on every render.

12. A 30-second micro-drama in one generation

This is the closest public example to a full ad script: two characters, three dialogue beats, 9:16 vertical, native audio, one call. The full prompt is below.

Vertical 9:16 micro-drama, one continuous episode, photorealistic, cinematic night grade, handheld-steady framing. Two characters, both from the reference image.
Character A: MAYA, the woman in the reference image. Same face, same straight dark shoulder-length hair, same rust-orange knit sweater over a white collared shirt, same thin silver necklace in every stage.
Character B: DANIEL, the man in the reference image. Same face, same short greying hair, same close-trimmed grey beard, same rimless glasses, same navy quarter-zip pullover in every stage.
Setting: a glass-walled meeting room inside an empty open-plan office at night, city lights behind the glass, cold overhead strip lighting, a plain unbranded silver laptop on the table. No visible brand logos anywhere. Stage 1: Initial state: DANIEL sits alone at the meeting room table, hands folded, looking at the dark window. Primary event: MAYA pushes the glass door open and crosses to the table, setting the laptop down hard in front of him. End state: MAYA stands over the table with both palms flat on it, looking down at DANIEL, who has not moved. MAYA says, tight and controlled: "The numbers do not reconcile. Someone moved them."
Stage 2: Initial state: MAYA turns the laptop screen toward DANIEL. Primary event: DANIEL does not look at the screen, he keeps his eyes on MAYA, and the camera pushes in slowly on her face as she registers that he is not surprised. End state: MAYA straightens up slowly, her hands leaving the table, her expression falling from anger into disbelief. DANIEL says, calm and quiet: "I know. I approved it."
Stage 3: Initial state: MAYA stands very still on the far side of the table. Primary event: she closes the laptop and picks it up, holding it against her chest, and DANIEL rises from his chair and steps between her and the glass door. End state: the two of them stand facing each other in the narrow gap by the door, neither moving. MAYA says, steady and low: "Then you already know what I have to do."
Audio: spoken English dialogue, two distinct voices, quiet office room tone with a faint air-conditioning hum, a low sustained tension drone underneath.
One generation, 30.04 seconds, 720x1280, two speaking characters, unedited. Turn the sound on. Credits: Segmind

Three things made it work.

Stages with a start, an event and an end. Each beat gets an initial state, one primary event and an end state. That's one thing happening per 10 seconds, with a clear frame to start from and a clear frame to land on. Multi-shot prompts crammed into six seconds failed in testing, while the same structure at 12 to 15 seconds worked.

Names in caps, repeated, and every line tagged. MAYA and DANIEL are always the subject of the sentence. No "she" across a stage boundary. Every line of dialogue says who speaks it. With two people in frame, an untagged line is a coin flip.

A reference image plus a written description. The faces came from one generated image holding both characters, and the prompt describes their hair, clothes and glasses again in words. The image carries the likeness. The text gives the model something to hold onto when the faces turn.

One thing we noticed watching it frame by frame. The prompt asks for "a plain unbranded silver laptop" and "No visible brand logos anywhere." The laptop in the returned file still has a logo on the lid. This fits ByteDance's guide, which only sanctions negatives for subtitles and audio. If a product must stay clean, describe what it looks like ("a matte silver laptop with a blank lid") and check every take.

For ad makers, swap the office drama for a problem, a turn and a payoff. Stage 1: the problem. Stage 2: the product. Stage 3: the result and the line that sells it.

13. The 30-second continuous workshop

This benchmark runs the same prompts on every model and publishes one try each, failures included. Its long-shot test is useful because it's built to catch drift: one woman, one blue apron, one red mug, one white cloth, 30 seconds.

One continuous 30-second take in a sunlit pottery workshop, no cuts. A woman wearing a plain blue apron stands at a wooden table with one red mug and a folded white cloth. First third: she picks up the red mug and turns it once in her hands. Middle third: she places it back on the table and wipes beside it with the white cloth. Final third: she folds the cloth, places it to the left of the mug, and looks toward the same window. The camera slowly tracks sideways throughout. Keep the same woman, apron, red mug, table and window. Natural workshop ambience, no music, no text.
30.09 seconds at 1280x720. The apron, mug and workshop stayed recognizable at every sampled checkpoint, and the action reached the final pose. Credits: Voyager

Look at how the prompt fights drift. It counts objects ("one red mug"). It uses "the same" twice. It ends with an explicit keep-list: "Keep the same woman, apron, red mug, table and window."

It also splits time in thirds instead of seconds. That's a softer form of the time budget, and it still gives each action its own window.

The keep-list is something you should copy into every product video. Put the product, the person and the setting in one line at the end.

14. The product turntable

Not every test went well, and the failures are useful too.

A matte black ceramic pour-over coffee dripper rotating slowly on a white turntable, soft studio light, seamless white background, product video, no text.
5 seconds, 720p. It opens on an extreme close-up of the dripper's rim before pulling out to the product. Credits: Voyager

The clip opened on an extreme close-up of the rim, so the first second reads as an abstract shape before the product appears. On a 5-second ad, that's 20% of your runtime with no recognizable product.

The prompt never says where the camera starts. Based on the dolly-path finding in example 7, the fix is to say it: "Open on a full product shot, the whole dripper centered in frame with space above and below. Hold the framing while the turntable rotates once." Our suggestion, untested.

The benchmark found three more problems worth knowing:

  • Their image-to-video run of the same dripper was refused as a possible copyright violation, even though the starting image was their own generated product shot.
  • A physics test of spilled water turned the water into a translucent haze.
  • Wait times for the same 5-second request ran from 170 seconds to 26 minutes.

Their overall take: on 5-second clips, Seedance 2.5 didn't look better than 2.0. The long-shot tests are where it earned its place.

What breaks, and how people worked around it

Put the hands-on reports side by side and the same problems keep showing up.

Old Seedance 2.0 prompts. In a short film made almost entirely in 30-second mode, an established 2.0 prompt produced "visibly broken, glitchy results" on 2.5. The model also resisted rapid, one-second-style cuts. Scenes that were allowed to "breathe" came out better. If you have a prompt library from 2.0, rewrite it into the four-block structure rather than pasting it over.

Accuracy after 15 seconds. Users on Reddit's r/generativeAI reported the model getting less accurate after the 15-second mark when a prompt had six shots. That fits the front-loading finding in example 9. Fewer, longer beats hold up better than many short ones.

Voice and accent drift. The model casts a voice based on what it thinks the character should sound like from the reference image. A character came out with a British accent nobody asked for. Writing "American" fixed it. Writing "American English" didn't, apparently because the word "English" pulled it back. Name the voice in the prompt, and test small wording changes.

Morphing in fast action. Testers saw visual "spikiness" and morphing in fast action scenes, even with structured prompts. Upscaling didn't remove it. Keep fast action short, or let a slower take carry the scene.

Odd prop scale. In one take, a book rendered as big as a suitcase. The workaround: a 30-second take usually holds several usable segments, so cut around the flaw rather than throwing the whole take away.

Reference order. When a set of first frames was fed in alongside references, the shots came back largely right but not always in the right order. Put the sequence in the prompt text. Don't count on upload order to set it.

Rotated vertical output. One 9:16 request in testing came back as a correct portrait file with the picture rotated 90 degrees inside it. No error. It was intermittent. If you publish vertical video at volume, check a frame from each clip.

Real faces and named IP. ByteDance's docs say reference images and videos with real human faces can't be uploaded. Seedance 2.0's launch drew cease-and-desist letters over recognizable characters, and you should expect prompts naming real people, franchises or studios to be refused. Every hands-on test above built its characters from generated portraits.

Storyboards that are too long. ByteDance's guide says multi-panel storyboards work best at 15 panels or fewer. Eighteen panels is their example of input that produces still frames or scrambled order.

6 Seedance 2.5 prompt templates for UGC ads

These are our templates, built from the patterns above for the ads EzUGC customers make most. We haven't run every one of them as written, so treat them as starting structures. Swap the brackets, keep the shape.

Each one follows the same rules: framing anchored to the body, seconds in every beat, sounds described, "no music," "no subtitles," and a keep-list at the end.

Template 1: Talking-head product review (15s, 9:16)

Image 1 is the creator. Image 2 is the product: [describe product, color, material, label].
A woman in her [age] reviews [product] in her [kitchen/bedroom/car], handheld selfie framing, natural daylight.
0–4s: She holds the product up next to her face, camera holds her from mid-chest to just above her head. She says: "[hook line under 10 words]."
5–10s: She turns the product in her hand to show the label, then uses it [describe the action]. She says: "[one benefit, specific]."
11–15s: She looks into the lens and smiles. She says: "[result or offer]." Product stays in frame, label toward camera.
Handheld selfie feel throughout, slight natural shake. American accent, warm, conversational. Quiet room tone, no music. No subtitles. Keep the same woman, outfit, product and room.

Template 2: Problem, turn, payoff (30s, 9:16)

Built on the staged structure from example 12.

Image 1 is JESS. Image 2 is the product: [describe].
Vertical 9:16, one continuous episode, natural light, handheld-steady framing. JESS: [hair, clothes, one distinguishing detail], same in every stage. Stage 1 (0–10s): Initial state: JESS [is frustrated by the problem]. Primary event: [the problem happens again, visibly]. End state: JESS looks at the camera. JESS says: "[problem line]."
Stage 2 (11–20s): Initial state: JESS picks up the product. Primary event: she [uses it, one clear action]. End state: [visible change]. JESS says: "[what it did]."
Stage 3 (21–30s): Initial state: [the after scene]. Primary event: JESS [relaxed action]. End state: JESS holds the product beside her face. JESS says: "[offer or CTA]."
Audio: spoken dialogue, one voice, [room tone and two specific sounds], no music. No subtitles. Keep the same JESS, outfit, product and setting.

Template 3: Product hero dolly (8s, 16:9 or 9:16)

Built on the dolly-path pattern from example 7.

A single continuous tracking shot moving [left to right / toward camera] across [surface]: past [feature 1], then [feature 2], then [feature 3], ending on the full [product] centered in frame, label facing camera. No cuts. [Lighting described as a source: window light from the left / soft overhead studio light]. [Two specific sounds], no music. No subtitles.

Template 4: Unboxing (12s, 9:16)

Image 1 is the product box: [describe]. Image 2 is the product inside: [describe].
Overhead shot of hands unboxing [product] on a [surface], first-person feel.
0–4s: Overhead shot, frame holds both hands and the whole box. Hands slide the lid off. Cardboard scrape.
5–8s: Hands lift the product out from Image 2 and turn it once toward camera. Tissue paper rustle.
9–12s: Push in to a close-up of the product in one hand, label readable. Soft daylight from above throughout. No music. No subtitles. Keep the same hands, box and product.

Template 5: Before and after (10s, 9:16)

Image 1 is the person. Image 2 is the product: [describe].
0–4s: Medium close-up, camera holds her from the shoulders to just above her
head. [Describe the "before" state specifically: dull skin, frizzy hair, messy
desk].
5–6s: She applies or uses the product, hands in frame.
7–10s: Same framing as 0–4s. [Describe the "after" state specifically]. She
smiles at the lens.
Same window light from the left throughout. [Sounds], no music. No subtitles.
Keep the same person, outfit, framing and room.

Template 6: Re-shoot a winning ad with new camera work

Built on ByteDance's camera-only edit.

Edit @Video 1. Keep the characters, actions, dialogue and visual style
unchanged. Adjust only the camera movement.
0–[x]s: [camera term from the official list], following [object or person].
[x]–[y]s: [camera term], following [object or person].
[y]–[end]s: push in on the product, then hold.
Keep the entire sequence smooth and continuous.

The last one is the most underrated. When an ad works, the expensive part is the script and the performance. A camera-only edit gives you a fresh-looking variant to test without touching either.

A checklist before you hit generate

Run your prompt against this before you spend a generation on it.

  1. Did you number every reference and give it a job? "Image 1 is the product" beats hoping the model guesses.
  2. Is there a one-line summary up top? Subject, place, event, style, camera.
  3. Does every beat have seconds? If you wrote `Shot 3:` without a duration, it'll get squeezed.
  4. Are timestamps whole seconds with no gaps? `0–4s`, `5–10s`, `11–15s`.
  5. Is the framing tied to the body? "From mid-chest to just above her head." Never "at eye level" alone.
  6. Is the camera route written in order? What it passes first, next and last.
  7. Did you name the sounds and end with "no music"? Room tone plus two specific sounds.
  8. Did you name the voice? Accent, tone, pace. "American," never "American English."
  9. Are your negatives only about subtitles and audio? Rewrite everything else as what you want.
  10. Is there a keep-list at the end? Person, outfit, product, setting.
  11. Are there real faces or brand names in your references? Swap in generated portraits and describe products in plain words.
  12. Is the clip longer than it needs to be? Thirty seconds is a ceiling. A 10-second hook doesn't get better at 30.

How to run Seedance 2.5 prompts in EzUGC

Seedance 2.5 is in the EzUGC dashboard alongside the other video models. Here's what the EzUGC version supports, as of October 5, 2026:

  • Durations: any whole number from 4 to 30 seconds.
  • Resolutions: 480p and 720p, in 16:9, 4:3, 1:1, 3:4, 9:16 and 21:9.
  • Native audio: dialogue, sound effects and ambience in the same generation.
  • Image inputs: first frame, last frame, and up to 30 reference images.
  • Prompt length: up to 10,000 characters, which is plenty for the four-block structure and a three-stage script.

A few workflow notes from the patterns above.

Make your creator first. Real faces can't be used as references, so generate your creator portrait with an image model like Nano Banana Pro in the AI influencer generator, then pass that image to Seedance 2.5 as a reference. Describe the person again in the prompt, as the micro-drama in example 12 did.

Use reference images to carry identity. A first frame locks the opening picture. A reference image passes the face or product without fixing the composition. For a creator who appears across three beats, reference is the safer choice.

Draft cheap, then commit. Use 480p to check the action and the wording. Then render the final at 720p, knowing the composition will be a new take.

Plan your Seedance usage. Seedance 2.5 sits inside the Seedance family allowance on each plan: 5 generations a cycle on Startup, 10 on Growth and 20 on Pro, inside the plan's overall 50, 100 or 500 videos. Failed generations don't count. Check the pricing page for current plans, and see Seedance 2.0 if a 15-second clip covers what you need.

If you're making a lot of ad variants, the AI commercial generator handles script, actor and product shots together, and you can drop into Seedance 2.5 when a scene needs a longer continuous take.

Sources

Every prompt in this post is quoted from its source. The videos are the outputs those sources published, rehosted here so they load fast. Credit goes to the people who ran the tests.

  • ByteDance Seed, Introducing Seedance 2.5, July 31, 2026. Examples 1 to 6.
  • ByteDance, Dreamina Seedance 2.5 prompt guide, BytePlus ModelArk doc 2607689, August 7, 2026. Four-block structure, camera vocabulary, negation rules, limits.
  • Credits: Segmind (examples 7 to 12), Voyager (examples 13 and 14), MindStudio (failure modes)

FAQ

What is the best prompt structure for Seedance 2.5?

ByteDance's official four blocks. Number your reference assets and give each a role. Write a one-sentence summary. Write the plot as shots or integer-second timestamps. Finish with notes that hold across the clip, like lighting, sound and camera position.

Does Seedance 2.5 follow timestamps?

Yes. ByteDance's guide says 2.5 supports integer-second timestamps, while 2.0 only followed shot numbers. Keep the timeline continuous, use whole seconds and don't use timestamps for rapid repeated actions.

How long can a Seedance 2.5 video be?

4 to 30 seconds per generation. You can extend a finished clip with another generation that continues the same characters and scene, which is how ByteDance's subway example reached 65 seconds.

Can I use negative prompts in Seedance 2.5?

There's no negative prompt field. ByteDance's guide sanctions negation for subtitles and audio only, like "No subtitles" or "No BGM." In the micro-drama in example 12, "No visible brand logos" didn't stop a logo appearing on the laptop. Describe what you want to see instead.

Why did my Seedance 2.5 video get blocked for copyright?

The safety check runs on the finished audio, so asking for music can get a fully rendered clip blocked. Describe the room tone and the sounds in the scene, and end the audio line with "no music." Add licensed music in the edit.

Can I use a real person's photo as a reference?

ByteDance's docs say reference images and videos containing real human faces can't be uploaded. Testers who kept characters consistent used generated portraits, then described the character again in the prompt.

Should I reuse my Seedance 2.0 prompts?

Rewrite them. In hands-on testing, 2.0 prompts produced glitchy output on 2.5, and the new model does better with longer, slower beats than with rapid cuts.

How many shots can I fit in 30 seconds?

ByteDance doesn't set a number. The hands-on reports point the same way: three stages of about 10 seconds held up well in the micro-drama test, while Reddit users saw accuracy fall after 15 seconds with six shots. Fewer, longer beats are safer.

Sources and citations

Frequently asked questions

Direct answers pulled into the page to improve answer-first relevance and scanability.

ByteDance's official four blocks. Number your reference assets and give each a role. Write a one-sentence summary. Write the plot as shots or integer-second timestamps. Finish with notes that hold across the clip, like lighting, sound and camera position.
Yes. ByteDance's guide says 2.5 supports integer-second timestamps, while 2.0 only followed shot numbers. Keep the timeline continuous, use whole seconds and don't use timestamps for rapid repeated actions.
4 to 30 seconds per generation. You can extend a finished clip with another generation that continues the same characters and scene, which is how ByteDance's subway example reached 65 seconds.
There's no negative prompt field. ByteDance's guide sanctions negation for subtitles and audio only, like "No subtitles" or "No BGM." In the micro-drama in example 12, "No visible brand logos" didn't stop a logo appearing on the laptop. Describe what you want to see instead.
The safety check runs on the finished audio, so asking for music can get a fully rendered clip blocked. Describe the room tone and the sounds in the scene, and end the audio line with "no music." Add licensed music in the edit.
ByteDance's docs say reference images and videos containing real human faces can't be uploaded. Testers who kept characters consistent used generated portraits, then described the character again in the prompt.
Rewrite them. In hands-on testing, 2.0 prompts produced glitchy output on 2.5, and the new model does better with longer, slower beats than with rapid cuts.
ByteDance doesn't set a number. The hands-on reports point the same way: three stages of about 10 seconds held up well in the micro-drama test, while Reddit users saw accuracy fall after 15 seconds with six shots. Fewer, longer beats are safer.
Tags:UGCAI

Written by