NEWYour next viral UGC ad is one click away. Meet Studio.Make an ad
PopcraftPopcraft
New model

Seedance 2.5 · ByteDance

Thirty seconds, in one generation

A whole ad, a product demo or a scene beat comes back as one thirty-second generation rather than five clips and a join — held together by up to fifty reference assets, and edited afterwards without a re-roll.

Live on Popcraft, beside Seedance 2.0.

30sin one generation
50reference assets per job
10+languages, natively
Against Seedance 2.0
Longest single generation
15s30s
Reference images
930
Reference clips, video + audio
310 + 10
Output aspect ratio
6 fixed0.4–2.5

Thirty seconds, twice

Two of ByteDance’s own demonstration films, at the length the model is built for. Each is one generation — picture, camera changes and sound arriving together, with nothing assembled afterwards.

Seedance 2.51920×108029.98sWith soundA street corner, end to end. One generation, with the music and the street sound arriving alongside the picture. The camera changes setup nine times inside it — street level, an overhead of the crossing, close work on faces, a backlit wide — and not one of those changes is a join between generations. This is the film running silently behind the headline at the top of the page. Here it has its sound.
Seedance 2.51280×72030.04sWith soundAn explainer with a beginning, a middle and an end. The history of money in one modelling-clay set — barter, coin, banknote, bank, a contactless tap — across nine staged beats that hold the material, the light and the scale the whole way.

What to look for in them

The four gains are the details that decide whether a shot reads as footage or as output.

Lighting

Light that obeys the room

Catchlights in the eyes read as glass with depth in them, and direction, falloff and bounce follow the physics of the scene instead of being painted on.

Camera

Moves that hold their line

Push, pull, pan and tilt trajectories come back smoother and more stable, which is what makes a thirty-second move survive being thirty seconds long.

Performance

Expression as a curve

Micro-expressions — a tremble at the eye, a press of the lips — carried as continuous movement across a whole beat rather than snapped between poses.

Realism

Skin, hair, and the uncanny gap

Skin texture, hair sheen and small facial movement closer to how a camera records them, with less of the plastic look that gives a generated face away.

Four things change, and only one of them is the length

2.5 steps forward on four axes: longer generations, richer references, precise editing, and multilingual creation. The length is the headline. Editing is the change a working day actually feels.

Length
30s one generation

Thirty seconds from one generation

Up to thirty seconds in a single generation, holding camera movement, continuity and a complete narrative arc across the whole clip. Seedance 2.0 tops out at fifteen. High-Fidelity Temporal Extension then carries a short clip forward, backward, or bridges two of them.

References
50 assets per job

Everything the shot needs, in one brief

Thirty images, ten video clips and ten audio clips can go into one generation — cast, product, location, camera behaviour, music and voice together. In 2.0 that ceiling was nine images and three clips.

Editing
Local rewrites

Change a detail, keep the clip

Controllable Video Editing holds the frame, the camera and the pacing and rewrites only what you name: a background, a garment, a product, an actor’s expression, the music. A flaw in a good thirty seconds stops being a reason to re-roll it.

Language
10+ languages

Brief it in your own words

Chinese, English, Spanish, Indonesian, Malay, Thai, Arabic, Portuguese, Vietnamese, Japanese and Korean among others, natively — the prompt no longer has to be translated before the model reads it. Instruction following is stronger across the board.

2.5 beside 2.0, row by row

Seedance 2.0 stays in the line-up, and at 4K it is still the right tool for a short shot.

CapabilitySeedance 2.0Seedance 2.5
Longest single generation15s30s
Reference images930
Reference video and audio clips310 video + 10 audio
Total reference duration15s30s video, 30s audio, counted apart
Audio-only referenceSupported
Timestamps in the promptShot numbers onlyWhole-second timestamps
Multi-view subject imagesNot advisedSupported
Output aspect ratioSix fixed ratiosAny ratio from 0.4 to 2.5, set by the input
MOV output, for editing and extensionSupported — colour and audio continuity across a join

Every asset arrives with a job

The count is the easy part. What makes it usable is that every asset is attached with an @ mark and given a stated role, so the model knows which picture is the actor and which is the light.

Images
30 per job

Up to 4K, 30 MB each. JPEG, PNG, WebP, BMP, TIFF, GIF, HEIC, HEIF. Multi-view sets of the same subject are supported in 2.5, where 2.0 advised against them.

Video clips
10 per job

480p to 4K, 200 MB each, 2–30s per clip, MP4 or MOV. The clips share a thirty-second total.

Audio clips
10 per job

15 MB each, 2–30s per clip, MP3 or WAV, with its own separate thirty-second total. Audio-only references are new in 2.5 — a music bed, a voice or an effects track can drive pacing, beat matching and lip-sync on its own.

What a reference can be asked to carry

Subject
The appearance or the voice of a person, animal, product, prop, location or invented character — the thing that has to stay itself for the whole clip.Use the woman in @image 1 as the lead, walking a city street at dusk.
Motion
Action, camera movement, timing and effects lifted off a clip and applied to a different subject. It transfers across species and styles, not only across takes.Follow the camera move and running rhythm in @video 1; the subject is @image 1.
White model
A rough untextured 3D blockout drives the camera choreography and the staging; the finished materials, lighting and atmosphere come from stills. Coarse geometry works better here than a detailed model.Render the motion in the white-model @video 1; subject @image 1, scene @image 2.
Style
Palette, grade and texture taken from a picture or a clip while the cast stays yours.Use the cold blue night grade of @video 1, keeping the lead from @image 1.
Audio
Music, melody, dialogue or vocal timbre. Cast a voice per character and the film comes back speaking in it.The father’s voice: deep, unhurried, middle-aged. Reference @audio 1.
Storyboard
A multi-panel board read as a plot outline. Fifteen panels or fewer, line art in preference to a rendered board, and no text drawn into the panels. The result follows the story, not the drawing frame for frame.Follow the four panels in @image 1; the character is @image 2.
Keyframes
One image as a first frame, two as first and last, or a set of independent keyframes. Unlike a storyboard, the output tracks these closely.@image 1 is the first frame, @image 5 the last; walk her from indoors to outdoors.
Wide street-level shot: the lead in a red cap and headphones crosses at the lights.Close shot from the side: the same woman passes a brick wall, her face in profile.Backlit wide: she walks towards the camera into low sun with others behind her.Final wide: she stands on the crossing at the end of the film with three other people.
One person, four setups, one generation. Four frames in order from the thirty-second street film above — a street-level wide, a close profile, a backlit approach, a final wide. Holding a face, a garment and a silhouette that steady across a whole clip is the job a subject reference is given when you brief one yourself.
A focused set beats a full one. Quality tracks the mix rather than the count: keep image subjects to eight or fewer, audio and video subjects to five, and reference clips in the five-to-ten-second range, where a two-second clip carries too little and a thirty-second one dilutes what matters. Where assets compete, the priority is core characters, then key products and props, then the environment, then style.

The re-roll stops being the only move

Both capabilities take a finished clip as @video 1 and leave everything you did not name alone. Editing rewrites part of what is already there. Extension writes what comes before or after it.

First frame of the street film: an empty crossing on a sunlit corner, a few people on the far pavement.Last frame of the street film: four people standing together on the same crossing at the end of the film.
The first and last frames of the thirty-second street film. Extension is asked for at one of these two edges — write backward into the minute before the empty crossing, or forward out of the moment the group is standing in. Editing works on everything between them, holding the frame and rewriting only what you name.
Controllable Video Editing

Rewrite a region, hold the rest

Add a subject, remove one, swap a garment or a scene against a reference image, or change the audio — a different music bed, an added effect, a re-voiced line. Timestamps narrow an edit to a window: 0:02–0:05 and nothing outside it moves.

The pattern that works is three parts long: point at the source, name the object and the change, then add a preservation clause so the model leaves the rest of the frame alone.

Replace the dark outfit on the man in @video 1
with the outfit in @image 2.
Keep everything else unchanged.
High-Fidelity Temporal Extension

Write what comes before and after

Extend forward from the last frame, backward into the moment before the first, or generate the missing bridge between two clips so they cut together. Character, scene and camera carry across the seam, and the transition can be asked to be seamless in both picture and sound.

Extension is how a strong six seconds becomes a finished thirty without regenerating the part that already works. MOV in and MOV out is the recommended path — it holds colour, brightness and audio consistency across the join.

Start from the final frame of @video 1 and extend
6 seconds: she leaves frame, the sky darkens,
streetlights come on one by one. Keep it seamless.

What the input decides for you

2.5 splits jobs by whether the source assets fix the shape of the output. 2.0 draws no such line, and this is the part most likely to surprise someone moving over.

TaskAspect ratioDuration
EditingLocked to the source clipFollows the source, within about a third of a second
First frame, or first and lastLocked to the first-frame imageYours to set
ExtensionLocked to the source clipYours to set
Reference generationOpenYours to set
Storyboards and keyframesOpenYours to set
Give a first and a last frame at the same dimensions. Where they differ, the last frame is stretched to match the first, and the stretch is the model doing what it was told rather than a fault to prompt around.

Where thirty seconds earns its keep

The model’s use cases fall into five families. Four of them are jobs that get shorter. The fifth only becomes possible once a generation runs this long.

Story and narrative

A scene that plays out

A plot beat that runs its full length in one generation, with an ensemble held together by a cast of reference stills. Continuity fixes — a visible rig, a stray object — become an edit instead of a re-shoot.

Advertising and commerce

One master, then every version

One master cut, then versions: swap the SKU, the model, the on-screen copy or the market, with the product’s appearance and the brand’s look locked by reference. Voiceover follows in the languages you need.

Knowledge and explainer

The whole idea in one take

One idea explained end to end in thirty seconds, a presenter swapped without re-recording, and the same set distributed in several languages.

Industrial

A procedure, start to finish

Assembly and operating steps shown end to end, with nameplates, spec panels and packaging text rewritten per product and the shot left standing.

What the extra length opens up

The work fifteen seconds refused

Game trailers and CG, virtual presenters, a short expanded into a series, a character aged or re-styled for a seasonal cut.

Previsualisation cuts across all five. A rough grey blockout — the kind a layout artist makes in an afternoon — goes in as the motion source, and the finished materials, lighting and atmosphere come from stills. The camera choreography you designed survives into the render, which is why coarse geometry works better here than a detailed model.

Write it like a shot list

Treat the model as a producer reading a structured brief. The shape below is the one that survives being handed to someone else.

Two clay figures barter a fleece for a pot beside a fire.A clay trader at a market stall holds up a string of shells.A hammer strikes a gold coin bearing a king’s head on a wooden block.A clay printing press pushes out a banknote.A balance weighs a note against a pan of coins between two figures.A banker sits at a desk with stacks of notes and a safe behind him.A figure pushes a card into a terminal beside a stack of gold bars.A figure taps a phone over a card reader against a neon skyline.The whole diorama pulled back: the set laid out end to end on a table, warm at one end and neon at the other.
Nine frames, in order, from the thirty-second modelling-clay film above. Barter, shell money, a struck coin, a printed note, the scales, the banker, the vault, a contactless tap, the whole set. This is what a shot list looks like once it has been generated: nine beats, each with its own staging, inside a single job.
The structure

Five parts, in this order

Subject, location, event, genre and style, camera movement. Then a shot sequence — whole-second timestamps or numbered shots, either is read — describing per segment what is seen, how the camera moves, what is said and what is heard. Close with the constants: the angle, the environment, the atmosphere, the sound that runs underneath.

Keep about a second of story per second of screen. Too little in a window and the model improvises; too much and it either cuts frantically or drops half of it.

The assets

Map every one, in the text

Number assets in upload order and bind each in the prompt: which is the subject, which the voice, which the action, which the scene. Writing a name onto the image and using it in the prompt is the reliable way to get two of the same character.

Where a reference is already accurate, say to follow it and stop. Re-describing the move it contains gives the model two briefs to reconcile.

img1–2 are character 1, voiced by @audio 1;
img3–4 are character 2, voiced by @audio 2.
Follow the spell-casting action in @video 1 and the
orbit move in @video 2.
Camera language
Standard terms go in as they are — shot sizes, push in, pull out, pan, track, orbit, low angle, one-take, dolly zoom, handheld, speed ramp. A niche term needs a plain-English gloss after it, and a transition needs both its trigger point and its method.
Action and expression
Describe most action broadly and reserve the detail for the two or three beats that have to land. Expressions read better written out than named with an idiom.
Positive over negative
Say what should be there. Negatives are honoured where they are about subtitles and audio — no subtitles, no background music — which are also the two things a model is most likely to add unasked.
Sound
Ambience, effects, dialogue and score can each be cast from an audio reference and described in the brief, including which of them should give way when someone speaks.

What a second of 2.5 costs

Seedance 2.5 is billed per second of output, at the resolution you pick. Plan-tier rates apply to it automatically — the discount belongs to the plan and applies on every generation.

List rate, 720p
42 credits / second

What a Seedance 2.5 generation costs at 720p on standard billing. 480p runs at 24 credits a second — the cheaper way to iterate before a final pass.

Pro plans
−40% on the rate, annual

−20% on monthly billing. The discount applies to every Seedance generation, not once.

Max plans
−60% on the rate, annual

−40% on monthly billing. The deepest rate, and the reason a thirty-second generation is worth running twice.

Current rates for every model live on the pricing page, which carries the figures your plan actually pays.

Questions

Yes — Seedance 2.5 is live in the Popcraft video model picker. Choose it beside Seedance 2.0, set a duration up to thirty seconds, and generate.

No. Seedance 2.0 and Seedance 2.0 4K remain in the model list. A six-second product shot does not need a thirty-second ceiling, and 4K output is a reason to pick 2.0 for it. What 2.5 adds is length, the reference count, and the ability to edit a clip you already have.

Fifty is a ceiling, and the guidance is to stay well under it: eight or fewer image subjects, five or fewer when the main references are audio or video, clips of five to ten seconds, and storyboards of fifteen panels or fewer. Past those, stability drops and generations start needing repeats.

Take a clip you already like and extend it. Extension is the cheapest test of the claim, because it starts from footage you have already judged. If the seam holds — the face, the grade, the camera speed — the rest of the model’s promises are worth spending a longer generation on. If it does not, you have learned that for the price of a few seconds.

Thirty seconds is the whole ad

The join between two fifteen-second clips is where a generated film gives itself away. Seedance 2.5 makes that join optional.

Keep exploring

Related models