
Seedance 2.5 Features: 30-Second Clips and 50 References
The Seedance 2.5 features ByteDance announced: 30-second single clips, 50 multimodal references, region editing, and why native 4K belongs to Seedance 2.0.
At the Volcano Engine FORCE conference in Beijing on June 23, 2026, ByteDance previewed Seedance 2.5, the next version of its Doubao video model. The headline Seedance 2.5 features are a single continuous 30-second clip generated directly, support for up to 50 multimodal reference materials in one generation, and region-level editing[1].
That June event was a preview, not the release. The status has since changed: as of August 17, 2026, Seedance 2.5 is callable on reAPI as doubao-seedance-2.5-face, with 480p, 720p and 1080p output and published production pricing.[3][4] The Seedance 2.5 release-status guide tracks the transition from announcement to API access.
The feature descriptions below still separate ByteDance's launch claims from independently measured results. API availability makes those claims testable, but it does not turn a vendor demo into an independent benchmark.
TL;DR
- Three announced pillars: 30-second single clips (up from 15), up to 50 multimodal references (up from 12), and region-level editing plus 3D previz[1].
- Native 4K is not one of them. The native-4K pipeline belongs to the Seedance 2.0 line announced at the same event. Several outlets merged the two[1].
- API access is live on reAPI: use model ID
doubao-seedance-2.5-face; output is 480p, 720p or (since August 2026) 1080p.[3][4] - The predecessor already leads. Seedance 2.0 sits at #1 on Artificial Analysis's Text-to-Video Arena with an Elo of 1,219 among models with audio[2].
- reAPI pricing is published: $0.118589/s at 480p and $0.266824/s at 720p without source video; source-video rates are $0.071153/s and $0.160094/s respectively.[3]
What ByteDance actually announced

| Capability | Seedance 2.0 | Seedance 2.5 |
|---|---|---|
| Single-clip duration | 15 seconds | 30 seconds, generated directly |
| Reference materials | 9 images + 3 videos + 3 audio clips | 30 images + 10 videos + 10 audio clips |
| Editing control | limited | region-level editing + 3D previz |
According to the keynote, the 30-second clip is produced as one native generation rather than several short clips joined at the seams.[1] The live reAPI contract documents the 50-reference ceiling as 30 images, 10 videos, and 10 audio clips.[4]
Why 30 seconds is harder than it sounds
Doubling maximum clip length from 15 to 30 seconds sounds incremental. In video generation it is one of the harder problems.
Most models hold quality for a few seconds and then drift. Characters subtly change appearance, lighting shifts, motion loses physical plausibility. The common workaround is generating several short clips and stitching them, which introduces visible seams and continuity errors at every join[1].
Generating a coherent 30-second take directly, if it holds outside curated demos, removes a real production tax. For the formats ByteDance named on stage, film and TV pre-visualization, advertising, and short-form animated drama, a continuous half-minute is often the difference between a usable shot and a clip that needs manual repair.
The honest caveat: a single 30-second clip is a ceiling. Sustained quality across that full window is exactly what independent testing has to confirm.
Fifty references is the consistency play

The jump to 50 reference inputs is the upgrade most relevant to professional work. A request can carry up to 30 images, 10 videos, and 10 audio clips, although a smaller, clearly assigned set is usually easier to control.[4]
In practice this is a bid for consistency. Feed the model your character turnarounds, product shots, brand colors, and style frames, and it has far more grounding to hold the same face, the same packaging, and the same look across a sequence.
For an advertiser, character and product fidelity across shots is the entire job. A model that drifts on the logo is unusable regardless of how cinematic the motion is.
Region editing and 3D previz
The third pillar is control. ByteDance demonstrated region-level editing: replacing a subject, background, or product inside an existing shot without changing the original motion, camera move, or lighting[1].
That matters for iteration. Instead of regenerating an entire clip and hoping the rest survives, you swap one element and keep everything already approved. For localizing a product into a different market or swapping a model, that targeted edit is often the whole job.
The keynote also showed 3D white-model previsualization, letting creators block out shots and camera moves in a rough 3D scene before committing to a full generation[1]. That brings storyboarding and camera blocking into the generation step rather than leaving them to trial and error.
The 4K confusion worth clearing up
Native 4K, including 4K 10-bit, was discussed at FORCE, but it belongs to the Seedance 2.0 line, not to the 2.5 headline slide, which leads with duration, references, and editing control[1].
Several outlets folded 4K into the 2.5 spec sheet. We keep them separate because that is what the official materials show, and because a spec that gets attributed to the wrong model is exactly the kind of error that survives into buying decisions. Our earlier piece on what we know about Seedance 2.5 flagged the same press-claimed 4K spec.
Where the shipping version stands
This article does not cite an independent Seedance 2.5 benchmark. For context, the July snapshot below records where Seedance 2.0 stood on a neutral blind human-preference leaderboard.

| Model | Text-to-Video Arena Elo |
|---|---|
| Seedance 2.0 (720p) | 1,219 |
| HappyHorse-1.0 | 1,124 |
| Kling 3.0 1080p Pro | 1,105 |
| Google Veo 3.1 | 1,094 |
Seedance 2.0 currently sits at #1 among models with audio, and leads the Image-to-Video board as well[2]. A first-place arena standing reflects aggregate human preference on sampled prompts, not a guaranteed win on every brief, and emphatically not a 2.5 number.
What it does tell you is that the baseline 2.5 builds on is already at the front of the field, which raises the bar for what the upgrade has to prove.
Pricing context
On reAPI, Seedance 2.5 costs $0.118589 per second at 480p and $0.266824 per second at 720p when no source video is attached. A five-second 720p prompt or image-reference generation therefore reserves 1,335 credits, or $1.335 after credit rounding.[3]
Source-video jobs use lower rates—$0.071153/s at 480p and $0.160094/s at 720p—but they run on a different clock. Billable seconds are the greater of output seconds plus the rounded-up total source-video duration, or ceil(5 / 3 × output seconds).[4] Compare the complete request shape rather than multiplying the discounted rate by output duration alone.
Who the upgrades are aimed at
These are not hobbyist features. They target people for whom consistency and iteration are the job.
Advertisers and brand teams are the clearest fit. The 50-reference ceiling exists so a campaign can hold a product, a logo, and a spokesperson on-model across every shot, and region editing turns localizing an ad into a targeted edit rather than a reshoot.
Pre-visualization teams get 3D white-model blocking and longer takes, which move storyboarding into the generation step.
Short-form animation studios need exactly the multi-shot narrative coherence a continuous 30-second clip enables.
Performance marketers producing volume benefit if the cost posture holds, because the economics of generating many on-brand variants shift decisively.
Consider the concrete case. A consumer-electronics brand needs the same 20-second hero spot in four markets. With a drift-prone short-clip model that is four near-from-scratch generations plus manual continuity fixes. With a reference stack holding the product fixed and region editing swapping only the talent and packaging copy, versions two through four become edits of the first. Same motion, same lighting, different market.
That workflow is why the boring upgrades, references and editing, may matter more than the headline 30 seconds.
What to test now that it is live
Three concrete checks matter because the launch demos were curated and production access now makes repeatable testing possible:
- Does the 30-second clip stay coherent end to end, or does drift appear at second 18 the way it does on shorter models pushed past their comfort zone?
- Do 50 references actually hold a character and a brand across shots, or does the marginal reference stop contributing well before the ceiling?
- Does region editing leave the rest of the frame untouched, including lighting and camera motion?
Until those checks have reproducible results, treat the capability ceilings as vendor specifications rather than quality guarantees.
FAQ
Is Seedance 2.5 available now?
Yes. On reAPI, the callable production model ID is doubao-seedance-2.5-face; it supports 480p, 720p and 1080p output.[3][4]
Does Seedance 2.5 really generate 30-second videos?
ByteDance specifies a single continuous 30-second clip, double the 15-second ceiling of Seedance 2.0.[1] The live reAPI route makes that duration available for testing, but output quality across the full take still depends on the prompt and references.[4]
Does Seedance 2.5 support native 4K?
Not on the current reAPI route. Seedance 2.5 supports 480p and 720p; use Seedance 2.0 for 1080p or 4K.[4]
How does it compare to Veo, Kling, and Sora?
This article does not cite a current independent Seedance 2.5 benchmark. Its comparison table records the cited July Seedance 2.0 leaderboard snapshot, not a 2.5 score.[2]
How many reference materials can Seedance 2.5 take?
Up to 50 on reAPI: 30 images, 10 videos, and 10 audio clips.[4]
What else did ByteDance announce at FORCE?
Seedream 5.0 for image, Seed-Audio 1.0 for audio, and Doubao 2.1 Pro for text, alongside platform figures including roughly 180 trillion tokens processed per day[1].
Can I use Seedance 2.5 through reAPI today?
Yes. Submit doubao-seedance-2.5-face to the video-generation endpoint. The live model page publishes both resolution tiers and the API documentation covers the request and billing rules.[3][4]
Reading the live release clearly
Seedance 2.5 is a confident set of claims from the team whose current model already tops the independent video arenas. The upgrades it leads with are the right ones for professional production, where consistency and iteration matter more than one cinematic clip.
The gap between a polished stage demo and a model holding quality across 30 seconds on your prompts is exactly what the live API can now reveal. Treat the published limits as constraints, run the test list above on representative footage, and choose Seedance 2.0 instead when 1080p or 4K output matters more than the 2.5 workflow.
References
- Volcano Engine. Seedance and the Doubao model family, announced at the FORCE conference. Retrieved July 2026 from volcengine.com
- Artificial Analysis. Text-to-Video Arena leaderboard. Retrieved July 2026 from artificialanalysis.ai/text-to-video/arena
- reAPI. Seedance 2.5 — live model page and production pricing. Retrieved August 17, 2026 from reapi.ai/models/seedance-2-5
- reAPI. doubao-seedance-2.5-face — model ID, supported resolutions, request schema, and source-video billing. Retrieved August 17, 2026 from reapi.ai/docs/seedance-2-5
Further reading
- reAPI. Seedance 2.5 release status. reapi.ai/blog/seedance-2-5-release-status
- reAPI. Seedance 2.5 status: API access, pricing, and limits. reapi.ai/blog/seedance-2-5-what-we-know-2026
- reAPI. Seedream 5.0 Pro and Seedance 2.5 workflow. reapi.ai/blog/seedream-5-0-pro-seedance-2-5-workflow
Author

Categories
More Posts

How to Use Codex to Generate Slide Decks From Research
How to use Codex to generate slides: the seven-step workflow, the outline checkpoint that decides deck quality, prompts that pin exact claims, and its limits.


MiniMax M3 API: 1M Context, Pricing, and Coding Guide (2026)
Use the MiniMax M3 API for coding, agents, and multimodal work. Compare official and reAPI pricing, 1M context behavior, thinking modes, and limits.


Seedance 2.5 Pricing: 53% More Per Token Than Seedance 2.0
Seedance 2.5 pricing is $10.70 per million tokens on BytePlus. See the per-second conversion, reference-video clock, and current reAPI rates.
