adds the lid and moves the drink to the pickup mat.
End state: the completed drink is centered on the pickup mat; both hands have left frame.
[Preserve]
Keep identity, apron, glass design, liquid color, prop ownership, screen direction,
counter geometry, camera axis, lighting, and room ambience consistent.
```
Notice what is absent: three actions squeezed into every second, a new camera move in every stage, and vague endings such as “the scene feels complete.” The model needs a state it can hand from one segment to the next.
## Three ways to use time
The official guide treats time as a pacing budget rather than frame-accurate editing.\[1]
1. **Time range:** `0–6s`, `6–14s`, `14–22s`. Use this to allocate story beats.
2. **Exact point:** “At 12 seconds, the practical light switches from blue to amber.” Use this for one transition.
3. **Relative timing:** “Two seconds after the lid closes, the pickup bell rings.” Use this when one event triggers another.
Keep ranges consecutive and non-overlapping. A range can drift slightly around its boundary, so do not use it to demand impossible event frequency. For camera-specific troubleshooting, see [Seedance 2.5 camera control](/blog/seedance-2-5-camera-control).
## How to prompt a video edit
Editing prompts fail when the requested change is clear but the protected content is not. Use a four-part contract:
```text
[Edit goal]
Edit @Video 1. From 5–8 seconds, change only the desk lamp from white to orange.
[Master]
@Video 1 is the sole editing master for people, action, composition, camera,
occlusion, audio, and event order.
[Edit scope]
Modify only the lamp body and the light it casts on the desk.
[Preserve]
Keep the person's identity, clothing, expression, position, hand motion,
room geometry, camera movement, dialogue, and ambience unchanged.
```
For subject replacement, add two more rules: keep the number of target objects fixed, and make the replacement inherit every appearance, occlusion, movement, and exit of the original object. For background replacement, exclude the subject silhouette from the edit scope.
Dreamina's editing modes are product workflows, not evidence of fields in a public Seedance 2.5 API. Keep the conceptual prompt separate from any future request schema.
## How to prompt video extension
An extension has one non-negotiable job: connect at the boundary.
### Forward extension
```text
@Video 1 is the source video to extend forward.
The first extension frame continues the last source frame. Preserve the cyclist's
pose and direction, bicycle position, road geometry, camera height, afternoon light,
and forward motion.
Then the cyclist passes beneath the bridge and slows beside the orange marker.
Keep one continuous cyclist and one bicycle; do not duplicate either subject.
```
### Backward extension
For a backward extension, describe the new preceding event first. Then define the source video's first frame as the required final state of the extension. Also list any character or effect that must not appear before the source video begins.
Boundary continuity does not mean pixel identity. Review the last frames before the join, the first frames after it, and the full extended sequence.
## First frame, last frame, and multiple keyframes
In multimodal reference mode, define anchor images separately:\[1]
```text
@Image 1 is the first frame. It defines the opening composition and object positions.
@Image 2 is the last frame. It defines the final composition and object positions.
@Image 3 defines the courier's identity and orange jacket only.
The courier carries one parcel from the van to the doorway in one continuous action.
Begin from @Image 1 and arrive naturally at @Image 2.
Preserve identity, parcel count, building geometry, lighting, and camera direction.
```
Use the same aspect ratio for first and last frames. If several images define intermediate stages, state that they are keyframes in order, define the visible state of each one, and ask for continuous transitions between them. Independent images are usually clearer than a dense storyboard collage.
## A practical prompt review checklist
Before generating, check eight things:
1. Is there one clear output goal?
2. Does every reference have one named role and an exclusion?
3. Are repeated views explicitly identified as the same subject?
4. Does each stage contain only one main state change?
5. Is every end state visible and testable?
6. Are dialogue language, speaker, and delivery explicit?
7. Does the preserve list cover identity, count, ownership, space, camera, and audio?
8. For edits or extensions, is the source video's control boundary unambiguous?
If a generation fails, remove ambiguity before adding adjectives. Reduce the reference set, narrow the edit scope, or split an overloaded stage.
## Reuse the prompt structure in an automated workflow
The reference manifest, staged timeline, end states, and preserve list are useful whether a creator generates manually or a product submits tasks through an API. For automation today, use [Seedance 2.0 on reAPI](/models/seedance-2-0), keep the model identifier configurable, and save successful prompts as an evaluation set for the future [Seedance 2.5 route](/models/seedance-2-5).
## FAQ
### What is the best Seedance 2.5 prompt structure?
Use goal, subject and event, scene, reference roles, staged timeline, camera, audio, end states, and a preserve list. Omit blocks that do not affect the result.
### How many references can Dreamina Seedance 2.5 use?
The official prompt guide lists up to 30 images, 10 videos, and 10 audio clips—50 materials in total. Video and audio inputs each have a combined 30-second limit.\[1]
### Should I upload all 50 references?
Usually not. Start with the smallest set that defines identity, product, scene, motion, and audio. More assets create more relationships that must be mapped and preserved.
### How do I write a 30-second prompt?
Use consecutive time ranges. Give each range one main event, an observable end state, and the state inherited from the previous range.
### Can I control an action at an exact second?
You can request an exact timestamp, but treat it as direction rather than frame-accurate editing. The official guide warns that actions may land slightly around a time boundary.\[1]
### How do I stop reference characters from swapping?
Map each character to a specific asset, give each a unique label, state which props belong to whom, and explicitly prohibit identity, clothing, position, action, and dialogue swaps.
### Is the Seedance 2.5 API available on reAPI?
Not yet. The page is a coming-soon preview. Use [Seedance 2.0 on reAPI](/models/seedance-2-0) for a production API and monitor [Seedance 2.5](/models/seedance-2-5) for a verified launch.
## The shorter prompt is often the more controlled prompt
Seedance 2.5 does not need more prose; it needs fewer unstated relationships. Label the inputs, give each stage one job, describe the state that survives the cut, and protect everything outside the requested change.
That is the difference between asking for a video and directing one. Keep the reference manifest and prompt template alongside the project files, then reuse them across revisions instead of rebuilding the brief from memory.
## References
1. ByteDance Dreamina. *Dreamina Seedance 2.5 Prompt Guide.* Modified July 31, 2026. Retrieved August 2, 2026 from [bytedance.larkoffice.com](https://bytedance.larkoffice.com/docx/A88jd0B47oAd8zxWp5ycZFMfnxh)
### Further reading
* reAPI. *Dreamina Seedance 2.5 User Guide.* [reapi.ai/blog/dreamina-seedance-2-5-user-guide](/blog/dreamina-seedance-2-5-user-guide)
* reAPI. *Seedance 2.5 camera control.* [reapi.ai/blog/seedance-2-5-camera-control](/blog/seedance-2-5-camera-control)
* reAPI. *Seedance 2.5 for ecommerce video.* [reapi.ai/blog/seedance-2-5-ecommerce-video](/blog/seedance-2-5-ecommerce-video)
* reAPI. *Seedream-to-Seedance handoff.* [reapi.ai/blog/seedream-seedance-handoff](/blog/seedream-seedance-handoff)
---
# Dreamina Seedance 2.5 User Guide: 30s & Long-Video Modes (https://reapi.ai/blog/dreamina-seedance-2-5-user-guide)
**Dreamina Seedance 2.5 is best understood as a production workflow, not a single text box.** You can build a complete scene of up to 30 seconds, extend eligible source videos into a longer sequence, plan an Ultra Long Video project of up to 180 seconds, direct events with timestamps, combine multimodal references, and repair selected parts without regenerating everything.\[1]
The practical skill is choosing the smallest workflow that solves the shot. If this is your first AI video, you can begin with the eight-second copy-and-change example below. If you already run professional productions, the same guide expands into beat sheets, asset manifests, continuity locks, editing passes, and scene-level review. Every major workflow includes a reusable example and a concrete way to fix common failures.
## TL;DR: choose the workflow before writing the prompt
* Use **standard generation** for one location, one main event, and roughly three to five consecutive beats inside 30 seconds.
* Use **Extend Video** when an existing clip already has the correct cast, framing, environment, and motion. Dreamina's guide says the source must be shorter than 30 seconds, and describes continued extensions up to a 60-second result; the extension prompt controls the newly added segment.\[1]
* Use **Ultra Long Video** for a structured project of up to 180 seconds. Plan it as connected scenes with shared continuity rules, not as one enormous paragraph.
* Assign one job to every image, video, and audio reference. Also state what each reference must not control.
* Use **time ranges** for sustained action, **exact moments** for a visible beat, and **relative timing** when one event triggers another.
* Choose **Smart Edit** for an outcome-level change, **Edit with Marks** for a specific object or region, and **Edit Video** for a change tied to source footage or time.
* For every edit, write three lists: **change**, **preserve**, and **validate**.
* Treat subtitles, dialogue, sound effects, ambience, and BGM as separate layers. Name the layer to remove and the layers to protect.
## Seedance 2.5 workflow at a glance

The six modules solve different production problems. **Create** establishes the shot. **Extend** protects continuity across a boundary. **Direct** assigns events to time. **Edit** changes the smallest possible target. **Localize** controls language and sound. **Block** supplies spatial structure through storyboards or Clay Renderer references.
## How to use this guide as a beginner or a professional
You do not need a production vocabulary to start. Beginners can copy the smallest example, change the subject and scene, and generate a short result. Experienced creators can use the same example as a base layer, then add references, timed beats, continuity locks, editing passes, and formal review criteria.
| Your experience | Start here | Add only after the baseline works | Skip at first |
| -------------------------- | ------------------------------------------------------------ | ------------------------------------------------------------ | ----------------------------------------------------- |
| First AI video | One subject, one action, one scene, 5–10 seconds | One camera instruction and one end state | Large reference packs and multi-scene stories |
| Some generation experience | 15–30-second beat sheet and 1–3 references | Audio, exact timing, first/last frames | Ultra-long projects before continuity is stable |
| Professional creator | Continuity sheet, asset manifest, scene cards, review rubric | Editing passes, Clay blocking, audio plan, transition design | Nothing by default—choose controls by production need |
### Beginner example: make one clear eight-second shot
Copy this complete example, then replace the subject, action, and location with your own.
**Input — complete prompt**
```text
Create an 8-second realistic video.
One golden retriever walks into a sunny kitchen, stops beside a blue food bowl,
and looks toward the camera.
Use a fixed eye-level medium shot. Keep the dog fully visible.
Natural morning room sound, no dialogue, no music.
End with the dog standing beside the bowl and looking at camera.
```
Why it works:
* **One subject:** the model does not need to track a cast.
* **One action chain:** enter, stop, look.
* **One location:** there is no scene transition to invent.
* **One camera rule:** a fixed medium shot reduces competing direction.
* **One end state:** you can immediately decide whether the result passed.
If the dog never reaches the bowl, shorten the entrance or remove “looks toward the camera.” Do not solve an overloaded action by adding decorative adjectives.
### Advanced version: build on the same shot
Once the baseline works, add controls instead of replacing the whole prompt:
**Input — complete prompt**
```text
Create a 12-second realistic video.
@Image 1 defines the same golden retriever's face, coat color, and red collar only.
@Image 2 defines the sunny kitchen layout and blue bowl only.
Do not copy the people, camera angle, or background objects from the references.
0–5 seconds: the dog enters from frame left and walks toward the blue bowl.
End state: all four paws stop beside the bowl.
5–9 seconds: the dog looks down at the bowl, then raises its head.
End state: the dog's head faces forward.
9–12 seconds: make one slow camera push-in as the dog looks at camera.
End state: the same dog and one blue bowl remain centered.
Preserve identity, red collar, bowl count, kitchen layout, morning light,
screen direction, floor contact, and natural room ambience.
```
The professional version adds identity ownership, reference exclusions, timing, camera movement, cardinality, and continuity. The story remains simple; only the control becomes more precise.
### Find the example that matches your task
| Task | Example in this guide |
| ------------------------------------- | ------------------------------------------------ |
| First text-to-video generation | Eight-second dog-and-bowl example above |
| Image-referenced character or product | Advanced dog example and asset-manifest section |
| Full 30-second scene | Travel-bottle timed prompt |
| Continue a successful clip | 10-second source plus 5-second athlete extension |
| Change one object | Marked suitcase replacement |
| Remove subtitles or BGM | Cleanup examples with protected layers |
| Two-person dialogue | Mina and Owen identity-and-voice example |
| Replace a green background | Rainy station composite example |
| Guide complex motion | Clay Renderer handoff example |
| Build a long story | 180-second designer scene-card example |
## Before generating: make a one-page production brief
A good production brief prevents more failures than a longer prompt. Write the information that must survive every generation, extension, or edit before uploading any references.
### Build a continuity sheet
| Continuity item | What to record | Example lock |
| --------------- | ----------------------------------------- | ------------------------------------------------------ |
| Character | Face, hair, clothing, age range, voice | Mina keeps the same short black hair and orange jacket |
| Prop | Shape, color, count, owner | One silver parcel, always carried by Mina |
| Space | Entrance, exits, relative positions | Counter remains left; pickup door remains frame right |
| Camera | Axis, height, direction, movement | Eye-level camera; no axis reversal |
| Light | Direction, color, time of day | Soft morning light from frame left |
| Audio | Dialogue language, voice, ambience, music | English dialogue, quiet store ambience, no BGM |
| Start state | What is visible before the action | Mina stands outside with the parcel in her right hand |
| End state | What must be visible after the action | Parcel is centered on the counter; both hands leave it |
Use observable locks. “Keep the scene cinematic” cannot be verified. “The orange jacket, one parcel, left-to-right movement, and morning light remain unchanged” can.
### Create an asset manifest
Dreamina's official prompt guide documents up to 30 images, 10 videos, and 10 audio clips, with 50 materials in total. Video references and audio references each have a combined limit of 30 seconds.\[2] The maximum is capacity, not a recommended starting point.
| Asset | Its only job | Inherit | Do not inherit |
| ---------- | ------------------ | -------------------------------------- | ----------------------------- |
| `@Image 1` | Character identity | Face, hair | Background, pose, clothing |
| `@Image 2` | Wardrobe | Orange jacket, black trousers | Model's face, studio lighting |
| `@Image 3` | Location | Counter layout, doorway, morning light | People, signs, products |
| `@Video 1` | Motion | Walking pace, parcel handoff | Performer, room, camera color |
| `@Audio 1` | Voice | Speaker tone and delivery | Words from the sample |
Start with the smallest set that fully defines the scene. Add another asset only when it resolves a specific failure.
### Define pass conditions before generation
Reviewers need a shared definition of “done.” A compact pass list might be:
1. The same character appears in every beat.
2. Exactly one parcel remains in the scene.
3. The action order matches the beat sheet.
4. The camera stays on the same side of the action.
5. Dialogue belongs to the correct speaker.
6. The last frame reaches the stated end condition.
These checks turn subjective rerolling into a repeatable review process.
## Start a project in Dreamina
1. Open the [official Dreamina creator site](https://dreamina.capcut.com/ai-tool/home) and enter **AI Video**.
2. Select Seedance 2.5, then choose the input path that matches the job: text, multimodal references, first and last frames, an existing source video, extension, or an editing workflow.
3. Set duration and aspect ratio before building references. Anchor images should use the same ratio as the intended output.
4. Upload the smallest useful reference pack and label each asset's role in the prompt.
5. Generate once, compare the result with the pass conditions, then choose between a prompt revision and a localized edit.
Do not respond to every failure by adding more text. If the identity is wrong, improve identity mapping. If the event order is wrong, simplify the beat sheet. If only one object is wrong, edit that object instead of regenerating the scene.
## Choose between 30 seconds, extension, and Ultra Long Video
The duration options represent different production methods.
| Workflow | Described ceiling | Best for | Planning unit | Main risk |
| ------------------- | ---------------------------------------------------------- | ---------------------------- | ------------------------------- | -------------------------------------- |
| Standard generation | Up to 30 seconds | One complete scene | Timed beats | Too many events for one timeline |
| Continued extension | Source under 30 seconds; continued result up to 60 seconds | Continuing a successful clip | Boundary state plus new segment | Visible jump at the join |
| Ultra Long Video | Up to 180 seconds | Multi-scene stories | Scene cards and transitions | Character, prop, light, or audio drift |
### Build one complete scene in 30 seconds
Thirty seconds works best when the viewer stays in one location and follows one main change. Begin with the final frame you need, then work backward into three or four stages.
**Input — complete 30-second prompt**
```text
[Goal]
Create a 30-second product story for one reusable travel bottle.
[References]
@Image 1 defines the bottle shape, matte-blue finish, and white logo placement only.
@Image 2 defines the kitchen layout and morning window light only.
Keep exactly one bottle. Do not copy people from either image.
[0–7 seconds]
One runner enters the kitchen and places the closed bottle at the center of the counter.
End state: the bottle stands upright; both hands leave it.
[7–16 seconds]
The runner opens the bottle, adds water, and closes the lid.
End state: the lid is fully closed; no water is spilled.
[16–24 seconds]
The runner picks up the same bottle and walks toward the door.
End state: the runner reaches the doorway with the bottle in the right hand.
[24–30 seconds]
Slow push-in as the runner turns the bottle label toward camera.
End state: one bottle fills the center third; the label is readable; the runner stops moving.
[Audio]
Natural kitchen ambience, lid click at 14 seconds, no dialogue, no BGM.
[Preserve]
Bottle geometry, matte-blue color, logo position, runner identity, wardrobe,
screen direction, kitchen layout, morning light, and camera axis.
```
This prompt gives every interval one main state change. If the model skips an event, reduce the number of actions before increasing the duration.
### Extend an existing video without a visible jump
Dreamina's guide describes a concrete Extend Video flow and continued results up to 60 seconds:\[1]
1. Choose the target clip in the generation stream and open its extension control.
2. Use an original clip shorter than 30 seconds; the guide marks that as an eligibility condition.
3. Set the added duration using the visible duration control, timeline gesture, or direct value available in the interface.
4. Write only the action required for the newly added portion. The source footage remains the master for the original segment.
5. Generate the combined result, then review the seconds immediately before and after the join.
Suppose the source is 10 seconds and the added segment is 5 seconds. The extension instruction should describe those new five seconds, not rewrite the first ten:
**Input — complete extension prompt**
```text
Extend @Video 1 forward by 5 seconds.
[Boundary state]
Continue from the final source frame: the athlete has just landed, knees bent,
right hand touching the floor, camera low and facing the athlete, blue arena light.
[New action]
The athlete rises, looks directly toward the camera, and takes one controlled breath.
Do not repeat the jump or add another landing.
[Preserve]
Same athlete, uniform, body position at the join, camera height, lens direction,
arena geometry, blue light, motion speed, crowd ambience, and source aspect ratio.
[End state]
The athlete stands still at center frame and looks at camera; both feet remain planted.
```
Review the join at normal speed and frame by frame. Check pose, screen position, movement direction, prop count, lighting, focus, camera velocity, ambience, and audio level. A good continuation begins by matching the boundary; new story information comes second.
### Plan an Ultra Long Video project of up to 180 seconds
Do not write a 180-second prompt as one uninterrupted paragraph. Divide the project into scene cards and give every card an entry state, one primary event, an exit state, and a transition.
| Scene card | Duration | Entry state | Primary event | Exit state | Transition |
| -------------- | -------: | ------------------------------ | --------------------------------- | --------------------------------- | --------------------------------------- |
| 1. Workshop | 0–25s | Designer enters empty workshop | Unpacks prototype | Prototype centered on table | Match cut on circular dial |
| 2. Street test | 25–60s | Dial fills frame | Product tested in rain | Product still working | Water droplet becomes window reflection |
| 3. Train | 60–105s | Reflection on train window | Designer reviews data | Green result appears | Screen glow becomes dawn light |
| 4. Lookout | 105–150s | Dawn light on face | Final field test | Designer smiles and packs product | Bag closes across frame |
| 5. End card | 150–180s | Dark frame after bag close | Product reveal and closing motion | Product centered, motion stopped | Hold final composition |
Prepare four reusable documents before generating:
* a **character bible** for face, hair, wardrobe, posture, and voice;
* a **prop ledger** for object appearance, count, owner, and condition;
* **location rules** for layout, time of day, weather, and camera direction;
* an **audio plan** for dialogue, ambience, effects, music, and transitions.
Test the hardest scene first. If a rain sequence, two-person exchange, or complex camera move cannot hold continuity in isolation, a longer timeline will magnify the problem.
## Use multimodal references without making them fight
Multimodal input is useful only when ownership is explicit. A face image, clothing image, motion clip, room image, and voice sample can work together because each controls a different layer.
### Set a priority order for conflicting references
When two assets contain overlapping information, state which one wins:
```text
Reference priority:
1. @Image 1 controls face and hair.
2. @Image 2 controls clothing only.
3. @Video 1 controls motion timing and camera movement only.
4. @Image 3 controls room layout and light only.
5. @Audio 1 controls voice identity and delivery only.
If any reference conflicts, follow this priority order.
```
A weak instruction says, “Use all references for the character and style.” A controlled instruction names the owner of every attribute and excludes the irrelevant material in each file.
### Keep first and last frames compatible
The first anchor establishes the opening composition and output ratio. The last anchor defines where the motion must arrive. Use matching aspect ratios, keep the subject count consistent, and explain what connects the two states. Additional identity references should not override the anchor compositions.\[2]
## Direct the timeline with three kinds of time instruction
Dreamina supports timestamp-based direction, but the timeline still needs a realistic action budget.\[1]
| Timing method | Use it for | Example |
| --------------- | -------------------------------- | ------------------------------------------------------------ |
| Time range | Sustained action or a story beat | `6–12s: she opens the case and removes one camera` |
| Exact moment | One visible or audible change | `At 18s, the red practical light turns blue` |
| Relative timing | Cause and effect | `Two seconds after the door closes, the train begins moving` |
### Build a four-track beat sheet
| Time | Picture | Dialogue | Effects and ambience | Music |
| ------ | -------------------------------------- | ---------------------------- | -------------------- | ----------------------- |
| 0–6s | Mechanic opens workshop | None | Door roll, room tone | None |
| 6–14s | Mechanic places one camera on bench | “Let's test the stabilizer.” | Case latch | Low pulse begins |
| 14–22s | Camera powers on; status light changes | None | Power chime | Pulse continues quietly |
| 22–30s | Mechanic demonstrates one smooth pan | “No shake.” | Motor sound | Music ends at 29s |
Then turn the table into prompt language and add a visible end state to every range. Avoid assigning simultaneous dialogue, a prop change, a location change, and a complex camera move to the same two seconds.
Treat timestamps as direction, not a promise of frame-accurate nonlinear editing. If a key action lands too early, simplify the preceding beat or express the timing relative to a clear trigger.
## Choose the right editing workflow
Dreamina lists Smart Edit, Edit with Marks, and Edit Video as distinct workflows.\[1] They share a preserve-first discipline, but they solve different problems.
| Workflow | Best suited to | Prompt must identify | Main review risk |
| --------------- | --------------------------------- | ---------------------------------------------- | ----------------------------- |
| Smart Edit | An outcome-level correction | Desired result and protected content | The change spreads too widely |
| Edit with Marks | One object or bounded region | Marked target, replacement, edges, occlusion | Flicker or damaged boundaries |
| Edit Video | A source-footage change over time | Source master, time range, action, audio locks | Timing or motion drift |
### Smart Edit: describe the result and the boundary
Use Smart Edit when the desired correction is easy to state but not tied to one tiny shape.
**Input — Smart Edit instruction**
```text
Change the afternoon scene to light rain.
Add rain only outside the café windows and on the exterior pavement.
Preserve the two people, faces, hair, clothing, table objects, indoor lighting,
dialogue, camera movement, and all reflections already inside the café.
Validate that no rain appears indoors and no person's appearance changes.
```
### Edit with Marks: control one object or region
Mark the smallest useful target. Describe the replacement, what should appear around its edges, and how it behaves under occlusion.
**Input — marked-object instruction**
```text
The marked red suitcase is the only editable object.
Replace it with one navy hard-shell suitcase matching @Image 2.
Keep its original size, path, wheel contact, hand contact, shadows, occlusion,
and time in frame. Preserve every person and all unmarked luggage.
```
### Edit Video: make the source the master
**Input — source-video instruction**
```text
@Video 1 is the sole master for composition, people, action order, camera, and audio.
From 8–12 seconds, change only the desk lamp body from white to orange.
Update the lamp's local light spill on the desk.
Preserve faces, clothing, hand motion, desk geometry, camera movement,
dialogue, room ambience, and every event outside 8–12 seconds.
```
The official prompt guide notes that editing preserves the source ratio and approximately preserves duration, with a possible small difference of about 0.3 seconds from frame handling.\[2] Review both edit boundaries instead of judging only the middle frame.
## Remove subtitles, BGM, and unwanted visual elements
Cleanup tasks fail when “audio” or “text” is treated as one layer. Separate the categories before editing.
### Remove irrelevant subtitles without deleting useful text
**Input — subtitle cleanup instruction**
```text
Remove the subtitle line at the bottom from 4–9 seconds only.
Reconstruct the floor texture and moving shadow behind the removed letters.
Preserve the store sign, product label, wall poster, faces, camera movement,
dialogue, sound effects, and all text outside the subtitle region.
```
Check for letter fragments, soft rectangles, repeating background texture, and flicker as the camera moves.
### Detach or remove BGM while protecting speech
Write an audio manifest before the edit:
**Input — audio-layer instruction**
| Layer | Action |
| ------------- | ----------------------------------------- |
| Dialogue | Preserve both speakers and timing |
| Sound effects | Preserve door, footsteps, glass placement |
| Ambience | Preserve quiet restaurant room tone |
| BGM | Remove throughout the clip |
After cleanup, listen for clipped consonants, pumping volume, missing ambience, or a sudden noise-floor change. “Remove all background audio” is too broad when the scene needs effects and room tone.
### Remove one object across time
For partial elimination, name the object, time range, newly revealed background, and occlusion behavior:
**Input — object-removal instruction**
```text
Remove the black microphone stand from 0–14 seconds.
Reconstruct the wooden stage floor and blue curtain behind it.
When the singer crosses the area, preserve the singer in front and rebuild only
the hidden portion of the stand. Keep the microphone in the singer's hand.
```
Review the complete motion path. A clean still frame can still hide a ghost, popping edge, or broken shadow in motion.
## Transfer an idea without copying unwanted details
“Transfer Ideas” is most useful when you isolate the layer worth borrowing: action logic, composition, camera rhythm, transition design, or story structure. Do not ask the source to control everything.
**Input — transfer instruction**
```text
[Transfer]
Use @Video 1 only for the sequence: reveal object, circle it once, then end on a top view.
[Replace]
Use the ceramic tea set from @Image 1, the quiet studio from @Image 2,
and the warm paper texture described below.
[Do not inherit]
Do not copy the source performer, brand marks, text, colors, room, product,
music, or voice.
```
Review the result for accidental source identities, logos, wording, color palettes, and props.
## Change spatial perspective without breaking geometry
A perspective edit needs spatial relationships, not an invented focal-length number. Describe the scene in layers:
* **foreground:** bicycle wheel crosses the lower-left edge;
* **midground:** courier and parcel remain centered;
* **background:** doorway stays behind the courier, two meters away;
* **view direction:** camera moves from front-left to side view without crossing behind the courier;
* **preserve:** body proportions, parcel shape, ground contact, horizon, and light direction.
After the edit, check scale, horizon, vanishing direction, occlusion order, contact shadows, and feet touching the floor. Objects that merely change viewpoint should not change size or ownership without a stated reason.
## Control tone references and multi-person scenes
Multi-person generation becomes much more stable when every person has a profile and every line names its speaker.
**Input — complete two-person prompt**
```text
= face from @Image 1 + orange jacket from @Image 2 + voice from @Audio 1.
= face from @Image 3 + grey shirt from @Image 4 + voice from @Audio 2.
0–6 seconds: Mina stands frame left; Owen stands frame right. Both remain still.
Mina says in calm English: {The test starts now.}
6–12 seconds: Owen presses one button with his right hand.
Owen says in a lower, measured voice: {Power is stable.}
12–18 seconds: Mina checks the display while Owen keeps both hands off the device.
No speaker, voice, clothing, position, or action swaps.
```
For a tone or voice reference, say whether the asset controls voice identity, pitch range, pace, emotion, or delivery. Keep the written dialogue separate so the sample does not accidentally supply the words.
## Use green-screen editing for cleaner composites
The subject silhouette is the protected asset. Preserve face, hair strands, semi-transparent edges, clothing, pose, subject scale, motion blur, and movement timing while changing only the background.
**Input — green-screen instruction**
```text
Replace the green background with a rainy station platform at dusk.
Keep the performer silhouette, hair edges, transparent raincoat, pose, scale,
walking motion, camera movement, and source duration unchanged.
Match the new background perspective and camera speed.
Add cool light from frame right, subtle wet-floor reflection, and a contact shadow
under both feet. Remove green spill without changing the raincoat color.
```
Inspect hair, fingers, motion-blurred edges, reflective clothing, feet, and any object crossing the silhouette. Edge shimmer is usually more visible during motion than on the preview frame.
## Use Clay Renderer for motion and spatial blocking
Clay Renderer turns a simple white-model or blocked 3D scene into a spatial guide. The block should control geometry, camera path, movement, contact, and occlusion; separate references should control final identity, wardrobe, materials, environment, and visual style.\[1]
**Input — Clay Renderer handoff**
```text
@Video 1 is the Clay Renderer blocking reference.
Inherit only the two-person motion, camera path, table position, hand contact,
and occlusion order from @Video 1.
Replace Character A with and Character B with .
Use the workshop environment from @Image 5 and product materials from @Image 6.
Do not inherit white clay materials, placeholder faces, untextured background,
or temporary lighting from the blocking reference.
Preserve the blocked walking paths, handoff timing, subject scale,
camera direction, table contact, and final positions.
```
Compare the result against the block for path, contact, collision, scale, camera direction, and occlusion. Clay Renderer is a reference workflow; it should not be described as a native Maya or Blender scene-file integration unless an official integration is documented.
## Build seamless transitions and multi-grid storyboards
A storyboard grid defines ordered visual states. It does not automatically explain the motion between them. Give each panel one job and write a transition sentence between every pair.
| Panel | Anchor state | Transition instruction |
| ----- | ------------------------------ | -------------------------------------------------------------- |
| 1 | Chef holds closed box at waist | Camera follows as the box rises toward the table |
| 2 | Box centered on table | Chef opens lid while camera moves closer without changing axis |
| 3 | Product visible inside box | Product rotates once as chef's hands leave frame |
| 4 | Product fills center frame | Light softens and movement settles into final hold |
At each join, define the shared state: character position, facing direction, velocity, hand pose, prop ownership, camera direction, lighting, and ambient sound. Useful transition patterns include continuous action, matched composition, foreground occlusion, and a shared shape. Treat them as directing methods, not guaranteed interface presets.
## Three end-to-end workflow examples
### Example 1: a 30-second product story
**Brief:** Show one compact speaker unfolding, powering on, and playing music on a hotel desk.
**Assets:** front, side, and rear product images; one hotel-room reference; one unfolding motion clip; one power-on chime.
**Sequence:** map each image to one product view, write three timed beats, lock the speaker count to one, and state the final label orientation. Generate the baseline before adding camera movement.
**Review:** product geometry, hinge direction, button location, one-speaker count, desk contact, chime timing, and final label visibility.
**Likely failure:** three product views become three speakers. **Correction:** state that all product images describe the same unit and keep exactly one speaker in every frame.
### Example 2: a 15-second result built from a 10-second source
**Brief:** Continue a successful athletic landing for five seconds.
**Assets:** one eligible 10-second source clip; no new identity reference unless the face already drifts.
**Sequence:** select Extend Video, add five seconds, restate the final pose and camera state, then describe the rise and final look. Do not rewrite the original ten seconds.
**Review:** one continuous athlete, no repeated landing, matching pose at the join, stable arena geometry, continuous camera motion, and unchanged ambience.
**Likely failure:** the extension begins with the athlete already standing. **Correction:** make the first extension frame continue the crouched landing pose before any new action.
### Example 3: a multi-scene 180-second brand film
**Brief:** Follow one designer testing a prototype across workshop, street, train, and outdoor locations.
**Assets:** character profile, wardrobe views, prototype views, four location references, motion clips for the hardest actions, voice sample, and storyboard anchors.
**Sequence:** create scene cards, lock the prototype owner and condition, define audio transitions, generate the most difficult scene first, then connect scenes using shared shapes or motion.
**Review:** face, clothing, prototype geometry, prop condition, time-of-day progression, screen direction, voice identity, BGM continuity, and every scene boundary.
**Likely failure:** wardrobe or product details drift after the second location. **Correction:** repeat the character and prop profiles in every scene card and reduce scene-specific references that contain conflicting people or products.
## Quality-control checklist before export
### Identity and objects
* [ ] Each person keeps the assigned face, hair, clothing, voice, and position.
* [ ] Object count, owner, geometry, label, and condition remain consistent.
* [ ] No reference-only person, logo, subtitle, or prop leaks into the result.
### Timing and story
* [ ] Every beat occurs in the intended order.
* [ ] Each time range reaches a visible end state.
* [ ] Dialogue, effects, and music support rather than compete with the action.
* [ ] The final frame satisfies the brief.
### Camera and space
* [ ] Camera axis, height, direction, and movement remain intentional.
* [ ] Scale, horizon, contact, shadows, reflections, and occlusion make sense.
* [ ] Extensions and transitions have no visible pose, light, or audio jump.
### Edits and cleanup
* [ ] Only the intended target changed.
* [ ] Edit boundaries do not flicker or drift.
* [ ] Removed text or objects reveal a plausible background.
* [ ] BGM cleanup preserves dialogue, effects, and ambience requested in the brief.
## Troubleshooting common Seedance 2.5 problems
| Symptom | Likely cause | Practical fix |
| ------------------------------------- | ----------------------------------------------- | ----------------------------------------------------------------------- |
| Timestamp ignored | Too many actions compete in one range | Reduce events and give the range one end state |
| Important action omitted | Timeline is overfilled | Split the action or move it into another stage |
| Character identity swaps | References are not mapped per person | Create named profiles and ban face, clothing, voice, and position swaps |
| References fight each other | Several assets control the same attribute | Add an explicit priority order and exclusions |
| Extension begins with a jump | Boundary pose or camera state is missing | Restate the final source frame before new action |
| Perspective edit distorts the subject | Spatial layers and protected geometry are vague | Define foreground, midground, background, scale, horizon, and contact |
| Removed object leaves a ghost | Hidden background and occlusion are unspecified | Describe what must be reconstructed through the full time range |
| Subtitle area flickers | Cleanup is judged on one frame | Review motion and require consistent texture reconstruction |
| BGM removal damages dialogue | Audio layers were grouped together | List dialogue, effects, ambience, and music separately |
| Green-screen edge shimmers | Fine edges and motion blur were not protected | Lock hair, fingers, transparent material, blur, and silhouette |
| Storyboard transition feels abrupt | Panels define states but not connecting action | Write a transition sentence between every pair |
| Long-video scenes drift | Global continuity rules are not repeated | Reuse character, prop, location, and audio profiles in every scene card |
## Using the workflow in a product pipeline
Dreamina is useful for hands-on creation and editing. For software automation, [Seedance 2.5 on reAPI](/models/seedance-2-5) is not yet callable; [Seedance 2.0](/models/seedance-2-0) is the current production option. Keep the model name configurable, save the prompts and pass conditions above as an evaluation set, and reuse the same continuity checks when a verified 2.5 endpoint becomes available.
## FAQ
### How much action should fit into a 30-second prompt?
Use one main event with roughly three to five beats. Give every beat one observable end state. If an action is repeatedly omitted, reduce the event count instead of adding more adjectives.
### How do I extend a scene without a visible jump?
Begin the extension by restating the source video's final pose, subject position, movement direction, camera state, light, and audio. Describe new action only after that boundary is locked.
### Can any source video be extended?
The official Dreamina guide says the original source must be shorter than 30 seconds for the described Extend Video workflow.\[1]
### Should I use all 50 reference slots?
Usually not. Use the smallest pack that defines identity, wardrobe, props, scene, motion, and sound. Every additional asset creates another relationship that must be mapped and preserved.
### What is the difference between Smart Edit and Edit with Marks?
Smart Edit is better for an outcome-level correction. Edit with Marks is better when you can point to one object or bounded region. In both cases, list what must remain unchanged.
### How do I stop two speakers from swapping identity or voice?
Create a named profile for each person, bind face, clothing, and audio references separately, name the speaker in every dialogue line, and specify who remains still during the other person's action.
### Can I remove BGM while keeping dialogue and effects?
Treat dialogue, sound effects, ambience, and music as four separate layers. Ask to remove BGM and explicitly preserve the other three, then listen for voice damage and level jumps.
### What should Clay Renderer control?
Use it for blocking: geometry, camera path, movement, contact, and occlusion. Use separate references for identity, clothing, material, environment, and final style.
### Why does a storyboard still need transition instructions?
Panels define anchor states, not all intermediate motion. A transition sentence explains how position, speed, camera, lighting, and props move from one anchor to the next.
## Build the shot, then protect what already works
Seedance 2.5 becomes easier to direct when each workflow has a narrow job. Build short scenes from timed end states. Extend from a documented boundary. Organize longer projects with scene cards. Map every reference. Use localized edits only after defining what cannot change.
The most useful habit is also the simplest: before every generation or edit, write **change**, **preserve**, and **validate**. That turns a feature list into a repeatable production process.
## References
1. ByteDance Dreamina. *Dreamina Seedance 2.5 User Guide.* Modified July 31, 2026. Retrieved August 2, 2026 from [bytedance.larkoffice.com](https://bytedance.larkoffice.com/wiki/NjnWwvf4BiFYFLk2RzrcEgaunGf)
2. ByteDance Dreamina. *Dreamina Seedance 2.5 Prompt Guide.* Modified July 31, 2026. Retrieved August 2, 2026 from [bytedance.larkoffice.com](https://bytedance.larkoffice.com/docx/A88jd0B47oAd8zxWp5ycZFMfnxh)
### Further reading
* reAPI. *Dreamina Seedance 2.5 Prompt Guide.* [reapi.ai/blog/dreamina-seedance-2-5-prompt-guide](/blog/dreamina-seedance-2-5-prompt-guide)
* reAPI. *Seedance 2.5 camera control.* [reapi.ai/blog/seedance-2-5-camera-control](/blog/seedance-2-5-camera-control)
* reAPI. *Seedance 2.5 for ecommerce video.* [reapi.ai/blog/seedance-2-5-ecommerce-video](/blog/seedance-2-5-ecommerce-video)
* reAPI. *Seedance 2.0 API documentation.* [reapi.ai/docs/seedance-2-0](/docs/seedance-2-0)
---
# Free AI API tiers in 2026: what each one actually limits (https://reapi.ai/blog/free-ai-api-limits-2026)
If you searched for a free AI API expecting to generate images or video without paying, Google's own pricing page has bad news. Every Nano Banana variant, all three Imagen 4 models, every Veo version and the Lyria music models list their Free Tier price as the same two words: "Not available"\[1]. The free tier is real, but it covers text.
That distinction is buried in a table most people skim past, and it explains a lot of frustration. Below is what each major provider's free tier actually permits, pulled from their own pricing pages in July 2026, plus the row in Google's table that almost nobody reads.
## TL;DR
* **Google's free tier has no image, video or music models.** Nano Banana, Nano Banana 2, 2 Lite, Pro, Imagen 4, Veo 2, Veo 3, Veo 3.1 and Lyria 3 all show "Not available" under Free Tier\[1].
* **Google marks free-tier input as used to improve its products.** Every model block carries the row "Used to improve our products", answered Yes for Free Tier and No for Paid Tier\[1]. Google does not spell out what "improve" covers.
* **Free-tier rate limits are no longer published.** Google says limits "can be viewed in Google AI Studio" and that "Specified rate limits are not guaranteed and actual capacity may vary"\[2].
* **OpenAI documents a free test request, but no ongoing allowance.** Its quickstart congratulates you on "running a free test API request"\[8]; the pricing page lists no free inference allowance beyond that\[3].
* **Replicate does have free limits; fal does not advertise any.** Replicate lets you "run select models on Replicate for free, but after a bit you'll be asked to set up billing"\[7]. fal's pricing page carries no free-tier language at all\[4].
* Paying a fraction of a cent per image is a different problem from paying nothing, and it is a much easier one to solve.
## What Google's free AI API tier actually covers
Google's Gemini API pricing page lists a Free Tier column next to a Paid Tier column for every model. Twenty models show "Free of charge" in that column: Gemini 3.6 Flash, 3.5 Flash, 3.5 Flash-Lite, 3.5 Live Translate, 3.1 Flash-Lite, 3.1 Flash Live Preview, 3.1 Flash TTS Preview, 3 Flash Preview, 2.5 Pro, 2.5 Flash, 2.5 Flash-Lite, 2.5 Flash Native Audio, 2.5 Flash Preview TTS, the Embedding models, Robotics-ER 1.6 Preview and Gemma 4\[1].
Every generative media model is on the other list. The table below is copied from that page.
| Model | Free Tier | Paid Tier |
| ----------------------------------------- | ------------- | --------------------------------- |
| Nano Banana (Gemini 2.5 Flash Image) | Not available | $0.039 per image |
| Nano Banana 2 Lite (3.1 Flash Lite Image) | Not available | $0.0336 per 1K image |
| Nano Banana 2 (3.1 Flash Image) | Not available | $0.067 per 1K image |
| Nano Banana Pro (Gemini 3 Pro Image) | Not available | $0.134 per 1K/2K image |
| Imagen 4 Fast / Standard / Ultra | Not available | $0.02 / $0.04 / $0.06 per image |
| Veo 3.1 Lite | Not available | $0.05 per second (720p) |
| Veo 3.1 Fast | Not available | $0.10 per second (720p) |
| Veo 3.1 Standard | Not available | $0.40 per second (720p and 1080p) |
| Veo 2 | Not available | $0.35 per second |
| Lyria 3 Clip / Pro | Not available | $0.04 / $0.08 per song |
Google is explicit about it in prose too. The Veo 3.1 description reads "available to developers on the paid tier of the Gemini API"\[1].
So when someone asks whether Nano Banana has a free API, the answer from Google's own documentation is no, on any of the four variants, for input or output. The Grounding with Google Search line on those models is also marked "Not available" for free tier\[1].
## The row nobody reads
Scroll to the bottom of any model block on that pricing page and there is a row labelled "Used to improve our products". For Free Tier it says Yes. For Paid Tier it says No\[1].
That row appears under every model, including the text models that genuinely are free of charge. Google does not define "improve our products" on that page, so the honest reading is the narrow one: on the free tier your prompts, and whatever images or documents you attach, are retained for Google's own use in a way the paid tier excludes.
For a weekend project that is a fine trade. For anything touching client work, internal documents or user-submitted content, it is worth reading twice before you paste a production prompt into a free-tier key. The paid tier flips that row to No, which is the actual product difference between the two tiers on many models.
## Rate limits you cannot look up before signing up
Until recently you could read Google's free-tier requests-per-minute and requests-per-day figures straight from the docs. That table is gone. The rate limits page now says limits "depend on a variety of factors (such as your usage tier) and can be viewed in Google AI Studio", followed by "Specified rate limits are not guaranteed and actual capacity may vary"\[2].
The usage-tier table is still there, and it is short. The Free tier qualification is "Active project or free trial" with no billing cap. Tier 1 requires you to "Set up and link an active billing account" and caps at $250. Tier 2 needs $100 paid plus three days, Tier 3 needs $1,000 paid plus thirty days\[2].
One free-tier ceiling is still published, and it is worth knowing: Search grounding on Gemini 2.5 Flash is "Free of charge, up to 500 RPD", with that daily limit shared with Flash-Lite\[1]. Everything else about free-tier throughput is per-project and only visible from inside your own account.
Two practical consequences. You cannot capacity-plan a free AI API before you have an account. And moving off the free tier is a billing-account step, not a plan upgrade, so the "just add a card later" migration is a little more involved than it sounds.
## What the other providers mean by free
The picture outside Google is simpler, and mostly it means there is no free AI API to discuss at all.
**OpenAI.** The quickstart ends with "Congrats on running a free test API request!" and then points you at billing\[8]. OpenAI does not publish how many such requests a new key gets, only that the walkthrough is free to complete. Beyond that the pricing page lists per-token and per-image rates with no free inference allowance; its only free quantities are storage, 1 GB for file search and 1 GB per account per month for ChatKit uploads\[3]. Free ChatGPT is a consumer product and does not come with API credits.
**fal.ai.** The pricing page is titled Pay-Per-Use. Reading it end to end, there is no mention of a free tier, a trial or starter credits\[4].
**Replicate.** This is the one real free tier in the group, and it is deliberately vague: "You can run select models on Replicate for free, but after a bit you'll be asked to set up billing"\[7]. No model list, no quota, no reset window. Past that point the pricing page applies: "You only pay for what you use on Replicate. Some models are billed by hardware and time, others by input and output"\[5] — and hardware-billed models charge for occupied GPU time, so a slow model costs more than a fast one at identical output.
Notice the shape of what is on offer. A documented free walkthrough call, or an unspecified allowance on an unspecified subset. Nobody in this market is giving away image or video generation at scale, because each generation has a hard compute floor underneath it. Text is cheap enough to subsidise for goodwill. A 4K image is not.
## What a fraction of a cent actually buys
Once you accept that generative media is not free anywhere, the question changes from "who is free" to "how small can the first bill be". That is a much better question, because the answer is now genuinely small.
Here are current per-generation rates on reAPI, where 1 credit is $0.001 and you are billed only for completed generations\[6]:
| Model and tier | Price |
| ---------------------------------------------------- | ----------------- |
| [GPT Image 2](/models/gpt-image-2), Stable Low | $0.005 per image |
| [Nano Banana 2 Lite](/models/nano-banana-2-lite), 1K | $0.02 per image |
| [GPT Image 2](/models/gpt-image-2), Self 1K | $0.03 per image |
| [GPT Image 2](/models/gpt-image-2), Self 4K | $0.08 per image |
| [Seedance 2.0 Mini](/models/seedance-2-0-mini), 480P | $0.046 per second |
| [Seedance 2.0 Fast](/models/seedance-2-0), 480P | $0.078 per second |
| [Seedance 2.0](/models/seedance-2-0), 480P | $0.095 per second |
| [Seedance 2.0 Mini](/models/seedance-2-0-mini), 720P | $0.098 per second |
| [Seedance 2.0 Fast](/models/seedance-2-0), 720P | $0.165 per second |
| [Seedance 2.0](/models/seedance-2-0), 720P | $0.205 per second |
Text models bill per token rather than per generation, and this is where Google's free tier is a genuine competitor. The reason to pay for text is that data row, not the capability\[1]:
| Model | Input | Output |
| ---------------------------------- | ------------------- | ----------------- |
| [Kimi K3](/models/kimi-k3) | $2.50 per 1M tokens | $12 per 1M tokens |
| [GPT-5.6 Sol](/models/gpt-5-6-sol) | $4 per 1M tokens | $24 per 1M tokens |
Two of the image rates are worth sitting with. Nano Banana 2 Lite at $0.02 per 1K image is below Google's own paid rate of $0.0336 for the same model\[1]\[6], and it does not require a linked billing account. And a low-tier GPT Image 2 render at half a cent means testing a prompt twenty times costs a dime.
Signing up on reAPI includes starter credits with no card required, which is enough to work through the [quickstart](/docs/api/quickstart) and see real output before you decide anything. Per-model rates for everything not listed above are on the [pricing page](/pricing), alongside each model's official list price for comparison.
One more constraint that has nothing to do with price: Google embeds a SynthID watermark in image output from its models. If your use case involves resale or downstream editing, read up on [what SynthID does to Nano Banana output](/blog/nano-banana-watermark-synthid) before you commit to a pipeline.
## FAQ
### Can I get an AI API for free?
For text, yes. Google's Gemini API offers around twenty text, audio and embedding models at no charge, including Gemini 2.5 Pro and the 3.x Flash family\[1]. For image, video or music generation, no major provider offers a free tier as of July 2026.
### Does Nano Banana have a free API?
No. All four Nano Banana variants list Free Tier as "Not available" for input, output and search grounding. Paid rates run from $0.0336 per 1K image on 2 Lite to $0.24 per 4K image on Pro\[1].
### Is there a free Veo 3.1 API?
No. Google describes Veo 3.1 as "available to developers on the paid tier of the Gemini API" and lists Free Tier as "Not available" for the Standard, Fast and Lite variants\[1].
### Does Google train on my free-tier API requests?
Google's wording is "Used to improve our products", not "train", and the page does not define the term. That row reads Yes on Free Tier and No on Paid Tier for every model\[1]. It applies to the models that are free of charge, since the media models have no free tier at all.
### What are the Gemini free-tier rate limits?
Per-model request limits are no longer published. The docs direct you to view your own limits in Google AI Studio and add that "Specified rate limits are not guaranteed and actual capacity may vary"\[2]. The one exception still printed is Search grounding at 500 requests per day on the free tier\[1]. Anyone quoting you a specific free-tier RPM figure for a model is reading a cached version of that page.
### Do OpenAI API keys come with free credits?
The quickstart explicitly describes its walkthrough call as "a free test API request"\[8], but OpenAI does not document a standing free allowance. The pricing page lists none; its only free quantities are 1 GB of file-search storage and 1 GB per month of ChatKit upload storage\[3].
### Is Replicate free to start?
Yes, within limits Replicate does not quantify: "You can run select models on Replicate for free, but after a bit you'll be asked to set up billing"\[7]. Which models and how much are both unstated, so it is a way to try something rather than something to plan against. After that, "You only pay for what you use"\[5].
### What is the cheapest way to test an image model?
Per-image rates are now low enough that testing is a rounding error. A Stable Low render on GPT Image 2 is $0.005, so twenty test images cost ten cents\[6]. That is a more predictable path than working around free-tier quotas you cannot see in advance.
## Picking a tier without getting stuck
If your workload is text, take Google's free tier. It is genuinely free and it covers capable models. The cost that is not measured in dollars is the "Used to improve our products" row, which matters for some projects and not others.
If your workload is images, video or music, skip the search for a free AI API and optimise for the smallest possible first bill instead. That means per-generation billing rather than subscriptions, no billing-account prerequisite, and rates you can read before you sign up. A free AI API that excludes the models you came for is not a starting point. Half a cent per image is.
## References
1. Google. *Gemini Developer API pricing — Free Tier and Paid Tier rates by model.* Retrieved July 2026 from [ai.google.dev/gemini-api/docs/pricing](https://ai.google.dev/gemini-api/docs/pricing)
2. Google. *Gemini API rate limits — usage tiers and tier qualification.* Retrieved July 2026 from [ai.google.dev/gemini-api/docs/rate-limits](https://ai.google.dev/gemini-api/docs/rate-limits)
3. OpenAI. *API pricing.* Retrieved July 2026 from [platform.openai.com/docs/pricing](https://platform.openai.com/docs/pricing)
4. fal.ai. *GenAI API pricing — pay-per-use.* Retrieved July 2026 from fal.ai/pricing
5. Replicate. *Pricing.* Retrieved July 2026 from replicate.com/pricing
6. reAPI. *Model pricing — GPT Image 2, Nano Banana 2 Lite, Seedance 2.0 Mini.* Retrieved July 2026 from [reapi.ai/models](/models)
7. Replicate. *Billing — free limits.* Retrieved July 2026 from replicate.com/docs/topics/billing
8. OpenAI. *API quickstart.* Retrieved July 2026 from [developers.openai.com/api/docs/quickstart](https://developers.openai.com/api/docs/quickstart)
### Further reading
* reAPI. *Nano Banana API free tier.* [reapi.ai/blog/nano-banana-api-free-tier](/blog/nano-banana-api-free-tier)
* reAPI. *What SynthID does to Nano Banana output.* [reapi.ai/blog/nano-banana-watermark-synthid](/blog/nano-banana-watermark-synthid)
---
# Gemini Omni API: Preview Specs, Pricing, and Limits (https://reapi.ai/blog/gemini-omni-api-preview-specs-pricing-2026)
**Google's direct Gemini Omni API is now available as a paid public preview under the model ID `gemini-omni-flash-preview`. It runs through the Gemini Interactions API, generates 3–10 second video at 720p and 24 FPS, and costs about $0.10 per second of video output.** It can create video from text or images, generate audio with the video, and edit a result across multiple turns by passing a `previous_interaction_id`.\[1]\[2]
Those specifications describe Google's direct Gemini Developer API. They do **not** describe every API carrying the name Gemini Omni. In particular, reAPI's `gemini-omni` is a supplier-routed video contract with a different endpoint, model name, request body, resolution tiers, and billing method. Treat the two as separate integrations.
## Gemini Omni API specs at a glance
| Item | Google direct public preview |
| -------------------- | -------------------------------------------------------------------- |
| Model ID | `gemini-omni-flash-preview` |
| API | Gemini Interactions API |
| REST endpoint | `POST https://generativelanguage.googleapis.com/v1beta/interactions` |
| Access | Paid Gemini API tier; no free tier |
| Input | Text, images, and video for editing |
| Output | Video with generated audio |
| Output length | 3–10 seconds |
| Output specification | 720p, 24 FPS |
| Aspect ratio | `16:9` or `9:16` |
| Context window | 1,048,576 tokens |
| Video output price | $17.50 per 1M output tokens, approximately $0.10 per second |
| Stateful editing | Yes, through `previous_interaction_id` |
| Lifecycle | Preview; interface and limits may change |
Google added the model to public preview on June 30, 2026. Its release notes and model card are the cleanest sources for the output limits: both identify 720p video between 3 and 10 seconds, while the model card also specifies 24 FPS and the 1,048,576-token context window.\[2]\[3]
## The correct model ID and endpoint
The direct model ID is:
```text
gemini-omni-flash-preview
```
Do not shorten it to `gemini-omni` in a direct Google request. That shorter identifier belongs to reAPI's supplier-routed contract, not the Gemini Developer API.
Google exposes Omni through the Interactions API rather than the older long-running video-generation pattern used by Veo. A minimum REST request looks like this:
```bash
curl -X POST \
"https://generativelanguage.googleapis.com/v1beta/interactions?key=$GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-omni-flash-preview",
"input": "A continuous handheld shot of a tabby cat crossing a sunny kitchen. Natural room audio, no dialogue."
}'
```
The direct REST response returns an interaction with a `steps` array. The generated MP4 appears as video content in a `model_output` step. Google's Python and JavaScript SDKs add an `output_video` convenience field, so SDK code and raw REST parsing are not identical.\[1]
For portrait output, add a response format:
```json
{
"response_format": {
"type": "video",
"aspect_ratio": "9:16"
}
}
```
Landscape `16:9` is the default. The public guide does not document a numeric `duration` request field like reAPI's route does. Google states that output falls within 3–10 seconds and shows natural-language timing and timecodes in prompts, such as `[0-3s]`, `[3-6s]`, and `[6-10s]`.\[1]
## How direct Google pricing works
Gemini Omni Flash Preview is available only on the paid tier. Google's current standard prices are:\[4]
| Meter | Direct Google price |
| --------------------------------------- | -------------------: |
| Input tokens, any listed modality | $1.50 per 1M tokens |
| Text output, including thinking tokens | $9.00 per 1M tokens |
| Video output, including thinking tokens | $17.50 per 1M tokens |
Google converts 720p video into output tokens at 5,792 tokens per second. At $17.50 per million tokens, that is about $0.101 per output second. A rough video-output estimate is therefore:
```text
3-second output ≈ $0.30
6-second output ≈ $0.61
10-second output ≈ $1.01
```
These are estimates for the video-output component, not guaranteed invoice totals. Input media, input text, text output, and thinking-token consumption can add to the charge. Retries and additional editing turns are also new billable interactions.
The important distinction is that Google does not publish a flat "$X per generation" price for this preview. It bills token consumption. If a third-party route quotes a fixed price for an 8-second or 10-second job, that is the third party's product contract.
## Multi-turn video editing with the Interactions API
Stateful editing is the strongest reason to use Google's direct interface. The first request creates a video and returns an interaction ID. A second request points to that state with `previous_interaction_id`:
```python
import base64
from google import genai
client = genai.Client()
first = client.interactions.create(
model="gemini-omni-flash-preview",
input="A woman playing violin outdoors in soft morning light.",
)
edited = client.interactions.create(
model="gemini-omni-flash-preview",
previous_interaction_id=first.id,
input="Remove the violin. Keep everything else the same.",
)
with open("edited.mp4", "wb") as file:
file.write(base64.b64decode(edited.output_video.data))
```
Each turn produces a new video. The model carries forward the video state and conversation, so the editing prompt can name only the requested change. Google's guidance favors short edit instructions and suggests adding "Keep everything else the same" when preserving the rest of the scene matters.\[1]
State has an operational cost. If you set `store: false`, the result cannot be edited later through `previous_interaction_id`. Persist the interaction ID alongside the asset record when later revisions are part of the product workflow. If you do not need editing, Google recommends `background=false`, `store=false`, and `stream=false` for a faster synchronous request.
## Input support and current preview limits
The model card lists text, image, and video input. Video input is capped at 10 seconds when used for editing, while generated output remains 3–10 seconds at 720p and 24 FPS.\[3]
The guide supports several request shapes:
* text-to-video;
* image-to-video from a starting image;
* multiple subject or style reference images;
* stateful editing of a previously generated result;
* editing an uploaded video through the Files API.
That broad list needs a preview-stage warning. Google's current limitations say uploaded audio references are unsupported, even though the model generates an audio track with video. Referencing multiple videos is unsupported. Video extension, interpolation between first and last frames, and voice editing are also unavailable. Generic video references up to three seconds can pass API schema validation but are not processed correctly by the model at present.\[1]
There are regional restrictions too. Editing uploaded videos is not currently available in the European Economic Area, Switzerland, or the United Kingdom, although users in those regions can edit video generated by the model. Image editing involving minors and certain recognizable people carries additional restrictions.
Other engineering limits worth planning around:
* no provisioned throughput;
* no system instructions, `temperature`, `top_p`, stop sequences, or dedicated negative-prompt parameter;
* English is fully supported, while other languages have not been formally evaluated;
* content safety filters apply to both prompts and generated media;
* every generated video carries an invisible SynthID watermark;
* large outputs should use URI delivery instead of inline base64.
The dedicated negative-prompt field is missing, but negative instructions can still be written in the ordinary prompt: "No dialogue," "No scene cuts," or "Do not add text."
## Google direct API vs reAPI's `gemini-omni`
The names are similar enough to cause implementation mistakes. The contracts are not interchangeable:
| Contract detail | Google direct preview | reAPI supplier route |
| ----------------- | ------------------------------------------------ | ---------------------------------------------------- |
| Model | `gemini-omni-flash-preview` | `gemini-omni` |
| Endpoint | `/v1beta/interactions` | `/api/v1/videos/generations` |
| Execution | Interaction response; inline or URI media | Async task submission and polling |
| Output resolution | Google documents 720p | Supplier delivery tiers: 720p, 1080p, 4K |
| Duration | 3–10 second output; prompt timing | Explicit `duration`: 4, 6, 8, or 10 |
| Images | Interactions multimodal content | Public `image_urls`, with 0, 1, or 3 entries |
| Video input | Direct editing workflow and preview restrictions | One public `video_urls` entry, supplier-route rules |
| Multi-turn state | `previous_interaction_id` | No equivalent field in the video-generation contract |
| Billing | Token-based; about $0.10 per output second | Per-job supplier-route pricing |
| Result shape | Interaction `steps` / SDK `output_video` | Task `output.video_urls` |
reAPI's [Gemini Omni API reference](/docs/gemini-omni) documents the supplier route. It accepts a prompt of up to 2,000 characters, explicit duration and resolution fields, `16:9` or `9:16`, public URL inputs, and task polling. Its 1080p and 4K options are supplier delivery tiers; they are not evidence that Google's direct public preview exposes those resolutions.
Here is the equivalent minimum shape on reAPI:
```json
{
"model": "gemini-omni",
"prompt": "A continuous handheld shot of a tabby cat crossing a sunny kitchen.",
"duration": 6,
"resolution": "1080p",
"aspect_ratio": "16:9"
}
```
The route returns a task ID, which the client polls until `output.video_urls` is available. It does not return a Google interaction ID and should not be treated as a wrapper around the direct Interactions API.
**Disclosure:** reAPI publishes this article and operates the reAPI supplier-routed endpoint described above. Direct-Google specifications and prices in this guide come from Google's documentation; reAPI contract details come from our published API reference. The distinction is stated because reAPI has a commercial interest in the routed product.
## Which API should you use?
Use Google's direct preview when the product specifically needs conversational editing, a stored interaction history, or a direct Google billing relationship. It is also the clearest route when a team wants to build against the newest Google interface and accepts preview lifecycle risk.
Use a supplier-routed contract when a fixed request schema, explicit resolution and duration tiers, public URL inputs, and a common async task pattern are more useful than Google interaction state. Validate the route's own pricing and media constraints rather than copying Google's settings into it.
This choice is separate from choosing a video model. If the actual question is whether to keep Veo in a production workflow, read the [Gemini Omni vs Veo 3.1 comparison](/blog/gemini-omni-vs-veo-3-1-2026). If the decision is about multimodal reference control and generation style, the [Gemini Omni vs Seedance 2.0 comparison](/blog/gemini-omni-vs-seedance-2-0-2026) covers that search intent. This page is only about the Gemini Omni API contract.
## Frequently asked questions
### Is Gemini Omni API available now?
Yes. Google released `gemini-omni-flash-preview` on the paid Gemini Developer API tier on June 30, 2026. It is a public preview, not a generally available production model.\[2]
### What resolution does the direct Gemini Omni API support?
Google's model card currently documents 720p output at 24 FPS. Do not infer direct 1080p or 4K support from a third-party Gemini Omni route.\[3]
### How much does Gemini Omni video cost?
Google charges $17.50 per 1M video output tokens and calculates 720p video at 5,792 tokens per second. That works out to approximately $0.10 per output second, before other input or output consumption.\[4]
### Can Gemini Omni edit the same video more than once?
Yes. Pass the prior interaction's ID as `previous_interaction_id` in the next Interactions API request. Keep storage enabled and persist the ID if the user may return for another edit.\[1]
### Does Gemini Omni accept an audio reference?
Not in the current direct API preview. It generates audio with video, but Google's limitation list says uploading audio references is unsupported.\[1]
### Is `gemini-omni` the Google model ID?
No. The direct Google model ID is `gemini-omni-flash-preview`. `gemini-omni` is the identifier used by reAPI's separate supplier-routed contract.
## Sources
1. Google AI for Developers. *Generate and edit videos with Gemini Omni Flash.* Updated July 30, 2026. [ai.google.dev/gemini-api/docs/omni](https://ai.google.dev/gemini-api/docs/omni)
2. Google AI for Developers. *Gemini API release notes — June 30, 2026.* [ai.google.dev/gemini-api/docs/changelog](https://ai.google.dev/gemini-api/docs/changelog)
3. Google AI for Developers. *Gemini Omni Flash model card.* [ai.google.dev/gemini-api/docs/models/gemini-omni-flash](https://ai.google.dev/gemini-api/docs/models/gemini-omni-flash)
4. Google AI for Developers. *Gemini Developer API pricing — Gemini Omni Flash Preview.* Retrieved July 30, 2026. [ai.google.dev/gemini-api/docs/pricing](https://ai.google.dev/gemini-api/docs/pricing)
---
# Gemini Omni vs Seedance 2.0: The 2026 Video Model Split (https://reapi.ai/blog/gemini-omni-vs-seedance-2-0-2026)
Google shipped Gemini Omni Flash on May 19, 2026 at I/O. ByteDance has held the Artificial Analysis Video Arena top spot with Seedance 2.0 since February. If you're picking Gemini Omni vs Seedance 2.0 right now, you're choosing between Google's first reasoning-and-editing-first video model and the model that benchmarks say is the best raw generator on the market.
The split is sharper than most "X vs Y" comparisons in this category. Seedance 2.0 throws a 1080p, multi-shot, audio-coupled clip back at you on one forward pass. Gemini Omni Flash gives you a 10-second clip you keep editing through conversation. Below is a capability-by-capability breakdown sourced to each vendor's own pages, with prices from the live reAPI listings.
## TL;DR
* **Release timing.** Seedance 2.0 launched February 12, 2026\[5]. Gemini Omni Flash launched May 19, 2026 at Google I/O\[1].
* **Benchmark gap.** Seedance 2.0 holds Elo 1,269 (text-to-video) and 1,351 (image-to-video) on the Artificial Analysis Video Arena, #1 in both categories\[6]. Gemini Omni was not on the leaderboard at launch.
* **Resolution ceiling.** Gemini Omni Flash supports 720p, 1080p, and 4K\[7]. Seedance 2.0 caps at 1080p\[8].
* **Duration.** Gemini Omni Flash outputs 4, 6, 8, or 10 seconds\[7]. Seedance 2.0 outputs 4 to 15 seconds with multi-shot cuts inside the same clip\[5].
* **References.** Gemini Omni Flash accepts 0, 1, or 3 image inputs\[9]. Seedance 2.0 accepts up to 9 images + 3 video clips + 3 audio clips per request\[10].
* **Editing model.** Gemini Omni is built around multi-turn conversational edits\[1]. Seedance 2.0 is single-pass with rich reference inputs.
* **The split.** Pick Gemini Omni when iteration on one clip matters more than peak raw quality. Pick Seedance 2.0 when you ship one polished clip and move on.
## Where each model comes from
Seedance 2.0 came out of ByteDance Seed on February 12, 2026\[5]. The launch went viral for photorealistic clips of named celebrities, and Disney sent ByteDance a cease-and-desist letter a day later\[11]. The model ships with C2PA watermarking by default. ByteDance positions it as a "unified multimodal audio-video joint generation architecture" that takes text, image, video, and audio as input, and generates lip-synced video with native audio across 8+ languages.
Gemini Omni Flash is the first model in a new Google DeepMind family announced at Google I/O on May 19, 2026. Sundar Pichai framed it on stage as part of Google's world-models push: "AI is moving from predicting text to simulating reality. Gemini Omni is the next step in that direction."\[4] Google's own product page says Gemini Omni will replace Veo in the Gemini app\[3]. Outputs are SynthID-watermarked, with verification available through the Gemini app, Chrome, and Google Search\[1].
Both models run behind reAPI's OpenAI-compatible `POST /api/v1/videos/generations`. You switch between them by changing the `model` field in the request body, no other infrastructure changes required.
## What each Gemini Omni vs Seedance 2.0 spec actually means
| Capability | Gemini Omni Flash | Seedance 2.0 |
| ------------------------------ | ------------------------------------------------------------------ | ------------------------------------------------------------------------------- |
| Text-to-video | yes | yes |
| Image-to-video (single) | yes (1 ref) | yes |
| Image-to-video (multi-ref) | up to 3 (fusion mode) | up to 9 images |
| First/last-frame interpolation | no | yes (`image_with_roles`) |
| Reference video | no | up to 3 clips, ≤15s combined\[10] |
| Reference audio | voice-reference only at launch\[1] | up to 3 clips, ≤15s combined\[10] |
| Native audio synthesis | yes | yes (joint generation, phoneme lip-sync)\[5] |
| Multi-shot in one output | no | yes, multiple cuts in one generation\[5] |
| Multi-turn conversational edit | yes\[1] | no |
| 4K output | yes\[7] | no (1080p ceiling)\[8] |
| Duration options | 4 / 6 / 8 / 10s\[7] | any 4–15s\[10] |
| Aspect ratios | 16:9, 9:16\[9] | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, adaptive\[10] |
| Watermarking | SynthID (Google)\[1] | C2PA (default on)\[5] |
| Avatar feature | yes (consumer-only at launch)\[1] | no |
Two cells do most of the work in the decision. Seedance 2.0 takes a 9+3+3 reference bundle in one request. Gemini Omni Flash takes 0/1/3 image references and one voice reference. If your pipeline relies on feeding a model "here is the product, here is the brand style clip, here is the music bed, now compose," Seedance 2.0's pipeline is built for it\[10]. Gemini Omni isn't built for that workflow yet.
The other one is the editing column. Gemini Omni's multi-turn conversational editing has no Seedance equivalent. "Make the violin invisible. Now change the camera angle to be over the violinist's shoulder. Now transport the violinist to the image environment" is a real prompt sequence from the Google blog\[1]. Each instruction builds on the last while keeping the character and scene coherent. Seedance 2.0 doesn't work that way. You write one prompt with one set of references, you get one clip.
## The benchmark gap nobody can close yet
Seedance 2.0 is #1 on the Artificial Analysis Video Arena leaderboard, with Elo 1,269 for text-to-video and 1,351 for image-to-video\[6]. Both scores are above Veo 3.1, Kling 3.0, and Sora 2 in the same arena.
Gemini Omni Flash launched four days before this post, and Google did not put it on the arena at launch. Early hands-on coverage is split. TechCrunch called the consumer demos genuinely impressive but flagged that several features were broken at I/O\[4]. Independent reviewers writing the day after launch said raw generation quality "trails" Seedance 2.0 on aggregate, while Omni's text rendering, physics intuition, and conversational edits opened new ground. Until the Arena gets enough votes on Omni Flash, the only honest read is: Seedance 2.0 is the verified raw-quality leader. Gemini Omni Flash is the most novel editing surface anyone has shipped.
*Same prompt run on Gemini Omni Flash and Seedance 2.0, side by side. Judge the raw-quality gap with your own eyes before you trust the Elo numbers.*
If a benchmark Elo decides your purchase, Seedance 2.0 wins today. If you're shipping a product where iterative refinement matters more than the single best frame, that Elo gap stops being the right axis.
## Editing through conversation, or compose-everything-up-front
Gemini Omni's headline is conversation. The Google product page is blunt: "Gemini Omni makes creating videos as easy as having a conversation."\[3] The blog goes further with a worked example: a violinist clip, refined across four prompts that each build on the last. Characters stay consistent, physics holds, the scene remembers what came before\[1]. This is the "Nano Banana for video" pitch, and it's the part of the model that no other 2026 video model competes with directly.
*Gemini Omni Flash leaning on Gemini's world knowledge to ground a single-prompt generation. The reasoning step is what separates Omni's outputs from a pure diffusion model run.*
Seedance 2.0 takes the opposite philosophy. Compose everything up front. Hand it text + 9 images + 3 video references + 3 audio references in one shot, and the model fuses them into one cohesive output\[10]. ByteDance's design assumes the user knows the spec, has the assets, and wants one finished clip with no back-and-forth. The reference pipeline is the editing surface. If you want a different result, you change the references and resubmit, you don't iterate.
For brand spots and product ads, Seedance 2.0's compose-up-front model maps cleanly to how creative briefs already work. For social experimentation, fan edits, or anything where the author doesn't know what they want until they see the first cut, Gemini Omni's conversation loop wins.
## Two different kinds of "multi"
Both models advertise "multi" as a differentiator. They mean different things.
Seedance 2.0's "multi" is **multi-shot inside one generation**: a single 15-second output contains multiple cuts and transitions, like an edited storyboard\[5]. You write a prompt that describes the scene progression, and the model emits one clip with the cuts already in it. This collapses what would otherwise be a 3-to-4 call workflow on Veo or Gemini Omni into a single call.
Gemini Omni's "multi" is **multi-turn refinement on one shot**: each conversational instruction reshapes the existing clip without losing thread\[1]. You don't get more shots, you get a more refined version of the same shot. The cost compounds across turns, but the consistency is the point. Refining the same scene through 5 turns is a different product entirely from generating 5 different scenes.
A pipeline that needs both is real. Generate the storyboard with Seedance 2.0's multi-shot, refine individual beats with Gemini Omni's multi-turn. Both models behind one endpoint makes that workflow a 30-line change instead of a vendor migration.
## Price math, 720p / 1080p / 4K with audio
Per-second pricing on Seedance 2.0 vs per-generation pricing on Gemini Omni Flash means the comparison flips by clip length. The table below is what each cheapest viable path costs on reAPI for the same output spec.
| Output spec | Gemini Omni Flash (per-gen) | Seedance 2.0 cheapest (per-sec, ref mode) |
| -------------------- | ------------------------------------------ | ------------------------------------------------------------------------ |
| 5s 720p with audio | n/a (Omni durations are 4 / 6 / 8 / 10) | $0.376 (Seedance 2.0 5s × $0.0752/s)\[8] |
| 6s 720p with audio | $0.204\[7] | $0.451 (5.97s × $0.0752/s)\[8] |
| 8s 1080p with audio | $0.216\[7] | $1.709 (8s × $0.2136/s Standard ref)\[8] |
| 10s 1080p with audio | $0.240\[7] | $2.136 (10s × $0.2136/s)\[8] |
| 10s 4K with audio | $0.480\[7] | not supported (1080p ceiling) |
Two things to call out. First, Gemini Omni Flash's per-generation rate is independent of duration in the same resolution bucket only by a small margin: 4s costs $0.18, 10s costs $0.24 at 720p/1080p\[7]. Seedance 2.0's per-second rate compounds linearly with duration, so the longer the clip, the larger the price gap in Omni's favor at the same resolution. Second, Seedance 2.0's reference-mode pricing (any of `image_urls`, `video_urls`, `audio_urls` set) is roughly 40% cheaper per second than text-only mode\[8], so the table above assumes ref mode.
For a 6-second 1080p clip with audio in May 2026, Gemini Omni Flash is the cheaper choice on reAPI. For multi-shot storyboards longer than 10 seconds or any output that needs reference video and audio composed in, Seedance 2.0 is the only model of these two that does it.
## When to actually pick which
**Pick Gemini Omni Flash when:**
* You're iterating on one clip and want multi-turn editing without re-running from scratch
* 4K output matters
* The 10-second cap is enough for your use case (Brichtova told TechCrunch this cap is a product decision, not a model limit\[4])
* Per-generation flat pricing simplifies your cost forecasting
* Avatar generation is something you want exposed (consumer-tier today, API surface coming\[1])
**Pick Seedance 2.0 when:**
* You ship one polished clip per call, no iteration
* Multi-shot storyboards in one 15-second output replace a 3-to-4 call workflow
* Reference video, reference audio, or both feed into the generation
* Lip-synced dialogue in non-English languages is a hard requirement
* 21:9 ultrawide or other non-standard aspect ratios are required
* The Arena Elo headline matters for your buyers
Neither is universally better. Gemini Omni vs Seedance 2.0 only resolves once you know your output spec and your team's iteration style.
## FAQ
### Is Gemini Omni an upgrade to Veo 3.1?
Google's product page says Gemini Omni will replace Veo in the Gemini app\[3]. Veo 3.1 remains available through the Vertex AI and Gemini API surfaces, and through aggregators. For consumer surfaces (Gemini app, Flow, YouTube Shorts), Omni Flash is the new default.
### Is Seedance 2.0 free to use?
Not at the API level. Seedance 2.0 is paid-tier on every provider that exposes it. ByteDance's consumer products (Dreamina, CapCut) include some Seedance 2.0 quota in their free tiers, but those quotas are not API-accessible.
### Does Gemini Omni Flash support multi-shot like Seedance 2.0?
No. Gemini Omni Flash outputs single continuous shots up to 10 seconds\[7]. Multi-shot sequences in one clip are a Seedance 2.0 capability\[5]. To get a Gemini Omni storyboard, you generate each shot separately, or use multi-turn conversational editing to refine one shot through phases.
### Can I use both behind one endpoint?
Yes. Both models run on reAPI's `POST /api/v1/videos/generations` with the same request envelope. Swapping between them means changing `model` from `gemini-omni` to `doubao-seedance-2.0` and adjusting fields that don't translate (Omni's `image_urls` accepts 0/1/3 entries, Seedance accepts up to 9; Omni has no `video_urls` or `audio_urls`).
### Which model has better physics?
Both vendors claim physics realism as a feature. Google's blog says Omni has "an improved intuitive understanding of forces like gravity, kinetic energy and fluid dynamics"\[1]. ByteDance's Seedance 2.0 paper covers complex motion and physical interaction. The Artificial Analysis Arena's image-to-video Elo (where physics often decides votes) currently puts Seedance 2.0 at 1,351, ahead of every model tested\[6]. Until Omni Flash gets enough Arena votes, "Seedance leads on physics" is the verified position.
### What about Hollywood IP risk?
Real risk on Seedance 2.0. ByteDance received a Disney cease-and-desist letter on February 13, 2026 over training-data concerns\[11]. Don't prompt named studio characters, films, or styles you don't have rights to. Gemini Omni Flash's outputs carry SynthID watermarks; Seedance 2.0's carry C2PA watermarks. Both make derivative work downstream-detectable.
### When does the Gemini Omni API open up?
Google said developer and enterprise API access will arrive "in the coming weeks" after the May 19 launch\[4]. reAPI exposed Gemini Omni on the standard videos endpoint at launch, so you can call it now without waiting on Google's direct API rollout. See the [Gemini Omni docs](/docs/gemini-omni) for request shape and the [model page](/models/gemini-omni) for live pricing.
## Routing both in one pipeline
For most teams shipping AI video in May 2026, Gemini Omni vs Seedance 2.0 isn't a decision you make once. It's a routing decision you make per request. Seedance 2.0 handles the polished one-shot output, the multi-reference compositions, and any time you need a multi-shot storyboard inside one clip. Gemini Omni Flash takes the 4K work, the with-audio clips at 1080p, and anything you need to iterate on through conversation. Both behind one OpenAI-compatible endpoint is a 30-line config change, not a vendor migration.
If forced to pick one model for everything, I'd pick Seedance 2.0 for commercial output that ships to paying customers today, on the strength of the verified Arena lead. I'd pick Gemini Omni Flash for any pipeline where the second draft matters more than the first, and let the next round of benchmarks decide the quality question. The Gemini Omni vs Seedance 2.0 split is the cleanest case I've seen in 2026 video where "pick one and stick with it" is actively the wrong answer.
## References
1. Google. *Introducing Gemini Omni.* Koray Kavukcuoglu, May 19, 2026. Retrieved May 2026 from [blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/)
2. Google DeepMind. *Gemini Omni — Model page.* Retrieved May 2026 from [deepmind.google/models/gemini-omni](https://deepmind.google/models/gemini-omni/)
3. Google. *Gemini Omni — Video Generation overview.* Retrieved May 2026 from [gemini.google/overview/video-generation](https://gemini.google/overview/video-generation/)
4. Rebecca Bellan. *Google's Gemini Omni turns images, audio, and text into video — and that's just the start.* TechCrunch, May 19, 2026. [techcrunch.com/2026/05/19/googles-gemini-omni-turns-images-audio-and-text-into-video-and-thats-just-the-start](https://techcrunch.com/2026/05/19/googles-gemini-omni-turns-images-audio-and-text-into-video-and-thats-just-the-start/)
5. ByteDance Seed. *Official Launch of Seedance 2.0.* February 12, 2026. [seed.bytedance.com/en/blog/official-launch-of-seedance-2-0](https://seed.bytedance.com/en/blog/official-launch-of-seedance-2-0)
6. Artificial Analysis. *Video Arena Leaderboard.* Retrieved May 2026 from [artificialanalysis.ai/video/arena](https://artificialanalysis.ai/video/arena)
7. reAPI. *Gemini Omni — Model page (live pricing).* Retrieved May 2026 from [reapi.ai/models/gemini-omni](/models/gemini-omni)
8. reAPI. *Seedance 2.0 — Model page (live pricing).* Retrieved May 2026 from [reapi.ai/models/seedance-2-0](/models/seedance-2-0)
9. reAPI. *Gemini Omni — API reference.* Retrieved May 2026 from [reapi.ai/docs/gemini-omni](/docs/gemini-omni)
10. reAPI. *Seedance 2.0 — API reference.* Retrieved May 2026 from [reapi.ai/docs/seedance-2-0](/docs/seedance-2-0)
11. Wikipedia contributors. *Seedance 2.0.* Retrieved May 2026 from [en.wikipedia.org/wiki/Seedance\_2.0](https://en.wikipedia.org/wiki/Seedance_2.0)
### Further reading
* Google. *Gemini Omni prompt guide.* [deepmind.google/models/gemini-omni/prompt-guide](https://deepmind.google/models/gemini-omni/prompt-guide/)
* reAPI. *Veo 3.1 vs Seedance 2.0: Picking a Video Model in 2026.* [reapi.ai/blog/veo-3-1-vs-seedance-2-0-2026](/blog/veo-3-1-vs-seedance-2-0-2026)
* reAPI. *Cheapest Veo 3.1 API in 2026.* [reapi.ai/blog/cheapest-veo-3-1-api-2026](/blog/cheapest-veo-3-1-api-2026)
---
# Gemini Omni vs Veo 3.1: Should You Migrate in May 2026? (https://reapi.ai/blog/gemini-omni-vs-veo-3-1-2026)
Google's Gemini Omni page says, in five words, "Gemini Omni will replace Veo in the Gemini app."\[3] If you only read the marketing, you'd assume Veo 3.1 is sunset. It isn't. Veo 3.1 still ships on Vertex AI, on the Gemini API, and on every aggregator that lists it. Google's own April 28, 2026 update to the Veo 3.1 Gemini API docs is more recent than Omni's launch\[4].
So when you're choosing Gemini Omni vs Veo 3.1 in May 2026, you're not choosing between a deprecated model and its replacement. You're choosing between a 7-month-old per-second model with five channels and a 2-day-old per-generation model with one. The right answer depends on which Veo 3.1 channel you'd otherwise be calling, and whether the workflow needs audio, 4K, first-and-last-frame control, or multi-turn editing.
## TL;DR
* **Scope of the "replacement."** Google replaced Veo with Omni in the Gemini app, Flow, and YouTube Shorts only\[3]. Veo 3.1 remains the documented Gemini API video model\[4].
* **API access timing.** Veo 3.1 Gemini API: live since October 2025. Gemini Omni public API: "coming weeks" from May 19, 2026\[5]. Through reAPI, both are callable today on the same endpoint.
* **Cheapest no-audio path.** Veo 3.1 Lite at $0.05 per 8-second 720p/1080p clip\[6] beats Gemini Omni's $0.216 for 8s 1080p\[7] by roughly 4x.
* **Cheapest with-audio path.** Gemini Omni Flash at $0.216 for 8s 1080p\[7] beats Veo 3.1 Fast Official at $1.20 (8s × $0.15/s)\[6] by 5.5x.
* **4K.** Veo 3.1 Fast alt at $0.30 per 8s clip wins on cost. Gemini Omni at $0.432 per 8s wins on bundled audio.
* **Duration flexibility.** Gemini Omni supports 4/6/8/10s\[8]. Veo 3.1 alt locks duration at 8s; Veo 3.1 official supports 4/6/8s.
* **First/last-frame interpolation.** Veo 3.1 Official only. Gemini Omni doesn't expose first/last anchoring.
* **Multi-turn conversational edit.** Gemini Omni only\[1].
* **The split.** Stay on Veo 3.1 Lite for cheap silent clips. Stay on Veo 3.1 Official for first/last-frame chains. Move to Gemini Omni for any with-audio workload at 1080p or 4K, or any flow where iteration matters more than one-shot quality.
## What Google actually replaced
The "Gemini Omni will replace Veo" claim runs across exactly three consumer surfaces: the Gemini app, Google Flow, and YouTube Shorts\[3]. Inside those, Omni Flash is now the default video model. Veo 3.1 disappears from the consumer UI in those products.
Outside those surfaces, Veo 3.1 is the same product it was on May 18. The Gemini API video docs at ai.google.dev still list Veo 3.1 as the supported text-to-video and image-to-video model\[4]. Vertex AI still exposes Veo 3.1 with `person_generation`, `resize_mode`, and other production controls. Aggregators routing through Google's Veo 3.1 weights are unaffected. Even Google's own April 28, 2026 announcement of new Veo 3.1 capabilities in the Gemini API (added safety filters, expanded prompt sizes) hit the public blog less than a month before Omni shipped\[4], hardly the cadence of a model being retired.
Nicole Brichtova, Google DeepMind's director of product management, told TechCrunch that Omni is "more than a Veo update" and described it as "the next step towards the progression of combining the intelligence of Gemini with the rendering capabilities of our media models."\[5] The framing is intentional. Omni is a different product category (reasoning + editing), not a drop-in upgrade for Veo's per-second video synthesis.
For the developer choosing Gemini Omni vs Veo 3.1, the takeaway is: don't migrate because the consumer UI did. Migrate when the capability or price math points at Omni for your specific workload.
## The capability matrix
| Capability | Gemini Omni Flash | Veo 3.1 (5 channels) |
| -------------------------------- | ------------------------------------------------------------------ | ---------------------------------------------------------------------------------------- |
| Text-to-video | yes | yes |
| Image-to-video (single) | yes (1 ref) | yes (Fast / Quality alt, Official tiers) |
| Multi-image reference | up to 3 (fusion mode)\[8] | up to 3 (Fast alt only)\[9] |
| First-frame anchoring | no | yes (Official tiers)\[9] |
| First + last-frame interpolation | no | yes (Official tiers)\[9] |
| Multi-turn conversational edit | yes\[1] | no |
| 4K output | yes ($0.432 per 8s with audio)\[7] | yes (4K-audio tier on Official; flat per-gen on alt)\[6] |
| Duration options | 4 / 6 / 8 / 10s\[8] | alt: 8s fixed; Official: 4 / 6 / 8s\[9] |
| Aspect ratios | 16:9, 9:16\[8] | 16:9, 9:16\[9] |
| Prompt length cap | 2,000 chars\[8] | 4,000 chars\[9] |
| Audio control | bundled (always on) | alt: no audio; Official: `generate_audio` toggle |
| `person_generation` control | not exposed | Official: `allow_adult` / `disallow`\[9] |
| Avatar feature (consumer) | yes\[1] | no |
| Watermarking | SynthID\[1] | SynthID |
| Billing model | per generation | alt: per generation; Official: per second |
Three asymmetries do most of the work in the decision.
**Veo 3.1's first/last-frame interpolation.** Omni doesn't have an equivalent. If your pipeline anchors clips on a known start frame and a known end frame (storyboard chaining, hero-clip retakes with locked composition), Veo 3.1 Official is the only model of these two that gives you that control.
**Omni's multi-turn conversational edit.** Veo doesn't have an equivalent. "Make the violin invisible. Now change the camera angle." Each instruction reshapes the existing clip without losing the scene\[1]. Veo can't iterate on its own output, you have to regenerate with a new prompt every time.
**Veo 3.1's audio toggle on Official.** Veo Official lets you generate silent video at a discount. Omni bundles audio into every generation at one price. If you don't want audio, you can save 33% on Veo Fast Official by setting `generate_audio: false`; Omni doesn't expose that switch.
## The five Veo 3.1 channels Omni doesn't replace
Veo 3.1 is five channels on reAPI, not one model, and each one has a different surface\[6]. Each maps to a different reason to stay or migrate.
**`veo3.1-lite`** is the budget channel. Per generation, 8-second duration locked, no audio, prompt-only (no image input). $0.05 per 8s 720p/1080p clip\[6]. Gemini Omni at $0.216 for 8s 1080p is roughly 4x more expensive on the same output spec. If your workload is "draft fast, throw most away," Lite is unbeaten.
**`veo3.1-fast`** is the workhorse alt channel. Per generation, accepts up to 3 image references, 8s fixed duration. $0.10 per 8s 720p/1080p, $0.30 per 8s 4K. Omni's three-image fusion mode covers a similar surface but at 2x the price for 1080p output.
**`veo3.1-quality`** is the high-fidelity alt channel. $0.75 per 8s 720p/1080p, $2.40 per 8s 4K. Used when face rendering, text legibility, or character consistency are make-or-break. Omni's raw quality is unconfirmed at this price tier; reviewers writing the day after launch said Omni's aggregate quality "trails" the leaderboard's frontier\[5].
**`veo3.1-fast-official`** unlocks the Vertex-grade controls: first/last-frame anchoring, `generate_audio` toggle, `person_generation` enum, 4 / 6 / 8s durations. Per second. $0.15/s for 720p/1080p with audio. None of these controls have an Omni equivalent.
**`veo3.1-quality-official`** is the production-grade tier. $0.40/s for 720p/1080p with audio. Same control surface as Fast Official, higher visual fidelity. Used for hero shots and paid creative.
When you're deciding Gemini Omni vs Veo 3.1, the right question isn't "which model wins" but "which of the five Veo 3.1 channels does this workload currently use, and does Omni's surface cover those controls?"
## Migrating the request body
On reAPI, both models run on `POST /api/v1/videos/generations`. The migration is a request-body edit, not an endpoint change. Here's the field-by-field translation from the most common Veo 3.1 setup to Omni.
**Veo 3.1 Fast (alt channel) request:**
```json
{
"model": "veo3.1-fast",
"prompt": "A neon city street in the rain, slow camera pan",
"image_urls": ["https://your-cdn.com/ref.jpg"],
"generation_type": "frame",
"aspect_ratio": "16:9",
"resolution": "1080p"
}
```
**Same request, migrated to Gemini Omni:**
```json
{
"model": "gemini-omni",
"prompt": "A neon city street in the rain, slow camera pan",
"image_urls": ["https://your-cdn.com/ref.jpg"],
"duration": 8,
"aspect_ratio": "16:9",
"resolution": "1080p"
}
```
Three changes to know about. First, `generation_type` doesn't exist on Omni; the mode is implicit from the `image_urls` count (0 text-to-video, 1 image-to-video, 3 fusion)\[8]. Second, Omni accepts `duration` as a first-class field, where Veo alt locks it to 8s. Third, `image_urls` length 2 is rejected by Omni with a 400 error code `20003`\[8]; Veo Fast accepts 1, 2, or 3.
If your source request was `veo3.1-fast-official`, the migration loses more. Drop `first_frame_image`, `last_frame_image`, `generate_audio`, `person_generation`, and `resize_mode` — Omni exposes none of these. If any of those were load-bearing, don't migrate.
## Price math, the channels that matter
For an 8-second 1080p clip on reAPI in May 2026:
| Channel | With audio | Without audio |
| ------------------------- | ------------------------------------------------------------------- | -------------------------------------------------------- |
| `veo3.1-lite` | not supported | $0.05 (alt, per-gen)\[6] |
| `veo3.1-fast` | not supported | $0.10 (alt, per-gen)\[6] |
| `veo3.1-quality` | not supported | $0.75 (alt, per-gen)\[6] |
| `veo3.1-fast-official` | $1.20 (8s × $0.15/s) | $0.80 (8s × $0.10/s)\[6] |
| `veo3.1-quality-official` | $3.20 (8s × $0.40/s) | $1.60 (8s × $0.20/s)\[6] |
| `gemini-omni` | $0.216 (per-gen, audio bundled)\[7] | $0.216 (no toggle)\[7] |
Two observations. First, if you need audio and you're not on Veo Official, Gemini Omni is the cheapest path on reAPI. Omni's $0.216 with-audio rate is 5.5x cheaper than Veo Fast Official's $1.20 and 14.8x cheaper than Veo Quality Official's $3.20. Second, if you don't need audio, Veo Lite at $0.05 stays unbeaten. Omni doesn't have a silent-mode discount because the price doesn't depend on audio.
For 4K with audio at 8s: Gemini Omni at $0.432 vs Veo Quality Official at $4.80 (8s × $0.60/s). Omni is 11x cheaper for 4K with audio. This is where the per-generation model lands hardest in Omni's favor.
## When to stay on Veo 3.1
* The workflow uses first-and-last-frame anchoring to chain shots into a longer sequence
* The workflow needs `person_generation` set to `disallow` for safety compliance
* The output spec is silent video and Veo Lite's $0.05 floor matters at scale
* The pipeline is shipping today and Omni's two-day-old release introduces uncomfortable risk
* Prompt length regularly exceeds 2,000 chars (Veo allows 4,000)
* The team has a Vertex AI contract with negotiated terms
## When to migrate to Gemini Omni
* You're paying for `veo3.1-fast-official` with audio at $1.20+ per 8s clip and the iterative-edit workflow doesn't need first-frame anchoring
* You need 4K with audio (Omni's $0.432 per 8s beats Veo Quality Official's $4.80 by 11x)
* The workflow is editing-heavy: regenerating from scratch for each tweak burns budget and breaks character consistency
* Duration flexibility (4 / 6 / 8 / 10s) matters and you're tired of the 8-second alt lock
* You want to consolidate down to one model behind one endpoint and the team can absorb the loss of Veo's official controls
## FAQ
### Is Veo 3.1 being deprecated?
No. Google removed Veo 3.1 from the Gemini app, Flow, and YouTube Shorts consumer surfaces and replaced it with Gemini Omni Flash\[3]. The Gemini API video docs continue to list Veo 3.1 as supported, with a documentation refresh as recent as April 28, 2026\[4]. Vertex AI's Veo 3.1 surface is unchanged.
### When does the Gemini Omni public API open?
Google said developer and enterprise API access would land "in the coming weeks" after the May 19, 2026 announcement\[5]. As of this post, no public Gemini API model ID for Omni has been documented. On reAPI, Gemini Omni is callable today at `gemini-omni` on the standard videos endpoint, ahead of Google's direct API rollout.
### Can I use both models behind one endpoint?
Yes. Both run on `POST /api/v1/videos/generations` on reAPI. Switching is a `"model"` field change. Field mappings are not identical (see the migration section above) — Omni drops `first_frame_image`, `last_frame_image`, `generate_audio`, `person_generation`, and `resize_mode`.
### Does Gemini Omni support first-frame and last-frame anchoring?
No. Omni exposes only image fusion at cardinality 0, 1, or 3\[8]. There is no concept of a "first" or "last" frame. If your pipeline depends on locking the opening or closing composition, Veo 3.1 Official is the only model that gives you that.
### Which model is cheaper for cheap iteration?
Depends on whether you need audio. Veo 3.1 Lite at $0.05 per 8s silent clip is the cheapest path for any kind of throw-away draft\[6]. If you need audio in every draft, Gemini Omni at $0.216 per 8s 1080p with audio is the cheapest with-audio rate on reAPI\[7].
### Will the Gemini Omni API undercut reAPI when it ships?
Possibly. Independent forecasters project Omni Flash at $0.20–$0.60 per second of video output on Google direct, plus token-based input pricing. reAPI's $0.216 per 8s 1080p clip is roughly $0.027/s flat, which sits below the low end of the Google direct forecast. The actual gap won't be known until Google publishes official pricing.
### What about Gemini Omni vs Veo 3.1 quality?
Omni Flash launched two days before this post and isn't on the Artificial Analysis Video Arena leaderboard yet\[10]. Veo 3.1 has been on the leaderboard for months and sits behind Seedance 2.0 but ahead of Sora 2 in aggregate. Early hands-on coverage of Omni called the consumer demos impressive but flagged that aggregate quality "trails" the leaderboard frontier\[5]. Treat Omni as the most novel editing surface, not the verified quality leader, until the Arena gets enough votes.
## Routing both for the next quarter
You don't decide Gemini Omni vs Veo 3.1 once for the whole pipeline in May 2026. You route per request. Send draft, silent, and first-frame-anchored workloads to the right Veo 3.1 channel. Send the with-audio, 4K, and multi-turn-editing workloads to Gemini Omni Flash. Both behind one OpenAI-compatible endpoint makes that routing a 30-line config change.
If forced to standardize on one model right now, I'd hold on Veo 3.1 for any pipeline that ships to paying customers, on the strength of seven months of production track record. I'd let Gemini Omni take the new workloads where its multi-turn editing and bundled audio matter. The "Gemini Omni replaces Veo" framing is true for Google's consumer UI and misleading for everywhere else. Read the Gemini Omni vs Veo 3.1 question as "what does each channel still earn its keep on" and the migration plan writes itself.
## References
1. Google. *Introducing Gemini Omni.* Koray Kavukcuoglu, May 19, 2026. Retrieved May 2026 from [blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/)
2. Google DeepMind. *Gemini Omni — Model page.* Retrieved May 2026 from [deepmind.google/models/gemini-omni](https://deepmind.google/models/gemini-omni/)
3. Google. *Gemini Omni — Video Generation overview.* Retrieved May 2026 from [gemini.google/overview/video-generation](https://gemini.google/overview/video-generation/)
4. Google. *Enhanced Veo 3.1 capabilities are now available in the Gemini API.* April 28, 2026. [blog.google/innovation-and-ai/technology/developers-tools/veo-3-1-gemini-api](https://blog.google/innovation-and-ai/technology/developers-tools/veo-3-1-gemini-api/)
5. Rebecca Bellan. *Google's Gemini Omni turns images, audio, and text into video — and that's just the start.* TechCrunch, May 19, 2026. [techcrunch.com/2026/05/19/googles-gemini-omni-turns-images-audio-and-text-into-video-and-thats-just-the-start](https://techcrunch.com/2026/05/19/googles-gemini-omni-turns-images-audio-and-text-into-video-and-thats-just-the-start/)
6. reAPI. *Veo 3.1 — Model page (live pricing across 5 channels).* Retrieved May 2026 from [reapi.ai/models/veo3-1](/models/veo3-1)
7. reAPI. *Gemini Omni — Model page (live pricing).* Retrieved May 2026 from [reapi.ai/models/gemini-omni](/models/gemini-omni)
8. reAPI. *Gemini Omni — API reference.* Retrieved May 2026 from [reapi.ai/docs/gemini-omni](/docs/gemini-omni)
9. reAPI. *Veo 3.1 — API reference.* Retrieved May 2026 from [reapi.ai/docs/veo3-1](/docs/veo3-1)
10. Artificial Analysis. *Video Arena Leaderboard.* Retrieved May 2026 from [artificialanalysis.ai/video/arena](https://artificialanalysis.ai/video/arena)
### Further reading
* Google. *Generate videos with Veo 3.1 in the Gemini API.* [ai.google.dev/gemini-api/docs/video](https://ai.google.dev/gemini-api/docs/video)
* reAPI. *Gemini Omni vs Seedance 2.0: The 2026 Video Model Split.* [reapi.ai/blog/gemini-omni-vs-seedance-2-0-2026](/blog/gemini-omni-vs-seedance-2-0-2026)
* reAPI. *Veo 3.1 vs Seedance 2.0: Picking a Video Model in 2026.* [reapi.ai/blog/veo-3-1-vs-seedance-2-0-2026](/blog/veo-3-1-vs-seedance-2-0-2026)
* reAPI. *Cheapest Veo 3.1 API in 2026.* [reapi.ai/blog/cheapest-veo-3-1-api-2026](/blog/cheapest-veo-3-1-api-2026)
---
# GLM-5.2 API Guide: 1M Context, Pricing, and Coding (2026) (https://reapi.ai/blog/glm-5-2-api-guide)
GLM-5.2 is Z.AI's open-weight flagship for long-horizon coding and agent work.
The **GLM-5.2 API** combines a one-million-token context window, 128K maximum
output, thinking that can be disabled, function calling, JSON mode, and an
OpenAI-compatible request shape.\[1]
Its headline price is competitive: Z.AI lists $1.40 per million input tokens,
$0.26 for cached input, and $4.40 per million output tokens. reAPI currently
lists $0.90 input and $3 output through its compatible gateway.\[2]
## TL;DR
* **Model id:** `glm-5.2`, with the dot.
* **Context:** 1M tokens; maximum output is 131,072 tokens.
* **Official price:** $1.40 input, $0.26 cached input, $4.40 output per MTok.
* **reAPI price:** $0.90 input and $3 output per MTok.
* **Thinking:** enabled by default but switchable off.
* **Reasoning effort:** accepts seven labels but collapses to three effective
behaviors; default is `max`.
* **Tools:** up to 128 functions, `tool_choice: "auto"`, streamed tool-call
arguments.
* **Main trap:** the page slug is `glm-5-2`, but API requests must send
`glm-5.2`.
## GLM-5.2 specifications
| Property | GLM-5.2 |
| ----------------- | --------------------------------- |
| Vendor | Z.AI |
| Distribution | Open weight |
| API model id | `glm-5.2` |
| Context window | 1,000,000 tokens |
| Maximum output | 131,072 tokens |
| Input/output | Text / text |
| Thinking default | Enabled |
| Reasoning default | `max` |
| Temperature | 0.0–1.0, default 1.0 |
| Top-p | 0.01–1.0, default 0.95 |
| Tools | Up to 128 functions |
| Structured output | JSON object mode, not JSON Schema |
Z.AI positions the model for long-horizon tasks. The long context supports
large codebases and document collections, but it should not be treated as a
reason to send every file on every turn. Search and context selection still
reduce latency and cost.
## GLM-5.2 API pricing
| Route | Input / MTok | Cached input | Output / MTok |
| ------------- | -----------: | -----------------------: | ------------: |
| Z.AI official | $1.40 | $0.26 | $4.40 |
| reAPI | **$0.90** | Not separately published | **$3.00** |
For a monthly workload with 100M uncached input tokens and 20M output tokens,
the simple standard-rate calculation is:
```text
Z.AI: 100 × $1.40 + 20 × $4.40 = $228
reAPI: 100 × $0.90 + 20 × $3.00 = $150
```
That example assumes the same token use and no cache hits. Z.AI's $0.26 cached
rate can materially change a workload with repeated system prompts, tool
definitions, or stable repository context.
## The reasoning-effort labels collapse into three behaviors
GLM-5.2 accepts more effort strings than it meaningfully executes. According
to Z.AI's request behavior, the values map as follows:

| Values sent | Effective behavior |
| ----------------------- | ------------------ |
| `none`, `minimal` | Skip thinking |
| `low`, `medium`, `high` | High reasoning |
| `xhigh`, `max` | Maximum reasoning |
The default is `max`, which is expensive for routine classification, extraction,
and simple code edits. If the task needs reasoning but not the maximum, send
`high`. If it is deterministic and simple, compare `none` with thinking
disabled.
`reasoning_effort` only matters while thinking is enabled. Do not expect `max`
to override `thinking: {"type": "disabled"}`.
## How conversation thinking is handled
GLM-5.2 separates visible `content` from `reasoning_content`. With the default
`clear_thinking: true`, prior reasoning is removed rather than preserved across
turns. Set it to `false` only when the complete reasoning history remains
relevant and your application can replay the full ordered messages correctly.\[3]
This differs from Kimi K3, where preserved thinking is mandatory. GLM-5.2 lets
you choose between a cleaner context and continuity of hidden work.
## Calling GLM-5.2 with the OpenAI Python SDK
```python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_REAPI_KEY",
base_url="https://api.reapi.ai/v1",
)
stream = client.chat.completions.create(
model="glm-5.2",
messages=[
{"role": "user", "content": "Refactor this module and add tests."}
],
max_tokens=8192,
stream=True,
extra_body={
"thinking": {"type": "enabled"},
"reasoning_effort": "high",
},
)
for chunk in stream:
delta = chunk.choices[0].delta
reasoning = getattr(delta, "reasoning_content", None)
if reasoning:
print(reasoning, end="")
if delta.content:
print(delta.content, end="")
```
The model id contains a dot. `glm-5-2` is the website slug and will produce an
unknown-model error if sent in the request body.
## Function calling and structured output
GLM-5.2 supports up to 128 function definitions and streams tool-call arguments
when `tool_stream` is enabled. Its `tool_choice` behavior is narrower than many
OpenAI models: `auto` is the supported choice. Design the prompt and tool
descriptions so the model can decide when to call them.
The model supports JSON object mode through:
```json
{
"response_format": {"type": "json_object"}
}
```
This is not JSON Schema enforcement. Validate the returned object in your
application and retry or repair when required fields are missing.
## Temperature, top-p, and stop sequences
Unlike several reasoning models, GLM-5.2 allows temperature and top-p tuning.
Z.AI recommends changing one rather than both, because two simultaneous
sampling changes make output behavior harder to attribute.\[1]
* Use the defaults for general agent work.
* Lower temperature for repeatable extraction or transformation.
* Use one stop string only; the API does not accept a long stop list.
* Keep `max_tokens` realistic. A 128K ceiling is not a target for every call.
For coding agents, reasoning effort and context selection usually affect cost
more than small sampling changes.
## When to use GLM-5.2
**Use it for large-codebase work.** The context window and tool support fit
repository navigation, migrations, test generation, and long-running repair.
**Use it for cost-sensitive agents.** Its official and reAPI rates sit well
below the flagship Claude and GPT tiers.
**Use it when open weights matter.** Teams can evaluate self-hosting or private
deployment rather than depend exclusively on a closed API.
**Avoid it when native image understanding is required.** The documented
GLM-5.2 API is text in and text out. Choose a vision model for image or video
analysis.
## Common integration mistakes
1. Sending `glm-5-2` instead of `glm-5.2`.
2. Leaving the default `max` effort on every low-value request.
3. Expecting `response_format` to enforce a JSON Schema.
4. Setting `tool_choice` to an unsupported forced-tool value.
5. Tuning temperature and top-p simultaneously.
6. Reading only `content` during streaming and discarding
`reasoning_content` without deciding whether it is needed.
7. Filling the 1M context window instead of retrieving relevant files.
## FAQ
### Is GLM-5.2 open source?
It is safer to describe GLM-5.2 as open weight unless the specific license and
release artifacts satisfy your definition of open source. Review the current
model license before redistribution or commercial deployment.
### How much does the GLM-5.2 API cost?
Z.AI lists $1.40 input, $0.26 cached input, and $4.40 output per million tokens.
reAPI currently lists $0.90 input and $3 output.
### Does GLM-5.2 support one million tokens?
Yes. Z.AI documents a 1M-token context window and 128K maximum output.
### Can thinking be disabled?
Yes. Set `thinking.type` to `disabled`, or use a no-thinking effort value where
supported. The default is enabled with maximum reasoning.
### Does GLM-5.2 support images?
No on the documented chat model surface. It accepts text and returns text.
### Is the GLM-5.2 API OpenAI compatible?
The reAPI route is compatible with OpenAI Chat Completions. Vendor-specific
fields such as `thinking` and `reasoning_effort` should be passed through the
SDK's extra-body mechanism.
## Conclusion
The GLM-5.2 API is compelling when a one-million-token context, open weights,
and low token prices matter more than native multimodality. The integration is
straightforward once three details are handled correctly: send `glm-5.2` with
the dot, lower the default `max` effort for routine traffic, and treat JSON
mode as syntax—not schema validation.
## References
1. Z.AI. *GLM-5.2 model guide and API behavior.* [docs.z.ai/guides/llm/glm-5.2](https://docs.z.ai/guides/llm/glm-5.2)
2. Z.AI. *Model pricing.* [docs.z.ai/guides/overview/pricing](https://docs.z.ai/guides/overview/pricing)
3. Z.AI. *Chat Completions API reference.* [docs.z.ai/api-reference/llm/chat-completion](https://docs.z.ai/api-reference/llm/chat-completion)
### Further reading
* reAPI. *GLM-5.2 API documentation.* [reapi.ai/docs/glm-5-2](/docs/glm-5-2)
* reAPI. *GLM-5.2 model page.* [reapi.ai/models/glm-5-2](/models/glm-5-2)
* reAPI. *Kimi K3 complete guide.* [reapi.ai/blog/kimi-k3-complete-guide](/blog/kimi-k3-complete-guide)
---
# GPT API Price Cut: 5 Models Now Cost 20% Less on reAPI (https://reapi.ai/blog/gpt-api-price-cut-2026)
The latest **GPT API price cut** is easy to misunderstand. OpenAI has not
announced a blanket reduction to its public API list prices. What changed for
reAPI users is the route: five current GPT models are now listed at **20% below
OpenAI's standard input and output token rates**.
The discount covers GPT-5.4, GPT-5.5, and the complete GPT-5.6 family—Sol,
Terra, and Luna. There is no new SDK or migration project. Existing
OpenAI-compatible code can use the lower rates by changing the base URL and,
where necessary, the model id.
This guide compares every rate, calculates the savings on a real workload,
and explains when OpenAI's Batch API or prompt caching can still change the
answer.
## TL;DR
* **Five GPT models are 20% below OpenAI's standard list price on reAPI:**
GPT-5.4, GPT-5.5, GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna.
* **This is a reAPI discount, not an OpenAI-wide price cut.** OpenAI's current
public prices remain unchanged.\[1]\[2]
* **The cheapest route is GPT-5.6 Luna** at $0.80 input and $4.80 output per
million tokens on reAPI.
* **Terra and GPT-5.4 share the same token price** on reAPI: $2 input and $12
output per million tokens. Benchmark both before choosing.
* **GPT-5.5 and GPT-5.6 Sol also share a price:** $4 input and $24 output per
million tokens on reAPI.
* **A 20% token discount produces a 20% bill reduction** when request mix and
token usage stay the same. A model that uses fewer output tokens can reduce
cost per completed task by more.
## GPT API prices before and after the discount
All prices below are USD per 1 million tokens. "OpenAI list" means standard
processing, not Batch, Flex, Priority, regional processing, or a negotiated
enterprise contract. reAPI figures are from its live public rate card on
August 1, 2026.\[3]
| Model | OpenAI input | reAPI input | OpenAI output | reAPI output | Reduction |
| ------------- | -----------: | ----------: | ------------: | -----------: | --------: |
| GPT-5.4 | $2.50 | **$2.00** | $15.00 | **$12.00** | 20% |
| GPT-5.5 | $5.00 | **$4.00** | $30.00 | **$24.00** | 20% |
| GPT-5.6 Sol | $5.00 | **$4.00** | $30.00 | **$24.00** | 20% |
| GPT-5.6 Terra | $2.50 | **$2.00** | $15.00 | **$12.00** | 20% |
| GPT-5.6 Luna | $1.00 | **$0.80** | $6.00 | **$4.80** | 20% |
OpenAI publishes a separate cached-input price for these models. The current
reAPI rate card does not expose a cached-input rate for the discounted default
routes, so this comparison does not invent one. Treat repeated prompt prefixes
as a separate benchmark rather than assuming the same 20% relationship.
## What the GPT API price cut means for a real bill
Suppose an agent processes 100 million input tokens and generates 20 million
output tokens each month. With identical token usage, the monthly calculation
is:
```text
monthly cost = input MTok × input rate + output MTok × output rate
```

| Model | OpenAI standard | reAPI | Monthly savings |
| ------------- | --------------: | -------: | --------------: |
| GPT-5.4 | $550 | **$440** | **$110** |
| GPT-5.5 | $1,100 | **$880** | **$220** |
| GPT-5.6 Sol | $1,100 | **$880** | **$220** |
| GPT-5.6 Terra | $550 | **$440** | **$110** |
| GPT-5.6 Luna | $220 | **$176** | **$44** |
The percentage is constant, but the dollar impact grows with spend. A team
currently paying $10,000 per month for the same mix of uncached standard
traffic would retain roughly $2,000 after moving it, before considering any
change in model behavior or token consumption.
That last condition matters. API cost is not merely price per token:
```text
cost per successful task
= price per token × tokens per attempt × attempts per success
```
A cheaper model that needs two retries can cost more than an expensive model
that succeeds once. A more concise model can beat the table by emitting fewer
billable output tokens. OpenAI says GPT-5.6 improves token efficiency, so a
carefully evaluated migration may save more than the headline 20% per-token
discount.\[1]
## Which discounted GPT model should you use?
The table creates two same-price pairs, which makes the buying decision more
interesting than simply choosing the newest model.
### GPT-5.6 Luna: the volume default
Luna is the lowest-priced model in this group. OpenAI positions it for
cost-sensitive, high-volume work, with a 1.05M-token context window and 128K
maximum output. It is the first candidate for classification, extraction,
ranking, routine support, and inexpensive agent subtasks.\[4]
At $0.80 input and $4.80 output per million tokens on reAPI, Luna is 60%
cheaper than Terra on both dimensions. Use the price gap to run a proper eval,
not to skip one: measure task completion, retries, latency, and total output
tokens.
### GPT-5.6 Terra versus GPT-5.4: same price, newer tier
Both cost $2 input and $12 output per million tokens on reAPI. Terra is the
natural first test because OpenAI describes it as the balanced GPT-5.6 tier,
competitive with GPT-5.5 while priced below the flagship.\[1]
GPT-5.4 remains useful when an existing production prompt has already been
validated against it. A same-price migration is not automatically free: model
behavior, reasoning defaults, cache behavior, and output length can change.
Run a shadow comparison before moving all traffic.
### GPT-5.6 Sol versus GPT-5.5: capability at the same rate
Both cost $4 input and $24 output per million tokens on reAPI. OpenAI's own
standard list price is also identical for the pair at $5 and $30. Sol is the
new flagship and adds the GPT-5.6 generation's updated reasoning, tool use,
and explicit prompt-caching controls, while GPT-5.5 is the stable baseline for
teams that have already tuned around it.\[1]\[2]
For new high-complexity deployments, start the evaluation with Sol. Keep
GPT-5.5 only when regression tests show that its established behavior wins on
your workload.
## When 20% off is not the cheapest option
A clean price table is useful, but three billing details can reverse a narrow
comparison.
### Batch and Flex processing
OpenAI offers Batch and Flex processing at 50% of standard rates for GPT-5.4
and GPT-5.5, and Batch processing for supported GPT-5.6 requests.\[2]\[5]
If a workload can wait and does not need an immediate response, compare those
asynchronous prices with reAPI's live rate instead of comparing standard
processing alone.
The 20% reAPI discount is most directly relevant to synchronous, standard API
traffic. A nightly offline enrichment job has different economics from a live
coding agent.
### Prompt caching
OpenAI lists cached reads for GPT-5.6 at 10% of uncached input price. GPT-5.6
also bills explicit cache writes at 1.25 times the uncached input rate, so the
net result depends on how often the prefix is reused.\[1]
For a long system prompt reused hundreds of times, cache hit rate can matter
more than a 20% reduction in ordinary input price. Track uncached input,
cache-write tokens, cache-read tokens, and output separately.
### Long-context requests
For GPT-5.4 and GPT-5.6 models with 1.05M context windows, OpenAI prices prompts
above 272K input tokens at 2x input and 1.5x output for the full request.\[4]\[6]
The reAPI public rate card shows base model ratios, not a complete long-context
invoice example. If your prompts cross that threshold, test a representative
request and verify billed usage before projecting savings.
## How to use the lower GPT API rates
reAPI exposes an OpenAI-compatible endpoint, so the Python client only needs a
different base URL. This example uses Terra, the balanced tier:
```python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_REAPI_KEY",
base_url="https://api.reapi.ai/v1",
)
response = client.chat.completions.create(
model="gpt-5.6-terra",
messages=[
{"role": "user", "content": "Review this pull request for data-loss risks."}
],
)
print(response.choices[0].message.content)
```
The available model ids are:
* `gpt-5.4`
* `gpt-5.5`
* `gpt-5.6-sol`
* `gpt-5.6-terra`
* `gpt-5.6-luna`
Do not replace the dots with hyphens in API requests. The reAPI website uses
hyphenated page slugs, but the wire model ids preserve OpenAI's dotted version
numbers.
## A safe migration checklist
Moving an API route is simple; proving that the new route is economical takes
a little more work.
1. **Export a representative sample.** Include short and long prompts, tool
calls, structured output, and known failure cases.
2. **Pin the model id.** Do not compare moving aliases if you need a stable
before-and-after result.
3. **Run both routes.** Measure answer quality, first-token latency, total
latency, input tokens, output tokens, retries, and errors.
4. **Calculate cost per accepted result.** Token price alone hides retries and
human review.
5. **Inspect feature compatibility.** Validate tools, structured outputs,
streaming, and image input used by the application.
6. **Move traffic gradually.** Start with a small percentage and keep the old
route available until error rates and bills match the test.
7. **Check the live rate card.** Pricing can change after an article is
published; treat the model page and API rate card as the source of truth.
## FAQ
### Did OpenAI lower GPT API prices?
Not across the five models in this comparison. OpenAI's current public list
prices remain $2.50/$15 for GPT-5.4, $5/$30 for GPT-5.5 and GPT-5.6 Sol,
$2.50/$15 for GPT-5.6 Terra, and $1/$6 for GPT-5.6 Luna, per million input and
output tokens. The 20% reduction described here is the current reAPI rate
relative to those standard prices.\[1]\[2]
### Which GPT models are 20% cheaper on reAPI?
GPT-5.4, GPT-5.5, GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna. The discount
applies to the published input and output token rates on reAPI's default
routes.\[3]
### What is the cheapest GPT-5.6 API model?
GPT-5.6 Luna. It costs $1 input and $6 output per million tokens at OpenAI's
standard rate, or $0.80 input and $4.80 output on reAPI's current default
route.\[3]\[4]
### Is GPT-5.6 Terra cheaper than GPT-5.4?
No. They have the same token prices in this comparison: $2/$12 on reAPI and
$2.50/$15 at OpenAI. Terra is newer, but the best value depends on task success,
token efficiency, and migration risk.
### Is GPT-5.6 Sol cheaper than GPT-5.5?
No. Both are $4/$24 per million input/output tokens on reAPI and $5/$30 at
OpenAI's standard rate. Sol offers the newer capability tier at the same
per-token price.\[1]
### Does reAPI offer discounted cached input?
The current public rate card does not publish a separate cached-input ratio
for these discounted default routes. Do not assume one. Compare a real cached
workload and check live billing before migration.
### Can I keep using the OpenAI Python SDK?
Yes. Set `base_url` to `https://api.reapi.ai/v1`, use a reAPI key, and keep the
canonical dotted GPT model id.
## The price cut is useful; cost per result is the decision
A uniform 20% reduction is unusually easy to budget. Keep the workload and
model fixed, and a $1,000 standard-token bill becomes roughly $800. The more
valuable opportunity is model routing: Luna for routine volume, Terra for the
balanced default, and Sol only where frontier capability improves the success
rate enough to pay for itself.
The honest headline is therefore narrower than "OpenAI cut prices" and more
useful to developers: **five current GPT models now run at 80% of OpenAI's
standard token rates on reAPI**. Verify your cache pattern, long-context usage,
and asynchronous options, then measure cost per accepted result before moving
production traffic.
## References
1. OpenAI. *GPT-5.6: Frontier intelligence that scales with your ambition.* Pricing and tier details. [openai.com/index/gpt-5-6](https://openai.com/index/gpt-5-6/)
2. OpenAI. *Introducing GPT-5.5.* API pricing and processing options. [openai.com/index/introducing-gpt-5-5](https://openai.com/index/introducing-gpt-5-5/)
3. reAPI. *Live public API pricing.* Retrieved August 1, 2026. [api.reapi.ai/api/pricing](https://api.reapi.ai/api/pricing)
4. OpenAI. *GPT-5.6 Luna model.* Specifications, pricing, caching, and long-context rules. [developers.openai.com/api/docs/models/gpt-5.6-luna](https://developers.openai.com/api/docs/models/gpt-5.6-luna)
5. OpenAI. *API pricing.* Standard and Batch processing rates. [openai.com/api/pricing](https://openai.com/api/pricing/)
6. OpenAI. *GPT-5.4 model.* Pricing and long-context rules. [developers.openai.com/api/docs/models/gpt-5.4](https://developers.openai.com/api/docs/models/gpt-5.4)
### Further reading
* reAPI. *How to use GPT-5.6.* [reapi.ai/blog/how-to-use-gpt-5-6](/blog/how-to-use-gpt-5-6)
* reAPI. *How to use GPT-5.5.* [reapi.ai/blog/how-to-use-gpt-5-5](/blog/how-to-use-gpt-5-5)
* reAPI. *GPT-5.6 model catalog.* [reapi.ai/models](/models)
---
# GPT Image 2 + Seedance 2.0: A Character Consistency Workflow (https://reapi.ai/blog/gpt-image-2-seedance-2-0-character-workflow)
**The most reliable GPT Image 2 + Seedance 2.0 workflow separates identity
design from motion.** Build and approve the character in GPT Image 2 first,
reuse that same visual package to create every shot's start frame, then ask
Seedance 2.0 to animate the frame instead of reinventing the person from text.
This reduces character drift because each model gets one job. GPT Image 2
controls face, silhouette, wardrobe, and framing. Seedance controls movement,
camera, timing, and audio. It is not a guarantee that a character “never
drifts”—no current API exposes a perfect identity lock—but it is a repeatable
production workflow with clear checkpoints.
## TL;DR
* **Lock the still before generating motion.** A weak or inconsistent start
frame becomes a weak video.
* **Create one canonical character package.** Keep a neutral portrait, full
body view, three-quarter view, and a short written identity block.
* **Generate each shot frame from those frozen references.** GPT Image 2 on
reAPI accepts up to 16 public image URLs in one image-to-image request.
* **Prompt Seedance for motion, not identity.** Pass the approved frame and
describe action, camera, pacing, and sound.
* **Reuse references without swapping them mid-project.** A “better” new face
halfway through a sequence creates a second definition of the character.
* **Chain clips with the returned last frame.** Seedance 2.0 can return a final
frame for the next segment, reducing visual jumps at joins.
## Why characters drift between AI video shots
Text descriptions are not identity records. “A woman with short black hair,
green eyes, and a red jacket” describes a category containing millions of
possible faces. Every fresh text-to-video request can sample another valid
person while still following the words.
Drift also enters through changes that look harmless to a human reviewer:
* a different crop hides the jawline or body proportions;
* dramatic lighting changes apparent eye, hair, and skin color;
* wardrobe prompts compete with facial details;
* a new reference introduces a different nose, age, or hairstyle;
* motion instructions ask for turns or occlusion the source frame cannot
support.
Seedance 2.0 accepts text, images, audio, and video as reference modalities,
with up to nine images, three videos, and three audio clips on its current open
platform.\[2] More references do not
automatically mean more consistency. They work only when they agree.
## The two-stage character consistency workflow

The pipeline has six gates. Do not move forward until the current gate passes,
because repairing identity after animation is slower than regenerating a still.
### 1. Write an identity contract
Define only features that must survive every shot. Keep the block short enough
to paste unchanged into image prompts.
```text
CHARACTER: Mara, fictional adult woman, early 30s
FACE: oval face, wide-set brown eyes, straight nose, small scar above left eyebrow
HAIR: chin-length black bob, center part
WARDROBE: burnt-orange field jacket, charcoal crew-neck shirt
ANCHOR: narrow silver compass pendant
DO NOT CHANGE: age, facial geometry, hair length, scar side, pendant shape
```
Avoid subjective labels such as “beautiful,” “cool,” or “cinematic.” They
invite the model to redesign the person. Physical anchors are easier to audit.
### 2. Generate the master portrait in GPT Image 2
Start with even light, a neutral expression, visible facial geometry, and a
simple background. GPT Image 2 supports both image generation and editing;
OpenAI documents text as input and image as both input and output.\[1]
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-image-2",
"prompt": "Character reference portrait. Fictional adult woman, early 30s, oval face, wide-set brown eyes, straight nose, small scar above left eyebrow, chin-length black bob with center part, neutral expression, even soft light, plain warm-gray background, realistic editorial photography, no text",
"size": "3:4",
"resolution": "2k"
}'
```
Save the returned URL as the canonical portrait. Do not choose a dramatic hero
image as the master; choose the face that is easiest to read.
### 3. Expand one portrait into a reference package
Use image-to-image requests with the approved portrait in `image_urls`.
Generate a full-body neutral stance, left and right three-quarter views, and a
wardrobe sheet. Ask the model to preserve identity rather than “make a similar
person.”
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-image-2",
"prompt": "Preserve the exact identity, age, facial geometry, scar position, hair cut, jacket, shirt, and compass pendant from the reference. Create a clean full-body character sheet with front and three-quarter views, neutral stance, even studio light, plain background, no labels, no extra people",
"image_urls": ["https://your-cdn.com/mara-master.png"],
"size": "16:9",
"resolution": "2k"
}'
```
The inexpensive `gpt-image-2` route returns one image per request and accepts
up to 16 public HTTP(S) reference URLs. It does not expose a seed or persistent
character ID.\[3] Consistency therefore comes
from repeatedly supplying the same approved pixels.
### 4. Create a controlled start frame for every shot
Write the shot list before generating more images. For each shot, feed the
master portrait and character sheet back into GPT Image 2, then change only the
scene, pose, lens, and composition.
```text
Preserve the exact identity and wardrobe from Image 1 and Image 2.
Medium-wide 16:9 frame. Mara stands beside a rain-streaked train window at
night, body at three-quarter angle, face fully visible, compass pendant visible.
Camera at eye level, 50mm lens, soft carriage practical light.
No additional people. Do not change facial geometry, age, hairstyle, scar side,
jacket color, or pendant.
```
Approve the still against the master before paying to animate it. This is the
cheapest point to reject a changed face, missing anchor, or incompatible pose.
### 5. Animate the frame with Seedance 2.0
Pass the approved shot frame first, followed by one or two canonical identity
references if needed. Seedance chooses image-to-video mode when `image_urls`
is present; there is no separate `mode` field on reAPI.\[4]
```bash
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "doubao-seedance-2.0-face",
"prompt": "@Image1 remains the same character. She turns slowly toward the rain-streaked window and exhales. Subtle blink, natural breathing, gentle train vibration. Slow 10% camera push-in. Preserve face, black bob, eyebrow scar, orange jacket, and silver compass pendant. No scene cut, no wardrobe change.",
"image_urls": [
"https://your-cdn.com/shot-01-start.png",
"https://your-cdn.com/mara-master.png"
],
"size": "16:9",
"resolution": "720p",
"duration": 5,
"generate_audio": true,
"return_last_frame": true
}'
```
Use a Face variant for identifiable real-person source material. Synthetic
characters can use an eligible non-Face or official route, but model
availability and safety rules still apply. Only upload likenesses you have the
right and consent to use.
### 6. Review, reject, and chain
Review the completed clip frame by frame. Check the face at the beginning,
middle, and end, not only the thumbnail.
| Check | Pass condition | If it fails |
| ------------------- | -------------------------------------------- | --------------------------------------------- |
| Facial geometry | Eyes, nose, jaw, and age remain stable | Regenerate from the same approved start frame |
| Hair silhouette | Length and part stay recognizable | Reduce head turns or wind motion |
| Identity anchors | Scar, pendant, and jacket remain present | Shorten the prompt and restate invariants |
| Hands and occlusion | Face is not rebuilt after a long obstruction | Reduce occlusion or split the action |
| Final frame | Clean enough to begin the next shot | Request another take before chaining |
When a clip passes, use `output.last_frame_url` as the first image of the next
Seedance request. Keep the canonical portrait in the remaining reference slots.
The last frame protects spatial continuity; the master portrait protects
identity.
## Prompting rules that reduce character drift
The order of instructions matters. A useful Seedance motion prompt follows
this sequence:
```text
identity binding → action → camera → environment motion → audio → invariants
```
For example:
```text
@Image1 remains the same character. She takes two measured steps toward the
door and looks over her left shoulder. Handheld camera tracks backward at chest
height. Curtains move slightly in the draft; quiet room tone and footsteps.
Preserve face, age, black bob, eyebrow scar, orange jacket, and pendant. No cut,
no new person, no costume change, no extreme motion blur.
```
Do not repeat a long visual description that conflicts with the supplied
frame. The frame already defines the scene. The prompt should spend most of its
tokens on change over time.
## What not to do
Five common shortcuts make drift worse.
1. **Do not generate every shot from text.** Each request samples a new person.
2. **Do not replace the master after shot three.** Fix a bad shot, not the
canonical identity.
3. **Do not mix references from different generations.** Minor disagreements
in face or clothing become competing instructions.
4. **Do not ask for extreme motion immediately.** Test a blink, small turn, or
short walk before flips, spins, crowds, and heavy occlusion.
5. **Do not use a drifted last frame as the only next reference.** Pair it with
the clean master or regenerate the transition.
The [Seedance 2.0 character consistency guide](/blog/seedance-2-0-character-consistency-guide)
covers reference ordering, voice reuse, multi-shot prompts, and longer chains
in more detail.
## Cost and iteration strategy
This workflow costs less when failures are caught in the still-image stage.
One GPT Image 2 request produces one candidate, while Seedance bills video by
output duration and resolution. Uploaded reference video can also add billable
seconds; image references do not.\[3]\[4]
A disciplined production loop is:
1. Generate low-count 1K or 2K character drafts.
2. Approve one master and stop exploring identity.
3. Generate and review every start frame as a still.
4. Animate five-second 720p tests with simple motion.
5. Increase duration or resolution only after the prompt passes.
Current rates can change, so calculate the budget from the live [GPT Image 2
model page](/models/gpt-image-2) and [Seedance 2.0 model page](/models/seedance-2-0)
rather than copying an old per-image or per-second figure into a production
spreadsheet. The [Seedance cost-per-second guide](/blog/seedance-2-0-cost-per-second)
explains the billing dimensions.
## A reusable folder structure
Treat character consistency as asset management, not prompt luck.
```text
project/
character/
identity-contract.txt
master-portrait.png
character-sheet.png
approved-refs.json
shots/
01-train-window/
start-frame.png
motion-prompt.txt
output.mp4
last-frame.png
02-platform/
start-frame.png
motion-prompt.txt
```
Version the identity contract and reference URLs. If the art director approves
a real redesign, create character package v2 and regenerate all dependent
shots instead of mixing versions.
## FAQ
### Does GPT Image 2 + Seedance 2.0 guarantee zero character drift?
No. It reduces drift by anchoring every shot to approved images, but neither
API exposes a perfect persistent character lock. Motion, occlusion, conflicting
references, and long clips can still change identity.
### How many character references should I use?
Start with the approved shot frame plus one master portrait. Add a character
sheet only when it contributes a missing angle or full-body information. The
API maximum is not a target.
### Should the GPT Image 2 master portrait be cinematic?
No. Use even light, a neutral expression, a visible face, and a simple
background. Build cinematic lighting into each shot frame after identity is
approved.
### Which Seedance 2.0 model should I use for a real person?
Use a Face variant that accepts identifiable real-person inputs and confirm
you have consent and usage rights. Official/non-Face routes may reject such
references.\[4]
### Can I reuse Seedance's last frame for the next clip?
Yes. Set `return_last_frame: true`, then pass the returned URL into the next
request. Keep the original master reference alongside it so small errors do not
compound across the sequence.
### Should I add negative prompts such as “no face drift”?
Brief invariants can help, but reference quality and achievable motion matter
more. A clean start frame plus a small, visible action is stronger than a long
list of negative terms.
## The character-consistent AI video workflow in one sentence
GPT Image 2 defines and stages the character; Seedance 2.0 animates approved
frames; a human review gate prevents one drifting shot from contaminating the
next. That is the practical GPT Image 2 + Seedance 2.0 character consistency
workflow—not a promise of perfect identity, but a controlled pipeline that can
be repeated, measured, and improved.
For a broader image-generation overview, read [How to use GPT Image 2](/blog/how-to-use-gpt-image-2).
For Seedance references beyond characters, see the [Seedance 2.0 API
documentation](/docs/seedance-2-0).
## References
1. OpenAI. *GPT Image 2 model documentation — image generation and editing, modalities, endpoints, and snapshots.* Retrieved August 2, 2026. [developers.openai.com/api/docs/models/gpt-image-2](https://developers.openai.com/api/docs/models/gpt-image-2)
2. Team Seedance et al. *Seedance 2.0: Advancing Video Generation for World Complexity.* April 2026. [arxiv.org/abs/2604.14148](https://arxiv.org/abs/2604.14148)
3. reAPI. *GPT Image 2 API documentation — reference limits, resolutions, and async task behavior.* Retrieved August 2, 2026. [reapi.ai/docs/gpt-image-2](/docs/gpt-image-2)
4. reAPI. *Seedance 2.0 API documentation — mode routing, reference limits, variants, last-frame return, and billing.* Retrieved August 2, 2026. [reapi.ai/docs/seedance-2-0](/docs/seedance-2-0)
5. Devenko AI. *GPT Image 2.0 + Seedance 2.0: The Only Workflow Where Your Character Never Drifts.* July 2026. [devenkoai.medium.com](https://devenkoai.medium.com/gpt-image-2-0-seedance-2-0-the-only-workflow-where-your-character-never-drifts-589ae53c79cd)
---
# Grok Imagine Video 1.5 API: Pricing, 1080p, and Limits (2026) (https://reapi.ai/blog/grok-imagine-video-1-5-api)
**Grok Imagine Video 1.5 is an image-to-video API, not a text-to-video model.**
You provide one source image plus a motion prompt, and the model animates that
first frame with optional native audio. xAI's current API offers 480p, 720p,
and 1080p pricing; the current reAPI route exposes 480p and 720p.\[1]
That distinction prevents the most common failed request: sending only text to
the 1.5 model. If you need text-to-video, reference-to-video, editing, or video
extension, xAI documents those capabilities under the broader
`grok-imagine-video` family rather than the 1.5 image-to-video endpoint.\[2]
## TL;DR
* **Input:** one image and a motion prompt.
* **Output:** a 1–15 second video through xAI; reAPI's recommended channel uses
6–15 seconds for fallback compatibility.
* **xAI list price:** $0.08/sec at 480p, $0.14/sec at 720p, and $0.25/sec at
1080p, plus $0.01 for the input image.\[1]
* **reAPI current recommended channel:** $0.03/sec at 480p and $0.06/sec at
720p; check the live model page before budgeting.
* **Native audio:** supported, with an `audio` toggle on the reAPI official
channel.
* **Maximum duration:** 15 seconds.
* **Main limitation:** Grok Imagine Video 1.5 does not accept text-only input
and the reAPI surface does not currently expose 1080p.
## Grok Imagine Video 1.5 specifications
| Property | xAI API | reAPI recommended channel |
| ------------- | ------------------------ | -------------------------------------- |
| Model ID | `grok-imagine-video-1.5` | `grok-imagine-video-1.5-official` |
| Input | Image + prompt | One public image URL + required prompt |
| Resolutions | 480p, 720p, 1080p | 480p, 720p |
| Duration | Up to 15 seconds | 6–15 seconds |
| Aspect ratios | Vendor-supported set | `1:1`, `16:9`, `9:16` |
| Audio | Native | Native, toggle available |
| Billing | Per second + image input | Per second |
| Execution | Asynchronous | Asynchronous task |
xAI describes the source image as the first frame. That makes the model useful
for product shots, illustrations, portraits, and keyframes where the starting
composition must be preserved.\[3]
## Grok Imagine Video 1.5 API pricing

### xAI direct pricing
| Resolution | Video rate | Input-image fee | 8-second example |
| ---------- | ---------: | --------------: | ---------------: |
| 480p | $0.08/sec | $0.01 | $0.65 |
| 720p | $0.14/sec | $0.01 | $1.13 |
| 1080p | $0.25/sec | $0.01 | $2.01 |
The example is `rate × 8 + $0.01`. A 15-second 1080p clip therefore reaches
$3.76 before retries. Resolution is the largest price lever.\[1]
### reAPI pricing and channel differences
reAPI currently exposes two channels. The recommended `official` route lists
$0.03/sec for 480p and $0.06/sec for 720p. The beta route has a broader aspect-
ratio set and an `nsfw_checker` field but a higher current rate. The official
route requires a prompt, supports 6–15 seconds, and provides an audio toggle.
For an 8-second clip, the current recommended route is approximately:
| Resolution | Per-second rate | 8-second clip |
| ---------- | --------------: | ------------: |
| 480p | $0.03 | $0.24 |
| 720p | $0.06 | $0.48 |
Rates are operational data, not permanent facts. Use the live
[Grok Imagine Video 1.5 model page](/models/grok-imagine-video-1-5) before
publishing a calculator or committing customer prices.
## How to call Grok Imagine Video 1.5 on reAPI
The endpoint is asynchronous. Submit a job, store the task id, and poll the
task endpoint until it completes.
```python
import requests
payload = {
"model": "grok-imagine-video-1.5-official",
"prompt": "Slow cinematic push-in, natural hair movement, warm sunset light",
"image_urls": ["https://your-cdn.com/keyframe.jpg"],
"aspect_ratio": "16:9",
"resolution": "720p",
"duration": 8,
"audio": True,
}
response = requests.post(
"https://reapi.ai/api/v1/videos/generations",
headers={
"Authorization": "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
json=payload,
timeout=30,
)
response.raise_for_status()
task = response.json()
print(task["id"])
```
Poll `GET /api/v1/tasks/{id}`. Completed results expose the generated video
URL, which should be copied to your own storage if the application needs a
durable asset.
## Writing prompts that preserve the source image
The strongest prompts describe motion and camera behavior rather than
redesigning the entire frame. The source image already supplies composition,
subject, clothing, palette, and background.
Use this order:
1. **Subject motion:** blink, turn, walk, fabric movement, product rotation.
2. **Camera motion:** push in, dolly left, locked camera, handheld drift.
3. **Environment motion:** wind, rain, traffic, water, particles.
4. **Audio:** ambience, dialogue, effects, or music direction.
5. **Invariants:** keep identity, product shape, logo placement, or framing.
Example:
```text
The camera makes a slow 10% push-in. The subject blinks once and turns slightly
toward the window. Curtains move in a light breeze. Preserve the exact face,
clothing, room layout, and morning color palette. Quiet room tone and distant
city ambience; no dialogue.
```
Avoid asking for ten shots in eight seconds unless rapid montage is the goal.
Short clips have little time to establish multiple camera moves and actions.
## 480p, 720p, or 1080p?
**Use 480p for motion tests.** It is the cheapest way to test identity drift,
camera interpretation, and prompt timing.
**Use 720p for most social and product drafts.** It provides a useful balance
between detail and retry cost and is the highest current reAPI option.
**Use 1080p through xAI when final delivery needs it.** The rate is more than
three times the 480p rate, so validate motion at lower resolution first. A
1080p generation is not exposed by the current reAPI route described here.
Upscaling a successful 720p clip can be cheaper than regenerating multiple
1080p attempts, but it cannot recover missing motion detail or fix identity
drift.
## Limits and failure modes
* **Text-only requests fail conceptually.** The 1.5 endpoint requires an image.
* **Temporary URLs expire.** Download completed videos promptly.
* **Audio adds another failure surface.** Check lip sync, unwanted dialogue,
clipping, and inconsistent ambience.
* **Longer clips multiply cost and drift risk.** Test the first six to eight
seconds before using the 15-second maximum.
* **Input images are billable on xAI direct.** Include the $0.01 image fee in
cost calculators.
* **Moderation applies.** Successful upload does not guarantee generation.
For identity-sensitive work, use a clean source image with one clear subject,
good lighting, and no conflicting motion cues.
## FAQ
### Does Grok Imagine Video 1.5 support text-to-video?
No. The current 1.5 model is image-to-video. Use a source image or select a
different Grok Imagine model that explicitly supports text input.
### Does it generate audio?
Yes. xAI documents native audio, and reAPI's official channel exposes an
`audio` toggle.
### How long can a video be?
Up to 15 seconds. The current reAPI recommended route uses a 6-second minimum.
### Does reAPI support 1080p?
Not on the current Grok Imagine Video 1.5 route. It exposes 480p and 720p. xAI
direct documents 1080p at $0.25 per second.
### How much does an eight-second 720p clip cost?
xAI direct lists approximately $1.13 including the input-image fee. The current
reAPI recommended route lists approximately $0.48. Check live pricing before
production use.
## Conclusion
The Grok Imagine Video 1.5 API is a focused keyframe animator: one image in,
up to 15 seconds of video out, with native audio and resolution-based pricing.
Test motion at 480p, move successful prompts to 720p, and use xAI's 1080p tier
only when the delivery requirement justifies the retry cost.
## References
1. xAI. *Grok Imagine Video 1.5 model and resolution pricing.* [docs.x.ai/developers/models/grok-imagine-video-1.5](https://docs.x.ai/developers/models/grok-imagine-video-1.5)
2. xAI. *Video generation capabilities and limitations.* [docs.x.ai/developers/model-capabilities/video/generation](https://docs.x.ai/developers/model-capabilities/video/generation)
3. xAI. *Imagine API overview and image-to-video quickstart.* [docs.x.ai/developers/model-capabilities/imagine](https://docs.x.ai/developers/model-capabilities/imagine)
### Further reading
* reAPI. *Grok Imagine Video 1.5 documentation.* [reapi.ai/docs/grok-imagine-video-1-5-official](/docs/grok-imagine-video-1-5-official)
* reAPI. *Grok Imagine Video 1.5 model page.* [reapi.ai/models/grok-imagine-video-1-5](/models/grok-imagine-video-1-5)
* reAPI. *AI video generation API pricing.* [reapi.ai/blog/ai-video-generation-api-pricing](/blog/ai-video-generation-api-pricing)
---
# Hailuo H3 Explained: MiniMax H3 Specs, Pricing, API, and Limits (2026) (https://reapi.ai/blog/hailuo-h3-minimax-h3-api-guide)
**Hailuo H3 is the search name people are using for MiniMax H3, the video
model MiniMax launched on July 31, 2026. Its API documentation also calls the
release Hailuo 03.** Those names refer to the same model: a multimodal system
that accepts text, images, video, and audio, then generates a 4–15 second video
with native stereo sound at up to 2K.\[1]\[2]
The naming is the easy part. The useful distinction is between the model's
three workflows: prompt-only text-to-video, first/last-frame image-to-video,
and reference-to-video with mixed image, video, and audio inputs. This guide
maps those modes to the exact API rules, separates MiniMax's direct price from
reAPI pricing, and shows the request shape that works on reAPI today.
## Hailuo H3 specs at a glance
| Item | MiniMax H3 / Hailuo 03 |
| -------------------- | ------------------------------------------------- |
| Release date | July 31, 2026 |
| Input | Text, images, video, and audio |
| Output | MP4 video with native stereo audio |
| Resolution | Up to 2K |
| Duration | 4–15 seconds, in whole-second increments |
| Generation modes | Text-to-video, image-to-video, reference-to-video |
| Reference capacity | Up to 9 images, 3 videos, and 3 audio clips |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16 |
| MiniMax direct price | $0.13/s at 2K; $0.09/s at 768p |
| reAPI model ID | `minimax-h3` |
MiniMax describes H3 as a general-purpose omni-modal generation model rather
than only a video renderer. The launch post says it can understand context
across the four input modalities and generate video, audio, images, and text,
although the public video API is the product surface documented in detail at
launch.\[1]
MiniMax also calls H3 an open model. Its launch article says model weights will
be released "in the coming days," subject to applicable laws and regulations.
Until the weights and license are actually published, treat self-hosting as an
announced next step rather than an available deployment path.
## Why Hailuo H3 and Hailuo 03 are the same model
The official release uses two naming layers. **MiniMax H3** is the model name
in MiniMax's launch article and the `model` value in its V2 generation API.
**Hailuo 03** appears in product and API descriptions for the same release.
“Hailuo H3” combines the product-family name with the model version, which is
why it has become a natural search term even though it is not the exact API
identifier.
Use the name that matches the context:
* Search and editorial copy can mention Hailuo H3, MiniMax H3, and Hailuo 03
together once, then settle on MiniMax H3.
* MiniMax's direct V2 endpoint expects `MiniMax-H3`.
* reAPI expects the lowercase model ID `minimax-h3`.
That last character-level difference matters. Model IDs are contract values,
not branding, so copying the display name into code will produce an invalid
request.
## One model routes three different video workflows

MiniMax H3 selects its generation mode from the media attached to the request.
The three routes share the same underlying model, but their validation rules
are not interchangeable.\[3]
### Text-to-video: prompt plus an explicit aspect ratio
Text-to-video is the simplest route. Send a prompt with no media references
and choose one of six aspect ratios. The prompt can describe the visuals,
camera movement, dialogue, music, ambience, and sound effects in the same
brief.
On reAPI, `aspect_ratio` is required for text-to-video and cannot be
`adaptive`. Use `16:9` for landscape, `9:16` for vertical social video, or one
of the four intermediate formats when the delivery surface needs it.
### Image-to-video: first frame, last frame, or both
Image-to-video accepts a first frame, a last frame, or both. Supplying two
frames gives the model a visual start and destination, making it useful for a
planned transition, a product reveal, or a shot that must land on a specific
composition.
The image controls the output orientation, so reAPI rejects `aspect_ratio` in
this mode. Crop the source frames to the intended delivery format before
submitting them. First/last-frame inputs also cannot be mixed with the separate
reference arrays in one request.
### Reference-to-video: assign identity, motion, and sound separately
Reference-to-video is the most flexible route. A request can carry up to nine
images, three video clips, and three audio clips. This makes it possible to
use one image for character or product identity, a video for motion and camera
language, and audio for voice or sound direction.
There are hard limits behind those headline counts:
* each reference video must be 2–15 seconds, with all video references
totaling no more than 15 seconds;
* each reference audio file must be 2–15 seconds, with all audio references
totaling no more than 15 seconds;
* audio cannot be the only reference—it must accompany at least one image or
video;
* reference mode can use an explicit aspect ratio or default to `adaptive`.
Do not send references as an undifferentiated pile. State each asset's job in
the prompt: identity from the portrait, motion from the clip, voice from the
audio, and lighting from the product frame. The API accepts many inputs, but a
clear division of responsibility gives the model fewer conflicting signals.
## Native stereo audio changes how the prompt is written
MiniMax H3 generates stereo audio with the video instead of adding a separate
audio track after the visual pass.\[1] The
prompt therefore needs to direct sound as deliberately as the camera.
A useful H3 prompt has four layers:
1. **Subject and action:** what is visible and what changes during the shot.
2. **Camera and timing:** framing, movement, focus, cuts, and pace.
3. **Spoken content:** exact dialogue and the intended speaker.
4. **Sound bed:** music, ambience, Foley, volume relationships, and silence.
For example:
> A six-second close-up of a brushed-metal espresso machine at sunrise. Slow
> dolly from left to right, warm window reflections, shallow depth of field.
> A soft switch click, rising steam, quiet café ambience, and one restrained
> piano chord. No dialogue.
Native audio removes an initial synchronization step; it does not guarantee a
final mix. Dialogue clarity, lip sync, music continuity, and unwanted effects
still need review on every accepted clip. For production budgeting, the
relevant metric is cost per usable result, not cost per generated second.
## MiniMax H3 pricing: direct API versus reAPI
MiniMax's direct pay-as-you-go table lists **2K output at $0.13 per second**
and **768p at $0.09 per second**.\[4] A
10-second 2K output therefore starts at $1.30 before billable input media.
Direct MiniMax billing adds three rules:
* audio input is free;
* the first five input images are free, then each additional image is $0.04;
* input video is billed by duration at the selected output-resolution rate.
reAPI has its own product contract. At publication, the
[MiniMax H3 model page](/models/minimax-h3) lists a flat **$0.1825 per second**
for 2K output. Reference-video seconds are added to generated-video seconds,
the first five reference images are free, and images six through nine cost
$0.055 each.\[5]\[6]
The reAPI formula is:
```text
cost = $0.1825 × (output seconds + reference-video seconds)
+ $0.055 × max(0, reference images - 5)
```
A 10-second output using a 6-second motion reference bills 16 seconds, or
$2.92. Adding seven reference images contributes two extra-image charges,
bringing that example to $3.03. Audio references add no charge.
These are two different price sheets for two different integrations. Do not
present the $0.13 MiniMax direct rate as the reAPI rate, or the reAPI rate as
MiniMax's official list price. Both can change after launch, so production code
should read the current model page rather than hard-code an article's number
into the UI.
## How to call the MiniMax H3 API on reAPI
reAPI exposes all three H3 modes through its standard asynchronous video
endpoint. A minimal text-to-video request is:
```bash
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer $REAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "minimax-h3",
"prompt": "A rain-soaked night market, slow handheld push forward, vendors calling softly, distant thunder and natural street ambience",
"aspect_ratio": "16:9",
"duration": 6
}'
```
Submission returns a task ID. Poll the shared task endpoint until the result is
complete:
```bash
curl https://reapi.ai/api/v1/tasks/TASK_ID \
-H "Authorization: Bearer $REAPI_API_KEY"
```
The completed response contains `output.video_urls`. Polling does not consume
credits, and failed generations are refunded automatically.\[5]
Change the request shape to select another mode:
```json
{
"model": "minimax-h3",
"prompt": "The character turns naturally and smiles as the camera pushes in",
"first_frame_url": "https://example.com/first-frame.jpg",
"last_frame_url": "https://example.com/last-frame.jpg",
"duration": 6
}
```
For reference-to-video, replace the frame fields with
`reference_image_urls`, `reference_video_urls`, and
`reference_audio_urls`. All media inputs must be public HTTP(S) URLs; base64
and `data:` URLs are rejected.
The full [MiniMax H3 API reference](/docs/minimax-h3) documents the media
formats, file-size limits, task response, errors, and billing behavior.
## Where Hailuo H3 fits—and what still needs testing
H3 is best evaluated as a **shot generator**, not a one-click long-form video
studio. Fifteen seconds can cover a product beat, a dialogue exchange, a social
clip, or one section of a commercial. A 60-second piece still needs multiple
generations, continuity planning, editing, and a final sound pass.
The mixed-reference route is the reason brand, character, and previsualization
teams should test it. The first/last-frame route is the reason motion designers
should care. Native stereo audio is the reason short-form teams may remove one
rough-cut step. None of those capabilities proves that H3 will hold a face,
logo, label, or voice across every brief.
MiniMax calls the model production-ready and publishes polished examples for
brand films, narrative content, social video, advertising, UI/UX, and gaming.
Those are vendor demonstrations, not independent benchmarks. A serious
evaluation should include hands touching products, small printed text, two
speakers in one frame, fast motion, long dialogue, and consistency across
separately generated shots.
The model launched one day before this article, so there is not yet enough
credible public testing for a quality ranking. That limitation is more useful
than an invented verdict: run a fixed prompt set, record accepted outputs per
ten attempts, and compare the total cost of approved shots against the model
already in your workflow.
## Frequently asked questions
### Are Hailuo H3, MiniMax H3, and Hailuo 03 the same model?
Yes. MiniMax H3 is the official launch and API model name; Hailuo 03 is the
product/API description for the same release. Hailuo H3 is the common hybrid
search term. On reAPI, use the model ID `minimax-h3`.
### How long can a MiniMax H3 video be?
Each generation can be 4–15 seconds in whole-second increments. Longer videos
must be split into shots and assembled in an editor.
### Does Hailuo H3 generate audio?
Yes. It generates native stereo audio with the picture and can direct dialogue,
music, ambience, and effects in the same prompt.
### Can MiniMax H3 use an audio reference by itself?
No. Reference audio must be paired with at least one reference image or video.
The audio-reference total cannot exceed 15 seconds.
### How much does MiniMax H3 cost?
MiniMax's direct API lists $0.13/s at 2K and $0.09/s at 768p. reAPI lists a
separate 2K rate of $0.1825/s, with input-video seconds billed at the same rate
and extra-image charges after the first five.
### Is MiniMax H3 open source?
MiniMax announced H3 as an open model, but the July 31 launch post said weights
would be released in the coming days. Check the current weights and license
before planning a self-hosted deployment.
## The practical reading of MiniMax H3
Hailuo H3 is a naming problem wrapped around a useful model design: one
MiniMax H3 endpoint covers prompt-only generation, first/last-frame animation,
and mixed-reference video with native stereo sound. Its 15-second ceiling keeps
it at the shot level, while the reference limits make those shots more
directable than a prompt-only workflow.
Start with the six-second default, assign every reference a clear role, and
judge H3 by accepted clips rather than launch demos. Developers can test the
same request shapes in the [MiniMax H3 playground](/models/minimax-h3) or move
directly to the [API documentation](/docs/minimax-h3).
**Disclosure:** reAPI publishes this article and operates the routed MiniMax H3
endpoint described above. MiniMax specifications and direct prices come from
MiniMax's public launch materials and documentation; reAPI contract details
come from our published model page and API reference.
## References
1. MiniMax. *MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities.* Published July 31, 2026. [minimax.io/blog/minimax-h3](https://www.minimax.io/blog/minimax-h3)
2. MiniMax API. *Video Generation Guide — MiniMax H3 capabilities, duration, resolution, and media inputs.* Retrieved August 1, 2026. [platform.minimax.io/docs/guides/video-generation](https://platform.minimax.io/docs/guides/video-generation)
3. MiniMax API. *Create a Video Generation Task V2 — request modes and validation rules.* Retrieved August 1, 2026. [platform.minimax.io/docs/api-reference/video-generation-v2-create](https://platform.minimax.io/docs/api-reference/video-generation-v2-create)
4. MiniMax API. *Pay as You Go Pricing — MiniMax H3 output and input-media rates.* Retrieved August 1, 2026. [platform.minimax.io/docs/guides/pricing-paygo](https://platform.minimax.io/docs/guides/pricing-paygo)
5. reAPI. *MiniMax H3 API reference — modes, parameters, billing, output, and errors.* Retrieved August 1, 2026. [reapi.ai/docs/minimax-h3](/docs/minimax-h3)
6. reAPI. *MiniMax H3 model page and live pricing.* Retrieved August 1, 2026. [reapi.ai/models/minimax-h3](/models/minimax-h3)
### Further reading
* reAPI. *MiniMax H3 vs Seedance 2.5.* [reapi.ai/blog/minimax-h3-vs-seedance-2-5](/blog/minimax-h3-vs-seedance-2-5)
* reAPI. *AI Video Generation API Pricing: Cost per Second Compared.* [reapi.ai/blog/ai-video-generation-api-pricing](/blog/ai-video-generation-api-pricing)
* reAPI. *Gemini Omni API: Preview Specs, Pricing, and Limits.* [reapi.ai/blog/gemini-omni-api-preview-specs-pricing-2026](/blog/gemini-omni-api-preview-specs-pricing-2026)
* reAPI. *Best Open-Source AI Video Models for Local GPUs.* [reapi.ai/blog/best-open-source-ai-video-models-local-gpu-2026](/blog/best-open-source-ai-video-models-local-gpu-2026)
---
# How Long Can Seedance Videos Be? 15s Now, 30s Next (https://reapi.ai/blog/how-long-can-seedance-videos-be)
Seedance videos can be 4 to 15 seconds long per generation. That is the official range for every Seedance 2.0 tier, set in ByteDance's model documentation, at 24 frames per second with a default of 5 seconds when you do not specify\[1]. The 15-second ceiling is a hard cap, not a plan limit: no platform, subscription, or API tier extends it, which is the answer to the "why does Seedance only make 15-second videos" complaints in every language.
The more useful questions sit around that number: what duration actually costs, how productions get past 15 seconds today, whether Runway's extend feature applies (short answer: no), and what changes when Seedance 2.5's announced 30-second generation ships. All four, sourced, below.
## TL;DR
* **The range is 4–15 seconds** per generation, 24fps, default 5s, on Seedance 2.0, Fast, and Mini alike\[1]. There is also a smart mode (duration −1) that lets the model pick a length to fit the prompt\[1].
* **You pay by the second**, so a 15-second clip costs exactly 3x a 5-second clip at the same tier; on reAPI that is anywhere from $0.45 (Mini, 480p, video ref) to $6.07 (Standard, 1080p, text mode) per 15s clip\[2].
* **Longer stories are chained, not generated**: the API returns the final frame of a clip on request, and feeding it forward as the next clip's first frame is the documented continuity mechanism\[1].
* **Runway does host Seedance 2.0, but its "Extend Video" is a Gen-3 feature**, unavailable for Seedance and retiring in July 2026 anyway\[3]\[4].
* **Seedance 2.5 officially claims 30-second single-pass generation**\[5], with a July release window; a 180-second beta figure circulates unverified.
## How the duration setting actually behaves
Duration on Seedance 2.0 is a request parameter, an integer from 4 to 15 seconds\[1]. Three behaviors worth knowing before you script against it.
First, the default is 5 seconds, and most platform UIs inherit it, which is why "my videos are always 5 seconds" is usually just an untouched setting. Second, ByteDance documents a smart-duration mode, passing −1, where the model chooses a length appropriate to the prompt\[1]; useful for dialogue beats, risky for anything that must cut on a beat grid. Third, when reference videos are involved, the input budget is separate from output duration: up to 3 reference clips totaling 15 seconds feed in regardless of how long the output runs\[1].
Frame budget: at 24fps\[1], 15 seconds is 360 frames, and that is the arena where long-generation quality lives or dies; temporal drift compounds per frame, which is exactly why vendors cap length rather than let quality decay advertise itself.
## What each extra second costs
Per-second billing makes duration a linear cost dial, and the tier matters far more than the length. On reAPI's current rates\[2], a maximum 15-second clip runs:
| Tier | 15-second clip |
| --------------------------- | -------------- |
| Mini, 480p, video reference | $0.45 |
| Fast, 480p, reference mode | $0.60 |
| Fast, 720p, text mode | $2.17 |
| Standard, 720p, text mode | $2.69 |
| Standard, 1080p, text mode | $6.07 |
The practical takeaway for anyone budgeting story content: fifteen seconds of Mini costs less than five seconds of Standard at 1080p. If your cut allows mixed tiers, spend Standard on the hero shots and Mini on connective footage, and the per-minute cost of a chained sequence drops by more than half. The [live pricing table](/models/seedance-2-0#pricing) has all cells current.
## Going past 15 seconds today
There is exactly one sanctioned mechanism and one hosted-platform myth to clear up.
The mechanism is last-frame chaining. The API can return a generation's final frame (`return_last_frame`)\[1]; hand that image to the next request as its first frame, keep your reference set identical, and the seam carries composition and lighting across clips. A 60-second sequence is four chained generations plus a crossfade-free edit. The technique, along with the reference discipline that keeps characters stable across the joins, is the core of our [character consistency guide](/blog/seedance-2-0-character-consistency-guide).
The myth is "Seedance extend on Runway." Runway does serve Seedance 2.0, including the Fast and Mini tiers, from its Standard plan up\[3], but the Extend Video feature people remember belongs to Runway's own Gen-3 models, was never available for third-party models, and is scheduled for retirement at the end of July 2026\[4]. On Runway, continuing a Seedance clip means the video-to-video route, which is a restyle, not an extension. Chaining via last frame remains the honest way to lengthen Seedance footage anywhere.
## The 30-second future, dated July
The 15-second ceiling is about to move, officially. ByteDance's promotional page for Seedance 2.5 headlines 30-second continuous single-pass generation with segment-level prompt control\[5], and Volcano Engine's documentation portal dates the release to July\[6]. A beta tester's claim of 180-second enterprise generations is a single unverified post; treat it accordingly.
Doubling the single-pass length halves the seams in narrative work, and segmented prompts aim to replace the shot-chaining workflow outright. Whether quality holds across 30 seconds is the open question no one outside the beta can answer; everything confirmed so far is collected in our [Seedance 2.5 pre-launch guide](/blog/seedance-2-5-what-we-know-2026).
## FAQ
### How long can a Seedance 2.0 video be?
4 to 15 seconds per generation, at 24fps, across all tiers; the default is 5 seconds\[1]. Longer sequences are made by chaining generations through the returned last frame.
### Why are my Seedance videos only 5 seconds?
That is the default duration parameter. Set it explicitly, up to 15, or use the smart mode (−1) to let the model choose\[1].
### Why does Seedance only create 15-second videos?
It is a model-level cap in ByteDance's documentation\[1], applied identically on every platform. Long single generations degrade temporal consistency, so vendors cap where quality holds; the successor raises the ceiling to a claimed 30 seconds\[5].
### Does a 15-second clip cost more than a 5-second one?
Exactly 3x at the same tier; billing is per output second\[2]. Tier choice moves cost far more than duration: 15 seconds of Mini ($0.45–$0.72) undercuts 5 seconds of 1080p Standard ($2.02)\[2].
### Does Runway have Seedance extend?
No. Runway hosts Seedance 2.0 for generation\[3], but Extend Video is a Gen-3-only feature that is being retired in July 2026\[4]. Lengthening Seedance footage means last-frame chaining, on any platform.
### Will Seedance 2.5 make longer videos?
That is its headline claim: 30-second continuous generation, officially stated on ByteDance's own promo page, with a July release window\[5]\[6].
## Duration is a workflow, not a setting
Fifteen seconds is the canvas, per-second billing is the meter, and last-frame chaining is how Seedance videos become as long as the story needs today. Set duration deliberately, mix tiers to protect the budget, and keep the reference set frozen across the chain. The [Seedance 2.0 API on reAPI](/models/seedance-2-0) exposes the whole loop, duration, references, and last-frame return, behind one endpoint, and when Seedance 2.5 doubles the canvas this month, the same pipeline simply makes fewer joins.
## References
1. Volcano Engine / BytePlus (ByteDance). *Seedance 2.0 — durations, frame rate, smart duration, reference limits, return\_last\_frame.* Retrieved July 2026 from [docs.byteplus.com/en/docs/ModelArk/1520757](https://docs.byteplus.com/en/docs/ModelArk/1520757)
2. reAPI. *Seedance 2.0 — model page and live pricing.* Retrieved July 2026 from [reapi.ai/models/seedance-2-0](/models/seedance-2-0)
3. Runway. *Third-party models on Runway — Seedance 2.0 availability.* Retrieved July 2026 from [help.runwayml.com/hc/en-us/articles/50488490233363](https://help.runwayml.com/hc/en-us/articles/50488490233363)
4. Runway. *Extend Video (Gen-3) — feature scope and retirement notice.* Retrieved July 2026 from [help.runwayml.com/hc/en-us/articles/30266515017875](https://help.runwayml.com/hc/en-us/articles/30266515017875)
5. Volcano Engine (ByteDance). *Doubao Seedance 2.5 — official promotional page (30-second generation).* Retrieved July 2026 from [ark.volcengine.com/promotion?modelName=seedance-2-5](https://ark.volcengine.com/promotion?modelName=seedance-2-5)
6. Volcano Engine (ByteDance). *Documentation portal — Seedance 2.5 July release notice.* Retrieved July 2026 from [volcengine.com/docs/search?q=Seedance 2.5](https://www.volcengine.com/docs/search?q=Seedance%202.5)
### Further reading
* reAPI. *Seedance 2.0 Character Consistency: References, Voice, Shots.* [reapi.ai/blog/seedance-2-0-character-consistency-guide](/blog/seedance-2-0-character-consistency-guide)
* reAPI. *Seedance 2.5: What We Know Before the Public Launch.* [reapi.ai/blog/seedance-2-5-what-we-know-2026](/blog/seedance-2-5-what-we-know-2026)
* reAPI. *Cheapest Seedance 2.0 in 2026: Real Prices, Compared.* [reapi.ai/blog/cheapest-seedance-2-0-2026](/blog/cheapest-seedance-2-0-2026)
---
# How to Get a Claude API Key, and When You Need One (https://reapi.ai/blog/how-to-get-claude-api-key)
Getting a Claude API key takes about a minute. Two settings you choose while creating it matter more than the key itself, and most people skip past both.
There is also a question worth asking before you start: whether you need an Anthropic key at all, or whether the thing you are building calls more than one vendor.
## TL;DR
* **Keys are created in the Console** at `platform.claude.com`, under Account Settings\[1].
* **You choose an expiration when you create the key.** That choice is made once, at creation\[1].
* **Workspaces segment keys and control spend** per use case\[1].
* **Auth is the `x-api-key` header**, or `Authorization`. SDKs read `ANTHROPIC_API_KEY` from the environment\[1]\[2].
* **The Workbench lets you try the API in a browser** before writing any code\[1].
* **If you are calling Claude alongside GPT or Gemini**, one gateway key covers all three instead of three vendor accounts.
## Creating the key
The API is made available through the web Console\[1].
1. Sign in at **`platform.claude.com`**.
2. Try the API in the browser first with the **Workbench** at `platform.claude.com/playground`, which needs no code.
3. Generate the key in **Account Settings**, at `platform.claude.com/settings/keys`.
4. **Pick an expiration** while creating it.
5. Optionally place it in a **workspace** at `platform.claude.com/settings/workspaces`.
Steps 4 and 5 are the ones worth slowing down for.
## The two settings people skip
**Expiration is chosen at creation.** You set each key's lifetime when you make it\[1]. A key with a long life is a key you will still be finding in an old `.env` next year. For anything that is not long-lived production infrastructure, a short expiry is the cheapest security control available.
**Workspaces segment keys and control spend by use case**\[1]. This is the feature that turns "someone's script ran overnight" from a billing surprise into a bounded one. One workspace per project, or per environment, means a runaway loop burns that workspace's budget rather than the account's.
Both are decisions you make once and live with. Neither is retrofittable to a key already in production without rotating it.
## Using the key
Two headers are accepted. Send **one** of them\[1]:
| Header | Value |
| --------------- | ------------------------- |
| `x-api-key` | Your API key from Console |
| `Authorization` | Bearer-style alternative |
The SDKs read the key from the environment automatically\[2]:
```bash
export ANTHROPIC_API_KEY="sk-ant-..."
```
```python
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
resp = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Summarize this changelog."}],
)
```
That is the whole integration for a single-vendor setup.
## When you do not need an Anthropic key

The Console flow above is the right answer when Claude is the only model you call.
It stops being the right answer the moment your product calls two vendors. Then you are maintaining an Anthropic account, an OpenAI account, and a Google Cloud project, three billing relationships, three sets of rate limits, and three key-rotation schedules, to run one feature. Adding a fourth model means a fourth of everything.
A gateway collapses that. reAPI exposes Claude through an OpenAI-compatible `/v1/chat/completions` endpoint, so the same client and the same key reach Claude, GPT, and Gemini, and switching between them is a model-string change:
```python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_REAPI_KEY",
base_url="https://api.reapi.ai/v1",
)
resp = client.chat.completions.create(
model="claude-opus-5",
messages=[{"role": "user", "content": "Summarize this changelog."}],
max_tokens=16000,
)
```
The native Anthropic `/v1/messages` surface is available too, so an SDK already written against Anthropic's format works after changing the base URL.
Rates are on [reapi.ai/models](/models), and the current Opus reference is at [reapi.ai/docs/claude-opus-5](/docs/claude-opus-5).
## One thing to get right on reAPI
There are **two separate credentials**, and mixing them up is the most common first-call failure.
| Key | Where you create it | What it calls |
| ----------- | ---------------------- | -------------------------------- |
| Gateway key | `api.reapi.ai` console | Chat models: Claude, GPT, Gemini |
| Media key | `reapi.ai` account | Image and video generation |
A gateway key sent to a media endpoint fails, and the reverse fails too. If a first call returns an auth error and the key looks correct, check which of the two you are holding.
## FAQ
### Where do I get a Claude API key?
In the Console at `platform.claude.com`, under Account Settings at `platform.claude.com/settings/keys`\[1].
### Can I try the API without writing code?
Yes. The Workbench at `platform.claude.com/playground` runs requests in the browser\[1].
### How do I authenticate a request?
Send your key in the `x-api-key` header, or use `Authorization`. Send one, not both\[1].
### What environment variable do the SDKs read?
`ANTHROPIC_API_KEY`\[2].
### Can I change a key's expiration later?
Expiration is chosen when the key is created\[1]. Changing it means creating a new key and rotating.
### How do I stop one project from spending the whole budget?
Use workspaces, which segment keys and control spend by use case\[1].
### Do I need an Anthropic key to use Claude through a gateway?
No. A gateway key replaces the per-vendor account when you are calling several models, and the endpoint is OpenAI-compatible.
### Why does my reAPI key return an auth error?
Most likely you are using the gateway key on a media endpoint or the media key on the gateway. They are separate credentials created in separate places.
## Deciding before you create
The mechanical part of how to get a Claude API key is four clicks in the Console. The part worth thinking about is the two settings attached to it, expiration and workspace, because both are chosen once and neither can be changed afterward without rotating the key.
And the question underneath all of it is how many vendors your product will end up calling. One vendor, one key, and the Console flow is exactly right. Three vendors and the arithmetic changes: three accounts, three invoices, three rate-limit regimes, and a fourth model means starting over. That is the point where a single gateway key stops being a convenience and starts being the simpler architecture.
## References
1. Anthropic. *API overview — Console, Workbench, key creation, expiration, workspaces, and authentication headers.* Retrieved July 2026 from [docs.claude.com/en/api/overview](https://docs.claude.com/en/api/overview)
2. Anthropic. *Get started — setting the API key and first request.* Retrieved July 2026 from [docs.claude.com/en/docs/get-started](https://docs.claude.com/en/docs/get-started)
### Further reading
* reAPI. *How to use Claude Opus 5.* [reapi.ai/blog/how-to-use-claude-opus-5](/blog/how-to-use-claude-opus-5)
* reAPI. *How to use Claude Code.* [reapi.ai/blog/how-to-use-claude-code](/blog/how-to-use-claude-code)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# How to Use Claude Code: Anthropic's Terminal Coding Agent (https://reapi.ai/blog/how-to-use-claude-code)
Learning how to use Claude Code starts with what it actually is: Anthropic's agentic coding tool, one that runs in your terminal, reads your entire codebase, and executes development tasks from natural-language instructions\[1]. It has accumulated 101,000 GitHub stars and 15,500 forks since general availability, which makes it one of the most widely adopted AI coding tools of 2026\[2].
The structural difference from earlier tools is worth stating plainly, because it decides whether the workflow suits you. Copilot-style autocomplete suggests the next line based on what you already wrote. Claude Code reads the whole project, plans an approach, edits across many files at once, runs the tests, handles the failures, iterates, and commits. You do not tell it which files matter.
This guide covers what it does, where it runs, how to install it, the five jobs it is disproportionately good at, how it compares to Cursor and Copilot, and what to do when you want that behavior inside your own product rather than in a terminal.
## TL;DR
* **1M token context window**, roughly 25,000 to 30,000 lines of code in one session without fragmentation\[1].
* **80.8% on SWE-bench Verified**, the highest publicly reported score among available tools at the time of writing\[1].
* **Plan Mode** shows you the file-by-file plan before it touches anything, which is what makes autonomous operation acceptable on a production repo\[1].
* **Runs almost everywhere**: terminal CLI, VS Code, Cursor, Windsurf, JetBrains, a desktop app, the web, and `@claude` on a GitHub issue\[1].
* **Access requires a Claude Pro or Max subscription, or a Claude Console account with API access**\[1].
* **MCP support** connects it to databases, internal APIs, ticketing, and observability, so it reasons about more than the repo\[1].
## What Claude Code is

Claude Code is a coding agent rather than a completion engine. Point it at a project and it navigates the structure the way a senior developer would: start at the entry points, follow the dependency graph, build a working model of how the components interact, then make changes that are architecturally coherent instead of locally correct and globally broken.
That distinction matters most on the tasks that actually consume engineering time. Debugging a production issue that behaves differently in staging. Implementing a feature that needs an ORM change, three endpoint updates, and a migration. Those are multi-file problems, and multi-file problems are where a full-context agent separates from autocomplete.
## The capabilities that matter
### A 1M token context window
Claude Code reads up to a million tokens in one context window, roughly 25,000 to 30,000 lines of code in a single session without chunking or retrieval augmentation\[1]. When a refactor touches 40 files across a 200,000-line codebase, you do not hand it a file list. It finds them.
### 80.8% on SWE-bench Verified
SWE-bench Verified evaluates a model against real GitHub issues from open-source repositories, scored by whether the proposed fix passes the repository's own test suite. Claude Code reaches 80.8%, the highest publicly reported figure among available tools\[1]. The benchmark measures genuine codebase understanding rather than pattern-matching on isolated snippets, which is why it tracks real usefulness better than most coding evals.
### Plan Mode
Before executing anything, Claude Code presents a structured plan: which files it will modify, what changes it will make, in what order. You review it, question it, adjust it, and only then approve\[1].
This is the feature that makes the rest of it usable on a production repo. Supervised autonomy keeps the developer in the loop at the one moment where being out of the loop is expensive.
### Agent Teams
Multiple Claude Code instances work on different parts of a problem simultaneously, coordinated by a lead agent that assigns subtasks and merges results\[1]. For a feature that decomposes into independent workstreams, a backend change, a frontend update, and a docs update running in parallel, this compresses calendar time rather than just token time.
## Where Claude Code runs

Broad adoption has a boring explanation: it meets developers where they already are\[1].
* **Terminal CLI** is the core experience. Navigate to the project, run `claude`, describe the task.
* **VS Code extension**, with visual diff review and interactive change selection. The same extension works in Cursor and Windsurf.
* **JetBrains plugin** covers IntelliJ IDEA, PyCharm, WebStorm, and the rest.
* **Desktop app** adds visual diff review across parallel sessions, session scheduling, and cloud execution for long-running jobs.
* **Web** at claude.ai/code, no local install, plus the Claude iOS app.
* **GitHub**: tag `@claude` on an issue and get a pull request back, without asking teammates to change how they work.

## How to use Claude Code: installation and first run
macOS and Linux:
```bash
curl -fsSL https://claude.ai/install.sh | bash
```
Or with Homebrew:
```bash
brew install --cask claude-code
```
Windows:
```powershell
irm https://claude.ai/install.ps1 | iex
```
Or with WinGet:
```powershell
winget install Anthropic.ClaudeCode
```
Then navigate to your project and run `claude`. It indexes the project structure and asks what you want to do.
One step matters more than the install method. Create a `CLAUDE.md` at the repository root and put your conventions in it: which test framework you use, which files must never be modified automatically, how you name things\[1]. It is read at the start of every session. Teams use it to encode the institutional knowledge a senior developer would otherwise pass on during onboarding, and it is the single highest-leverage configuration in the tool.
## Five jobs it is disproportionately good at
**Large-codebase refactoring.** When a refactor touches dozens of files and needs consistent changes to signatures, imports, and tests at once, coordination is the hard part. Claude Code builds a dependency map, finds every call site, makes coordinated changes, and runs the suite before showing you the result. Rakuten's team documented it working autonomously for seven hours on an activation-vector extraction method inside a large multi-language library, producing 99.9% numerical accuracy against the reference implementation\[1].
**Cross-file debugging.** When a bug surfaces in user behavior but originates in the interaction between three modules written by different people at different times, full-context navigation lets it follow the execution path, hypothesize failure modes, write diagnostics, and fix every place the broken assumption is made rather than the most obvious one.
**Test generation for untested code.** It reads the implementation, infers intended behavior from the code and whatever docs exist, and generates tests covering normal paths, edge cases, and error conditions. It tends to find bugs while doing it.
**Git workflow automation.** Describe the commit in plain language, including what to leave out, and it handles staging, the commit message, the push, and a pull request with a structured description. With the GitHub MCP integration it can read open issues, implement fixes, run tests, and open the PR with full context.
**Architecture analysis.** Point it at an unfamiliar codebase and ask it to explain the architecture. You get a structured walkthrough of the organization, each major component, the data flow between services, and the critical paths. Useful for onboarding, technical due diligence, and writing the documentation nobody ever wrote.
## MCP integration
Claude Code supports the Model Context Protocol, Anthropic's open standard for connecting models to external tools and data\[1]. That extends its reasoning past the repository: query the production database to understand the data model before writing a migration, read the ticket to understand acceptance criteria before implementing, check the observability platform for error patterns before debugging.
The plugins directory carries a growing library of community integrations, so most common tools are already covered.
## Claude Code, Cursor, or Copilot
| | Claude Code | Cursor | GitHub Copilot |
| ------------------ | --------------------------------- | ----------------------- | --------------------------- |
| Interface | Terminal CLI + IDE extensions | Custom VS Code fork | IDE extension |
| Context window | 1M tokens | \~128K tokens | \~128K tokens |
| SWE-bench score | 80.8% | Not reported | Not reported |
| Agentic operation | Full: plan, execute, test, commit | Partial (composer mode) | Partial (Copilot Workspace) |
| Multi-file editing | Native, coordinated | Yes | Limited |
| Pricing | Claude Pro/Max, $20 to $100/mo | $20/mo | $10 to $39/mo |
Source for the comparison rows is the same guide referenced throughout\[1].
The honest read is that these are not really competing for the same minutes. Cursor is a VS Code fork with AI integrated throughout the editor: fast inline completion, multi-model support, a visual IDE. Copilot is the most widely deployed because it ships inside enterprise Microsoft subscriptions, and it handles routine completion well without operating as a full agent on complex multi-file work.
Most developers who use both use Cursor for daily editing and Claude Code for the tasks that need deep codebase understanding: large refactors, architecture changes, security audits, subtle cross-file bugs.
## When you want this behavior inside your own product
Claude Code is Anthropic's agent, and it is the right tool when the work happens in your terminal. It is the wrong shape for a different job: putting the same behavior inside software you ship. A code-review bot in your CI, a migration agent for your customers, an internal tool that reads a repo and answers questions about it. None of those are a CLI on a developer's laptop.
That job needs the models directly. Access to Claude Code itself runs through a Claude Pro or Max subscription, or a Claude Console account with API access\[1], and when you build your own agent you are working at that second layer.
reAPI exposes the Claude models through an OpenAI-compatible `/v1/chat/completions` endpoint, so one client and one key cover Claude alongside GPT and Gemini:
```python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_REAPI_KEY",
base_url="https://api.reapi.ai/v1",
)
resp = client.chat.completions.create(
model="claude-opus-5",
messages=[{"role": "user", "content": "Review this diff and report issues with severity."}],
max_tokens=16000,
stream=True,
)
```
The native Anthropic `/v1/messages` surface is available too, so SDKs written for either format work unchanged. Model pages and rates are at [reapi.ai/models](/models), and the endpoint reference for the current Opus generation is at [reapi.ai/docs/claude-opus-5](/docs/claude-opus-5).
Use Claude Code for your own repo. Use the API when the agent has to run in someone else's.
## FAQ
### What model powers Claude Code?
It runs on Anthropic's current frontier Claude models, and the default tracks whatever Anthropic has most recently shipped rather than staying pinned to one version. Check Anthropic's documentation for the model behind your plan at any given time.
### Is Claude Code free?
No. It requires a Claude Pro or Max subscription, priced between $20 and $100 per month, or a Claude Console account with API access\[1].
### How big a codebase can Claude Code handle?
The context window is 1M tokens, roughly 25,000 to 30,000 lines in a single session without fragmentation\[1]. Larger repositories still work, because it navigates the dependency graph rather than loading everything at once.
### What is `CLAUDE.md` for?
Persistent per-project instructions read at the start of every session: coding conventions, test framework, files that must never be modified automatically. It is where teams encode onboarding knowledge\[1].
### Is Claude Code better than Cursor?
They solve different problems. Cursor is a full IDE experience with fast inline editing; Claude Code is a terminal agent with a much larger context window and full plan-execute-test-commit operation. Many developers run both\[1].
### Can Claude Code open pull requests?
Yes. It integrates with git directly, and tagging `@claude` on a GitHub issue returns a pull request. With the GitHub MCP integration it can read the issue, implement, test, and submit with full context\[1].
### What is Plan Mode?
A review gate. Claude Code presents the file-by-file plan before making changes, and waits for approval\[1].
### How do I build a coding agent of my own?
Call the models through an API rather than the CLI. Point an OpenAI-compatible client at `https://api.reapi.ai/v1` and set the model string, or use the native Anthropic `/v1/messages` surface.
## Picking the right layer
Claude Code earns its adoption on one property: it holds an entire project in view and acts on it, rather than guessing at the next line. The million-token window, the 80.8% SWE-bench score, and Plan Mode are three expressions of the same design decision, and they land hardest on refactors, cross-file debugging, and unfamiliar codebases.
Know which layer your problem lives at. If the work is in your repo and your terminal, install it and write a good `CLAUDE.md` before anything else. If the work is inside a product you ship to other people, you want the model API underneath, not the CLI on top of it. Learning how to use Claude Code well mostly means recognizing which of those two you are actually doing.
## References
1. Anthropic. *Claude Code — overview, capabilities, surfaces, and configuration.* Retrieved July 2026 from [platform.claude.com/docs/claude-code](https://platform.claude.com/docs/claude-code)
2. Anthropic. *Claude Code — official repository.* Retrieved July 2026 from [github.com/anthropics/claude-code](https://github.com/anthropics/claude-code)
3. Anthropic. *Claude Code documentation.* Retrieved July 2026 from [docs.claude.com/en/docs/claude-code/overview](https://docs.claude.com/en/docs/claude-code/overview)
### Further reading
* reAPI. *How to use Claude Opus 5.* [reapi.ai/blog/how-to-use-claude-opus-5](/blog/how-to-use-claude-opus-5)
* reAPI. *How to use Claude Fable 5.* [reapi.ai/blog/how-to-use-claude-fable-5](/blog/how-to-use-claude-fable-5)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# How to Use Claude Fable 5: Refusals, Fallback, and Cost (https://reapi.ai/blog/how-to-use-claude-fable-5)
Anthropic released Claude Fable 5 on June 9, 2026, the first publicly available model from its Mythos family, at $10 per million input tokens and $50 per million output\[1]. It is the most capable model Anthropic has shipped to the general public, and the steepest list price in the Claude lineup.
Learning how to use Claude Fable 5 means learning one thing the other Claude models do not make you handle: it can decline a request. Fable 5 ships with safety classifiers that Mythos 5 does not have, and a declined request comes back as a successful HTTP 200 with `stop_reason: "refusal"` rather than an error\[1]. Your integration needs three new behaviors for that: refusal handling, a fallback path to another model, and an understanding of how refusals are billed.
This guide covers what Fable 5 actually does, how it differs from its restricted sibling Mythos 5, the refusal-and-fallback contract in detail, the export-control suspension that took it offline for eighteen days, and why the arrival of Claude Opus 5 changed the answer to "should I be paying for this?"
## TL;DR
* **List price is $10 / $50 per million tokens**, twice Claude Opus 4.8\[3]. On reAPI it runs $8.00 / $40.00, which is 80% of Anthropic's published rate.
* **It can refuse.** Fable 5 carries safety classifiers covering cybersecurity, biology and chemistry, and distillation attempts. A refusal is an HTTP 200 with `stop_reason: "refusal"`, and the response reports which classifier declined\[1].
* **Refusals are not billed** when nothing was generated, and fallback credit refunds the prompt-cache cost of retrying on another model\[1].
* **Thinking cannot be disabled.** Adaptive thinking is always on; `thinking: {"type": "disabled"}` is unsupported. Use `effort` to control depth\[1].
* **30-day data retention is mandatory.** Fable 5 and Mythos 5 are Covered Models and are not available under zero data retention\[1].
* **Claude Opus 5 changed the math.** Anthropic describes Opus 5 as delivering frontier intelligence at half the cost of Fable 5\[4].
## What Claude Fable 5 is

Claude Fable 5 is Anthropic's most capable widely released model, built for the most demanding reasoning and long-horizon agentic work\[1]. The API identifier is `claude-fable-5`.
The launch is two models, and conflating them is the easiest mistake to make.
| | Claude Fable 5 | Claude Mythos 5 |
| -------------------- | ------------------- | ---------------------------------------------- |
| API model ID | `claude-fable-5` | `claude-mythos-5` |
| Availability | Generally available | Limited release through Project Glasswing only |
| Safety classifiers | Yes | No |
| Context / max output | 1M tokens / 128k | Same |
| List price | $10 / $50 per MTok | Same |
They share the same underlying capabilities and the same specs\[1]. The difference is the classifier layer. So when a benchmark number is attributed to "Mythos 5 / Fable 5" on a cybersecurity or biology row, read it carefully: that ceiling belongs to the model without the brakes. The public Fable 5 you can actually call will decline on exactly those questions by design.
## Benchmarks
Anthropic published a head-to-head against Mythos Preview, Opus 4.8, GPT-5.5, and Gemini 3.1 Pro. The standout is agentic coding, where the lead is large rather than marginal\[1].

| Benchmark | Mythos 5 / Fable 5 | Mythos Preview | Opus 4.8 | GPT-5.5 | Gemini 3.1 Pro |
| -------------------------------------------- | ------------------ | -------------- | -------- | ------- | -------------- |
| Agentic coding (SWE-Bench Pro) | **80.3%** | 77.8% | 69.2% | 58.6% | 54.2% |
| Agentic coding (FrontierCode Diamond, xhigh) | **29.3%** | n/a | 13.4% | 5.7% | n/a |
| Agentic coding (Terminal-Bench 2.1) | **88.0%**\* | n/a | 82.7% | 83.4%† | 70.7%† |
| Knowledge work (GDPval-AA) | **1932** | n/a | 1890 | 1769 | 1314 |
| Knowledge work vision (GDP.pdf, no tools) | **29.8%** | n/a | 22.5% | 24.9% | 16.7% |
| Spatial reasoning (Blueprint-Bench 2) | **38.6%** | n/a | 14.5% | 36.2% | 26.5% |
| Tool use (AutomationBench) | **17.4%** | n/a | 15.5% | 12.9% | 9.6% |
| Computer use (OSWorld-Verified) | 85.0% | **85.4%** | 83.4% | 78.7% | 76.2% |
| Legal (Legal Agent Benchmark) | **13.3%** | n/a | 10.4% | 2.1% | 0.0% |
| Reasoning (HLE, no tools) | **59.0%**\* | 56.8% | 49.8% | 41.4% | 44.4% |
| Reasoning (HLE, with tools) | 64.5%\* | **64.7%** | 57.9% | 52.2% | 51.4% |
| Biology (BioMysteryBench, hard) | **46.1%**\* | 29.6% | 40.0% | n/a | n/a |
| Biology (BioMysteryBench, human-solved) | **83.9%**\* | 82.6% | 80.4% | n/a | n/a |
| Cybersecurity (ExploitBench, Cap%) | **78.0%**\* | 69.0% | 40.0% | 34.0% | n/a |
| Health (HealthBench Professional) | **66.0%**\* | 64.7% | 56.9% | 51.8% | n/a |
† Terminal-Bench figures for GPT-5.5 and Gemini 3.1 Pro were run through Codex CLI and Gemini CLI respectively.
Anthropic's own methodology note under that table is the most important thing on it: reported scores land within a one to three percentage point difference for Mythos 5 and Fable 5, and the table shows the higher of the two. Starred rows show a larger gap "due to our blocking safeguards for cybersecurity and biology-related questions," and on those "Claude Fable 5 performs closer to Claude Opus 4.8 due to fallbacks"\[1].
Read that carefully before quoting any starred number. The 78.0% cybersecurity figure and the 66.0% health figure are not what the public model delivers on those questions.
The unstarred rows are the gains the public model keeps. SWE-Bench Pro at 80.3% is more than ten points above Opus 4.8 and over twenty above GPT-5.5 and Gemini 3.1 Pro.
FrontierCode is the row that explains the price.

On that hardest subset, Fable 5 climbs from roughly 11% at low effort to about 31% at max, while Opus 4.8 tops out near 13% at xhigh and comes back down at max, and GPT-5.5 stays flat around 5% to 6% across its whole ladder. Fable 5 converts extra compute into real accuracy on hard problems in a way the others do not. You are paying for a model designed to keep working a problem longer.
Partner results from the launch are specific enough to be falsifiable: Stripe reported completing a 50-million-line Ruby codebase migration in roughly one day against a two-month manual estimate; Hex said Fable 5 was the first model to clear 90% on its core analytics benchmark; Hebbia recorded its highest finance-benchmark score. On vision, the model completed Pokémon FireRed using vision alone.
## How to use Claude Fable 5 when it refuses
This is the part of Fable 5 that changes your code, and it is worth being precise because it is easy to describe wrongly.
Fable 5 does not silently reroute a risky request to a smaller model. When a classifier declines, the Messages API returns `stop_reason: "refusal"` as a **successful HTTP 200 response, not an error**, and the response reports which classifier declined it\[1]. Code that indexes the first content block unconditionally will break on this. Check the stop reason before reading content.
Retrying on another Claude model is your call, and there are three supported paths\[1]:
* **Server-side.** Pass the `fallbacks` parameter and let the API retry for you, either in its `"default"` mode using Anthropic's recommended models or with a list you name. In beta on the Claude API.
* **Client-side.** Use the SDK middleware to retry from the client on any platform.
* **Manual.** Build the retry yourself, in any language.
The billing rules matter as much as the mechanics. You are **not billed for a request refused before any output is generated**, and when you retry on another model, fallback credit refunds the prompt-cache cost of switching so you do not pay it twice\[1].
Three domains are fenced: cybersecurity, including exploitation and offensive work; biology and chemistry; and distillation attempts aimed at extracting Fable 5's capabilities to train another model.
## What the export-control suspension revealed
Fable 5 has already been taken offline once, and the reason is worth knowing before you build a product on it.
On Friday, June 12, three days after launch, the US government applied export controls to Fable 5 and Mythos 5, requiring Anthropic to restrict access to foreign nationals whether inside or outside the United States. Because the order took effect immediately and Anthropic had no reliable way to verify nationality in real time, it **suspended access to both models for all users**\[2].
The controls were lifted on June 30, and Fable 5 returned globally on July 1 across the Claude Platform, Claude.ai, Claude Code, and Claude Cowork, with AWS, Google Cloud, and Microsoft Foundry re-enabled afterward\[2]. Mythos 5 access was restored to a set of US organizations following government approval on June 26\[2].
Eighteen days of total unavailability is a real operational fact, not a footnote. If Fable 5 is the only model on a production path, that path had no service for eighteen days. The practical lesson is the same one the refusal contract already teaches: this model needs a fallback, and the fallback should be configured before you need it.
## The retention requirement
Fable 5 and Mythos 5 carry mandatory 30-day data retention and are **not available under zero data retention**, because both are designated Covered Models\[1]. Organizations that previously held zero-retention agreements do not keep them for this model.
If you have data-residency or retention commitments to your own customers, this is the clause to read before adopting Fable 5. It is not negotiable per-account, and no amount of gateway configuration changes it.
## Pricing, and what Opus 5 changed
Fable 5 lists at $10 per million input tokens and $50 per million output, twice Opus 4.8's $5 / $25\[3]. That was the calculus in June.
It is not the calculus now. When Anthropic shipped Claude Opus 5 on July 24, it described the model as delivering frontier intelligence **at half the cost of Claude Fable 5**\[4], and Opus 5 lists at $5 / $25, unchanged from Opus 4.8\[3]. Opus 5 also has no classifier-refusal layer of Fable 5's kind, no mandatory 30-day retention, and a full effort ladder.
On reAPI the spread is wider still:
| Model | reAPI rate per 1M tokens |
| --------------- | --------------------------- |
| Claude Fable 5 | $8.00 input / $40.00 output |
| Claude Opus 5 | $2.40 input / $12.00 output |
| Claude Opus 4.8 | $4.00 input / $20.00 output |
Fable 5 costs more than three times what Opus 5 costs on the same gateway. That does not make Fable 5 pointless. It does mean the burden of proof moved: you now need a specific reason to reach for it, measured on your own tasks.
## Fable 5, Opus 5, or Opus 4.8?
**Reach for Fable 5** when the task is long-horizon and genuinely hard, and when you have measured that it wins: repo-scale migrations, multi-file refactors against a trustworthy test suite, overnight agent loops, dense analytical work where being wrong is expensive. The FrontierCode curve is the evidence, since accuracy keeps climbing with effort where other models flatten.
**Start with Opus 5** for most new work. Anthropic's own framing puts it close to Fable 5's frontier intelligence at half the price, and it does not carry the refusal layer or the retention requirement.
**Stay on Opus 4.8** for the broad middle where your evals already characterize its behavior, and where its cheaper fast mode matters.
One operational note that is easy to miss: a security or life-sciences team that needs full capability is **worse off** paying Fable 5 rates, because those are exactly the domains where the classifiers decline. Calling Opus 4.8 or Opus 5 directly, or applying to Project Glasswing for Mythos 5, is the honest path.
## Autonomy and file-based memory
The capability Anthropic stresses most is duration. Fable 5 and Mythos 5 can work autonomously for longer than any previous Claude models, and the supporting examples are deliberately long-horizon: with file-based memory, Mythos 5's performance on Slay the Spire improved three times more than Opus 4.8's, and in life-sciences work it ran autonomously over a week to assemble single-cell data for millions of cells across 138 species.
Supported features at launch include effort, task budgets in beta, the memory tool, code execution, programmatic tool calling, tool-result clearing through context editing, compaction, and vision\[1].
Two request-surface rules follow from the always-on thinking. `thinking: {"type": "disabled"}` is unsupported, so `effort` is the only lever for thinking depth. And the raw chain of thought is never returned: `thinking.display` defaults to `"omitted"`, with `"summarized"` returning a readable summary instead\[1].
## Calling Claude Fable 5 alongside your other models
Read the last four sections together and one requirement falls out: **Fable 5 needs a second model wired up next to it**, not as an optimization but as a correctness requirement. It refuses on three domains by design. It was fully unavailable for eighteen days. And a model at a third of its price now covers most of what people bought it for.
reAPI exposes Claude Fable 5 through an OpenAI-compatible `/v1/chat/completions` endpoint, so the same client that calls Opus 5, GPT, and Gemini calls it too:
```python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_REAPI_KEY",
base_url="https://api.reapi.ai/v1",
)
resp = client.chat.completions.create(
model="claude-fable-5",
messages=[{"role": "user", "content": "Plan and execute this multi-file refactor."}],
max_tokens=16000,
stream=True,
)
```
Rates on the gateway are $8.00 per million input tokens and $40.00 per million output, which is 80% of Anthropic's published $10 / $50. The gateway also exposes the native Anthropic `/v1/messages` surface, so SDKs written for that format work unchanged. Details are on [reapi.ai/docs/claude-fable-5](/docs/claude-fable-5) and [reapi.ai/models/claude-fable-5](/models/claude-fable-5).
Switching the model string when a refusal comes back is a three-line change when both models sit behind one key and one client. It is a project when they do not.
## FAQ
### What happens when Claude Fable 5 refuses a request?
The Messages API returns a successful HTTP 200 with `stop_reason: "refusal"`, and reports which classifier declined it. It is not an error response, so check the stop reason before reading the content\[1].
### Am I billed for a refused request?
No, not when the request is refused before any output is generated. If you retry on another model, fallback credit refunds the prompt-cache cost of the switch\[1].
### What is the difference between Claude Fable 5 and Claude Mythos 5?
Same capabilities, same specs, same price. Mythos 5 ships without the safety classifiers and is available only through Project Glasswing to approved customers. Fable 5 is the generally available version\[1].
### Can I turn thinking off on Claude Fable 5?
No. Adaptive thinking is always on and `thinking: {"type": "disabled"}` is unsupported. Control depth with the `effort` parameter instead\[1].
### Does Claude Fable 5 support zero data retention?
No. Fable 5 and Mythos 5 are designated Covered Models, carry 30-day retention, and are not available under zero data retention\[1].
### Why was Claude Fable 5 unavailable in June?
The US government applied export controls on June 12, 2026, requiring restricted access for foreign nationals. With no way to verify nationality in real time, Anthropic suspended access for all users. The controls were lifted June 30 and Fable 5 returned July 1\[2].
### Should I use Claude Fable 5 or Claude Opus 5?
Start with Opus 5 unless you have measured Fable 5 winning on your own tasks. Anthropic describes Opus 5 as delivering frontier intelligence at half Fable 5's cost, without the refusal layer or the retention requirement\[4].
### How do I call Claude Fable 5 with the OpenAI SDK?
Point the OpenAI client at `https://api.reapi.ai/v1` and set `model` to `claude-fable-5`. The endpoint is OpenAI-compatible, and the native Anthropic `/v1/messages` surface is available too.
## Buying the top of the lineup on purpose
Fable 5 is a genuinely different class of model on the row that matters most for agentic work, and the FrontierCode curve is the honest argument for it: on hard problems, extra effort keeps buying accuracy instead of flattening out. If your workload is a repo-scale migration or an overnight loop where a failed run costs more than the tokens, that curve is worth paying for.
Everything else about it is a constraint you have to design around. It declines three domains by contract, it carries mandatory 30-day retention with no zero-retention option, it went fully dark for eighteen days under export controls, and a model at a third of the price now sits close behind it. None of that is disqualifying, and all of it argues the same thing: know how to use Claude Fable 5 with a fallback already wired up, and reach for it deliberately rather than by default.
## References
1. Anthropic. *Introducing Claude Fable 5 and Claude Mythos 5 — capabilities, refusals, fallback, billing, retention, and supported features.* Retrieved July 2026 from [platform.claude.com/docs/en/about-claude/models/introducing-claude-fable-5-and-claude-mythos-5](https://platform.claude.com/docs/en/about-claude/models/introducing-claude-fable-5-and-claude-mythos-5)
2. Anthropic. *Redeploying Fable 5 — export-control timeline and restoration.* Retrieved July 2026 from [anthropic.com/news/redeploying-fable-5](https://www.anthropic.com/news/redeploying-fable-5)
3. Anthropic. *Pricing — token, cache, and batch rates.* Retrieved July 2026 from [platform.claude.com/docs/en/about-claude/pricing](https://platform.claude.com/docs/en/about-claude/pricing)
4. Anthropic. *What's new in Claude Opus 5.* Retrieved July 2026 from [platform.claude.com/docs/en/about-claude/models/whats-new-opus-5](https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5)
5. Anthropic. *Refusals and fallback — response shapes and retry paths.* Retrieved July 2026 from [platform.claude.com/docs/en/build-with-claude/refusals-and-fallback](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback)
### Further reading
* reAPI. *How to use Claude Opus 5.* [reapi.ai/blog/how-to-use-claude-opus-5](/blog/how-to-use-claude-opus-5)
* reAPI. *How to use Claude Opus 4.8.* [reapi.ai/blog/how-to-use-claude-opus-4-8](/blog/how-to-use-claude-opus-4-8)
* reAPI. *Claude Fable 5 endpoint reference.* [reapi.ai/docs/claude-fable-5](/docs/claude-fable-5)
---
# How to Use Claude Opus 4.8: Coding, Honesty, and Cost (https://reapi.ai/blog/how-to-use-claude-opus-4-8)
Anthropic released Claude Opus 4.8 on May 28, 2026, forty-one days after Opus 4.7 and at exactly the same rate: $5 per million input tokens and $25 per million output\[1]\[3]. Knowing how to use Claude Opus 4.8 well is mostly a question of what changed underneath a price tag that did not move.
Three things did. Agentic coding went up again, SWE-bench Pro climbing from 64.3% to 69.2%\[1]. Fast mode got three times cheaper. And Anthropic spent most of its announcement on something that is not a capability number at all: the model is roughly four times less likely than Opus 4.7 to let a flaw in code it wrote pass without flagging it\[1]. For anyone running agents unattended, that last one changes the risk profile more than the benchmark does.
## TL;DR
* **Price unchanged, capability up.** $5 / $25 per million tokens, same as Opus 4.7 and 4.6\[3]. On reAPI the same model runs $4.00 / $20.00, which is 80% of Anthropic's published rate.
* **The honesty gain is the headline.** About four times less likely than Opus 4.7 to let its own code flaws pass unremarked, and roughly seventeen times less likely than Sonnet 4.6 to produce a dishonest summary of its own agentic work\[1].
* **Fast mode dropped to a third.** $10 / $50 per million tokens for roughly 2.5x output speed, against the $30 / $150 fast mode on previous Claude models\[1].
* **The default effort level moved from `medium` to `high`.** Harnesses that relied on the old default get deeper reasoning, higher latency, and more output tokens unless they set it explicitly\[1].
* **Fixed thinking budgets were removed.** Code passing `budget_tokens` has to move to `thinking: {"type": "adaptive"}`\[1].
* **It does not win everywhere.** GPT-5.5 still takes Terminal-Bench 2.1, 78.2% against 74.6%, and GPQA Diamond slipped slightly from Opus 4.7\[1].
## What Claude Opus 4.8 is

Claude Opus 4.8 is the frontier model in Anthropic's Claude 4 family and a direct replacement for Opus 4.7 on every surface Anthropic ships to. The API identifier is `claude-opus-4-8`, with no date suffix\[2]. It reads up to 1M tokens in one request and returns up to 128K, or up to 300K on the Batch API with a beta header\[2].
If you want the specification-level tour, we wrote one: [what is Claude Opus 4.8](/blog/what-is-claude-opus-4-8) covers the context window, modalities, and capability list. This guide is about the operational side, which is where the interesting decisions live.
## How Opus 4.8 performs on coding and agentic benchmarks

Anthropic's headline comparison puts Opus 4.8 against Opus 4.7, GPT-5.5, and Gemini 3.1 Pro across the evaluations it considers representative of real agentic work\[1].
| Benchmark | Opus 4.8 | Opus 4.7 | GPT-5.5 | Gemini 3.1 Pro |
| --------------------------------------------- | --------- | -------- | --------- | -------------- |
| Agentic coding (SWE-bench Pro) | **69.2%** | 64.3% | 58.6% | 54.2% |
| Agentic terminal coding (Terminal-Bench 2.1) | 74.6% | 66.1% | **78.2%** | 70.3% |
| Multidisciplinary reasoning (HLE, no tools) | **49.8%** | 46.9% | 41.4% | 44.4% |
| Multidisciplinary reasoning (HLE, with tools) | **57.9%** | 54.7% | 52.2% | 51.4% |
| Agentic computer use (OSWorld-Verified) | **83.4%** | 82.8% | 78.7% | 76.2% |
| Knowledge work (GDPval-AA, Elo) | **1890** | 1753 | 1769 | 1314 |
| Agentic financial analysis (Finance Agent v2) | **53.9%** | 51.5% | 51.8% | 43.0% |
Opus 4.8 leads six of the seven rows. The exception is Terminal-Bench 2.1, where GPT-5.5 wins 78.2% to 74.6%\[1].
Three caveats belong next to that table, because they matter for an actual buying decision. SWE-bench Verified, reported separately, moves to 88.6% from 87.6%, a small step at a level where the benchmark is starting to saturate. GPQA Diamond went the wrong way, slipping to 93.6% from Opus 4.7's 94.2%. And these are vendor-run evaluations, so read the whole table as a directional signal rather than a neutral audit. Anthropic also reports 96.7% on USAMO 2026 and a rise to 84.3% on single-agent BrowseComp\[1].
The honest summary is "broadly ahead, not ahead everywhere."
## What partners report from production
Benchmarks overstate real-world gains often enough that the early-access partner statements in Anthropic's launch note are worth more than another eval row. Several are specific enough to act on\[1].
* **Databricks** reported a step change in agentic reasoning at roughly 61% lower token cost than Opus 4.7. This is the one that matters for production economics, because a capability gain that also cuts spend is rare in a point release.
* **Cursor** reported meaningfully more efficient tool calling, completing end-to-end tasks in fewer steps at the same level of intelligence.
* **Cognition**, which builds Devin, reported clean tool use and instruction-following consistent enough to keep autonomous engineering workloads running unattended.
* **Harvey** recorded its highest score on its Legal Agent Benchmark, and the first model to clear 10% on its strictest all-pass standard.
Treat the 61% figure as a hypothesis to test on your own traffic rather than a number to put in a budget. Token cost is workload-dependent, and the default effort change described below pushes in the opposite direction.
## The honesty upgrade
The most distinctive part of this release is a behavior change, not a capability number. On Anthropic's internal misaligned-behavior evaluation, scored 1 to 10 where lower is better, Opus 4.8 lands close to Claude Mythos Preview and well below both Opus 4.7 and Sonnet 4.6\[1]. In practical terms Anthropic reports it is about four times less likely than Opus 4.7 to let flaws in its own code pass unremarked, and about seventeen times less likely than Sonnet 4.6 to produce a dishonest summary of its agentic work\[1].

Two things keep that in perspective. These are rate reductions, not eliminations. A model four times less likely to hide a flaw still occasionally hides one, so human review of agentic code does not go away. And the honesty metrics come from Anthropic's own evaluations rather than third-party benchmarks, so they are the vendor's framing until independent testing catches up.
The direction still matters. In a long unattended loop, a model that proactively says it is unsure is materially safer to deploy than one that confidently ships a broken result.
## Dynamic Workflows
The flagship new feature ships as a research preview in Claude Code and addresses a problem every agentic-coding user eventually hits: some jobs are too large for one agent and one context window. A framework migration across hundreds of thousands of lines. A dependency upgrade touching every service. A repo-wide refactor.
The design is orchestrator-and-workers. Opus 4.8 plans the work, splits it into independent segments, and dispatches them to hundreds of subagents running in parallel. Each subagent plans, executes, and verifies its own slice, and the orchestrator merges the verified results. Anthropic frames the repository's existing test suite as the gate for what counts as done, and describes the system as carrying a codebase-scale migration from kickoff to merge\[1].
This is an orchestration feature that the model's improved long-horizon coherence makes viable, which is also why the honesty work and the workflow work belong in the same release. Fanning out hundreds of agents is only useful if each one reliably flags what it could not finish.
## Effort, fast mode, and the cost equation
Two changes here, and the first one is a silent bill increase if you miss it.
**The API default effort level moved from `medium` to `high`**, with `xhigh` and `max` available above it\[1]. Any harness that relied on the old default now reasons deeper, runs slower, and emits more output tokens without a line of code changing. If you are comparing Opus 4.8 cost against Opus 4.7 cost, match the effort levels first or the comparison is meaningless.
**Fast mode dropped to a third of its former price.** It runs the same model at roughly 2.5x output speed for $10 per million input tokens and $50 per million output, against $30 / $150 on previous Claude models\[1]. That is still double the standard rate, but it moves fast mode from "too expensive to justify" into range for latency-sensitive products.
Here is the full official rate card\[3]:
| Item | Price per MTok |
| ------------------------ | -------------- |
| Input | $5 |
| Output | $25 |
| Cache hits and refreshes | $0.50 |
| Cache write (5-minute) | $6.25 |
| Cache write (1-hour) | $10 |
| Batch input | $2.50 |
| Batch output | $12.50 |
| Fast mode input | $10 |
| Fast mode output | $50 |
The Batch API halves both dimensions, and prompt caching makes repeated context cheap to reread at $0.50 per million tokens\[3]. Between a flat standard price, a cheaper fast tier, and an effort default that moved, the question stopped being whether Opus is affordable and became which effort level, in which mode, for which class of task.
## How to use Claude Opus 4.8: migrating from Opus 4.7
Anthropic describes the move as largely non-breaking, but three changes are worth catching before production traffic flips\[1].
1. **The default effort level changed from `medium` to `high`.** Set it explicitly or accept deeper reasoning, higher latency, and more output tokens.
2. **Fixed extended-thinking token budgets were removed.** Code passing a `budget_tokens` value migrates to `thinking: {"type": "adaptive"}`.
3. **The Messages API now accepts system entries mid-task without breaking the prompt cache.** Not a breaking change, but a behavior to know about if your agent updates instructions mid-run.
The checklist that holds for any model upgrade: replay 20 to 50 real tasks from your production trace at matched effort levels, compare end-to-end cost and quality rather than per-token rates, re-tune prompts that pin effort or thinking budgets, and pilot Dynamic Workflows on one genuinely large job before wiring it into automation. The token-cost math is empirical. Measure it on your own workload.
## Who should care right now
**Teams already running Opus 4.7.** The SWE-bench Pro gain plus the reported token-cost reduction make a migration pilot worth the time, provided you budget for the new `high` effort default and benchmark cost on real tasks first.
**Teams with codebase-scale jobs.** Dynamic Workflows is the reason to upgrade if you have a migration or repo-wide refactor that stalled because it was too big for one agent. Start with a single bounded job where the test suite is trustworthy.
**Anyone running unattended agentic loops.** The honesty improvements matter most where no human reviews each step in real time: overnight runs, autonomous pipelines, long research tasks.
**Teams that should probably skip it.** If you are starting fresh rather than migrating, look at [Claude Opus 5](/blog/how-to-use-claude-opus-5) before committing to 4.8. It lists at the same $5 / $25, and on reAPI it runs cheaper than Opus 4.8 does. Opus 4.8 remains the right answer when you need a model whose behavior your evals already characterize, or when you specifically want thinking off at high effort levels, which Opus 5 no longer permits.
## Calling Claude Opus 4.8 alongside your other models
Everything above is a routing problem wearing a migration costume. The effort default moved, so cost per task moved. Fast mode became viable for a class of routes it was priced out of. GPT-5.5 still wins terminal coding. Opus 5 exists now at the same list price. None of that resolves into one setting, and all of it changes again at the next release.
reAPI exposes Claude Opus 4.8 through an OpenAI-compatible `/v1/chat/completions` endpoint, so the client you already use for GPT and Gemini calls it too:
```python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_REAPI_KEY",
base_url="https://api.reapi.ai/v1",
)
resp = client.chat.completions.create(
model="claude-opus-4-8",
messages=[{"role": "user", "content": "Refactor this module and add tests."}],
max_tokens=16000,
stream=True,
)
```
Rates on the gateway are $4.00 per million input tokens and $20.00 per million output, which is 80% of Anthropic's published $5 / $25. The endpoint reference is at [reapi.ai/docs/claude-opus-4-8](/docs/claude-opus-4-8), and current rates live on [reapi.ai/models/claude-opus-4-8](/models/claude-opus-4-8).
Swapping a model string is the easy part. Being able to swap it without a second SDK, a second key, and a second invoice is the part worth building once.
## FAQ
### When was Claude Opus 4.8 released?
May 28, 2026, forty-one days after Opus 4.7, at unchanged pricing\[1].
### How much does Claude Opus 4.8 cost?
$5 per million input tokens and $25 per million output, with cache reads at $0.50 and the Batch API at half price\[3]. Fast mode is $10 / $50\[1].
### What is the biggest thing to catch when migrating from Opus 4.7?
The API default effort level moved from `medium` to `high`. If your harness relied on the old default, you get deeper reasoning, higher latency, and more output tokens without changing any code\[1].
### Does Claude Opus 4.8 beat GPT-5.5?
On five of the six rows where Anthropic compares them directly, yes. GPT-5.5 wins agentic terminal coding on Terminal-Bench 2.1, 78.2% against 74.6%\[1].
### What are Dynamic Workflows?
A Claude Code research preview in which Opus 4.8 acts as an orchestrator, splits a large job into independent segments, dispatches them to hundreds of parallel subagents, and merges verified results using the repository's own test suite as the gate\[1].
### Is the honesty improvement independently verified?
Not yet. The four-times and seventeen-times figures come from Anthropic's internal evaluations rather than third-party benchmarks\[1]. They are rate reductions, not eliminations, so human review of agentic output still applies.
### Should I use Opus 4.8 or Opus 5?
Opus 5 lists at the same $5 / $25 and is a step-change rather than an incremental upgrade. Opus 4.8 stays relevant when your evals already characterize its behavior, or when you need thinking disabled at high effort levels, which Opus 5 rejects.
### How do I call Claude Opus 4.8 with the OpenAI SDK?
Point the OpenAI client at `https://api.reapi.ai/v1`, set `model` to `claude-opus-4-8`, and send a normal chat request. The endpoint is OpenAI-compatible, so no SDK rewrite is needed.
## Matching the effort level to the job
The useful way to read this release is that Anthropic held the price flat and spent the release on things that only show up in production: a model that flags its own mistakes, an orchestration primitive for jobs too big for one context window, and a fast mode that finally costs what latency-sensitive products can pay. The benchmark table is real but modest. The behavior change is the part that alters how much supervision an agent needs.
If you are migrating, the work is small and specific. Pin the effort level explicitly rather than inheriting the new `high` default, move any `budget_tokens` code to adaptive thinking, replay real production tasks at matched effort before trusting any cost comparison, and pilot Dynamic Workflows on one bounded job with a test suite you trust. That is how to use Claude Opus 4.8 without discovering the effort default in your invoice.
## References
1. Anthropic. *Introducing Claude Opus 4.8 — benchmarks, honesty evaluations, Dynamic Workflows, effort defaults, and fast-mode pricing.* Retrieved July 2026 from [anthropic.com/news/claude-opus-4-8](https://www.anthropic.com/news/claude-opus-4-8)
2. Anthropic. *Models overview — context, output, modality, and capabilities.* Retrieved July 2026 from [platform.claude.com/docs/en/about-claude/models/overview](https://platform.claude.com/docs/en/about-claude/models/overview)
3. Anthropic. *Pricing — token, cache, and batch rates.* Retrieved July 2026 from [platform.claude.com/docs/en/about-claude/pricing](https://platform.claude.com/docs/en/about-claude/pricing)
### Further reading
* reAPI. *How to use Claude Opus 5.* [reapi.ai/blog/how-to-use-claude-opus-5](/blog/how-to-use-claude-opus-5)
* reAPI. *What is Claude Opus 4.8?* [reapi.ai/blog/what-is-claude-opus-4-8](/blog/what-is-claude-opus-4-8)
* reAPI. *Claude Opus 4.8 endpoint reference.* [reapi.ai/docs/claude-opus-4-8](/docs/claude-opus-4-8)
---
# How to Use Claude Opus 5: Benchmarks, Effort, and Cost (https://reapi.ai/blog/how-to-use-claude-opus-5)
Anthropic launched Claude Opus 5 on July 24, 2026, describing it as a step-change improvement over Claude Opus 4.8\[6]. It lists at exactly the same price as the model it succeeds: $5 per million input tokens and $25 per million output\[3]. So the question of how to use Claude Opus 5 well has almost nothing to do with the list price, which did not move, and almost everything to do with a dial most teams never touch.
That dial is effort. Instead of publishing one number per benchmark, Anthropic published score-versus-cost curves across all five effort levels\[1]. The spread between the cheapest and the most expensive rung on those curves is wider than anything a prompt rewrite will recover, and the official guidance is to sweep the ladder against your own evals rather than turn everything up\[2].
## TL;DR
* **Same list price, different model.** $5 / $25 per million tokens, unchanged from Opus 4.8\[2]. On reAPI the same model runs $2.40 / $12.00, which is 48% of Anthropic's published rate.
* **Thinking is on by default now.** On Opus 4.8, omitting the `thinking` field meant no thinking. On Opus 5 the same request thinks\[2]. Since `max_tokens` caps thinking plus response text together, tight budgets carried over will truncate.
* **Disabling thinking is capped at effort `high`.** Sending `thinking: {"type": "disabled"}` with `xhigh` or `max` returns a 400, checked per request\[2].
* **Opus 5 does not win everything.** Anthropic's own table has it losing agentic coding on DeepSWE v1.1 to GPT-5.6 Sol, 68.8% against 72.7%\[1].
* **Do not inherit effort settings.** Anthropic says `low` and `medium` now produce strong quality at a fraction of the tokens and latency, and tells you to start at the default `high` and move in both directions against your own evals\[2].
* **Priority Tier does not cover it**, and it sits in its own rate-limit bucket rather than the shared Opus pool\[5]\[4].
## What Claude Opus 5 is
Anthropic positions Claude Opus 5 for complex agentic coding and enterprise work\[2]. The API identifier is `claude-opus-5`, with no date suffix to append.
| Property | Claude Opus 5 |
| -------------------- | ------------------------------------------------------- |
| API model ID | `claude-opus-5` |
| Context window | 1M tokens, both the default and the maximum |
| Max output | 128k tokens |
| List price | $5 / $25 per million input / output tokens |
| Fast mode | $10 / $50, Claude API only |
| Thinking | On by default, adaptive |
| Effort levels | `low`, `medium`, `high`, `xhigh`, `max`, default `high` |
| Prompt cache minimum | 512 tokens, down from 1,024 |
There is no smaller context variant\[2]. The 1M window is what you get. Beyond the Claude API, the model is available on Amazon Bedrock as `anthropic.claude-opus-5`, on Google Cloud as `claude-opus-5`, and on Microsoft Foundry\[2]. Fast mode is the exception: it runs on the Claude API only, not on any of the three cloud platforms\[2].
## The full official benchmark table

This is Anthropic's published comparison, transcribed in full, including the rows Opus 5 does not win\[1]. Where the Fable 5 column carries a different model name, that is Anthropic's own annotation.
| Benchmark | Opus 5 | Fable 5 | Opus 4.8 | GPT-5.6 Sol |
| --------------------------------------------- | --------- | -------------------- | -------- | ----------- |
| Agentic terminal coding (Frontier-Bench v0.1) | **43.3%** | 33.7% | 21.1% | 34.4% |
| Knowledge work (GDPval-AA v2, Elo) | **1861** | 1747 | 1593 | 1736 |
| Novel problem-solving (ARC-AGI-3) | **30.2%** | — | 1.5% | 7.8% |
| Agentic search (BrowseComp) | **90.8%** | 87.4% | 84.3% | 90.4% |
| Multidisciplinary reasoning (HLE, no tools) | 56.3% | **56.5%** | 49.8% | — |
| Multidisciplinary reasoning (HLE, with tools) | **64.7%** | 63.9% | 57.9% | — |
| Computer use (OSWorld 2.0) | **70.6%** | 66.1% | 55.7% | 62.6% |
| Agentic coding (DeepSWE v1.1) | 68.8% | 69.7% | 59.0% | **72.7%** |
| Agentic coding (FrontierCode v1.1, Main) | 53.4% | **53.5%** | 46.5% | 47.5% |
| Business workflows (AutomationBench) | **26.0%** | 17.4% | 17.0% | 18.1% |
| Legal (Legal Agent Benchmark, held-out) | 11.7% | **13.3%** | 10.4% | 2.5% |
| Health (HealthBench Professional) | 59.8% | **66.0%** (Mythos 5) | 57.4% | 60.5% |
| Biology (BioMysteryBench, hard) | **49.4%** | 46.5% | 42.4% | — |
| Biology (BioMysteryBench, human-solved) | **90.1%** | 89.0% (Mythos 5) | 88.5% | — |
One methodology note before anyone quotes these numbers. On the companion chart that breaks Frontier-Bench down by effort level, Anthropic states that those results come from an internal run on the mini-SWE-agent harness with a GKE backend, averaging reward over five attempts per task, and that Opus 4.8 served as the fallback when safety classifiers refused a request from Opus 5 or Fable 5\[1]. These are vendor-run evaluations. Read them as a directional signal, not a neutral audit.
## Where it wins and where it loses
Opus 5 takes nine of the fourteen rows. The margins on the wins are not subtle. Frontier-Bench goes from 21.1% on Opus 4.8 to 43.3%, slightly more than double. ARC-AGI-3 goes from 1.5% to 30.2%, and the nearest competitor in that row is GPT-5.6 Sol at 7.8%. AutomationBench sits at 26.0% against a cluster of 17% to 18%\[1].
It loses the other five, and those are more useful for picking a model than the wins are.
* **DeepSWE v1.1**, agentic coding: 68.8% against GPT-5.6 Sol's 72.7%. This is the clearest loss on the board, and it lands on the workload Opus 5 is marketed for. Anthropic left the row in.
* **HealthBench Professional**: 59.8%, behind both Mythos 5 at 66.0% and GPT-5.6 Sol at 60.5%.
* **Legal Agent Benchmark**: 11.7% against Fable 5's 13.3%. Both numbers are low enough that the benchmark is mostly measuring distance from the goal.
* **HLE without tools** (56.3% against 56.5%) and **FrontierCode Main** (53.4% against 53.5%) are ties inside the noise, but they do rule out the claim that Opus 5 dominates Fable 5 across the board.
The pattern: Opus 5 wins where a task runs long, uses tools, and requires holding state across many steps. It ties or loses on single-shot hard questions.
## The effort ladder decides your bill
Anthropic publishes the full score-versus-cost ladder rather than one number per benchmark, and reading it changes how you configure the model\[1].

Three things follow from it.
**The top rung is not the recommended starting point.** Anthropic's instruction is to start at the default, `high`, and adjust in either direction based on your evals: step down where quality holds to save tokens and latency, step up for the most demanding work\[2]. Note where that sentence starts. The recommended entry point is the middle of the ladder, not the top of it, and effort now carries more weight because Opus 5 converts extra effort into better results more reliably than any earlier Opus model\[2].
**Low and medium are stronger than they used to be.** Anthropic explicitly calls out efficiency at lower effort, saying `low` and `medium` produce strong quality at a fraction of the tokens and latency of higher settings\[2]. Effort defaults carried over from Opus 4.8 are the wrong starting point.
**Effort is the cost lever, not the prompt.** The official guidance is to start at the default `high`, then move in either direction against your own evals: step down where quality holds to save tokens and latency, step up for the most demanding work\[2]. At `xhigh` or `max`, set a large `max_tokens` so the model has room to think and act across subagents and tool calls\[2].
## Two breaking changes under the hood
The request surface is mostly inherited from Opus 4.8. Sampling parameters are still rejected, fixed thinking budgets are still rejected, and last-assistant-turn prefill still fails. Two things did change, and both break code carried over unmodified.
**Thinking runs when you omit the parameter.** On Opus 4.8, a request without a `thinking` field ran without thinking. On Opus 5, the same request thinks, with the model deciding when and how much on each turn\[2]. The wire value did not change; `thinking: {"type": "adaptive"}` is still valid and equivalent to the default. What changed is what happens when you say nothing. Because `max_tokens` is a hard limit on thinking plus response text together, Anthropic tells you to revisit it for any workload that ran without thinking before\[2]. A budget tuned for visible output alone will now truncate mid-answer.
**Turning thinking off requires effort `high` or below.** `thinking: {"type": "disabled"}` paired with `xhigh` or `max` returns a 400. The check runs on every request, so a session that succeeded at `high` starts failing the moment a later call raises effort while thinking is still disabled\[2].
Anthropic also documents what goes wrong with thinking off: the model can write a tool call into its text output instead of emitting a `tool_use` block, or leak internal XML tags into the visible response\[2]. The first failure mode is the dangerous one in an agent loop, because the turn ends cleanly and the call never executes. The official recommendation is to keep thinking on and control cost with a lower effort level instead.
Three additions are worth knowing about:
* **Mid-conversation tool changes** (beta header `mid-conversation-tool-changes-2026-07-01`) let you add or remove tools between turns while preserving the prompt cache, instead of pinning a tool list for the life of a session\[2].
* **A `"default"` fallbacks mode** (beta header `server-side-fallback-2026-07-01`) applies Anthropic's recommended fallback models by refusal category rather than a list you maintain\[2].
* **The prompt cache minimum dropped to 512 tokens** from 1,024, so prompts previously too short to cache now create cache entries with no code changes\[2].
## What actually moves the bill
Here is the complete official rate card\[3]:
| Item | Price per MTok |
| ------------------------ | -------------- |
| Input | $5 |
| Output | $25 |
| Cache hits and refreshes | $0.50 |
| Cache write (5-minute) | $6.25 |
| Cache write (1-hour) | $10 |
| Batch input | $2.50 |
| Batch output | $12.50 |
| Fast mode input | $10 |
| Fast mode output | $50 |
The Batch API runs at half the standard rate on both dimensions, and fast mode doubles both\[3]. Anthropic describes Opus 5 as delivering frontier intelligence at half the cost of Claude Fable 5\[2], which checks out against the rate card: Fable 5 lists at $10 / $50\[3].
Three levers matter more than the per-token rate.
**Effort.** The published cost-per-attempt spread across the ladder is wide enough that moving one route from `xhigh` down to `medium` saves more than any prompt rewrite\[1].
**Verbosity.** Anthropic states that default responses and written deliverables run longer on Opus 5\[2]. Thinking tokens bill as output, and so does the extra prose. Lowering effort reduces thinking but does not reliably shorten visible output, so write an explicit brevity instruction.
**Verification scaffolding you can now delete.** Opus 5 verifies its own work without being told to, and Anthropic says instructions carried over from earlier models, like "include a final verification step" or "use a subagent to verify", cause over-verification\[2]. Deleting them lowers token spend with no quality tradeoff, which is a rare kind of optimization.
## Rate limits and Priority Tier
Two operational facts that catch teams during rollout.
Priority Tier does not cover Claude Opus 5. Anthropic's service-tier documentation states that Priority Tier is supported on all available Claude models except Claude Mythos 5, Claude Mythos Preview, Claude Opus 5, and Claude Sonnet 5\[5]. If your capacity planning assumed Priority Tier coverage, it does not apply here.
Opus 5 also has its own rate-limit bucket. The combined Opus limit applies to traffic across Opus 4.8, 4.7, 4.6, and 4.5; Opus 5 is not part of that pool\[4]. That is good news during a migration, since Opus 5 traffic does not eat the budget your existing Opus 4.8 routes depend on. Fast mode has dedicated limits again separate from standard Opus limits, and returns a 429 with a `retry-after` header when exceeded\[4].
## How to use Claude Opus 5: a migration checklist
Anthropic's migration guidance reduces to a model string change plus a review of the two behavior changes\[2]. In practice the list is a little longer.
1. Change the model string to `claude-opus-5`.
2. Audit every route that disables thinking. Either keep thinking disabled and drop effort to `high` or below, or keep the effort level and remove the `thinking` field.
3. Raise `max_tokens` on routes that never set a `thinking` field. They think now, and the budget covers both.
4. Re-run an effort sweep against your own evals. Start at the default `high` and move in both directions rather than inheriting Opus 4.8 settings.
5. Delete verification instructions from prompts and verification steps from harnesses.
6. Add an explicit brevity instruction to user-facing routes.
7. Recheck short prompts you had written off as uncacheable against the new 512-token minimum.
8. Handle refusals as responses, not errors. Safety classifiers can decline a request and return a stop reason rather than an HTTP error, so code that reads the first content block unconditionally will break.
## Calling Claude Opus 5 alongside your other models
Everything above points at the same operational problem. This model has five effort levels, its own rate-limit bucket, no Priority Tier, and a benchmark table where it loses agentic coding to GPT-5.6 Sol and health to Mythos 5. The correct configuration is not one setting. It is a routing decision per workload, and it changes every time a vendor ships.
That is the layer worth building deliberately. A code-review pass and a bulk extraction job have no business running at the same effort, and a route that Opus 5 loses on has no business staying on Opus 5 out of habit. Handling that with one SDK client per vendor, one key per vendor, and one billing surface per vendor is where the cost of multi-model work actually accumulates.
reAPI exposes Claude Opus 5 through an OpenAI-compatible `/v1/chat/completions` endpoint, so the client you already use for GPT and Gemini calls it too:
```python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_REAPI_KEY",
base_url="https://api.reapi.ai/v1",
)
resp = client.chat.completions.create(
model="claude-opus-5",
messages=[{"role": "user", "content": "Refactor this module and add tests."}],
max_tokens=16000,
stream=True,
)
```
Rates on the gateway are $2.40 per million input tokens and $12.00 per million output, which is 48% of Anthropic's published $5 / $25. The full request surface, including `output_config.effort` and the thinking rules above, is documented on the [reapi.ai/docs/claude-opus-5](/docs/claude-opus-5) page, and current rates live on [reapi.ai/models/claude-opus-5](/models/claude-opus-5).
None of that removes the routing decision. It removes the integration tax on making one.
## FAQ
### Is Claude Opus 5 better than Claude Fable 5?
On nine of the fourteen rows in Anthropic's published table, yes, and at half Fable 5's list price\[3]. But Fable 5 still takes HLE without tools, FrontierCode Main, and the Legal Agent Benchmark, and the HealthBench Professional row goes to Mythos 5\[1].
### Does Claude Opus 5 cost more than Opus 4.8?
No. Both list at $5 per million input tokens and $25 per million output\[3]. Anthropic shipped the capability jump without a price change\[2].
### What is the biggest breaking change from Opus 4.8?
Thinking is on by default. A request that omitted the `thinking` field ran without thinking on Opus 4.8 and thinks on Opus 5, which changes both cost and `max_tokens` behavior. Second is that `thinking: {"type": "disabled"}` now returns a 400 at effort `xhigh` or `max`\[2].
### Which effort level should I use?
Start at the default, `high`, then sweep in both directions against your own evals\[2]. Anthropic frames `max` as the top tier for the deepest possible reasoning rather than a better default, and separately notes that `low` and `medium` hold up at a fraction of the tokens and latency on this model\[2].
### What is the Claude Opus 5 context window?
1M tokens, which is both the default and the maximum. There is no smaller context variant. Maximum output is 128k tokens\[2].
### Is Claude Opus 5 covered by Priority Tier?
No. Priority Tier covers all available Claude models except Mythos 5, Mythos Preview, Opus 5, and Sonnet 5\[5]. Opus 5 also sits in its own rate-limit bucket rather than the combined Opus 4.x pool\[4].
### Can I still turn thinking off?
Yes, at effort `high` or below. Anthropic advises against it, documenting that the model may write tool calls into visible text or leak internal tags when thinking is disabled, and recommends lowering effort instead\[2].
### How do I call Claude Opus 5 with the OpenAI SDK?
Point the OpenAI client at `https://api.reapi.ai/v1`, set `model` to `claude-opus-5`, and send a normal chat request. The endpoint is OpenAI-compatible, so no SDK rewrite is needed.
## Picking an effort level and moving on
The interesting thing about this release is that the pricing page is the least informative part of it. The list price did not move, so the whole delta is in behavior: thinking on by default, a 400 that did not exist before, a cache floor cut in half, and an effort ladder whose recommended entry point is the middle rung rather than the top one. Anthropic also published five benchmark rows where its new flagship loses, which is worth more than the nine where it wins.
If you are migrating, the work is small and specific. Change the model string, audit the routes that disable thinking, raise `max_tokens` where thinking is now implicit, delete the verification instructions you wrote for an older model, and sweep effort against your own evals rather than inheriting a setting. That is how to use Claude Opus 5 without paying for the top rung of a ladder that stops climbing.
## References
1. Anthropic. *Introducing Claude Opus 5 — benchmark comparison table and effort-level cost curves.* Retrieved July 2026 from [anthropic.com/news/claude-opus-5](https://www.anthropic.com/news/claude-opus-5)
2. Anthropic. *What's new in Claude Opus 5 — behavior changes, new features, availability, and migration.* Retrieved July 2026 from [platform.claude.com/docs/en/about-claude/models/whats-new-opus-5](https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5)
3. Anthropic. *Pricing — token, cache, batch, and fast-mode rates.* Retrieved July 2026 from [platform.claude.com/docs/en/about-claude/pricing](https://platform.claude.com/docs/en/about-claude/pricing)
4. Anthropic. *Rate limits — per-model buckets and fast-mode limits.* Retrieved July 2026 from [platform.claude.com/docs/en/api/rate-limits](https://platform.claude.com/docs/en/api/rate-limits)
5. Anthropic. *Service tiers — Priority Tier model coverage.* Retrieved July 2026 from [platform.claude.com/docs/en/api/service-tiers](https://platform.claude.com/docs/en/api/service-tiers)
6. Anthropic. *Release notes — Claude Opus 5 launch entry dated July 24, 2026.* Retrieved July 2026 from [platform.claude.com/docs/en/release-notes/overview](https://platform.claude.com/docs/en/release-notes/overview)
### Further reading
* reAPI. *What is Claude Opus 4.8?* [reapi.ai/blog/what-is-claude-opus-4-8](/blog/what-is-claude-opus-4-8)
* reAPI. *Claude Opus 5 endpoint reference.* [reapi.ai/docs/claude-opus-5](/docs/claude-opus-5)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# How to Use Claude Sonnet 5: Effort, Cost, and Limits (https://reapi.ai/blog/how-to-use-claude-sonnet-5)
Anthropic released Claude Sonnet 5 on June 30, 2026 with an unusually direct pitch: near-Opus intelligence on coding and agentic work, at Sonnet pricing\[1]. The official benchmarks back up most of it.
Knowing how to use Claude Sonnet 5 well is mostly about two things the launch post does not lead with. The effort ladder is the lever that makes the cost claim real, and a denser tokenizer means the same prompt now costs about 30% more tokens than it did on Sonnet 4.6, so a migration that only swaps the model string will quietly change your bill and can truncate your output\[1].
This guide covers the benchmarks, where Sonnet 5 ties and where it trails Opus 4.8, the four API changes that break carried-over code, the pricing including the September increase, and who should actually run it.
## TL;DR
* **Near-Opus on knowledge work.** GDPval-AA v2 Elo of 1618 against Opus 4.8's 1615, a statistical tie\[1].
* **Still trails on the hard tail.** SWE-bench Pro 63.2% against 69.2%, and USAMO 2026 79.5% against 96.7%\[1].
* **Introductory pricing runs through August 31, 2026**: $2 input / $10 output per million tokens, rising to $3 / $15 on September 1\[1].
* **First Sonnet with the full effort ladder**: `low`, `medium`, `high` (default), `xhigh`, `max`\[1].
* **The tokenizer is about 30% denser.** Per-token prices are unchanged, so an identical request costs more than on Sonnet 4.6. Re-baseline before you migrate\[1].
* **`budget_tokens` and sampling parameters now return 400**\[1].
## What Claude Sonnet 5 is
Claude Sonnet 5 (`claude-sonnet-5`) succeeds Sonnet 4.6 as a drop-in replacement. Anthropic calls it the most agentic Sonnet model yet, built to make plans, drive browsers and terminals, and run autonomously for long stretches at a level that recently required larger and more expensive models\[1].
It shipped as the default model for Free and Pro users on claude.ai, and simultaneously on the Claude API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry in preview. It carries a 1M-token context window by default and 128K maximum output\[1].
## Benchmarks

| Benchmark | Sonnet 5 | Sonnet 4.6 | Opus 4.8 (reference) |
| ----------------------------------- | -------- | ---------- | -------------------- |
| Agentic coding (SWE-bench Pro) | 63.2% | 58.1% | **69.2%** |
| Agentic coding (Terminal-Bench 2.1) | 80.4% | 67.0% | **82.7%** |
| Reasoning (HLE, no tools) | 43.2% | 34.6% | **49.8%** |
| Reasoning (HLE, with tools) | 57.4% | 46.8% | **57.9%** |
| Computer use (OSWorld-Verified) | 81.2% | 78.5% | **83.4%** |
| Knowledge work (GDPval-AA v2, Elo) | **1618** | 1395 | 1615 |
Two things stand out. Against Sonnet 4.6 this is a clean sweep with real jumps: Terminal-Bench climbs 13.4 points and the knowledge-work Elo rises 223. And on knowledge work Sonnet 5 nominally edges Opus 4.8, 1618 against 1615, which is a tie, while landing within roughly two to six points of the flagship on agentic coding and computer use\[1].
That "within a few points, at a third of the cost" gap is the entire value proposition.
Deeper numbers from the system card: 85.2% on the classic 500-problem SWE-bench Verified subset, 78.3% on SWE-bench Multilingual across nine languages, and 38.8 on Cognition's FrontierCode v1, more than double Sonnet 4.6's 15.1. Cursor's independent measurement puts it at 61.2% on CursorBench against Sonnet 4.6's 49% and Opus 4.8's 63.8%. BrowseComp reaches 84.7% single-agent and 86.6% multi-agent\[1].
One honest soft spot: on USAMO 2026 Sonnet 5 scores 79.5%, a big jump over Sonnet 4.6's 55.0% but far behind Opus 4.8's 96.7%\[1]. Heavy formal-proof mathematics is not this model's job.
## The effort ladder is the cost lever
Sonnet 5 is the first Sonnet-tier model to expose the full ladder: `low`, `medium`, `high` (the default), `xhigh`, and `max`\[1]. This is not a cosmetic knob. It is what makes the Opus-quality-at-Sonnet-cost claim real.

The BrowseComp curve is the clearest illustration. Sonnet 5's accuracy climbs from roughly 60% at `low` to about 85% at `max`, and it reaches Opus-4.8-class accuracy at a markedly lower cost per task. Opus 4.8 still tops out slightly higher, but Sonnet 5 gets most of the way there for far less money, which is exactly what a high-volume agent loop needs\[1].
Computer use tells a more sober story.

On OSWorld-Verified, Opus 4.8 stays ahead at every effort level. Sonnet 5 still beats Sonnet 4.6 comfortably and remains far cheaper, but if pixel-accurate GUI control is your core workload the flagship keeps a genuine edge\[1].
The practical rule: `xhigh` for the hardest long-horizon coding and agentic tasks, `high` for most work, `medium` or `low` for latency-sensitive or simple jobs. Tune it per route rather than defaulting to `max` everywhere.
## Against the cost-comparable tier
The system card publishes competitor numbers against GPT-5.5 and Gemini 3.5 Flash, which is the cost-comparable tier rather than the absolute frontier. Read it as "beats the models at its price point," not "beats everything"\[1].
| Evaluation | Sonnet 5 | GPT-5.5 | Gemini 3.5 Flash |
| ------------------------ | -------- | -------- | ---------------- |
| SWE-bench Pro | **63.2** | 58.6 | 55.1 |
| Terminal-Bench 2.1 | 80.4 | **83.4** | 76.2 |
| HLE (no tools) | **43.2** | 41.4 | 40.2 |
| HLE (with tools) | **57.4** | 52.2 | n/a |
| OSWorld-Verified | **81.2** | 78.7 | 78.4 |
| FrontierCode v1 | **38.8** | 25.5 | n/a |
| GDPval-AA v2 (Elo) | **1618** | 1509 | 1357 |
| AutomationBench | 13.5 | 12.9 | **14.5** |
| HealthBench Professional | **57.8** | 51.8 | n/a |
Sonnet 5 wins most head-to-heads against GPT-5.5 and loses the one that gets quoted most: Terminal-Bench 2.1, where GPT-5.5 driving the Codex CLI scores 83.4 against 80.4. Against Gemini 3.5 Flash the sweep is broader, with AutomationBench the single exception. Note that AutomationBench scores are low across the board; these are unsaturated benchmarks, not solved ones\[1].
## Where it ties Opus 4.8 and where it trails
**Ties or edges ahead:** GDPval-AA v2 knowledge work (1618 vs 1615), agentic search on BrowseComp at comparable accuracy for a given task cost, Real-World Finance v2 (1219 vs 1222), and the AA-Briefcase professional-work Elo (1393 vs 1352)\[1].
**Trails:** SWE-bench Pro (63.2 vs 69.2), Terminal-Bench 2.1 (80.4 vs 82.7), OSWorld (81.2 vs 83.4), HLE no-tools (43.2 vs 49.8), CursorBench (61.2 vs 63.8), Toolathlon (54.3 vs 59.9), and USAMO 2026 (79.5 vs 96.7)\[1].
The pattern is consistent. On open-ended knowledge work and agentic search, Sonnet 5 is effectively Opus-class. On the hardest coding, computer use, and formal math, Opus 4.8 keeps a real but modest lead. For workloads running many agent turns where cost dominates, Sonnet 5 is the rational default and Opus is the escalation path for the hard tail.
## How to use Claude Sonnet 5: four changes that break old code
Sonnet 5 adopts the Opus 4.7/4.8 request surface, so migrating is more than a string swap\[1].
**Adaptive thinking is on by default.** Sonnet 4.6 ran without thinking unless asked. Sonnet 5 thinks adaptively out of the box, deciding per request how much reasoning a task warrants. You can still disable it explicitly when you need raw speed.
**`budget_tokens` is removed.** Manual thinking budgets now return a 400. Control depth through effort levels instead.
**Sampling parameters are rejected.** Non-default `temperature`, `top_p`, and `top_k` return a 400. Steer with prompting.
**The tokenizer is about 30% denser.** Sonnet 5 shares the tokenizer introduced with Opus 4.7, which turns the same text into roughly 30% more tokens. Per-token pricing is unchanged, so an equivalent request can cost more than it did on Sonnet 4.6, and a `max_tokens` limit tuned for the old tokenizer can now truncate output mid-thought.
That last one is the expensive surprise. Re-measure real spend on your own prompts rather than assuming the sticker price maps cleanly from Sonnet 4.6.
One capability addition worth noting: Sonnet 5 is the first Sonnet-tier model with a high-resolution image tier, up to 2576 pixels on the long edge and 4784 visual tokens, against the old 1568-pixel cap. It activates automatically, which matters for dense document, chart, and screenshot understanding\[1].
## Pricing
| Item | Introductory (through Aug 31, 2026) | Standard (from Sep 1, 2026) |
| ---------- | ----------------------------------- | --------------------------- |
| Input | $2 / MTok | $3 / MTok |
| Output | $10 / MTok | $15 / MTok |
| Cache read | \~$0.20 / MTok | \~$0.30 / MTok |
| Batch | $1 / $5 | 50% of standard |
Against Opus 4.8 at $5 / $25, Sonnet 5 is roughly a third of the cost at introductory rates and about 60% at standard rates\[1]. Layer in the effort dial and prompt caching and the gap widens for a well-tuned agent workload.
Two caveats. The denser tokenizer means you should recompute real spend rather than trusting the sticker comparison. And Priority Tier is not available on Sonnet 5\[1].
## Safety behavior your code has to handle
Sonnet 5 is the first Sonnet-tier model shipping with real-time cybersecurity safeguards. Requests touching prohibited or high-risk cyber and biology topics can be refused, returned as a **successful HTTP 200 with a refusal stop reason rather than an error**\[1].
If you are building security tooling or life-sciences workflows where benign-but-adjacent requests might trip a classifier, handle that stop reason explicitly instead of assuming every 200 carries usable content. Anthropic also reports lower hallucination and sycophancy rates than Sonnet 4.6, which matters wherever output is trusted downstream without a human check\[1].
## Who should use Claude Sonnet 5
It is the rational default for high-volume agentic workloads where cost per turn compounds: coding agents, browser and terminal automation, retrieval and research loops. It fits product builders embedding an LLM into a shipping feature, and long-context document work where the 1M window and low per-token price make bulk processing viable.
Reach for Opus 4.8 or Opus 5 instead when the task is in the hard tail: the most demanding coding, pixel-accurate computer use, or formal mathematics, where the flagship's few-point lead is worth the premium.
## Running a Sonnet-tier model on reAPI
To be straightforward about coverage: **Claude Sonnet 5 is not on the reAPI gateway today.** The Sonnet-tier model available is Claude Sonnet 4.6, and the Claude models on the gateway are Opus 5, Opus 4.8, Opus 4.7, Sonnet 4.6, and Fable 5.
That matters less than it sounds like, because of where the prices actually land. Sonnet 5 lists at $3 / $15 from September. On reAPI, **Claude Opus 5 runs $2.40 / $12.00** and Sonnet 4.6 runs the same. So the tier-above model costs less on the gateway than Sonnet 5's standard list price, and Opus 5 is the model Anthropic describes as delivering frontier intelligence at half the cost of Fable 5.
```python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_REAPI_KEY",
base_url="https://api.reapi.ai/v1",
)
resp = client.chat.completions.create(
model="claude-opus-5",
messages=[{"role": "user", "content": "Plan and run this multi-step refactor."}],
max_tokens=16000,
stream=True,
)
```
The endpoint is OpenAI-compatible and the native Anthropic `/v1/messages` surface is available too. Rates and request shapes are on [reapi.ai/models](/models), and the Opus 5 reference is at [reapi.ai/docs/claude-opus-5](/docs/claude-opus-5).
If your evaluation specifically requires Sonnet 5's behavior rather than its price point, call it through Anthropic directly.
## FAQ
### Is Claude Sonnet 5 better than Opus 4.8?
On open-ended knowledge work and agentic search they are effectively tied, and Sonnet 5 nudges ahead on GDPval-AA v2 (1618 vs 1615). On the hardest coding, computer use, and formal math, Opus 4.8 keeps a modest lead\[1].
### How much does Claude Sonnet 5 cost?
$2 per million input tokens and $10 per million output through August 31, 2026, rising to $3 / $15 on September 1. Prompt caching cuts cached-context reads by roughly 90%, and the Batch API halves the rate\[1].
### What is the `xhigh` effort level?
An effort setting between `high` and `max`, tuned for long-horizon agentic and coding tasks. Sonnet 5 is the first Sonnet-tier model to support the full ladder\[1].
### Why did my costs go up after migrating from Sonnet 4.6?
The tokenizer is about 30% denser, so the same text becomes roughly 30% more tokens at unchanged per-token prices. Recount your prompts and raise `max_tokens` so long outputs are not truncated\[1].
### Does Claude Sonnet 5 support a 1M-token context window?
Yes, 1M by default at the standard per-token rate with no long-context surcharge, and up to 128K output tokens\[1].
### What breaks when migrating from Sonnet 4.6?
`budget_tokens` returns a 400, non-default sampling parameters return a 400, adaptive thinking is now the default, and the denser tokenizer changes both cost and truncation behavior\[1].
### Can Claude Sonnet 5 control a computer?
Yes. It supports tool use, web search, web fetch, code execution, and computer use, scoring 81.2% on OSWorld-Verified against Opus 4.8's 83.4%\[1].
### Is Claude Sonnet 5 available on reAPI?
Not currently. The gateway carries Claude Opus 5, Opus 4.8, Opus 4.7, Sonnet 4.6, and Fable 5. Opus 5 on reAPI costs less per token than Sonnet 5's standard list rate.
## Choosing between a tier and a price point
Claude Sonnet 5 delivers on its central promise. For coding and agentic work it lands close enough to Opus 4.8 that, at a third of the cost during the introductory window, it becomes the sensible default and reserves the flagship for the hard tail. It wins most head-to-heads against GPT-5.5 and Gemini 3.5 Flash, with honest exceptions on Terminal-Bench, AutomationBench, and formal math.
Two things decide whether that arithmetic survives contact with your workload. The September price step takes it from a third of Opus to about 60%, and the denser tokenizer means the sticker comparison understates real spend. Measure both on your own prompts before committing. That is how to use Claude Sonnet 5 without discovering the tokenizer in your invoice, and it is also why the tier label matters less than the number you actually pay per completed task.
## References
1. Anthropic. *Introducing Claude Sonnet 5.* Retrieved July 2026 from [anthropic.com/news/claude-sonnet-5](https://www.anthropic.com/news/claude-sonnet-5)
2. Anthropic. *Models overview — specs and pricing for all current Claude models.* Retrieved July 2026 from [platform.claude.com/docs/en/about-claude/models/overview](https://platform.claude.com/docs/en/about-claude/models/overview)
3. Anthropic. *Pricing — token, cache, and batch rates.* Retrieved July 2026 from [platform.claude.com/docs/en/about-claude/pricing](https://platform.claude.com/docs/en/about-claude/pricing)
### Further reading
* reAPI. *How to use Claude Opus 5.* [reapi.ai/blog/how-to-use-claude-opus-5](/blog/how-to-use-claude-opus-5)
* reAPI. *How to use Claude Opus 4.8.* [reapi.ai/blog/how-to-use-claude-opus-4-8](/blog/how-to-use-claude-opus-4-8)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# How to Use Codex to Generate Slide Decks From Research (https://reapi.ai/blog/how-to-use-codex-to-generate-slides)
Codex is most useful for slide decks when the input is structured. Not "make me a PowerPoint," but a folder holding research notes, competitor analysis, product positioning, screenshots, and an existing brand template, turned into a deck that follows the same rules every time.
OpenAI documents this directly: Codex can manipulate `.pptx` files, generate visuals, apply repeatable layout rules, update existing decks, and build new ones, using the Slides skill for PowerPoint editing and the ImageGen skill for illustrations and cover art\[1].
Learning how to use Codex to generate slides is mostly about one discipline: making it propose the outline before it writes a single slide, and telling it which text it is not allowed to rewrite.
## TL;DR
* **Start from an existing brand deck**, and make Codex summarize the visual rules before generating anything\[1].
* **Demand an outline first.** Slide-by-slide title, core message, source evidence, recommended visual, presenter note. Approve it before any PPTX gets written.
* **Pin the claims it must not touch.** Approved messaging, pricing language, legal claims. Tell it to mark gaps as "needs verification" rather than fill them in.
* **Keep text editable and charts native.** Do not let it rasterize slides.
* **Render and inspect before delivery.** A deck can be content-correct and visually broken.
* **Use ImageGen for cover art and concepts**, never for product UI, real logos, or data charts.
## The seven-step workflow

### 1. Collect structured inputs
Codex works from source material, so the quality of the deck tracks the organization of the folder. Product brief, research notes, brand template, screenshots, competitor notes, interview summaries, campaign metrics, messaging framework, approved claims, logo assets.
### 2. Start from a brand template
OpenAI's guidance is to start from an existing deck where possible\[1]. Make the model read it first:
```text
First inspect the existing brand deck. Identify the slide size, title style,
body-text style, color palette, logo placement, image treatment, chart style,
and spacing rules. Summarize these rules before generating new slides.
```
### 3. Get the outline before the deck
This is the step people skip, and it is the one that decides whether the deck argues anything.
```text
Before creating the deck, propose a slide-by-slide outline. For each slide,
include the title, core message, source evidence, recommended visual, and
presenter note. Do not generate the PPTX until I approve the outline.
```
Narrative flow is not generic. A market research deck, a launch deck, and a sales deck should not share a structure, and you find out whether the model understood that at the outline stage, not after it has rendered forty slides.
### 4. Pin the claims it may not rewrite
Decks carry approved messaging, legal claims, and pricing language. A model that paraphrases those has created a compliance problem, not a draft.
```text
Preserve all approved product claims exactly as written. Do not invent customer
quotes, pricing, performance numbers, or competitor claims. If evidence is
missing, mark it as needs verification instead of filling it in.
```
The instruction that does the work is the last clause. Given a gap, a model will fill it plausibly unless told what to do instead.
### 5. Generate editable slides with native charts
```text
Create the deck as an editable PPTX. Keep text as editable PowerPoint text.
Use native charts for simple bar, line, pie, and histogram visuals when
practical. Do not rasterize full slides. Use image assets only for screenshots,
illustrations, and complex visuals.
```
Teams revise wording many times. A rasterized slide is a dead end.
### 6. Use ImageGen deliberately
Define one visual direction and reuse it, rather than letting each slide drift.
```text
Define a consistent visual direction for this deck. Use clean SaaS-style
illustrations, realistic product context, and a polished B2B visual tone.
Save the image prompts so future slides can match the same direction.
```
**Generate**: cover slides, product-concept illustrations, market-trend visuals, customer-journey diagrams, section dividers.
**Do not generate**: exact product UI screenshots, legal or compliance claims, real customer logos, or data charts that should stay native and editable.
### 7. Render and validate
```text
Render the deck to slide images. Review every slide for clipped text, overlap,
inconsistent spacing, unreadable screenshots, font substitution, layout drift,
and visual-style mismatch. Fix all issues before saving the final PPTX.
```
A deck can be entirely correct and still look broken. Rendering to images is how the model sees what you would see.
## Deck types this works well for
The pattern fits any deck whose structure is stable while the content changes\[1].
**Market research.** Market overview, growth drivers, customer segments, buyer pain points, competitor landscape, pricing patterns, adoption barriers, opportunity areas, recommendations.
**Product launch.** Product overview, customer problem, target audience, positioning statement, key features, differentiation, launch narrative, messaging pillars, campaign plan, timeline, call to action.
**Product demo.** Demo objective, user problem, workflow overview, step-by-step screenshots, feature highlights, before-and-after, customer value, next steps.
**Competitive analysis.** Competitor overview, feature comparison, pricing comparison, messaging comparison, strengths and weaknesses, positioning map, differentiation strategy, recommended talk tracks.
**Sales enablement.** Buyer persona, pain points, discovery questions, value proposition, objection handling, competitor responses, proof points, demo flow, closing talk track.
## A complete worked prompt
```text
Use the Slides and ImageGen skills to generate a 12-slide market research deck.
Inputs:
- the attached brand PowerPoint template
- market research notes
- competitor analysis
- customer interview summaries
- product screenshots
First inspect the brand template and summarize the visual rules.
Then propose a slide-by-slide outline before generating the deck.
Include market overview, customer segments, pain points, competitor landscape,
positioning gaps, product opportunity, and recommendations.
Preserve approved claims exactly. Mark any missing evidence as needs
verification rather than filling it in. Keep all text editable and use native
charts. Render to images and fix layout issues before saving the final PPTX.
```
Every clause in that prompt maps to a failure this workflow is designed to prevent.
## Where the approach breaks down
Worth stating plainly, because the failure modes are predictable.
**It will not invent a strategy.** Codex arranges and renders the argument you supply. Given thin inputs it produces a well-formatted deck that says nothing, which is harder to spot than an ugly one.
**Unverified numbers are the main risk.** A model asked for a market-size slide without a source will produce a confident figure. The "mark as needs verification" instruction is the mitigation, and it only works if you actually read for those markers.
**Visual consistency degrades across long decks** unless you pin the direction once and reference it, which is what step 6 is for.
**Brand nuance stays human.** The model can match a palette and a type scale. Whether the deck sounds like your company is a judgment call.
## Building this into your own product
The workflow above runs inside Codex, on your machine, against your folder. That is the right shape when the deck is yours.
It is the wrong shape for a different job: putting deck generation inside software you ship. A tool that turns a customer's report into slides, an internal service that assembles weekly reviews, a feature that drafts sales collateral from a CRM record. None of those are a developer running a CLI.
That job needs the models directly. The reasoning that plans the outline and the image model that produces the cover art are both API calls, and reAPI exposes both behind one key:
```python
from openai import OpenAI
client = OpenAI(api_key="YOUR_REAPI_KEY", base_url="https://api.reapi.ai/v1")
outline = client.chat.completions.create(
model="gpt-5.6-terra",
messages=[{"role": "user", "content": "Propose a 12-slide outline from this research. Mark unsupported claims."}],
max_tokens=16000,
)
```
Cover art and section visuals come from the image side of the same account, at [reapi.ai/models/gpt-image-2](/models/gpt-image-2), which handles in-image text well enough for titles and labels. The full catalog is at [reapi.ai/models](/models).
Use Codex for your own decks. Use the API when the deck belongs to someone else.
## FAQ
### Can Codex edit an existing PowerPoint file?
Yes. OpenAI documents Codex manipulating `.pptx` files directly, updating existing presentations as well as building new ones, through the Slides skill\[1].
### How do I keep Codex from inventing statistics?
Instruct it to preserve approved claims verbatim and to mark missing evidence as needing verification rather than filling it in. Then read for those markers before the deck ships.
### Should I let Codex generate charts?
Use native PowerPoint charts for simple bar, line, pie, and histogram visuals so they stay editable. Reserve generated images for illustrations and complex visuals\[1].
### Why ask for an outline first?
Because narrative structure is where decks fail, and fixing it at the outline stage costs one message instead of a full regeneration.
### What should ImageGen not be used for?
Exact product UI, legal or compliance claims, real customer logos without permission, and data charts that should remain native and editable.
### Does this work for non-marketing decks?
Yes. The workflow is structure-driven, so it applies to any deck type whose format is stable while the content changes.
### How do I build slide generation into my own application?
Call the models through an API rather than the CLI. Outline planning is a chat completion; cover art and diagrams are image generations.
## Making the outline the checkpoint
The temptation with a capable agent is to describe the deck and accept what comes back. That produces a plausible artifact whose argument nobody chose.
The discipline that makes this workflow reliable is small: read the brand rules back before generating, approve an outline before rendering, pin the sentences that cannot change, and look at rendered images before calling it done. Learning how to use Codex to generate slides is mostly learning to put a human checkpoint at the outline, where changing your mind is still cheap.
## References
1. OpenAI. *Platform documentation — Codex, skills, and model capabilities.* Retrieved July 2026 from [platform.openai.com/docs](https://platform.openai.com/docs)
2. OpenAI. *Codex — open-source repository and usage documentation.* Retrieved July 2026 from [github.com/openai/codex](https://github.com/openai/codex)
### Further reading
* reAPI. *How to use GPT Image 2.* [reapi.ai/blog/how-to-use-gpt-image-2](/blog/how-to-use-gpt-image-2)
* reAPI. *How to use GPT-5.6.* [reapi.ai/blog/how-to-use-gpt-5-6](/blog/how-to-use-gpt-5-6)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# How to Use Gemini 3.6 Flash: Speed, Price, and Limits (https://reapi.ai/blog/how-to-use-gemini-3-6-flash)
On July 21, 2026, Google shipped three Gemini models in one announcement: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a restricted security specialist called Gemini 3.5 Flash Cyber\[1]. No Pro model, no keynote, no claim of a new intelligence frontier. The entire release argues one thing: teams running agents in production need fewer tokens and lower latency per task, not a bigger model.
That framing tells you who it is for. Learning how to use Gemini 3.6 Flash well is mostly about the compounding math, because Google cut the output price and the model also emits fewer tokens to do the same work. Those two effects multiply. This guide covers the official benchmark data, where 3.6 Flash honestly loses to GPT-5.6 Luna and Claude Sonnet 5, the four API changes that break carried-over code, and which of the three models actually fits your workload.
## TL;DR: the July 2026 Gemini Flash lineup
| Model | List price per 1M tokens | Positioning |
| -------------------------- | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| **Gemini 3.6 Flash** | $1.50 input / $7.50 output | Workhorse for coding, agents, and knowledge work. Roughly 17% fewer output tokens than 3.5 Flash\[4] |
| **Gemini 3.5 Flash-Lite** | $0.30 input / $2.50 output | Fastest model in the 3.5 series at 350 output tokens per second\[4]. High-volume pipelines |
| **Gemini 3.5 Flash Cyber** | Not publicly priced | Vulnerability finder inside DeepMind's CodeMender agent. Governments and vetted partners only\[3] |
Two roadmap notes hide in the same post: Gemini 3.5 Pro is still in partner testing with no public date, and Google confirmed it has started its most ambitious pre-training run yet, for Gemini 4\[1].
On reAPI, Gemini 3.6 Flash runs $1.20 input and $6.00 output per million tokens, which is 80% of Google's published rate. Flash-Lite and Flash Cyber are not on the gateway.
## The efficiency play

Gemini 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output, with cached input at $0.15 per million\[1]. Its predecessor charged $9 per million output, so Google cut the output price by roughly 17% while improving benchmark scores across the board.
The sticker price is the smaller half of the story. Artificial Analysis measured 3.6 Flash consuming about 17% fewer output tokens than 3.5 Flash on identical workloads\[4], and Google reported token reductions of up to 65% on Datacurve's DeepSWE benchmark\[1]. The model also takes fewer reasoning steps and fewer tool calls to finish multi-step work.
For an agent that loops dozens of times per task, those effects stack: cheaper tokens, fewer tokens, fewer loops. A per-token price comparison understates the gap.
### Official benchmarks against its own predecessors

Google's launch chart compares 3.6 Flash against 3.5 Flash and the older 3.1 Pro on four agentic benchmarks\[1].
| Benchmark | Gemini 3.1 Pro | Gemini 3.5 Flash | Gemini 3.6 Flash |
| ------------------------------------------------ | -------------- | ---------------- | ---------------- |
| DeepSWE v1.1 (long-horizon software engineering) | 12% | 37% | **49%** |
| MLE-Bench (machine learning engineering) | 42.6% | 49.7% | **63.9%** |
| GDPVal-AA v2 (knowledge work, Elo) | 965 | 1349 | **1421** |
| OSWorld-Verified (computer use) | 76.2% | 78.4% | **83.0%** |
Three more gains over the prior generation sit in Google's evaluation methodology: SWE-Bench Pro at 58.7% against 55.1%, Terminal-Bench 2.1 at 78.0% against 76.2%, and the GDM-MRCR v2 long-context retrieval test at full 1M-token depth, where 3.6 Flash roughly doubles both predecessors at 54.0% against under 27%\[5].
From the model card: the knowledge cutoff moved to March 2026 from January 2025, the context envelope stays at 1M input and 64K output tokens, and the model accepts text, images, audio, and video natively\[2].
## Where 3.6 Flash wins and where it loses
Google's chart only compares Gemini against Gemini. Here is the cross-vendor picture using mid-2026 published results\[5].
| Benchmark | Gemini 3.6 Flash | GPT-5.6 Luna | Grok 4.5 | Claude Sonnet 5 |
| ------------------ | ---------------- | ------------ | --------- | --------------- |
| DeepSWE v1.1 | 49% | **67%** | n/a | n/a |
| Terminal-Bench 2.1 | 78.0% | **84.7%** | n/a | n/a |
| SWE-Bench Pro | 58.7% | n/a | **64.7%** | n/a |
| MLE-Bench | 63.9% | n/a | n/a | **66.9%** |
| GDPVal-AA v2 (Elo) | 1421 | n/a | n/a | **1607** |
| OSWorld-Verified | **83.0%** | n/a | n/a | n/a |
No model sweeps the board, and 3.6 Flash is not trying to. It leads on OSWorld-Verified computer use, on both CharXiv chart-reasoning variants, and on both GDM-MRCR long-context tests\[5]. Those are the multimodal and long-context lanes Google has consistently owned.
It clearly trails GPT-5.6 Luna on agentic coding, on both DeepSWE and Terminal-Bench. It trails Claude Sonnet 5 on ML engineering and knowledge work, and Sonnet 5's 1607 Elo on GDPVal-AA is a different weight class from 3.6 Flash's 1421\[5].
The independent read agrees. Artificial Analysis scores 3.6 Flash at 50 on its Intelligence Index, rank 21 of 187 models tracked. That is not frontier intelligence, and Google is not claiming it is. The same evaluation ranks it first on output speed in its class at 303.6 tokens per second and describes its token usage as fairly concise, while noting that time to first token, at 11.54 seconds under high thinking effort, sits at the slow end of its tier\[4].
One methodology caveat before anyone quotes the cross-vendor table: vendor scores are generally reported at each model's maximum thinking configuration, and agent harnesses differ between vendors. Treat small gaps as noise and only large ones as signal.
The practical read. If the job is raw coding-agent horsepower, GPT-5.6 Luna and Claude Sonnet 5 still hold the top of the board, at several times the price. If the job is computer use, document-heavy multimodal work, or long-context retrieval at scale, 3.6 Flash offers the best score per dollar in its bracket.
## Gemini 3.5 Flash-Lite
The second release targets the opposite end of the curve. Flash-Lite runs at 350 output tokens per second as measured by Artificial Analysis, priced at $0.30 per million input and $2.50 per million output, which is one-fifth and one-third of 3.6 Flash's rates\[1]\[4].

The generational jump is larger here than anywhere else in the release\[1]:
* Terminal-Bench 2.1: 54.0% against 31.0% for 3.1 Flash-Lite, a 23-point move in agentic terminal coding.
* GDPVal-AA v2: 1140 Elo against 642, close to double on real-world knowledge work.
* GDM-MRCR v2 long context: 72.2% against 60.1%.
The comparison worth pausing on is against the older, bigger 3 Flash: Flash-Lite now wins SWE-Bench Pro at 54.2% against 49.6% and OSWorld-Verified at 74.0% against 65.1%\[1]. The cheapest current model beats the previous generation's mid-tier on coding and computer use.
Google positions it for high-throughput work: agentic search, bulk document processing, classification, subagent fan-out, with thinking levels configurable from minimal, the default, up to high\[1].
## Gemini 3.5 Flash Cyber
The third model is the one Google will not sell you. Flash Cyber is a version of 3.5 Flash fine-tuned to find, validate, and patch security vulnerabilities, and it ships exclusively inside CodeMender, DeepMind's code-security agent, in a limited pilot for governments and trusted partners\[3].
The fine-tuning leans on assets few labs have: the OSV.dev database of more than 700,000 open-source vulnerabilities, over a decade of OSS-Fuzz results, and security-workflow data from Google-scale codebases including Chromium\[3]. Inside CodeMender, multiple Flash Cyber sub-agents analyze different code paths and merge findings into one report, which is the architectural bet: a cheap model called up to five times competing with single calls to much larger ones.

| System | CyberGym score |
| -------------------------------------------------- | -------------- |
| GPT-5.5-Cyber in OpenAI agent | **85.6%** |
| Claude Mythos 5 in Anthropic agent | 83.8% |
| GPT-5.6 Sol in OpenAI agent | 83.6% |
| Gemini 3.5 Flash Cyber in CodeMender (max 5 calls) | 83.2% |
| Claude Mythos Preview in Anthropic agent | 83.1% |
Read that honestly: Flash Cyber does not top the chart. OpenAI's dedicated GPT-5.5-Cyber leads by 2.4 points and Mythos 5 edges it by 0.6. Google's claim is competitive performance from a Flash-sized model, which is a cost claim rather than a crown\[3].
Where it does win is in Google's own deeper evaluations. On a V8 JavaScript engine test it confirmed 55 unique real issues against 47 for stock 3.5 Flash and 36 for Claude Opus 4.6, including 10 that no other model found\[3]. Google's Cloud Vulnerability Research team reports using it to find a memory-corruption vulnerability in a production service and build a working exploit chain, bypassing ASLR and W^X, within two hours\[3].
Which explains the leash. Google states the capability is dual-use, limits access to vetted defenders, deploys it only inside CodeMender's reporting workflow, and requires human approval before patches ship\[3].
## How to use Gemini 3.6 Flash: four API changes that break old code
The model card describes 3.6 Flash as based directly on 3.5 Flash, the same natively multimodal sparse reasoning foundation and the same 1M/64K envelope, with the gains coming from post-training rather than a new base model\[2]. The request surface is where the work is.
1. **Model IDs are `gemini-3.6-flash` and `gemini-3.5-flash-lite`**\[1].
2. **`thinking_budget` is replaced by a `thinking_level` string enum.** 3.6 Flash defaults to `medium`, Flash-Lite to `minimal`\[1].
3. **The classic sampling knobs are gone.** `temperature`, `top_p`, and `top_k` were removed and should be deleted from request payloads\[1].
4. **`candidate_count` is unsupported**, and prefilled model turns, meaning conversations that end on a non-empty model role, are rejected\[1].
Both public models expose the full built-in tool suite, including computer use, search grounding, and code execution\[1]. Coming from 3.5 Flash, the migration is a model-string change plus deleting deprecated parameters. There is no new SDK surface to learn.
## Which model should you pick
A routing heuristic straight off the published numbers:
* **Default to 3.6 Flash** for agent loops involving coding, computer use, document parsing, or anything multi-step. The token-efficiency gains compound exactly where agent costs explode.
* **Route bulk, latency-sensitive work to Flash-Lite**: extraction, classification, search summarization, subagent fan-out. Raise `thinking_level` only where sampled quality checks say you must.
* **Keep frontier coding on frontier models.** If DeepSWE-class autonomy is the job, the honest numbers say GPT-5.6 Luna and Claude Sonnet 5 still finish tasks 3.6 Flash cannot, and the premium may be worth it on low-volume, high-stakes runs.
* **Flash Cyber is not a choice you get to make** unless you are a government or an invited partner.
## Calling Gemini 3.6 Flash alongside your other models
Notice what that routing list actually says. Three of the four bullets send work somewhere other than Gemini. The honest conclusion from Google's own benchmark table is that 3.6 Flash should win some of your traffic and lose the rest, and the split moves every time a vendor ships.
reAPI exposes Gemini 3.6 Flash through an OpenAI-compatible `/v1/chat/completions` endpoint, so the client already calling GPT and Claude calls it too. Note the wire model id keeps Google's dot:
```python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_REAPI_KEY",
base_url="https://api.reapi.ai/v1",
)
resp = client.chat.completions.create(
model="gemini-3.6-flash",
messages=[{"role": "user", "content": "Summarize this report and pull out the numbers."}],
max_tokens=16000,
stream=True,
)
```
Rates on the gateway are $1.20 per million input tokens and $6.00 per million output, which is 80% of Google's published $1.50 / $7.50. The gateway also exposes a native Gemini endpoint type for this model, so SDKs written against Google's own format work without a rewrite. Current rates and the request surface live on [reapi.ai/models/gemini-3-6-flash](/models/gemini-3-6-flash) and [reapi.ai/docs/gemini-3-6-flash](/docs/gemini-3-6-flash).
To be clear about coverage: Flash-Lite and Flash Cyber are not on the gateway. Flash Cyber is not available to anyone outside Google's pilot regardless of how you call it.
## FAQ
### Is Gemini 3.6 Flash actually cheaper than 3.5 Flash?
Yes, on output: $7.50 against $9 per million tokens, roughly a 17% cut. Because the model also emits about 17% fewer tokens for the same work, effective cost per completed task drops more than the sticker price suggests. Cached input is $0.15 per million\[1]\[4].
### What is the Gemini 3.6 Flash context window?
1M input tokens and 64K output tokens, with text, image, audio, and video accepted natively\[2].
### Do I need to change code to upgrade from 3.5 Flash?
Minimally. Swap the model string to `gemini-3.6-flash`, replace `thinking_budget` with `thinking_level`, and remove `temperature`, `top_p`, `top_k`, and `candidate_count` from requests\[1].
### Is Gemini 3.6 Flash better than GPT-5.6 Luna or Claude Sonnet 5?
Not at agentic coding. Luna leads DeepSWE 67% to 49% and Terminal-Bench 84.7% to 78.0%, and Sonnet 5 leads MLE-Bench and GDPVal-AA\[5]. Gemini 3.6 Flash wins computer use, chart reasoning, and long-context retrieval, at a fraction of the price.
### Can I access Gemini 3.5 Flash Cyber?
Not unless you are a government agency or an invited partner in the CodeMender pilot\[3].
### What happened to Gemini 3.5 Pro?
It remains in partner testing with no announced date, while Google has confirmed that Gemini 4 pre-training is already underway\[1].
### How fast is Gemini 3.6 Flash?
Artificial Analysis ranks it first on output speed in its class at 303.6 tokens per second, but time to first token at high thinking effort is 11.54 seconds, which is slow for its tier\[4]. Stream anything a person watches.
### How do I call Gemini 3.6 Flash with the OpenAI SDK?
Point the OpenAI client at `https://api.reapi.ai/v1` and set `model` to `gemini-3.6-flash`, keeping the dot in the id. The endpoint is OpenAI-compatible, so no SDK rewrite is needed.
## Reading a release that competes on unit economics
This launch is worth understanding as a pricing move rather than a capability move. Google shipped no Pro tier, made no frontier claim, and put its engineering into the two numbers that decide what an agent costs: price per token and tokens per task. Both went down about 17%, and they multiply. Against that, the cross-vendor table shows 3.6 Flash losing agentic coding to GPT-5.6 Luna and knowledge work to Claude Sonnet 5, which is the correct trade for a model in this bracket rather than a flaw.
The practical version: default agent loops to it, send bulk extraction lower, keep frontier coding on frontier models, and measure effective cost per completed task instead of per-token rates, because that is the only number in which the token-efficiency gain shows up. That is how to use Gemini 3.6 Flash without paying frontier prices for work it already does well, or expecting it to win the rows it plainly loses.
## References
1. Google. *Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber.* Retrieved July 2026 from [blog.google/technology/google-deepmind](https://blog.google/technology/google-deepmind/)
2. Google DeepMind. *Gemini 3.6 Flash model card.* Retrieved July 2026 from [deepmind.google/models/gemini](https://deepmind.google/models/gemini/)
3. Google DeepMind. *Introducing Gemini 3.5 Flash Cyber and CodeMender.* Retrieved July 2026 from [deepmind.google/discover/blog](https://deepmind.google/discover/blog/)
4. Artificial Analysis. *Gemini 3.6 Flash — intelligence, performance, and price analysis.* Retrieved July 2026 from [artificialanalysis.ai/models](https://artificialanalysis.ai/models)
5. OfficeChai. *Gemini 3.6 Flash beats Gemini 3.5 Flash and Gemini 3.1 Pro on most benchmarks.* Retrieved July 2026 from [officechai.com](https://officechai.com/)
### Further reading
* reAPI. *How to use Claude Opus 5.* [reapi.ai/blog/how-to-use-claude-opus-5](/blog/how-to-use-claude-opus-5)
* reAPI. *Gemini 3.6 Flash endpoint reference.* [reapi.ai/docs/gemini-3-6-flash](/docs/gemini-3-6-flash)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# How to Use GPT-5.5: Agentic Strengths and Its One Flaw (https://reapi.ai/blog/how-to-use-gpt-5-5)
OpenAI released GPT-5.5 on April 23, 2026, describing it as its first fully retrained base model since GPT-4.5\[1]. Within twenty-four hours it took the top spot on the Artificial Analysis Intelligence Index.
Knowing how to use GPT-5.5 well starts with an uncomfortable number that did not make the launch post. On the AA-Omniscience benchmark it posts the highest accuracy of any model tested at 57%, and an **86% hallucination rate on the questions it gets wrong**\[2]. Claude Opus 4.7 hallucinates at 36% on the same measure. When GPT-5.5 is wrong, it is confidently wrong.
That single fact determines how you should deploy it, and this guide is organized around it.
## TL;DR
* **Trained end-to-end as an agent**, not a chat model with tools bolted on. OpenAI concentrated training on agentic coding, computer use, knowledge work, and early scientific research\[1].
* **Terminal-Bench 2.0 at 82.7%**, roughly 13 points ahead of Claude Opus 4.7 and Gemini 3.1 Pro\[1].
* **The price rise is smaller than it looks.** List went from $2.50/$15 to $5/$30, but roughly 40% fewer output tokens on Codex tasks puts the effective per-task rise near 20%\[1].
* **On reAPI it runs $4.00 / $24.00** per million tokens, 80% of the published rate.
* **The hallucination profile is the catch.** Highest accuracy, and the highest confidence when wrong\[2].
* **Coding is still a draw.** SWE-bench Pro goes to Opus 4.7, 64.3% against 58.6%\[1].
## What GPT-5.5 is
GPT-5.5 is an agentic model first and a chat model second. It ships in two variants sharing a 1M-token context window and text-plus-vision input\[1]:
* **GPT-5.5**, with reasoning effort levels `xhigh`, `high`, `medium`, `low`, and non-reasoning.
* **GPT-5.5 Pro**, a higher-compute variant pushing harder on long-horizon tasks.
## Where it leads

| Benchmark | GPT-5.5 | GPT-5.4 | Claude Opus 4.7 | Gemini 3.1 Pro |
| --------------------- | --------- | ------- | --------------- | -------------- |
| Terminal-Bench 2.0 | **82.7%** | 75.1% | 69.4% | 68.5% |
| Expert-SWE (internal) | **73.1%** | 68.5% | n/a | n/a |
| GDPval (wins or ties) | **84.9%** | 83.0% | 80.3% | 67.3% |
| OSWorld-Verified | **78.7%** | 75.0% | 78.0% | n/a |
| BrowseComp | 84.4% | 82.7% | 79.3% | **85.9%** |
| FrontierMath Tier 1–3 | **51.7%** | 47.6% | 43.8% | 36.9% |
| CyberGym | **81.8%** | 79.0% | 73.1% | n/a |
Three things worth flagging\[1].
**Terminal-Bench 2.0 is the decisive win.** It is the closest public proxy for finishing a real engineering task end to end without hand-holding, and a 13-point lead at this end of the curve is unusual.
**GDPval matters more than it looks.** It tests output quality across 44 occupations against industry professionals, and winning or tying 84.9% of the time is the actual commercial pitch: knowledge-work parity with a domain expert.
**SWE-bench Pro is a loss, and the community disputes the benchmark.** GPT-5.5 scores 58.6% against Opus 4.7's 64.3%, with open questions about training-set memorization on this particular test. Treat the Opus lead with some skepticism, but do not assume the coding gap closed.

## What is genuinely new
The architectural detail that matters is not in any benchmark table: GPT-5.5 was trained end-to-end as an agent, with native expectations around multi-step tool use, screen reading, and verification\[1].
**Multi-step tool calls without supervision.** OpenAI demonstrated the model completing over 1,000 sequential tool calls without intervention. That is where the Terminal-Bench and OSWorld numbers come from.

**Self-checking before submission.** The model verifies its own output before returning it. CodeRabbit reported expected issue detection jumping from 58.3% to 79.2% in their review benchmark, with shorter responses and a stronger bias toward small workable changes\[1].
**Computer use in Codex.** Screen reading is exposed through Codex, letting the model see and interact with arbitrary desktop applications.
## The hallucination problem
This is the section the launch coverage buried, and the one that should shape your deployment.
On AA-Omniscience, GPT-5.5 posts the highest accuracy of any model tested at 57%, with an 86% hallucination rate on the questions it gets wrong. Claude Opus 4.7 sits at 36% and Gemini 3.1 Pro Preview at 50%\[2].
In practice: when GPT-5.5 is wrong, it does not flag uncertainty. For agentic workflows where the model takes real actions on real systems, that is a non-trivial failure mode.
The likeliest explanation is that OpenAI optimized for agentic confidence, training the model to act rather than defer. That trade is rational for tool-heavy workflows and costly for pure question answering.
Two mitigations actually work:
**Raise reasoning effort.** Use `xhigh` for anything where output quality matters more than latency. Higher effort visibly reduces hallucination on the tasks tested.
**Ground it in tools.** File reads, test runs, search results. GPT-5.5's strength is tool use, so use the tools to check it. A claim the model verified against a file is a different kind of claim from one it produced from memory.
This is also where Opus 4.7 still wins for production. If the cost of a confidently wrong answer is large, a lower hallucination rate matters more than five points on a benchmark.
## Pricing
| Model | Input per 1M | Output per 1M |
| ------------------ | ------------ | ------------- |
| GPT-5.5 | $5 | $30 |
| GPT-5.5 Pro | $30 | $180 |
| GPT-5.4 (previous) | $2.50 | $15 |
The naive read is that GPT-5.5 doubled the price. The honest read is that it is meaningfully more token-efficient, using roughly 40% fewer output tokens on Codex tasks for equivalent work, which puts the effective per-task rise closer to 20% than 100%\[1]. Factor in a lower retry rate and agentic workloads come out ahead.
GPT-5.5 Pro at $30 / $180 is for one narrow case: the long-horizon, high-stakes task where five to ten benchmark points justify a six-fold cost premium. That is a small slice of real traffic.
On reAPI, GPT-5.5 runs **$4.00 input and $24.00 output** per million tokens, which is 80% of the published rate.
One thing worth knowing before you pick it: on the same gateway, **GPT-5.6 Sol costs exactly the same $4.00 / $24.00**. GPT-5.6 shipped in July with the `max` and `ultra` reasoning modes and a higher Terminal-Bench score. Unless your evals are already calibrated against GPT-5.5's behavior, there is no price argument for staying.
## How to use GPT-5.5: where it fits and where it does not
**Use it for** long-horizon agentic coding, computer-use tasks in Codex, research workflows that browse and synthesize, mathematical reasoning, and multi-occupation knowledge work where the 84.9% GDPval result is a credible signal\[1].
**Do not use it for** tasks where hallucination cost is high and you cannot wrap the call in verification, pure coding workloads where Opus still leads on SWE-bench Pro, or price-sensitive high-volume traffic where cheaper models undercut it by five to ten times per token.
## Calling GPT-5.5 on reAPI
The endpoint is OpenAI-compatible, so the model string is the only change. Note the wire id keeps the dot:
```python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_REAPI_KEY",
base_url="https://api.reapi.ai/v1",
)
resp = client.chat.completions.create(
model="gpt-5.5",
messages=[{"role": "user", "content": "Read the failing test, find the bug, and patch it."}],
max_tokens=16000,
stream=True,
)
```
Given the hallucination profile, the integration detail that matters most is not the request shape. It is whether the surrounding harness gives the model something to check itself against. Rates are on [reapi.ai/models/gpt-5-5](/models/gpt-5-5), and switching to `gpt-5.6-sol` or `gpt-5.6-terra` is a one-string change on the same key.
## FAQ
### Is GPT-5.5 worth the price increase over GPT-5.4?
For agentic and knowledge-work tasks, yes. The per-task cost is roughly 20% higher rather than 100%, because of token-efficiency gains, and the quality jump is real. For simple chat or low-stakes generation, the older model was better value\[1].
### How does GPT-5.5 compare to Claude Opus 4.7 for coding?
Opus 4.7 leads SWE-bench Pro by about six points with significantly lower hallucination. GPT-5.5 leads on Terminal-Bench, agentic tool use, and long-horizon tasks. For code review and refactoring, Opus is safer. For agentic loops with tools, GPT-5.5 wins\[1].
### Why does GPT-5.5 hallucinate more than Opus 4.7?
The likeliest explanation is optimization for agentic confidence: the model was trained to take action rather than defer. Mitigate with `xhigh` reasoning plus tool-grounded verification\[2].
### Can GPT-5.5 handle 1M tokens reliably?
Yes for retrieval. For complex reasoning over the full window, expect degradation past roughly 200K, as with other 1M-context models. Chunk or use retrieval for high-stakes long-context work.
### Should I use GPT-5.5 or GPT-5.6?
On reAPI they cost the same, and GPT-5.6 added the `max` and `ultra` reasoning modes with a higher Terminal-Bench score. Stay on GPT-5.5 only if your evaluations are already calibrated to its behavior.
### Does GPT-5.5 generate images?
No. It is text and vision input only. Image generation in OpenAI's stack is [GPT Image 2](/blog/how-to-use-gpt-image-2).
### What is GPT-5.5 Pro for?
Long-horizon, high-stakes tasks where a five-to-ten-point benchmark gain justifies a six-fold cost premium at $30 / $180 per million tokens\[1].
### How do I set reasoning effort?
Through the reasoning-effort parameter, with levels `xhigh`, `high`, `medium`, `low`, and non-reasoning. Raise it when correctness matters more than latency\[1].
## Deploying a confident model carefully
The headline metric of this release, index leadership, is the least useful thing about it. The real story is a model trained from the ground up as an agent, whose dominant benchmarks all map to multi-step, tool-using, real-world task completion. That is a different center of gravity from a precision instrument for high-stakes single-step work.
The deployment rule follows directly from the hallucination number. Give it tools and let it check itself, raise reasoning effort when correctness beats latency, and never let a confidently-worded answer reach a user or a production system without something the model can be checked against. That is how to use GPT-5.5 well: not as an oracle, but as a very capable agent that needs a verification loop around it.
## References
1. OpenAI. *Models — GPT-5.5 specifications, benchmarks, and reasoning-effort levels.* Retrieved July 2026 from [platform.openai.com/docs/models](https://platform.openai.com/docs/models)
2. Artificial Analysis. *Intelligence Index and AA-Omniscience hallucination measurements.* Retrieved July 2026 from [artificialanalysis.ai/models](https://artificialanalysis.ai/models)
3. OpenAI. *Pricing — per-token rates by model and tier.* Retrieved July 2026 from [platform.openai.com/docs/pricing](https://platform.openai.com/docs/pricing)
### Further reading
* reAPI. *How to use GPT-5.6.* [reapi.ai/blog/how-to-use-gpt-5-6](/blog/how-to-use-gpt-5-6)
* reAPI. *How to use GPT Image 2.* [reapi.ai/blog/how-to-use-gpt-image-2](/blog/how-to-use-gpt-image-2)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# How to Use GPT-5.6: Sol, Terra, and Luna Tiers Compared (https://reapi.ai/blog/how-to-use-gpt-5-6)
OpenAI began a limited preview of GPT-5.6 on June 26, 2026, and the strangest thing about that launch was that almost nobody could use it. Access went to roughly twenty vetted partners through the API and Codex, in a phased rollout OpenAI says the U.S. government requested\[1].
That gate has since lifted. Knowing how to use GPT-5.6 is now a practical question rather than a waiting game, and the answer is mostly about picking one of three tiers: **Sol** the flagship, **Terra** the balanced workhorse, and **Luna** the budget tier. Each is a durable capability tier under one generation number, not a temporary variant.
This guide covers the three-tier design, the one benchmark OpenAI chose to lead with and what its four asterisks mean, the new `max` and `ultra` reasoning modes, the pricing that makes Terra the story, and the cybersecurity review that shaped the rollout.
## TL;DR
* **Three durable tiers.** Sol (flagship), Terra (balanced, roughly 2x cheaper than GPT-5.5), Luna (fast and low cost)\[1].
* **Sol tops Terminal-Bench 2.1** at 88.8% single-agent, or 91.9% in `ultra` mode. The single-agent lead over Claude Mythos 5 is 0.8 points, which is inside the noise band\[1].
* **`ultra` is not one agent.** It orchestrates subagents in parallel, so its score is not a like-for-like comparison against single-agent baselines\[1].
* **Terra is the commercial story.** GPT-5.5-class coding at half the price\[1].
* **On reAPI all three run at 80% of OpenAI's published rates**: Sol $4.00 / $24.00, Terra $2.00 / $12.00, Luna $0.80 / $4.80 per million tokens.
* **The evaluation window was deliberately narrow.** OpenAI published coding, biology, and cybersecurity only, holding back SWE-bench, GDPval, math, and hallucination numbers\[1].
## What GPT-5.6 is
GPT-5.6 succeeds GPT-5.5, the fully retrained agentic model from April 2026. Where GPT-5.5 was one model with reasoning-effort levels, GPT-5.6 is a family of three tiers under one generation number. OpenAI's framing is explicit: the number identifies a generation, while Sol, Terra, and Luna identify durable capability tiers that can advance on their own cadence\[1].
That sentence is the most consequential design decision in the release. It decouples "how smart" from "how new." A future generation could ship a new Luna without touching Sol, and a Sol upgrade does not force everyone on Terra to re-test their pipelines.
| Tier | Positioning |
| ----------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| **GPT-5.6 Sol** | The flagship, and the only tier exposing the new `max` and `ultra` reasoning modes. The cybersecurity and long-horizon-agent model |
| **GPT-5.6 Terra** | The balanced workhorse. Competitive with GPT-5.5 while being about 2x cheaper. Most production workloads belong here |
| **GPT-5.6 Luna** | Fast and affordable, aimed at high-volume routine tasks |
Multiple third-party reports put Sol's context window at roughly 1.5 million tokens, about 43% larger than the GPT-5.5 generation, though OpenAI's preview post focused on safety and evaluations rather than the spec sheet\[1]. Treat that figure as reported rather than officially confirmed.
## Terminal-Bench 2.1, and its four asterisks
OpenAI published only three benchmark areas in the preview, coding, biology, and cybersecurity, explicitly holding back the rest for general availability\[1]. The one it wants you to remember is Terminal-Bench 2.1, which tests command-line agentic workflows requiring planning, iteration, and tool coordination.

| Model | Terminal-Bench 2.1 |
| -------------------------- | ------------------ |
| GPT-5.6 Sol (`ultra` mode) | **91.9%** |
| GPT-5.6 Sol | 88.8% |
| Claude Mythos 5 | 88.0% |
| GPT-5.6 Terra | 84.3% |
| Claude Fable 5 | 84.3% |
| GPT-5.5 | 83.4% |
| GPT-5.6 Luna | 82.5% |
| Claude Opus 4.8 | 78.9% |
| Gemini 3.1 Pro Preview | 70.7% |
Read carefully, this table is more honest than the "new state of the art" headline suggests, in four ways\[1].
**The 91.9% record uses `ultra`, which is not a single agent.** It comes from Sol orchestrating subagents. Against single-agent baselines it is not apples to apples. The fair comparison is Sol's 88.8%.
**Single-agent Sol beats Claude Mythos 5 by 0.8 points.** Two runs of the same model can disagree by more than that. OpenAI holds a real lead at the top of the coding curve, but on this benchmark it is narrow, not a blowout.
**Terra ties Claude Fable 5 and barely edges GPT-5.5.** 84.3% matches Fable 5 exactly and sits less than a point above the model it replaces. Paired with the price cut, that is the actual commercial story: same coding ability for half the money, not a capability leap in the mid-tier.
**Gemini 3.1 Pro Preview trails by about 21 points.** That gap is real but benchmark-specific. Gemini leads elsewhere on long-context retrieval and multimodal work not shown here.
On biology, OpenAI reports Sol achieving stronger GeneBench v1 results than GPT-5.5 while using fewer tokens, without publishing raw numbers, so that is a directional claim rather than a verifiable score\[1].
## `max` and `ultra`
The most important addition is not in any benchmark cell. GPT-5.6 introduces two reasoning settings, both exclusive to Sol\[1].

**`max`** is the familiar lever taken one notch further: a single agent given the most time to reason deeply on one hard problem.
**`ultra`** is the genuinely new primitive. It goes beyond a single agent by spawning subagents that split complex work into parallel pieces and recombine the results. This is OpenAI productizing multi-agent orchestration inside a single model endpoint, the pattern developers have been hand-rolling with frameworks for the past year.
The practical implication is real: for a multi-file refactor or an end-to-end audit, you no longer build the orchestration layer yourself. The catch is cost and latency, because subagents mean more total tokens and more wall-clock time. Reserve `ultra` for the high-stakes task where a few points justify the spend.
It also means benchmark comparisons need care. A score produced by an agent that quietly spawns subagents is a different kind of result from a single forward pass.
## Pricing

| Tier | OpenAI list (input / output per 1M) | On reAPI (input / output per 1M) |
| ------------- | ----------------------------------- | -------------------------------- |
| GPT-5.6 Sol | $5 / $30 | **$4.00 / $24.00** |
| GPT-5.6 Terra | $2.50 / $15 | **$2.00 / $12.00** |
| GPT-5.6 Luna | $1 / $6 | **$0.80 / $4.80** |
Two things stand out in OpenAI's own numbers. Sol holds GPT-5.5's exact $5 / $30 pricing while adding capability and the new modes, which is a rare case of more for the same money. And Terra delivers GPT-5.5-class coding at half the price, which is the change most teams will actually feel\[1]. Layer Sol's token-efficiency gains on top, since it reaches comparable results with fewer output tokens, and the effective drop in cost per task is larger than the headline rates imply.
Every tier runs at 80% of the list rate on reAPI, so the relative economics between tiers are unchanged. Terra is still the default and Luna is still the volume play.
GPT-5.6 also overhauls prompt caching, adding explicit cache breakpoints and a 30-minute minimum cache life. Cache writes bill at 1.25x the uncached input rate while cache reads keep the 90% discount\[1]. For agents that re-send a large system prompt on every step, predictable caching can dominate the bill.
## The review that shaped the rollout
This part is worth understanding even now that access has opened, because it is a precedent rather than a footnote.
OpenAI restricted the initial release to about twenty vetted partners at the U.S. government's request, after previewing the models' capabilities ahead of launch\[1]. The trigger was cybersecurity capability. Sol meaningfully advances vulnerability research, and in evaluations involving Chromium and Firefox it identified bugs and exploitation primitives, the building blocks of an exploit, without autonomously producing a functional full-chain exploit under the conditions tested. OpenAI states Sol does not cross the Cyber Critical threshold in its Preparedness Framework, and that distinction is the entire basis for releasing it\[1].
Notably, OpenAI did not endorse the arrangement as permanent, saying this kind of government access process should not become the long-term default because it keeps the best tools from the defenders who need them\[1].
The safety stack shipped alongside it is worth knowing because it previews how high-capability models will ship generally: trained-in refusals, real-time cyber and biology classifiers that can pause generation for review by a larger model, account-level review of flagged activity, and differentiated access to the most sensitive capabilities. OpenAI dedicated over 700,000 A100-equivalent GPU hours to automated red-teaming for universal jailbreaks\[1].
OpenAI is candid that these safeguards misfire: legitimate security researchers may hit refusals on dual-use requests.
## Where the case is strong and where it is thin
**Strong.** Agentic command-line coding, where single-agent Sol tops the benchmark and `ultra` extends it. Token efficiency, which compounds for high-volume agents. Mid-tier value, where Terra's price cut is the most broadly useful change in the release. And native multi-agent orchestration, which removes a layer developers used to build by hand\[1].
**Thin.** The headline is a 0.8-point win on one benchmark. The evaluation window was deliberately small, with no SWE-bench, GDPval, math, or hallucination numbers published, so whether GPT-5.6 fixed the confident-hallucination problem that dogged GPT-5.5 is still unknown. And everything is vendor-reported under OpenAI's own harness\[1].
Until an expanded suite and independent scoring land, judgment on general capability should stay reserved.
## How to use GPT-5.6: which tier to pick
**Terra for most production workloads.** It matches GPT-5.5-class capability at roughly half the cost, and the benchmark says its coding is a tie with Claude Fable 5.
**Sol for long-horizon coding, security work, and anything justifying `ultra`.** The subagent orchestration is worth real money on tasks where a few points change the outcome, and worth nothing on routine calls.
**Luna for high-volume, latency-sensitive, routine traffic.** At 82.5% on Terminal-Bench it is not far behind GPT-5.5's 83.4%, at a fraction of the cost.
## Calling GPT-5.6 on reAPI
All three tiers sit behind one OpenAI-compatible `/v1/chat/completions` endpoint, so switching tiers is a model-string change rather than an integration. Note that the wire ids keep OpenAI's dots:
```python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_REAPI_KEY",
base_url="https://api.reapi.ai/v1",
)
resp = client.chat.completions.create(
model="gpt-5.6-terra",
messages=[{"role": "user", "content": "Refactor this module and add tests."}],
max_tokens=16000,
stream=True,
)
```
Swap `gpt-5.6-terra` for `gpt-5.6-sol` or `gpt-5.6-luna` to move tiers. Because the same key and client also reach Claude and Gemini, the tier decision and the vendor decision become the same kind of decision, which is the point when a benchmark table has your default model losing one row and winning another.
Current rates are on [reapi.ai/models/gpt-5-6-sol](/models/gpt-5-6-sol), [reapi.ai/models/gpt-5-6-terra](/models/gpt-5-6-terra), and [reapi.ai/models/gpt-5-6-luna](/models/gpt-5-6-luna).
## FAQ
### What is the difference between `max` and `ultra` mode?
`max` is a single agent reasoning longer and deeper on one problem. `ultra` orchestrates multiple subagents that split complex work in parallel and recombine it. Both are exclusive to Sol, and `ultra` produced the 91.9% Terminal-Bench headline\[1].
### Is GPT-5.6 Sol better than Claude for coding?
On Terminal-Bench 2.1, single-agent Sol at 88.8% edges Claude Mythos 5 at 88.0% by 0.8 points, a real but narrow lead. With `ultra` Sol reaches 91.9%, but that uses subagents and is not a like-for-like single-agent comparison\[1].
### Which GPT-5.6 tier should I use?
Terra for most production workloads, since it matches GPT-5.5-class capability at about half the cost. Sol for long-horizon coding, security work, or tasks that justify `ultra`. Luna for high-volume routine calls\[1].
### Why was the U.S. government involved in the release?
Sol meaningfully advances cybersecurity capability, including vulnerability research, without crossing OpenAI's Cyber Critical threshold. At the government's request OpenAI ran a phased preview to about twenty vetted partners while a cyber framework was developed\[1].
### Does GPT-5.6 fix GPT-5.5's hallucination problem?
Unknown. OpenAI published only coding, biology, and cybersecurity evaluations in the preview and held back hallucination and knowledge-work numbers\[1]. Pair the model with source-grounded verification for high-stakes output.
### How big is GPT-5.6 Sol's context window?
Third-party reports put it around 1.5 million tokens, roughly 43% larger than the GPT-5.5 generation, but OpenAI's preview post did not confirm a spec sheet\[1].
### What changed in prompt caching?
Explicit cache breakpoints and a 30-minute minimum cache life. Cache writes bill at 1.25x the uncached input rate; cache reads keep the 90% discount\[1].
### How do I call GPT-5.6 with the OpenAI SDK?
Point the client at `https://api.reapi.ai/v1` and set `model` to `gpt-5.6-sol`, `gpt-5.6-terra`, or `gpt-5.6-luna`, keeping the dots in the id.
## Reading a three-tier release
GPT-5.6 is two releases wearing one announcement. One is an ordinary, well-executed family update: a sensible naming reset, a mid-tier that halves cost without losing capability, and an `ultra` mode that bakes multi-agent orchestration into the endpoint. The other is a frontier model whose initial availability was set in coordination with a government, shipped behind a 700,000-GPU-hour safety stack.
For anyone actually building, the useful summary is smaller than either story. The benchmark lead at the top is 0.8 points and vendor-reported. The change you will feel is Terra's price. That is how to use GPT-5.6 sensibly: default to the middle tier, reserve Sol for the tasks that pay for `ultra`, and keep source-grounded verification in place until the expanded evaluation suite tells you whether the hallucination problem moved.
## References
1. OpenAI. *Models — the GPT-5.6 family, reasoning modes, and specifications.* Retrieved July 2026 from [platform.openai.com/docs/models](https://platform.openai.com/docs/models)
2. OpenAI. *Pricing — per-token rates by model and tier.* Retrieved July 2026 from [platform.openai.com/docs/pricing](https://platform.openai.com/docs/pricing)
### Further reading
* reAPI. *How to use Claude Opus 5.* [reapi.ai/blog/how-to-use-claude-opus-5](/blog/how-to-use-claude-opus-5)
* reAPI. *How to use Claude Fable 5.* [reapi.ai/blog/how-to-use-claude-fable-5](/blog/how-to-use-claude-fable-5)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# How to Use GPT Image 2: Text Rendering and Thinking Mode (https://reapi.ai/blog/how-to-use-gpt-image-2)
OpenAI announced ChatGPT Images 2.0 on April 21, 2026, powered by the new `gpt-image-2` model, and began rolling it out the next day\[1]. The headline change is not resolution. It is that generated text inside images now works.
Not "mostly works, with a handful of charming typos." Text at roughly 99% accuracy across Latin, Japanese, Korean, Hindi, Bengali, Arabic, and a dozen other scripts\[1]. For two years the standard workflow was to generate the image, then composite real text on top in Figma. That step is largely gone.
Learning how to use GPT Image 2 well is mostly about two things: writing briefs precise enough to exploit the text rendering, and knowing when to turn Thinking Mode off.
## TL;DR
* **Text rendering at \~99% accuracy**, including non-Latin scripts at near-native quality\[1].
* **Thinking Mode** lets the model browse the web for references, plan multiple variants, and self-verify its output before returning\[1].
* **2K native resolution** with 16:9 support, and a December 2025 knowledge cutoff\[1].
* **It replaces DALL-E 3**, which retired May 12, 2026, and the interim GPT Image 1.5\[1].
* **A completely new architecture**, single-pass rather than the prior two-stage pipeline, giving sub-three-second typical generation\[1].
* **Not the photorealism leader.** Nano Banana 2 stays ahead on pure photographic realism and 4K output; GPT Image 2 leads on text fidelity and prompt adherence\[1].
## What GPT Image 2 is
GPT Image 2 is OpenAI's flagship image model, shipping on three surfaces: inside ChatGPT as Images 2.0, inside the Codex coding agent for inline asset generation, and as the `gpt-image-2` API\[1].
The architecture is a clean break from the DALL-E line. OpenAI has declined to confirm whether it is diffusion-based, autoregressive, or hybrid, saying only that it is a completely new architecture moving from a two-stage inference pipeline to a single-pass generator. The practical consequences are faster generation and much tighter coupling between prompt semantics and rendered output\[1].
## The text rendering breakthrough

The specific improvements are worth separating, because they unlock different workflows\[1]:
* **Latin-script spelling at \~99%**, including dense menu text, packaging labels, nutrition panels, long editorial headlines, and UI copy.
* **Non-Latin scripts at near-native accuracy** for major world languages. Previous models treated these as decoration: visually plausible, semantically meaningless.
* **Small text and iconography** in dense compositions, including form fields, badge labels, and legal fine print, rendered readably.
* **Stylistic consistency between text and design language**, so a Brooklyn café sign and a Tokyo subway poster read as different designers rather than the same generic model.

The bookstore example is the one to study. Each title is rendered in a distinct Indic script, and each is coherent text a reader of that language can identify\[1]. That is the capability that opens localized content pipelines: regional ad creative, multilingual packaging, Indic-language e-commerce thumbnails, at a speed that was previously impossible.
## Thinking Mode

Thinking Mode is the feature that changes how you interact with the model. With standard prompting it does what any image model does: read prompt, generate pixels. With Thinking Mode on, it also\[1]:
**Browses the web** for reference material when the prompt names real brands, products, or time-sensitive content.
**Plans multiple output variants** before committing to pixels, which is what makes one prompt produce an ad campaign in three sizes or a comic across five panels.
**Self-verifies before returning**, checking that rendered text is spelled correctly and that layout hierarchy survived.
The demonstration OpenAI leads with is a single prompt asking for a poster about merch currently on its own website, which produced an accurate composition because the model went and looked. No reference upload, no manual product list\[1].
For briefs containing phrases like "in the style of their current hero page," this removes the manual reference-gathering step. For bulk generation of simple visuals, turn it off, because it costs more and takes longer.
## Commercial-grade output

The advertisement above is a fully synthetic single-prompt generation. It carries a plausible product photograph, a coherent brand wordmark in two scripts, secondary typographic detail, a legible announcement with location and social handle, and stylistic consistency between photograph, type, texture, and palette\[1].
Two years ago this was the classic generative-image failure case. The model could produce the photograph, but any real text in the composition devolved into garbled glyphs.
The implication is worth naming directly: for most small-business marketing the bottleneck is no longer visual production. It is the brief.
## Deeper structural understanding
OpenAI's positioning line is "image generation with a point of view," meaning the model has structural understanding of what it renders rather than statistical coverage of a corpus. In practice\[1]:
* **Instruction following** improves on compositional requests: left-aligned headline, product at right, brand mark bottom-right.
* **World knowledge** is stronger. The model knows which camera a product shot should look like it was taken on, which paper stock a vintage cover should emulate.
* **Character consistency** holds across multi-panel sequences.
* **Iconographic literacy** improved, so periodic tables, anatomical diagrams, and architectural plans render correctly rather than as plausible-looking nonsense.
This is the axis where it most clearly differs from purely photographic models. On a photorealistic still life, Nano Banana 2 is competitive or ahead. On a composition requiring the model to *know* something and render it correctly, GPT Image 2 leads\[1].
## Where it does not fit
The honest limitations\[1]:
* **Not a replacement for designers on identity-critical brand work.** First-draft quality is excellent; a creative director should still direct.
* **Not the photorealism leader.** Nano Banana 2 and specialist photo-real models retain edges on hyper-realistic still life and portraiture.
* **Thinking Mode is slower and pricier.** Disable it for bulk generation of simple visuals.
* **Content policy applies.** Real public figures, copyrighted characters, and certain commercial references are restricted.
## How to use GPT Image 2 on reAPI
reAPI exposes the model on an async task endpoint. Submit returns a `task_id`; poll until ready.
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-image-2",
"prompt": "Editorial poster for a Kyoto coffee shop opening, hand-lettered headline in cream serif, warm cinematic light, legible address line at the bottom",
"size": "16:9",
"resolution": "2k"
}'
```
Three things to know before writing the integration.
**Resolution is a priced dimension.** The gateway exposes 1K, 2K, and 4K, each with its own rate, so a resolution choice is a cost choice. Current rates are on [reapi.ai/models/gpt-image-2](/models/gpt-image-2), which is the canonical source rather than any figure quoted in an article.
**Media inputs are public URLs only.** No base64, no `data:` payloads, on any model on the platform.
**Prompt precision pays here more than on older models.** Briefs that worked with DALL-E 3, vague and visually focused, underperform relative to what this model does with explicit layout, typographic, and reference instructions. The ceiling is higher and the floor requires more precise input\[1].
Full request and response shapes are in the [reapi.ai/docs/gpt-image-2](/docs/gpt-image-2) reference.
## FAQ
### What is GPT Image 2?
OpenAI's flagship image generation model, announced April 21, 2026, replacing DALL-E 3 and GPT Image 1.5. It ships in ChatGPT as Images 2.0, in Codex, and as the `gpt-image-2` API\[1].
### How accurate is GPT Image 2 at rendering text?
Roughly 99% on Latin script, with near-native accuracy on major non-Latin scripts including Japanese, Korean, Hindi, Bengali, and Arabic\[1].
### What is Thinking Mode?
A mode in which the model browses the web for references, plans multiple variants before generating, and self-verifies rendered text and layout before returning output. It costs more and takes longer, so disable it for bulk simple generation\[1].
### What resolution does GPT Image 2 support?
2K native with 16:9 support alongside existing aspect ratios\[1]. On reAPI, 1K, 2K, and 4K are exposed as separately priced options.
### Is GPT Image 2 better than Nano Banana 2?
They lead on different axes. Nano Banana 2 is ahead on pure photographic realism and 4K output. GPT Image 2 leads on text fidelity, prompt adherence, and integrated Thinking Mode\[1].
### What happened to DALL-E 3?
It retired on May 12, 2026. GPT Image 2 replaces it and the interim GPT Image 1.5\[1].
### Does GPT Image 2 accept base64 image input on reAPI?
No. Every model on the platform takes public http(s) URLs only for media inputs.
### How should I write prompts for GPT Image 2?
More specifically than for DALL-E 3. Give explicit layout, typographic, and reference instructions rather than a vague visual description, because the model now acts on structural detail that older models ignored\[1].
## Updating the brief, not just the model string
The useful summary of this release is that it moves the bottleneck. When generated text was unusable, the constraint was the model. Now that a café advertisement with readable Japanese secondary type comes out of a single prompt, the constraint is how precisely you can describe what you want.
That has a practical consequence for anyone migrating a pipeline. Prompt templates written for the previous generation are tuned for a model that ignored structural instruction, and they will underuse this one. Rewrite them with explicit layout and typography, keep Thinking Mode for briefs that reference the real world and off for volume, and pick the resolution deliberately since it is a priced dimension. That is how to use GPT Image 2 without paying for capability your briefs never ask for.
## References
1. OpenAI. *GPT Image 2 — model card, capabilities, and API reference.* Retrieved July 2026 from [platform.openai.com/docs/models](https://platform.openai.com/docs/models)
2. OpenAI. *Pricing — per-image and per-token rates by model.* Retrieved July 2026 from [platform.openai.com/docs/pricing](https://platform.openai.com/docs/pricing)
### Further reading
* reAPI. *How to use Nano Banana 2 Lite.* [reapi.ai/blog/how-to-use-nano-banana-2-lite](/blog/how-to-use-nano-banana-2-lite)
* reAPI. *GPT Image 2 endpoint reference.* [reapi.ai/docs/gpt-image-2](/docs/gpt-image-2)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# How to Use Nano Banana 2 Lite: Speed, Price, and Limits (https://reapi.ai/blog/how-to-use-nano-banana-2-lite)
Google shipped Nano Banana 2 Lite on June 30, 2026, the fastest and cheapest tier of its Nano Banana image line\[1]. Its official API name is Gemini 3.1 Flash-Lite Image, and the entire point is speed: a 1K image in about four seconds, at roughly half the cost of the full Nano Banana 2.
Knowing how to use Nano Banana 2 Lite well comes down to one number that is easy to miss. On image generation it gives up about 19 Elo points to base Nano Banana 2, roughly a 1.5% quality gap, while running five times faster and costing half as much\[1]. That trade is excellent for most work and wrong for a specific kind of work, and this guide is about telling those apart.
One naming trap first: the text-only `gemini-3.1-flash-lite` is a different model and does not generate images. The image model is the separate `-image` variant\[1].
## TL;DR
* **About four seconds per 1K image**, against 20 seconds for base Nano Banana 2\[1].
* **Roughly 3.4 cents per 1K image** at Google's list rate, or about 1.7 cents in batch mode. On reAPI it is 20 credits per image, which is $0.020, about 41% below Google's published rate.
* **1K is the ceiling.** No 2K, no 4K. Those need base Nano Banana 2 or Nano Banana Pro\[1].
* **Generation quality nearly matches the workhorse tier** (1251 vs 1270 Elo) but **editing does not** (1308 vs 1387)\[1].
* **Ten aspect ratios**, including 16:9, 9:16, and 21:9\[1].
* **Rated "Low reasoning" by Google**, so complex multi-constraint scenes belong on a higher tier\[1].
## The three Nano Banana tiers
Nano Banana 2 Lite replaces the first-generation Nano Banana, which Google now labels legacy, so the lineup is a clean three-tier stack\[1].

| Tier | Underlying model | Latency | Cost | Visual quality | Reasoning |
| ---------------------- | --------------------------- | ------- | ------ | -------------- | --------- |
| **Nano Banana 2 Lite** | Gemini 3.1 Flash-Lite Image | Low | Low | Medium | Low |
| Nano Banana 2 | Gemini 3.1 Flash Image | Medium | Medium | High | Medium |
| Nano Banana Pro | Gemini 3 Pro Image | High | High | High | High |
Note the version detail. Lite and base Nano Banana 2 both run on Gemini 3.1 Flash, while Pro runs on the Gemini 3 Pro Image line, a different and more capable model built for reasoning-heavy generation\[1]. Lite is explicitly the medium-quality, low-reasoning tier. You trade fidelity and complex-instruction handling for the lowest latency and cost in the family.
## Benchmarks
Google published a four-panel chart comparing Lite against base Nano Banana 2, the legacy model, and three competitors. Elo scores are sourced to LMArena, latency to Artificial Analysis\[1].

| Model | Generation Elo | Editing Elo | Latency (1K) | List price (1K) |
| ---------------------- | -------------- | ----------- | ------------ | --------------- |
| **Nano Banana 2 Lite** | 1251 | 1308 | **4.0s** | $0.034+ |
| Nano Banana 2 | **1270** | **1387** | 20.0s | $0.067+ |
| Nano Banana (legacy) | 1151 | 1295 | 7.0s | $0.039 |
| Flux 2 Klein 9B | 1069 | 1224 | 4.4s | **$0.015** |
| Grok Imagine Image | 1174 | 1329 | 6.4s | $0.020 |
| Seedream v5 Lite | 1132 | 1294 | 45.1s | $0.035 |
The trade is favorable. On generation, Lite gives up about 19 Elo to base Nano Banana 2 while running five times faster at half the price. On editing the gap is wider, 1308 against 1387, so edit-heavy workloads still favor the full model. Against the outside field, Lite beats Flux 2 Klein 9B and Seedream v5 Lite on both quality axes, and beats Seedream on speed by more than ten times. It trails Grok Imagine Image slightly on editing while leading it on generation\[1].
Two caveats belong with that table. Google picked its own comparison set and did not include GPT Image, so no official Lite-versus-GPT-Image number exists. And Google reports Elo values without a published leaderboard rank, so any "Lite is number one" claim is unverified\[1].
## Capabilities and limits
* **Resolution:** 1K (1024px) and 512px only. No 2K or 4K, which are exclusive to base Nano Banana 2 and Pro. On reAPI the model is exposed as a single 1K output, so there is no resolution field to set.
* **Aspect ratios:** 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, and 21:9. That covers social, web, and widescreen including the 16:9 slide ratio\[1].
* **Core features:** text-to-image, image editing, prompt adherence, character consistency, and legible in-image text, all inherited from the family. Multi-image composition works, though exact character and object limits are not clearly documented\[1].
* **Text rendering languages:** Google did not enumerate a language list for Lite specifically. Treat non-Latin scripts and long strings as something to test rather than assume\[1].
The honest framing is that Lite is a good-enough-for-volume model. It handles clean single-subject generation, straightforward edits, and standard aspect ratios at speed, and steps back from high fidelity, complex multi-constraint scenes, and dense factual infographics, which is what the low-reasoning rating is telling you.
## Pricing and the cost math
This is where the model earns its place, and it is worth stating carefully because secondary coverage got it wrong.
Google bills Lite at $0.25 per million input tokens and $30 per million output tokens. A 1K image is 1,120 output tokens, which works out to $0.0336, about 3.4 cents per **single image**, not per thousand images as some outlets reported. Batch mode halves the output rate to $15 per million, dropping a 1K image to roughly 1.7 cents. There is no free tier\[1].
Within the family, base Nano Banana 2 bills output at $60 per million (a 1K image is about $0.067, 2K about $0.10, 4K about $0.15) and Pro at $120 per million (4K about $0.24)\[1].
The detail that makes the math simple: Lite and base Nano Banana 2 use the same 1,120 tokens for a 1K image. The entire saving comes from the cheaper per-token output rate, not from a smaller image\[1]. For 1K work, Lite is a straight 2x discount over the workhorse tier.
On reAPI the rate is **20 credits per image**, which is $0.020, roughly 41% below Google's published $0.034. One image per request, one flat rate, no resolution tiers to reason about.
## How it compares to competitors
Against the field Google chose to show, Lite is quality-competitive but not the cheapest. Flux 2 Klein 9B at $0.015 and Grok Imagine Image at $0.020 both undercut it. Seedream v5 Lite is close on price at $0.035 but eleven times slower\[1].
Where Lite wins is the balance of three axes at once: strong Elo, four-second latency, and mid-tier pricing. If your only metric is cost per image, cheaper models exist. If you need decent quality and real-time latency together, Lite is the sweet spot.
## How to use Nano Banana 2 Lite: prompting that works
The low-reasoning rating means it rewards clear, concrete prompts over clever multi-constraint ones.
* **Lead with subject and style, then details.** "A minimalist product hero shot of a ceramic coffee mug, soft studio lighting, warm neutral background" beats a paragraph of layered conditions.
* **State the aspect ratio explicitly.** All ten are supported; saying so up front avoids awkward crops.
* **Keep in-image text short.** Legible text works, but quality on long strings and non-Latin scripts is undocumented. Headlines and single labels are safe; dense paragraphs are risky.
* **Iterate cheaply, then finalize.** At four seconds and a couple of cents per image, generate a dozen variations here, then re-render only the winner on a higher tier if you need 2K or 4K.
* **Use editing for controlled changes.** For same-scene background swaps, Lite's 1308 editing Elo is capable, though the full model's 1387 is stronger for demanding edits.
## Lite, base Nano Banana 2, or Pro
Three questions settle it.
**Do you need output above 1K?** If yes, skip Lite. It caps at 1K.
**Is the task edit-heavy or reasoning-heavy?** Lite's editing Elo trails base Nano Banana 2 by a real margin, and complex multi-constraint infographics belong on Pro. For straightforward generation, Lite is nearly as good at half the cost.
**Is latency or volume the constraint?** Lite is the only tier that delivers sub-five-second turnaround. Base Nano Banana 2's 20-second latency rules it out of interactive, high-volume loops.
The healthy pattern is two-tier: draft on Lite, finalize on a larger model only for the specific assets that earn the extra cost. For most teams the overwhelming majority of generated images never need that upgrade, which is exactly why a fast, cheap tier changes the economics of shipping custom visuals at all.
## Calling Nano Banana 2 Lite on reAPI
reAPI exposes the model on an async task endpoint. Submit returns a `task_id`; poll until it is ready.
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nano-banana-2-lite",
"prompt": "a hand-illustrated travel postcard of Kyoto at dusk, glowing paper lanterns, warm cinematic light",
"aspect_ratio": "16:9"
}'
```
Editing uses the same endpoint. Add `image_urls` with up to ten reference images and the request becomes a prompt-based edit rather than a text-to-image generation. There is no `mode` field to set; the presence of `image_urls` is what switches it.
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nano-banana-2-lite",
"prompt": "replace the background with a quiet stone lane at dusk, keep the subject unchanged",
"image_urls": ["https://example.com/source.png"],
"aspect_ratio": "16:9"
}'
```
Two platform rules worth knowing before you write the integration. Media inputs are **public http(s) URLs only**, with no base64 or `data:` accepted on any model. And billing is a flat 20 credits per image with one image per request, so cost forecasting is multiplication rather than token estimation.
Full request and response shapes are in the [reapi.ai/docs/nano-banana-2-lite](/docs/nano-banana-2-lite) reference, and current rates are on [reapi.ai/models/nano-banana-2-lite](/models/nano-banana-2-lite).
## FAQ
### What is Nano Banana 2 Lite's official name?
Gemini 3.1 Flash-Lite Image. "Nano Banana 2 Lite" is Google's marketing name for that image model. The text-only `gemini-3.1-flash-lite` is a different model that does not generate images\[1].
### How much does Nano Banana 2 Lite cost?
About 3.4 cents per 1K image at Google's list rate, 1,120 output tokens at $30 per million, or roughly 1.7 cents in batch mode. On reAPI it is 20 credits, which is $0.020. Ignore the "per 1,000 images" figure some outlets published; the rate is per single image\[1].
### How fast is Nano Banana 2 Lite?
Around four seconds per 1K image, against 20 seconds for base Nano Banana 2, roughly five times faster\[1].
### What resolutions does it support?
1K (1024px) and 512px only. No 2K or 4K; those need base Nano Banana 2 or Nano Banana Pro. On reAPI the model is exposed as a single 1K output\[1].
### Can Nano Banana 2 Lite edit images?
Yes. Send `image_urls` alongside the prompt and the same endpoint performs a prompt-based edit, with up to ten reference images. Its editing Elo of 1308 trails base Nano Banana 2's 1387, so demanding edits still favor the larger model\[1].
### How does it compare to GPT Image?
There is no official head-to-head. Google did not include GPT Image in its comparison set\[1].
### Does it accept base64 images?
Not on reAPI. Every model on the platform takes public http(s) URLs only for media inputs.
### What aspect ratios are supported?
Ten: 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, and 21:9\[1].
## Picking the tier that matches the job
Nano Banana 2 Lite is a well-judged addition rather than a headline model, and that is the point. It keeps almost all of base Nano Banana 2's generation quality while running five times faster and costing half as much, in exchange for a 1K ceiling, medium fidelity, and weaker editing and reasoning.
For high-volume, latency-sensitive, cost-sensitive work it is the right default, with the larger tiers reserved for assets that genuinely need 2K or 4K output or maximum quality. If you are wiring it into a product, remember the corrected math and the platform rules: about 3.4 cents per single 1K image at list, 20 credits on reAPI, public URLs only for reference images, and one image per request. That is how to use Nano Banana 2 Lite without paying workhorse-tier prices for draft-quality work.
## References
1. Google. *Gemini API — image generation, models, and capabilities.* Retrieved July 2026 from [ai.google.dev/gemini-api/docs/image-generation](https://ai.google.dev/gemini-api/docs/image-generation)
2. Google. *Gemini Developer API pricing.* Retrieved July 2026 from [ai.google.dev/gemini-api/docs/pricing](https://ai.google.dev/gemini-api/docs/pricing)
3. Google DeepMind. *Gemini image generation models.* Retrieved July 2026 from [deepmind.google/models/gemini-image](https://deepmind.google/models/gemini-image/)
### Further reading
* reAPI. *Nano Banana 2 Lite endpoint reference.* [reapi.ai/docs/nano-banana-2-lite](/docs/nano-banana-2-lite)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# How to Use Seedance 2.0 for Free: What Actually Works (https://reapi.ai/blog/how-to-use-seedance-2-0-for-free)
You can use Seedance 2.0 for free today, and I mean the real model, not a lookalike site. There are exactly two working routes I could verify in July 2026: ByteDance's own Dreamina app, which hands logged-in users daily credits\[1], and Krea's free tier, which allows a little under one video a day\[2]. Everything else that markets itself as "free Seedance 2.0" is either a trial wrapped around a subscription, a withdrawn promotion, or a fake site reselling generations.
The honest version of this guide has three parts: how to use it for free right now, what the free routes deliberately hold back, and the point where free stops being worth your time, because on a per-clip basis the paid floor is about twenty cents.
## TL;DR
* **Dreamina is the official free route**: sign in free, spend daily credits on Seedance 2.0 or 2.0 Fast\[1]. Real-person faces are blocked on 2.0 there, and outputs carry invisible watermarking\[3].
* **Krea's free tier** includes Seedance 2.0 access at under one clip per day; its $9/mo Basic plan covers roughly 20 videos if you outgrow that\[2].
* **"Unlimited Seedance" annual deals are gone or hollow.** OpenArt, the platform most cited for one, does not even list Seedance 2.0 anymore, only 1.5 Pro\[11].
* **"Free Seedance token" sites are scams.** The top search results for the model are unaffiliated domains by their own terms pages; none of them can hand out real capacity\[4].
* **There is no free local option.** The weights are closed; nothing to download, nothing to run offline\[5].
* **The paid floor is lower than you think**: Seedance 2.0 Mini runs from $0.03/s on reAPI, so about $1 buys roughly five 5-second clips\[6].
## The two free Seedance 2.0 routes that actually work
**Dreamina (dreamina.capcut.com).** This is ByteDance's own consumer app and the closest thing to an official free tier. A free login gets daily credits you can point at the model or its Fast variant in the video tool\[1]. Credit amounts shift with promotions and region, and ByteDance does not publish a fixed public number, so treat whatever your account shows as the offer. The important part: this is the actual model, same as the API serves.
The two-minute version: create a free account, open the video tool, switch the model picker to 2.0 (or 2.0 Fast, which stretches the same credits further), write a prompt or attach references that contain no real faces, and generate at 720p to conserve credits. Outputs land in your asset library with download enabled. Credits refresh on a daily cadence rather than accumulating, so an unused day is a spent day; plan one deliberate test per session instead of burning the allowance on prompt typos.
**Krea (krea.ai).** Krea's free tier includes the whole lineup, 2.0 through Mini, rate-limited to just under one video per day\[2]. That cadence is enough to evaluate whether the model fits your use case, which is what a free tier is for.
That is the whole list. Higgsfield's Seedance 2.0 access starts at its mid subscription tier, not free\[9]; Pollo's entry plan is $15/mo with watermarked output\[10]; and the official BytePlus API is enterprise-billed with no free video allowance\[7].
A regional note, since "how to use Seedance 2.0 for free in India" is a surprisingly common search: the restriction that exists is on the United States, where ByteDance's consumer rollout stayed paused after the copyright disputes\[8]. India is not specially restricted; the Dreamina and Krea routes above are the same there.
## What free deliberately holds back
Free access is real but shaped. Three limits to know before you build any plan around it.
**Volume.** Daily credits on Dreamina and sub-one-a-day limits on Krea are evaluation budgets, not production budgets. A content channel posting daily burns through a free allowance before breakfast.
**Faces.** On Dreamina, Seedance 2.0 and 2.0 Fast carry a flat "Real human faces are not supported" restriction on uploads\[1]\[3]. If your free experiments involve photos of people, they will bounce; our [not-eligible guide](/blog/seedance-not-eligible-explained) covers why and what the sanctioned routes are.
**Watermarking.** ByteDance embeds invisible watermarking in Dreamina outputs by design\[3], and cheap subscription tiers elsewhere add visible marks (Pollo's Lite plan, for instance\[10]). Free output is fine for testing prompts; it is not clean commercial deliverable.
## The "unlimited Seedance" deals, honestly
A whole family of searches asks about a platform that "gave unlimited Seedance 2.0 for a year" on a pro plan. Whatever was briefly true during launch-season promotions, here is the July 2026 state: OpenArt, the platform most associated with that pitch, no longer lists the model at all; its newest Seedance option is 1.5 Pro\[11]. Video generation is priced per second of GPU time everywhere upstream, which is why "unlimited" offers either cap you with fair-use clauses, degrade you to lighter models, or quietly disappear. If an offer says unlimited, read the cap.
And then there is the scam tier. The top organic results for "seedance" are domains unaffiliated with ByteDance by their own terms of service, several charging $15 to $100 a month\[4]. "Free Seedance tokens" sites are the same operation with a different hook. Nothing outside ByteDance's surfaces and known platforms has real capacity to give away; details and receipts are in our [platform guide](/blog/what-is-seedance-2-0-and-how-to-use-it).
One more non-route: running it locally for free. The model's weights have never been released, so there is nothing to install\[5]. Anything labeled a local Seedance build is a different model wearing the name.
## When free stops being worth it
Here is the math that ends most "how do I stay free" quests. On reAPI, the Mini tier bills $0.048/s at 480p text-to-video, and $0.03/s with a video reference\[6]. A 5-second clip is 15 to 24 cents. One dollar of credit is four to six clips, at API quality, with no daily gate, no queue for a credit refresh, and no watermark added by us. And if the project needs the full-quality Standard tier, 720p with references is $0.1086/s, about 54 cents per 5-second clip.
So the free tiers are the right place to answer "does the model fit my use case," and precisely the wrong place to grind out a project one rationed clip a day. The crossover point is embarrassingly low: if your time is worth anything, a few dollars of per-second API credit outperforms a week of babysitting free allowances. Signup credits on reAPI let you exercise the API itself before putting money in, and the [live pricing table](/models/seedance-2-0#pricing) shows exactly what each tier costs before you commit anything.
## FAQ
### Is Seedance 2.0 completely free anywhere?
Free with limits, yes: Dreamina's daily credits\[1] and Krea's sub-one-a-day free tier\[2]. Unlimited free does not exist anywhere, because upstream generation is metered per second.
### How do I use Seedance 2.0 for free without a watermark?
You mostly don't. Dreamina embeds invisible watermarking by design\[3], and budget subscription tiers add visible marks. Clean commercial output is what per-second API billing is for.
### Can I use Seedance 2.0 for free in India?
Yes, the same two routes: Dreamina and Krea's free tier. The notable regional restriction is the United States, where ByteDance's consumer rollout remains paused\[8].
### Are "free Seedance 2.0 token" generators real?
No. The squatter domains topping search results disclaim any ByteDance affiliation in their own terms\[4], and no third party has free capacity to distribute. Do not enter payment details on them.
### Did a platform really offer unlimited Seedance 2.0 for a year?
Launch-window promotions existed, but the flagship example, OpenArt, no longer carries the model at all\[11]. Treat any surviving "unlimited" claim as capped fine print.
### Can I run Seedance 2.0 locally for free?
No. The weights are closed and have never been distributed\[5]. Open-weight video models exist, but they are different models.
### What is the cheapest paid way to use Seedance 2.0?
Seedance 2.0 Mini from $0.03/s with a video reference, $0.048/s text-to-video at 480p, on reAPI's per-second billing\[6]. Full route-by-route numbers are in our [cheapest Seedance 2.0 comparison](/blog/cheapest-seedance-2-0-2026).
## Free is a test drive, not a plan
Use Seedance 2.0 for free the way the free routes are designed: log into Dreamina, spend the daily credits, confirm the model does what your project needs, and keep your expectations about faces and watermarks calibrated. The moment the answer is "yes, this works," stop rationing. At a twenty-cent floor per clip, the paid tier costs less than the coffee you drink while waiting for tomorrow's free credits, and that is the honest end of every guide to using Seedance 2.0 for free.
## References
1. Dreamina (CapCut). *AI video generation — Seedance model picker and credit system.* Retrieved July 2026 from [dreamina.capcut.com/ai-tool/home](https://dreamina.capcut.com/ai-tool/home?type=video)
2. Krea. *Video generation — Seedance models and plan limits.* Retrieved July 2026 from krea.ai/video
3. CapCut Newsroom. *Dreamina Seedance 2.0 global rollout — face restrictions and invisible watermarking.* Retrieved July 2026 from [capcut.com/newsroom/dreamina-seedance-2](https://www.capcut.com/newsroom/dreamina-seedance-2)
4. reAPI. *What Is Seedance 2.0 and How to Use It — squatter-site documentation.* Retrieved July 2026 from [reapi.ai/blog/what-is-seedance-2-0-and-how-to-use-it](/blog/what-is-seedance-2-0-and-how-to-use-it)
5. ByteDance Seed. *Seedance 2.0 — official model page (no weight distribution).* Retrieved July 2026 from [seed.bytedance.com/en/seedance2\_0](https://seed.bytedance.com/en/seedance2_0)
6. reAPI. *Seedance 2.0 Mini — model page and live pricing.* Retrieved July 2026 from [reapi.ai/models/seedance-2-0-mini](/models/seedance-2-0-mini)
7. BytePlus. *ModelArk — Seedance pricing (token billing, resource packs).* Retrieved July 2026 from [docs.byteplus.com/en/docs/ModelArk/1544106](https://docs.byteplus.com/en/docs/ModelArk/1544106)
8. Reuters. *ByteDance suspends launch of video AI model after copyright disputes.* Retrieved July 2026 from [reuters.com/technology/bytedance-suspends-launch-video-ai-model](https://www.reuters.com/technology/bytedance-suspends-launch-video-ai-model-after-copyright-disputes-information-2026-03-14/)
9. Higgsfield. *Pricing — Seedance availability by plan.* Retrieved July 2026 from higgsfield.ai/pricing
10. Pollo.ai. *Pricing.* Retrieved July 2026 from pollo.ai/pricing
11. OpenArt. *Video models — Seedance lineup.* Retrieved July 2026 from openart.ai/video/t2v
### Further reading
* reAPI. *Cheapest Seedance 2.0 in 2026: Real Prices, Compared.* [reapi.ai/blog/cheapest-seedance-2-0-2026](/blog/cheapest-seedance-2-0-2026)
* reAPI. *Seedance 2.0 "Not Eligible": Why It Happens, What Works.* [reapi.ai/blog/seedance-not-eligible-explained](/blog/seedance-not-eligible-explained)
* reAPI. *Seedance 2.0 API documentation.* [reapi.ai/docs/seedance-2-0](/docs/seedance-2-0)
---
# Is GPT-5.6 Out Yet? Yes — and Here Is What Shipped (https://reapi.ai/blog/is-gpt-5-6-out-yet)
Yes. GPT-5.6 is out, and it has been for a while.
Search volume has not caught up: "when is GPT-5.6 coming out" and its variants still pull tens of thousands of searches a month, all of them asking about a release that already happened. If you arrived here from one of those, the short version is below, followed by what actually shipped and what it costs.
## TL;DR
* **GPT-5.6 is released.** OpenAI previewed it on June 26, 2026 and general availability followed\[1].
* **It is three models, not one**: Sol, Terra, and Luna, as durable capability tiers under one generation number\[1].
* **The launch was unusual**: initial access went to roughly twenty vetted partners at the U.S. government's request, before broader release\[1].
* **Terra is the tier most teams want**: GPT-5.5-class coding at roughly half the price\[1].
* **Sol added two reasoning modes**, `max` and `ultra`, the second of which orchestrates subagents\[1].
* **On reAPI all three run at 80% of list**: Sol $4.00/$24.00, Terra $2.00/$12.00, Luna $0.80/$4.80 per million tokens.
## What "released" means here
The rollout is why the search question persisted longer than usual.
OpenAI previewed GPT-5.6 on June 26, 2026, but did not open it broadly. Access went to about twenty vetted partners through the API and Codex, in a phased rollout OpenAI says the U.S. government requested after reviewing the models' cybersecurity capabilities ahead of launch\[1].
So for a few weeks there was a model with published benchmarks, published pricing, and almost no one able to use it. That gap is what generated the wave of "is it out yet" searches, and it is also why answers from that period contradict each other.
That gate has since lifted. The models are generally available.
## The three tiers

| Tier | Positioning | List price per 1M (in / out) | On reAPI |
| ----------------- | -------------------------------------------------------------- | ---------------------------- | -------------- |
| **GPT-5.6 Sol** | Flagship. The only tier with `max` and `ultra` reasoning modes | $5 / $30 | $4.00 / $24.00 |
| **GPT-5.6 Terra** | Balanced workhorse, about 2x cheaper than GPT-5.5 | $2.50 / $15 | $2.00 / $12.00 |
| **GPT-5.6 Luna** | Fast and low cost, for high-volume routine work | $1 / $6 | $0.80 / $4.80 |
The naming is deliberate. OpenAI's framing is that the number identifies a generation while Sol, Terra, and Luna identify durable capability tiers that advance on their own cadence\[1]. A future generation can ship a new Luna without touching Sol.
For anyone who was waiting: **Terra is the release**. It matches GPT-5.5-class coding at roughly half the price, which is the change most workloads actually feel\[1].
## What actually shipped
**Terminal-Bench 2.1 at 88.8% single-agent**, or 91.9% in `ultra` mode. The single-agent lead over Claude Mythos 5 is 0.8 points, which is inside the noise band, so read it as a narrow win rather than a generational gap\[1].
**`ultra` is not one agent.** It orchestrates subagents in parallel, so its headline score is not a like-for-like comparison against single-agent baselines\[1].
**Sol holds GPT-5.5's exact pricing** while adding capability, and prompt caching was overhauled with explicit cache breakpoints and a 30-minute minimum cache life\[1].
**The evaluation window was deliberately narrow.** OpenAI published coding, biology, and cybersecurity results only, holding back SWE-bench, GDPval, math, and hallucination numbers\[1]. Whether GPT-5.6 fixed GPT-5.5's confident-hallucination problem is still unknown.
The full breakdown, including the four asterisks on that benchmark table, is in our [GPT-5.6 guide](/blog/how-to-use-gpt-5-6).
## If you were waiting to migrate
The practical answer for most teams is short.
**Move to Terra**, not Sol. The price cut is the release's real story, and Terra ties Claude Fable 5 on Terminal-Bench while sitting less than a point above the GPT-5.5 it replaces\[1].
**Reserve Sol for work that justifies `ultra`.** Subagent orchestration costs more tokens and more wall-clock time. It earns that on a multi-file refactor or an end-to-end audit, and wastes it on routine calls.
**Keep verification in place.** Until the expanded evaluation suite lands, pair the model with source-grounded checks for anything high-stakes.
```python
from openai import OpenAI
client = OpenAI(api_key="YOUR_REAPI_KEY", base_url="https://api.reapi.ai/v1")
resp = client.chat.completions.create(
model="gpt-5.6-terra",
messages=[{"role": "user", "content": "Refactor this module and add tests."}],
max_tokens=16000,
stream=True,
)
```
Swap the model string for `gpt-5.6-sol` or `gpt-5.6-luna` to change tiers. Note the wire ids keep the dot. Rates are on [reapi.ai/models](/models).
## FAQ
### Is GPT-5.6 out yet?
Yes. It was previewed on June 26, 2026 with restricted access, and is now generally available across all three tiers\[1].
### Why did it take so long to become available?
OpenAI ran a phased preview to roughly twenty vetted partners at the U.S. government's request, following a review of the models' cybersecurity capabilities\[1].
### What is the difference between Sol, Terra, and Luna?
Sol is the flagship and the only tier with `max` and `ultra` reasoning modes. Terra is the balanced default at roughly half GPT-5.5's price. Luna is the low-cost tier for high-volume routine work\[1].
### Which GPT-5.6 tier should I use?
Terra for most production workloads. Sol for long-horizon coding, security work, or anything that justifies `ultra`. Luna for high-volume, latency-sensitive traffic\[1].
### How much does GPT-5.6 cost?
List is $5/$30 for Sol, $2.50/$15 for Terra, and $1/$6 for Luna per million tokens. On reAPI each runs at 80% of that\[1].
### Is GPT-5.6 better than GPT-5.5?
On Terminal-Bench, yes, though Terra sits less than a point above GPT-5.5 there. The larger change is the price: Terra delivers GPT-5.5-class capability for about half the cost\[1].
### What is `ultra` mode?
A reasoning setting exclusive to Sol that spawns subagents to split complex work in parallel, rather than a single agent reasoning longer\[1].
### Did GPT-5.6 fix the hallucination problem?
Unknown. OpenAI held back hallucination and knowledge-work evaluations in the preview\[1].
## Answering the question the search was really asking
"When is GPT-5.6 coming out" was the right question for about three weeks in mid-2026, and the answer stopped being a date some time ago. What most people asking it actually want to know is whether it is worth switching, and that answer is more specific than a release note.
It is worth switching to **Terra**, because GPT-5.5-class coding at half the price is a real change to unit economics. It is worth reaching for **Sol** only where `ultra` earns its cost. And it is worth keeping verification around either, because the evaluation OpenAI published was deliberately narrow and the numbers that would settle the hallucination question were not in it.
## References
1. reAPI. *How to use GPT-5.6 — tiers, the Terminal-Bench table, reasoning modes, pricing, and the phased rollout.* [reapi.ai/blog/how-to-use-gpt-5-6](/blog/how-to-use-gpt-5-6)
2. OpenAI. *Models — GPT-5.6 specifications and reasoning modes.* Retrieved July 2026 from [platform.openai.com/docs/models](https://platform.openai.com/docs/models)
3. OpenAI. *Pricing — per-token rates by model and tier.* Retrieved July 2026 from [platform.openai.com/docs/pricing](https://platform.openai.com/docs/pricing)
### Further reading
* reAPI. *How to use GPT-5.6.* [reapi.ai/blog/how-to-use-gpt-5-6](/blog/how-to-use-gpt-5-6)
* reAPI. *How to use GPT-5.5.* [reapi.ai/blog/how-to-use-gpt-5-5](/blog/how-to-use-gpt-5-5)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# Kimi K3: The Complete Guide to Moonshot's 2.8T Flagship (https://reapi.ai/blog/kimi-k3-complete-guide)
Moonshot AI shipped **Kimi K3** in mid-July 2026, and it did so backwards: the API documentation, the pricing page, and a live `kimi-k3` model ID appeared before any benchmark table, tech blog post, or weight release\[1]. A week of leak threads filled the vacuum, and most of what circulated was wrong in at least one dimension. This guide is the version assembled from Moonshot's own documentation — what is confirmed, what is deliberately fixed, and what is still unpublished.
It is also written from a specific vantage point: we serve Kimi K3 on the [reAPI gateway](/models/kimi-k3), so the sections below cover not just what the model is, but what calling it actually looks like — the parameters that will reject your request, the caching behavior that decides your bill, and the migration traps waiting for K2.x code.
## TL;DR
* **Kimi K3 is a 2.8-trillion-parameter flagship** built on Kimi Delta Attention (a hybrid linear attention mechanism) plus Attention Residuals, with a Stable LatentMoE that activates 16 of 896 experts per token\[1]\[3].
* **The context window is 1,048,576 tokens with flat pricing** — no long-context tier — and output defaults to 131,072 tokens, raisable to the full window\[1]\[2].
* **Reasoning is always on.** The only lever is a three-rung `reasoning_effort` dial (`low` / `high` / `max`, no `medium`), and it defaults to `max`\[4].
* **Sampling is locked**: `temperature` 1.0, `top_p` 0.95, `n` 1, both penalties 0. Sending anything else returns an error\[1].
* **Moonshot's published rate is $3.00 input / $15.00 output per 1M tokens** (cache-hit input $0.30). On reAPI the same model runs $2.50 / $12.00 — below the published rate on both dimensions\[2]\[8].
* **Vision is native** (image and video input), but public image URLs are rejected — base64 or an uploaded file reference only\[5].
## Kimi K3 at a glance
Everything in this table comes from Moonshot's API documentation, retrieved July 26, 2026:
| Spec | Kimi K3 (official) |
| ------------------------ | ---------------------------------------------------------------------------------------------------------------- |
| Model ID | `kimi-k3` |
| Total parameters | 2.8 trillion |
| Architecture | Kimi Delta Attention (hybrid linear attention) + Attention Residuals; Stable LatentMoE, 16 of 896 experts active |
| Context window | 1,048,576 tokens (1M) |
| Max output | 131,072 tokens default, raisable to 1,048,576 |
| Modalities | Text, image, and video input; text output |
| Reasoning | Always on; `reasoning_effort` with `low` / `high` / `max`, default `max` |
| Input price (cache miss) | $3.00 per 1M tokens |
| Input price (cache hit) | $0.30 per 1M tokens |
| Output price | $15.00 per 1M tokens |
| Context tiering | None — flat per-token rate at any context length |
Two numbers deserve a second look. The parameter count — 2.8T — puts K3 in a class no open-weight model has occupied before, and Moonshot has committed to releasing weights, which would make it the first open-source model in the 3-trillion class\[3]. And the pricing: at $3.00 / $15.00, K3 costs three to four times what the K2 line did. Moonshot is pricing this as a frontier flagship, not the value play the Kimi name used to imply — while still undercutting the closed frontier models on both sides of the ledger.
## The launch: API first, paper trail later
The rollout order is worth understanding because it explains the conflicting information still circulating. In the days before launch, community trackers caught a promo page flashing on Moonshot's platform and beta selectors surfacing in the Kimi app; the consensus leak settled on "\~2.5T parameters, 1M context." Close, but the official figure is 2.8T. A wordless teaser video from Moonshot's official account followed, and then the product outpaced the press cycle: the API docs, quickstart, and pricing page went live the same day the first financial reporting on the launch appeared\[1]\[2].
That sequencing — API and docs before benchmarks and weights — is the reverse of how the K2 flagships launched, each of which arrived with a full tech blog and open weights on day one. The practical consequence: for the first week of K3's life, the documentation was the only reliable source, and everything else was inference.
## Architecture: Kimi Delta Attention, LatentMoE, Attention Residuals
The most technically interesting line in the K3 docs is that the model is "built on Kimi Delta Attention, a hybrid linear attention mechanism, and Attention Residuals"\[3]. This is not marketing vocabulary — it is a production deployment of the **Kimi Linear** research line Moonshot published in late 2025, and understanding it explains both the 1M-token window and the flat pricing that makes the window usable\[6].
**Kimi Delta Attention (KDA)** is a linear attention mechanism — a refinement of Gated DeltaNet with finer-grained gating. Standard softmax attention grows its key-value cache linearly with sequence length, which is what makes million-token contexts ruinously expensive to serve. Linear attention maintains a fixed-size recurrent state instead. The historical cost was quality: pure linear attention loses precise recall over long documents.
**The hybrid layout** is how the Kimi Linear paper resolved that trade-off: interleave KDA layers with full-attention layers at a 3:1 ratio, so three of every four layers use the cheap linear mechanism while every fourth retains exact global attention. In the published results this configuration outperformed full attention on quality benchmarks while cutting KV-cache memory by up to 75% and decoding up to 6x faster at 1M-token contexts\[6].
**Stable LatentMoE** is the sparse mixture-of-experts configuration: 896 experts with 16 active per token\[3]. That sparsity is what keeps a 2.8T model servable at all — the active compute per token is a small fraction of the total parameter count.
**Attention Residuals** is the one component with no published paper behind it yet; the term first appears in the K3 documentation itself. Until a technical report lands, the honest position is that we know the name and not the mechanism.
Why this matters practically: a 1M-token context is only as good as its economics. Most providers that offer very long contexts either tier the pricing upward or quietly degrade. K3's docs state there is no tiering — the same per-token rate applies whether you send 5,000 tokens or 900,000\[2]. That pricing decision is only rational if the marginal cost of long context really has collapsed, which is exactly what KDA is for. The architecture is the pricing story.
## What the API actually ships
The K3 API surface is opinionated in ways that will surprise teams migrating from K2.x or from OpenAI-style models. These are documented constraints, not observations\[1]:
* **Reasoning is always on.** There is no non-thinking mode — Moonshot's own FAQ answers "how do I turn off the chain of thought" with a flat *you can't*. The lever is `reasoning_effort`, a top-level field with three rungs: `low`, `high`, and `max`. There is no `medium`, and the default is the most expensive rung, `max`\[4].
* **Sampling is fixed.** `temperature` is locked at 1.0, `top_p` at 0.95, `n` at 1, and both penalty parameters at 0. Passing any other value returns an error, so requests should simply omit them. If your pipeline tunes temperature per task, that lever is gone.
* **Output ceilings are enormous.** `max_completion_tokens` defaults to 131,072 and can be raised to the full 1,048,576 — a single call can, in principle, emit a million tokens. Long-form generation (full codebases, book-length drafts, multi-file refactors) is clearly a design target.
* **Streaming separates reasoning from answers.** Streamed responses carry `reasoning_content` deltas distinct from `content` deltas, so a UI can render the thinking trace separately from the final answer.
* **Preserved Thinking is mandatory.** In multi-turn and tool-call loops, the complete assistant message — including `reasoning_content` and `tool_calls` — must go back into `messages` unchanged. Trimming to `content` alone, the habit most OpenAI-shaped code has, silently degrades the model.
* **Vision has sharp edges.** Image and video input are native, but public image URLs are not accepted — you send base64 data URIs or upload files and reference them by file ID\[5].
* **Structured output is first-class.** A JSON Schema with `strict: true` constrains the final answer field (not the reasoning, which is prose).
### The new tool-calling stack
Two capabilities debut with K3, and both target the same problem: agents with large tool inventories burning context on tool definitions\[1].
**`tool_choice: "required"`** forces at least one tool call on a turn — the way to stop an agent answering from memory when it was supposed to look something up. The K2 models reject this value; K3 accepts it.
**Dynamic tool loading** is the more novel one: place a complete tool definition inside a `system` message mid-conversation (a `tools` field with no `content`), and the tool becomes available from that point onward. Instead of front-loading 80 tool schemas into every request, an orchestrator injects tools exactly when a workflow phase needs them. One caveat: the server does not retain a dynamically loaded tool — keep that system message in your request history or the tool disappears next turn.
A warning from the docs worth repeating: Moonshot's built-in `web_search` tool is flagged as being updated and is not recommended for production use right now. If your agent needs search, bring your own.
### Automatic context caching
K3's context caching requires no cache IDs, TTL management, or extra parameters — keep a long prefix stable across requests and the platform attempts a cache hit on its own, billed at $0.30 per 1M tokens instead of $3.00\[1]. That is a 90% discount on repeated context, and it is the difference between the 1M window being a demo feature and a working pattern: load a corpus once as a stable prefix, then iterate against it at one-tenth the price.
Two conditions decide whether the hit happens. The previous request's prompt tokens must exceed 256 — below that, nothing is cached. And the `reasoning_effort` rung must not change: switching effort mid-conversation invalidates the prefix cache\[4]. Pick the rung before the conversation starts.
Early third-party telemetry suggests the pattern works as advertised: OpenRouter's live stats for K3 have shown a cache hit rate above 75% across traffic, pulling the weighted average input price well under a third of list\[7].
## Pricing: a deliberate break from the value playbook
Every Kimi release until now competed primarily on cost. K3 does not. Here is where it sits against its own family and the frontier competition (per 1M tokens, cache-miss input rates):
| Model | Input | Output | Cache-hit input | Context |
| ----------------------------- | --------- | ---------- | --------------- | ------- |
| Kimi K2.6 | $0.95 | $4.00 | $0.16 | 256K |
| DeepSeek V4 Pro | $1.74 | $3.48 | $0.145 | 128K |
| **Kimi K3 (Moonshot direct)** | **$3.00** | **$15.00** | **$0.30** | **1M** |
| **Kimi K3 on reAPI** | **$2.50** | **$12.00** | — | **1M** |
| Claude Opus 4.8 | $5.00 | $25.00 | — | 1M |
| GPT-5.5 | $5.00 | $30.00 | — | 400K |
The positioning is legible at a glance: K3 costs three to four times its own siblings and 40–50% less than the closed frontier. Moonshot is betting that a 1M-token window, native video input, and frontier-class capability justify a premium tier within the open-weight world.
Speed is the honest asterisk on that bet. Early third-party telemetry shows K3 producing roughly 28 tokens per second with a time-to-first-token around four seconds\[7] — unsurprising for an always-on reasoning model at this scale, but slow. K3 is built for depth-per-call, not calls-per-minute; latency-sensitive interactive loops belong on a different model.
One housekeeping note: Moonshot is pruning the lineup aggressively alongside the launch. The legacy `moonshot-v1` series and older K2.x variants are closed to new users, with a full platform sunset announced for the end of August 2026. Moonshot wants everyone on K3.
## Migrating from K2.x: a practical checklist
If you run K2.6 or K2.7-Code today, the migration surface is small but real:
1. **Swap the model ID** to `kimi-k3` — the API remains OpenAI-compatible.
2. **Replace the `thinking` parameter** with top-level `reasoning_effort`. The K2-era `thinking` object is not a K3 parameter and will error.
3. **Strip sampling parameters.** Remove `temperature`, `top_p`, and penalty settings from K3 requests — they are fixed server-side and any other value rejects.
4. **Set effort deliberately.** The default is `max`, the most expensive rung, and reasoning tokens bill as output. Drop to `high` or `low` where the task does not earn the deliberation — and do not switch rungs mid-conversation, because it invalidates your prefix cache.
5. **Re-audit cost models.** Output tokens cost 3.75x the K2.6 rate at list price, and always-on reasoning inflates output counts. Claw back spend by structuring prompts for the automatic cache: stable prefix first, variable suffix last.
6. **Rework vision ingestion** if you passed public image URLs — K3 accepts only base64 or uploaded file references\[5].
7. **Echo assistant messages back complete**, including `reasoning_content` and `tool_calls`, in every multi-turn loop.
## Who should use K3 today
Based strictly on what is confirmed, K3 makes sense right now for **long-horizon work where per-task quality dominates per-token cost**: multi-hour coding agents navigating large repositories, million-token document corpora queried repeatedly against a cached prefix, and multimodal reasoning jobs that need images, video, and code in one context.
It is the wrong choice today for latency-sensitive interactive products (28 tokens/sec is a different tool for a different job), for pipelines that need tunable sampling (everything is locked), and for pure cost-optimization plays — that is what the cheaper tiers are for.
## Running Kimi K3 through reAPI
K3 is live on the reAPI gateway as a drop-in OpenAI-compatible endpoint: swap the base URL to `https://api.reapi.ai/v1`, send your gateway key as a bearer token, and set the model string to `kimi-k3`. The same `openai` SDK you already use works unchanged.
```bash
curl https://api.reapi.ai/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k3",
"messages": [
{ "role": "user", "content": "Refactor this module and add tests." }
],
"reasoning_effort": "high",
"stream": true
}'
```
Two reasons to route K3 through reAPI rather than direct:
* **Price.** reAPI bills K3 at $2.50 input / $12.00 output per 1M tokens — below Moonshot's published rate on both dimensions, a full fifth lower on output\[8]. On a model whose always-on reasoning inflates output counts, the output rate is the number that matters.
* **One key, many models.** The same gateway key reaches Claude, GPT, Gemini, DeepSeek, GLM, and the Kimi line, which makes K3 a routing decision rather than a vendor commitment — put it behind the same fallback and load-balancing logic as everything else.
The full request surface — the fixed parameters, the caching conditions, the vision part format, the error catalog — is documented in our [Kimi K3 API docs](/docs/kimi-k3), and current rates are on the [model page](/models/kimi-k3).
## FAQ
### Is Kimi K3 open source?
Moonshot describes K3 as an open-source release and has committed to publishing weights, which would make it the first open model in the 3-trillion-parameter class\[3]. As of this writing the weights drop is imminent but the primary access path remains the API.
### Can I turn off or reduce K3's reasoning?
You cannot turn it off — reasoning is always on. You can reduce it: `reasoning_effort` accepts `low`, `high`, and `max`, with `max` as the default\[4]. `low` exists precisely for the case where the reasoning is costing more than the answer is worth.
### Does the 1M context cost extra?
No. Pricing is flat per token regardless of context length — a 900K-token prompt bills at the same rate as a 9K one\[2]. On Moonshot direct, cache-hit input drops to $0.30 per 1M tokens.
### Why did my K3 request return an error about temperature?
`temperature`, `top_p`, `n`, and both penalty parameters are fixed server-side, and sending any other value is an error rather than an override\[1]. Delete them from the request. The same applies to the K2-era `thinking` object — use top-level `reasoning_effort` instead.
### Can I use K3 in Claude Code, Codex, or Cline?
Yes — the API is OpenAI-compatible, so any tool that accepts a custom base URL and key can drive it, including Claude Code, Codex, Cline, RooCode, and OpenCode. Point the tool at `https://api.reapi.ai/v1` with a gateway key and the model id `kimi-k3`.
### How does K3 compare to Claude Opus 4.8 or GPT-5.5 on benchmarks?
Moonshot has published no official evaluation table for K3 as of this writing — no SWE-Bench, no Terminal-Bench, nothing independently verifiable. Treat specific numbers circulating on social media as unverified until the benchmark table and the technical report land.
## References
1. Moonshot AI. *Kimi K3 Quickstart — official API documentation: model ID, fixed sampling parameters, output ceilings, tool calling, context caching.* Retrieved July 2026 from [platform.kimi.ai/docs/guide/kimi-k3-quickstart](https://platform.kimi.ai/docs/guide/kimi-k3-quickstart)
2. Moonshot AI. *Kimi K3 pricing — $3.00 / $15.00 per 1M tokens, $0.30 cache-hit input, no context tiering.* Retrieved July 2026 from [platform.kimi.ai/docs/pricing/chat-k3](https://platform.kimi.ai/docs/pricing/chat-k3)
3. Moonshot AI. *Models overview — 2.8T parameters, Kimi Delta Attention + Attention Residuals, Stable LatentMoE (16 of 896 experts), open-source commitment.* Retrieved July 2026 from [platform.kimi.ai/docs/api/models-overview](https://platform.kimi.ai/docs/api/models-overview)
4. Moonshot AI. *Using reasoning\_effort — three rungs (low / high / max), default max, cache invalidation on switching.* Retrieved July 2026 from [platform.kimi.ai/docs/guide/use-reasoning-effort](https://platform.kimi.ai/docs/guide/use-reasoning-effort)
5. Moonshot AI. *Using Kimi vision models — image and video input, base64 and file references, no public URLs.* Retrieved July 2026 from [platform.kimi.ai/docs/guide/use-kimi-vision-model](https://platform.kimi.ai/docs/guide/use-kimi-vision-model)
6. Moonshot AI. *Kimi Linear: An Expressive, Efficient Attention Architecture (arXiv:2510.26692) — KDA, 3:1 hybrid layout, KV-cache and decoding results.* Retrieved July 2026 from [arxiv.org/abs/2510.26692](https://arxiv.org/abs/2510.26692)
7. OpenRouter. *MoonshotAI: Kimi K3 — live telemetry: cache hit rate, weighted input price, throughput and TTFT.* Retrieved July 2026 from openrouter.ai/moonshotai/kimi-k3
8. reAPI. *Kimi K3 — model page, API documentation, and current rates.* Retrieved July 2026 from [reapi.ai/docs/kimi-k3](/docs/kimi-k3)
### Further reading
* reAPI. *Kimi K3 API Documentation.* [reapi.ai/docs/kimi-k3](/docs/kimi-k3)
* reAPI. *Kimi K3 Model Page — current pricing.* [reapi.ai/models/kimi-k3](/models/kimi-k3)
* reAPI. *What Is Claude Opus 4.8? Anthropic's New Model Explained.* [reapi.ai/blog/what-is-claude-opus-4-8](/blog/what-is-claude-opus-4-8)
* reAPI. *What Is reAPI?* [reapi.ai/blog/what-is-reapi](/blog/what-is-reapi)
---
# Kimi K3 vs Claude Opus 5: Open Weights or Managed Reliability? (https://reapi.ai/blog/kimi-k3-vs-claude-opus-5)
**Choose Kimi K3 when open weights, native multimodal input, and deployment
control justify serious infrastructure. Choose Claude Opus 5 when you want a
managed model with a mature enterprise API and stronger published evidence for
long-running agentic work.** That is the useful Kimi K3 vs Claude Opus 5 split.
It is tempting to turn this comparison into one benchmark leaderboard. There
is no credible apples-to-apples table yet. Moonshot's K3 report predates Opus
5 in its comparison set, while Anthropic's Opus 5 launch table does not include
K3.\[1]\[2] Combining scores from different harnesses would look precise and answer the wrong question.
## TL;DR
* **Kimi K3 is open weight.** Its 2.8T-parameter MoE checkpoint is available
under the Kimi K3 License; 104B parameters activate per token.\[1]
* **Claude Opus 5 is managed only.** You get a mature API, multiple cloud
platforms, adaptive thinking, and no model-serving infrastructure.
* **Both have roughly 1M context and 128K-class default output capacity.** K3
can expose a much larger completion ceiling, but using it is expensive.
* **Hosted price does not follow openness.** On reAPI, K3 is $2.50 input / $12
output while Opus 5 is $2.40 / $12 per million tokens.
* **Self-hosting K3 is a cluster decision.** Native MXFP4 weights alone have a
raw floor around 1.4TB before runtime overhead.
* **There is no honest universal winner.** K3 wins deployment control; Opus 5
wins operational simplicity and published agent reliability.
## Kimi K3 vs Claude Opus 5 at a glance
| Category | Kimi K3 | Claude Opus 5 |
| -------------------- | ---------------------------------------- | ---------------------------------- |
| Distribution | Open weights under Kimi K3 License | Proprietary managed API |
| Architecture | 2.8T MoE, 104B active per token | Undisclosed |
| Context | 1,048,576 tokens | 1M tokens |
| Output | 128K default; API can allow up to 1M | 128K maximum |
| Input | Text, image, native visual understanding | Text and image |
| Thinking | Always on; `low`, `high`, `max` | Adaptive; `low` through `max` |
| Sampling | Fixed temperature/top-p | Vendor-controlled with effort dial |
| Official token price | $3 / $15 | $5 / $25 |
| reAPI token price | $2.50 / $12 | $2.40 / $12 |
| Self-hosting | Permitted under license; very demanding | Not available |
## Open weights change control, not just price
Kimi K3 gives teams the checkpoint. Moonshot describes it as a 2.8T-parameter
sparse Mixture-of-Experts model with 104B activated parameters, 896 routed
experts, native MXFP4 weights, and a one-million-token context window.\[1]
That enables private deployment, adaptation, infrastructure experiments, and
research that a closed API cannot support. It does not automatically mean
cheap local inference. Four bits across 2.8 trillion parameters produce a raw
weight floor of roughly 1.4TB decimal before scales, indexes, activations, and
runtime state. The active 104B parameters reduce computation, not the full
checkpoint that must remain accessible.
Claude Opus 5 makes the opposite trade. Anthropic runs the serving stack,
controls updates and safeguards, and exposes the model through Claude API,
Amazon Bedrock, Google Cloud, and Microsoft Foundry.\[3] You give up checkpoint access and gain an operational surface that can be adopted without a multi-node inference project.

## Why benchmark claims need restraint
Moonshot's official K3 report compares K3 with Claude Fable 5, Claude Opus 4.8,
GPT-5.6 Sol, GPT-5.5, and GLM-5.2. Opus 5 is absent.\[1] Anthropic's Opus 5 launch compares against Fable 5, Opus 4.8, and GPT-5.6 Sol, but not K3.\[2]
There are common benchmark names across external leaderboards, but harness,
effort, tool access, fallback rules, and dates differ. A score from Kimi Code
at `max` is not automatically comparable with a Claude Code or mini-SWE-agent
run published by another vendor.
What can be said safely:
* K3's own report positions it near the frontier on coding, knowledge work,
multimodal tasks, and browser agents.
* Anthropic's launch data positions Opus 5 strongly on long-horizon coding,
computer use, automation, and self-verification.
* Neither source establishes a controlled K3-versus-Opus-5 winner.
A production comparison should run both models with the same tools, timeouts,
task set, and success rubric.
## Reasoning and conversation state
K3 always reasons. Its top-level `reasoning_effort` accepts `low`, `high`, and
`max`, defaulting to `max`. Preserved Thinking is always active, which means a
multi-turn application must send the complete assistant message back—including
`reasoning_content` and tool calls. Changing effort mid-conversation also
invalidates prefix-cache reuse.\[4]
Opus 5 also turns thinking on by default, but its ladder has five levels from
`low` through `max`, defaulting to `high`. Thinking can be disabled at `high`
or below, although Anthropic recommends keeping it adaptive and lowering
effort to control cost.\[3]
The practical difference is flexibility. Opus lets one application reserve
maximum reasoning for hard cases and run routine work lower. K3 provides a
smaller ladder and requires the application to preserve reasoning state exactly.
## Multimodal input and long context
K3 is natively multimodal and Moonshot documents image and video understanding.
Its API does not accept arbitrary public image URLs: visual input must use a
base64 data URI or Moonshot file reference.\[5]
Opus 5 accepts image input and has a 1M context window, but it is not presented
as a native video-understanding model. Video workflows normally extract frames,
transcripts, or structured events before sending them to Claude.
For long codebases and document sets, both windows are large enough that
retrieval strategy matters more than the last 48K tokens. Sending an entire
repository on every turn can be slower and more expensive than combining
search, summaries, and targeted file loading.
## API pricing: open weight is not automatically cheaper
| Route | Input / MTok | Cached input | Output / MTok |
| ----------------------- | -----------: | -----------------------: | ------------: |
| Moonshot Kimi K3 | $3.00 | $0.30 | $15.00 |
| Anthropic Claude Opus 5 | $5.00 | $0.50 cache hit | $25.00 |
| reAPI Kimi K3 | **$2.50** | Not separately published | **$12.00** |
| reAPI Claude Opus 5 | **$2.40** | Not separately published | **$12.00** |
Direct vendor pricing makes K3 clearly cheaper. On reAPI the difference nearly
disappears: Opus input is ten cents cheaper and output costs the same. That is
a useful reminder that model licensing and hosted inference pricing are
different markets.
Self-hosting changes the calculation again. K3 eliminates per-token vendor
fees but adds accelerators, storage, networking, inference engineering,
monitoring, and underutilized capacity. It wins economically only when control
or sustained scale pays for that stack.
## Which model should you choose?
### Choose Kimi K3 when
* weights must run inside your own controlled environment;
* native image and video understanding matters;
* you can operate a large MoE inference stack;
* license-level modification and deployment control are requirements;
* a lower direct vendor token rate matters more than managed ecosystem depth.
### Choose Claude Opus 5 when
* you need production API reliability without serving the model;
* long-running coding, computer use, and self-verification are core workloads;
* deployment across Anthropic, Bedrock, Vertex AI, or Foundry matters;
* five effort levels and managed fallbacks fit the application;
* time-to-production is more valuable than checkpoint control.
### Evaluate both when
The workload is a hosted coding agent. On reAPI their output price is identical,
so a controlled A/B test can decide on completion rate, latency, and token use
without separate integrations.
## Calling Kimi K3 and Opus 5 on reAPI
```python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_REAPI_KEY",
base_url="https://api.reapi.ai/v1",
)
def ask(model, prompt):
return client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": prompt}],
max_tokens=12000,
)
kimi = ask("kimi-k3", "Find the root cause and propose a minimal patch.")
opus = ask("claude-opus-5", "Find the root cause and propose a minimal patch.")
```
For K3 multi-turn sessions, preserve the complete returned assistant message.
Do not reconstruct it from visible `content` alone.
## FAQ
### Is Kimi K3 better than Claude Opus 5?
No controlled official comparison establishes that. K3 is better when open
weights and deployment control are requirements; Opus 5 is better when managed
reliability and long-horizon agent tooling matter more.
### Is Kimi K3 open source?
It is safest to call K3 open weight. Moonshot publishes the weights under the
Kimi K3 License, which permits broad use and modification subject to its terms.
That is not the same claim as publishing every training component and dataset.
### Can Kimi K3 run on a consumer GPU?
Not as a normal local deployment. Its native four-bit weights have a raw floor
around 1.4TB, so practical self-hosting is a cluster-scale project.
### Which model has the larger context window?
They are effectively tied for ordinary planning: K3 has 1,048,576 tokens and
Opus 5 has 1M. Application-side retrieval and compaction usually matter more.
### Which model is cheaper on reAPI?
Opus 5 is slightly cheaper on input at $2.40 versus K3's $2.50 per million
tokens. Both are $12 per million output tokens at the current listed rate.
## Verdict
Kimi K3 vs Claude Opus 5 is a control-versus-operations decision before it is
a benchmark fight. K3 opens the checkpoint and asks you to carry a formidable
serving stack. Opus 5 closes the weights and removes that stack from your job.
For hosted API workloads, test both; for private deployment, K3 is the only one
of the pair that gives you the weights.
## References
1. Moonshot AI. *Kimi K3 official repository and technical report.* [github.com/MoonshotAI/Kimi-K3](https://github.com/MoonshotAI/Kimi-K3)
2. Anthropic. *Introducing Claude Opus 5.* [anthropic.com/news/claude-opus-5](https://www.anthropic.com/news/claude-opus-5)
3. Anthropic. *What's new in Claude Opus 5.* [platform.claude.com/docs/en/about-claude/models/whats-new-opus-5](https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5)
4. Moonshot AI. *Kimi K3 Quickstart and reasoning effort.* [platform.kimi.ai/docs/guide/kimi-k3-quickstart](https://platform.kimi.ai/docs/guide/kimi-k3-quickstart)
5. Moonshot AI. *Using Kimi vision models.* [platform.kimi.ai/docs/guide/use-kimi-vision-model](https://platform.kimi.ai/docs/guide/use-kimi-vision-model)
### Further reading
* reAPI. *Kimi K3 complete guide.* [reapi.ai/blog/kimi-k3-complete-guide](/blog/kimi-k3-complete-guide)
* reAPI. *Can Kimi K3 run on a 4GB GPU?* [reapi.ai/blog/unbelievable-run-kimi-k3-2-8-trillion-parameters-on-a-single-4gb-gpu](/blog/unbelievable-run-kimi-k3-2-8-trillion-parameters-on-a-single-4gb-gpu)
* reAPI. *How to use Claude Opus 5.* [reapi.ai/blog/how-to-use-claude-opus-5](/blog/how-to-use-claude-opus-5)
---
# Mammouth.ai Alternatives in 2026: 5 Tools Compared (https://reapi.ai/blog/mammouth-ai-alternatives)
Mammouth.ai is a French app that puts GPT, Claude, Gemini, Mistral, Grok, DeepSeek, and a stack of image and video models behind one subscription and one chat window\[1]. It is the "stop juggling six AI tabs" pitch, aimed at people who want frontier models without six separate logins. If you are looking at **Mammouth.ai alternatives**, you usually want the same thing through a different door: a free way to reach the same models, an API instead of a chat box, or open-source software you run yourself.
This guide compares five Mammouth.ai alternatives on what separates them in practice: which models you get, what you can do besides chat, and how you actually use the thing, whether that is a polished app, a developer API, or self-hosted software. Four are products in their own right. The fifth is reAPI, which we build, and I will be straight about it: reAPI is not a consumer chat app, so if you want a finished interface to log into and start typing, Mammouth or Poe is the honest answer, not us. Where reAPI fits is the developer who wants those same models behind one OpenAI-compatible API to put inside their own product. Everything below came from each vendor's own pages, checked in June 2026.
## TL;DR
* **Poe** is the closest match: thousands of models and user-built bots in one app, group chat, very long context, plus an OpenAI-compatible API\[4]\[5].
* **Perplexity** is a web-grounded answer engine with citations, plus a Sonar and Agent API that routes across OpenAI, Anthropic, Google, and xAI models\[6]\[7].
* **OpenRouter** is one OpenAI-compatible API (and a chat) to 400+ models across 60+ providers, with provider routing and bring-your-own-key\[8]\[9].
* **HuggingChat and LibreChat** are the open-source route: a free hosted chat over open-weight models, or a self-hosted app you point at any provider you like\[10]\[12].
* **reAPI** puts 200+ models behind one OpenAI-compatible API with a deep image and video catalog, for builders who want the models in their own product rather than a chat UI.
## What Mammouth.ai does well, and where it leaves gaps
Mammouth's strength is packing a lot of AI into one tidy app.
Where it is strong:
* **Many models, one window.** GPT, Claude, Gemini, Mistral, Grok, DeepSeek, Kimi, and more, switchable mid-conversation, with a reprompt feature that sends the same question to another model so you can compare\[1].
* **More than chat.** Image and video generation (including Sora 2, Veo 3.1 Lite, Kling, Grok Imagine), text-to-speech and voice chat, file uploads, web search, custom assistants, and a coding CLI\[1].
* **An API on the side.** Beyond the app, Mammouth exposes an OpenAI-compatible endpoint, so the same account works from code and from tools like n8n or Cursor\[2].
* **European data posture.** A Paris company that says it does not train on your data and keeps processing inside the EEA\[3].
Where teams hit walls:
* **One app's walls.** You get Mammouth's interface and its model menu. You cannot mix in a provider it has not added, and the app installs as a browser PWA rather than a native app or extension\[1].
* **Closed, frontier-first.** The catalog is curated commercial models. If you want open-weight models you can run or fine-tune yourself, that is not the center of gravity here.
* **A product to use, not a platform to build on.** The API exists, but Mammouth is shaped around a person typing in a chat window, not a team wiring models into their own software.
## How to evaluate a Mammouth.ai alternative
Five questions sort the field:
* **Which models?** Frontier commercial, open-weight, or both.
* **App or API?** A ready interface you log into, or an endpoint you call from code.
* **More than text?** Image, video, voice, web search, agents.
* **Hosted or self-run?** Someone else's servers, or your own machine.
* **Open or closed?** Open-source software you can inspect and host, or a closed product.
## Five Mammouth.ai alternatives worth comparing
### 1. Poe: the closest multi-model chat alternative
Poe, from Quora, is the most direct match. It runs on the same "every model in one app" idea, at a larger scale\[4].
* **What it does:** chat across thousands of models and user-built bots in one interface, generate images, video, and audio, run web search, talk to several bots at once, build your own bots, and sync across devices, with context windows up to 2M tokens\[4].
* **Models:** GPT-5.x, Claude Opus and Sonnet, Gemini, Grok, GLM, DeepSeek, Qwen, and Kimi, plus image models (Nano-Banana-2, GPT-Image-2, FLUX, Ideogram) and video models (Veo 3.1, Sora 2, Kling, Runway)\[4].
* **Developer access:** an OpenAI-compatible API at `api.poe.com` supporting both Chat Completions and Responses formats, plus a creator platform for building and monetizing bots\[5].
* **For:** anyone who wants the widest model and bot selection in a single consumer app, with an API on top.
* **Vs Mammouth:** the same one-app, many-models concept, with a broader catalog and a bot ecosystem, and also an OpenAI-compatible API. The trade is that Mammouth's interface is tighter and more curated.
### 2. Perplexity: for web-grounded answers and research
If most of your Mammouth use is asking questions and getting sourced answers, Perplexity is built around exactly that\[6].
* **What it does:** real-time, web-wide research and Q\&A with citations, using its own Sonar models plus a deep-research mode\[6].
* **Models:** the Sonar family (Sonar, Sonar Pro, Sonar Reasoning Pro, Sonar Deep Research), and an Agent API that routes to OpenAI, Anthropic, Google, and xAI models\[6]\[7].
* **Developer access:** a Sonar API and an OpenAI-compatible Agent API with built-in web search and tools\[7].
* **For:** research-heavy work where grounded, cited answers matter more than a wide model menu.
* **Vs Mammouth:** narrower as a general model switchboard, deeper on web-grounded answering. Mammouth actually bundles Perplexity's Sonar for its own search, so picking Perplexity is going to the source.
### 3. OpenRouter: for one API key across every model
OpenRouter is the alternative for people who would rather have an API than an app\[8].
* **What it does:** a unified, OpenAI-compatible interface to 400+ models across 60+ providers, with provider routing and fallback, bring-your-own-key, data-policy routing, and a web chat on top\[8]\[9].
* **Models:** 400+, including Claude Opus 4.8, GPT-5.5, Gemini 3.1 Pro, and most open-weight families\[8].
* **Developer access:** the OpenAI API specification works out of the box, with one key and aggregated billing across providers\[9].
* **For:** developers and power users who want maximum model reach behind one key, with a chat available when they want it.
* **Vs Mammouth:** far more models and explicit routing control, but API-first. The chat is a utility, not a polished consumer product the way Mammouth's is.
### 4. HuggingChat and LibreChat: the open-source route
If what you like about Mammouth is "many models in one place" but you want it free or self-hosted, the open-source world has two answers\[10]\[12].
* **HuggingChat:** a free hosted chat over open-weight models (Llama, Qwen, DeepSeek, Gemma, GLM, gpt-oss, Kimi), with an "Omni" router that picks a model for you, MCP tool support, image generation, and voice input. The code is Apache-2.0 and self-hostable\[10]\[11].
* **LibreChat:** a self-hosted, MIT-licensed app that unifies Anthropic, OpenAI, Google, Ollama, Mistral, OpenRouter, Perplexity, and any OpenAI-compatible endpoint, with agents, a code interpreter, web search, memory, artifacts, and file RAG\[12].
* **For:** people who want open-weight models for free (HuggingChat) or a private app they run on their own infrastructure with their own keys (LibreChat).
* **Vs Mammouth:** you trade Mammouth's packaged, frontier-first product for open-source control. HuggingChat leans toward open-weight models; LibreChat is bring-your-own-everything.
### 5. reAPI: for the same models behind one developer API
reAPI is a developer platform, not a consumer chat app. Up front: there is no polished interface to log into and start chatting the way Mammouth gives you, and no custom-assistant builder. What it gives you is 200+ models behind one OpenAI-compatible API, with a deep image and video catalog under the same key.
* **What it does:** chat completions for frontier LLMs (GPT, Claude, Gemini), image models (GPT-Image-2, Gemini 3 Pro Image), and a deep video catalog (Veo 3.1, Seedance, Wan, Kling), OpenAI-compatible for chat and on REST endpoints for media.
* **Developer access:** one base URL and a bearer key. The OpenAI SDK works by changing the base URL.
* **For:** developers who want Mammouth's breadth of models inside their own app or workflow, not in someone else's chat UI.
* **Vs Mammouth:** both speak the OpenAI format, and Mammouth even offers its own API. reAPI's difference is being API-first across 200+ models with a deep media catalog, rather than a consumer subscription app. If you want a finished chat product to use, Mammouth or Poe is the fit. If you want the models as an endpoint to build on, that is reAPI.
## Mammouth.ai vs. the alternatives at a glance
| Tool | Models | Beyond chat | How you use it | Open-source |
| ----------------------- | ----------------------------------------------- | --------------------------------- | --------------------------- | ---------------------- |
| Mammouth.ai | Frontier commercial (GPT, Claude, Gemini, more) | Image, video, voice, search, CLI | Web app (PWA) + API | No |
| Poe | Thousands, plus user bots | Image, video, audio, search, bots | Web + iOS/Android app + API | No |
| Perplexity | Sonar, plus routed GPT/Claude/Gemini/xAI | Web search, research, citations | App + API | No |
| OpenRouter | 400+ across 60+ providers | Routing, BYOK, chat utility | API-first + chat | No (service) |
| HuggingChat / LibreChat | Open-weight / any provider | Search, tools, agents, media | Hosted chat / self-hosted | Yes (Apache-2.0 / MIT) |
| reAPI | 200+, incl. deep media | Image and video under one key | API (no consumer app) | No |
Capability claims are from each vendor's official pages as of June 2026; features change, so confirm before you commit.
## The split that actually decides it
Two questions sort this list faster than any feature table.
First, do you want an app or an API? Mammouth, Poe, and HuggingChat are finished apps you log into. OpenRouter and reAPI are APIs first, where OpenRouter bolts on a chat and reAPI stays endpoints. LibreChat sits in the middle: an app, but one you host yourself.
Second, open or closed, hosted or self-run? If you want open-weight models for free, HuggingChat. If you want a private app on your own server, LibreChat. If you want the widest closed-model reach in one consumer app, Poe. And if you want those models built into your own product, OpenRouter or reAPI.
## Calling reAPI from the OpenAI SDK
Because reAPI speaks the OpenAI format, moving a chat call over is a base-URL change, and image and video run on REST endpoints under the same key.
```python
from openai import OpenAI
client = OpenAI(
base_url="https://reapi.ai/api/v1",
api_key="rk_live_YOUR_REAPI_KEY",
)
resp = client.chat.completions.create(
model="claude-opus-4-8",
messages=[{"role": "user", "content": "Summarize this thread in three bullets."}],
)
```
The honest way to choose is to run a real workload through one or two of these and compare model coverage and output for yourself, rather than trusting any single capability table, this one included.
## FAQ
### Does Mammouth.ai have an API?
Yes. Alongside the app, Mammouth exposes an OpenAI-compatible API at `api.mammouth.ai/v1`, so you can call its models from code or from tools like n8n, Cursor, and VS Code\[2].
### Which Mammouth.ai alternative is free?
HuggingChat is a free hosted chat over open-weight models, and LibreChat is open-source and self-hosted, where you pay only for your own provider keys and hosting\[10]\[12].
### Which Mammouth.ai alternative is closest to it?
Poe. It runs the same many-models-in-one-app idea at a larger scale, adds a bot ecosystem, and also exposes an OpenAI-compatible API\[4]\[5].
### Which alternative has the most models?
OpenRouter lists 400+ models across 60+ providers, and Poe lists thousands once you count user-built bots\[8]\[4].
### Is reAPI a chat app like Mammouth?
No. reAPI is a developer API for 200+ models, OpenAI-compatible for chat and REST for media. For a ready chat interface, use Mammouth, Poe, or HuggingChat; for an endpoint to build on, use reAPI.
### Which Mammouth.ai alternative can I self-host?
LibreChat (MIT) and HuggingChat's chat-ui (Apache-2.0) are both open-source and self-hostable, so you can run a multi-model chat on your own infrastructure\[11]\[12].
## Picking a Mammouth.ai alternative
Mammouth.ai does a clear job: it packs frontier models, image and video generation, and voice into one curated app with a European data posture. The reason to look at a Mammouth.ai alternative is usually a specific one. Poe for the widest model and bot selection in a consumer app, Perplexity for web-grounded research, OpenRouter for one API key across every model, and HuggingChat or LibreChat when you want open-source and free or self-hosted. If what you really want is those same models behind one OpenAI-compatible API to build into your own product, reAPI is the alternative aimed at developers rather than chat users, with the honest caveat that it is not a finished app you log into. Run a real task through two of them, and let model coverage and the way you actually work decide which Mammouth.ai alternative fits.
## Further reading
* [reapi.ai/models](/models) lists the full catalog of LLM, image, and video models behind one key.
* [reapi.ai/docs/api/quickstart](/docs/api/quickstart) covers the OpenAI-compatible chat endpoint in a few lines.
* [Best CometAPI alternatives](/blog/best-cometapi-alternatives) runs the same comparison for a unified API gateway.
## References
1. Mammouth.ai. *Product overview and available models.* Retrieved June 2026 from mammouth.ai
2. Mammouth.ai. *API quick start — OpenAI-compatible endpoint.* Retrieved June 2026 from info.mammouth.ai/docs/api-quick-start
3. Mammouth.ai. *About privacy — data handling and EEA processing.* Retrieved June 2026 from info.mammouth.ai/docs/about-privacy
4. Poe. *About — models, bots, and capabilities.* Retrieved June 2026 from [poe.com/about](https://poe.com/about)
5. Poe. *Creator platform — OpenAI-compatible API.* Retrieved June 2026 from [creator.poe.com/docs/external-applications/openai-compatible-api](https://creator.poe.com/docs/external-applications/openai-compatible-api)
6. Perplexity. *Getting started — overview and models.* Retrieved June 2026 from [docs.perplexity.ai/getting-started/overview](https://docs.perplexity.ai/getting-started/overview)
7. Perplexity. *Agent API — multi-provider models and tools.* Retrieved June 2026 from [docs.perplexity.ai/getting-started/pricing](https://docs.perplexity.ai/getting-started/pricing)
8. OpenRouter. *The unified interface for LLMs.* Retrieved June 2026 from openrouter.ai
9. OpenRouter. *FAQ — OpenAI compatibility, routing, and keys.* Retrieved June 2026 from openrouter.ai/docs/faq
10. HuggingChat. *Chat with open-source AI models.* Retrieved June 2026 from [huggingface.co/chat](https://huggingface.co/chat/)
11. Hugging Face. *chat-ui — the open-source codebase powering HuggingChat.* Retrieved June 2026 from [github.com/huggingface/chat-ui](https://github.com/huggingface/chat-ui)
12. LibreChat. *The open-source AI platform.* Retrieved June 2026 from [librechat.ai](https://www.librechat.ai/)
---
# Mammouth AI Pricing: Plans, Limits, and API Access (2026) (https://reapi.ai/blog/mammouth-ai-pricing-plans-limits-2026)
**Mammouth AI has three monthly chat plans, but the currency and tax display depends on location.** In a US-region view checked on July 30, 2026, Starter, Standard, and Expert appeared at $12, $24, and $72 before VAT; European visitors could see €12, €24, and €72 with VAT included. The plans are not unlimited: quotas fully renew after three hours, with Standard receiving three times the Starter quota and Expert receiving ten times the Starter quota.\[1]\[2]
The API is a separate meter. Each subscription includes a small monthly API credit, but developers can also fund Mammouth's OpenAI-compatible API on a pay-as-you-go basis without buying a chat subscription.\[3] That distinction is the key to understanding Mammouth AI pricing: the subscription pays for the Mammouth app and its renewable usage allowance; API calls consume a dollar-denominated API balance.
## Mammouth AI plans at a glance
| Plan | Monthly price | Chat quota | Included API credit | Best fit |
| -------- | -------------------------: | --------------: | ------------------: | ---------------------------------------------------- |
| Starter | $12 in the checked US view | Reference quota | $2/month | Regular individual use |
| Standard | $24 in the checked US view | 3× Starter | $4/month | Frequent reasoning, document, or image work |
| Expert | $72 in the checked US view | 10× Starter | $10/month | Heavy professional use across models and media tools |
The dollar prices above are the US-region display checked on July 30, 2026; Mammouth's official page localizes currency and tax treatment, so verify the amount shown for your location. The included API credits come from Mammouth's API documentation.\[1]\[3] Mammouth also offers team workspaces at the selected plan's per-user price, with centralized billing and shared projects; enterprise pricing requires a quote.
All three individual plans include the same broad product: a hosted chat interface with access to model families from providers such as OpenAI, Anthropic, Google, Mistral, xAI, DeepSeek, and others. The app also lists web research, image and video generation, projects, voice features, and coding tools. Paying for a higher plan primarily buys more usage, not a completely different application.
## How the three-hour quota works
Mammouth does not publish a single “messages per month” number. Its quota documentation says limits are defined per session and fully renew after three hours.\[2] The amount consumed by one exchange depends on what that exchange asks the service to do.
Long prompts and responses cost more quota than short ones. Large documents add input. A long conversation gets progressively more expensive because previous context is sent again. Web search and document-generation tools add work, and a more capable model consumes more than a lighter one.
That makes the headline multipliers more useful than trying to translate each plan into an exact message count:
* **Starter is the reference allowance.** It suits regular chat and occasional use of expensive tools.
* **Standard provides three times the Starter quota.** It is the safer middle plan for daily document work, reasoning models, or repeated image generation.
* **Expert provides ten times the Starter quota.** It is aimed at users who regularly run the expensive models or media tools.
At the displayed 12/24/72 price points, Expert costs six times as much as Starter but carries ten times the relative quota. That is a better quota-to-price ratio, although it only saves money if the extra allowance is actually used.
### Hitting a quota does not always stop the chat
Mammouth uses model thresholds. If a user exhausts the allowance for a demanding model, the app can switch the conversation to a lighter model until the quota renews. Its documentation gives a Claude example: Opus can fall back to Sonnet, then Sonnet to Haiku.\[2]
This keeps a session running, but it also means “I can still send messages” is not proof that the original model remains active. Anyone choosing a plan for a specific premium model should watch the model shown in the interface during a long session and test the workload across a full three-hour window.
## Subscription quota and API credit are different
The easiest pricing mistake is to treat the included API credit as a cash version of the app quota. It is not.
The chat subscription covers use inside Mammouth's hosted product, subject to the renewable quota. The included API credit is a separate monthly balance for calls made with an API key:
| Subscription | Monthly API credit included |
| ------------ | --------------------------: |
| Starter | $2 |
| Standard | $4 |
| Expert | $10 |
Mammouth's API guide states that all subscribers receive those credits. It also allows pay-as-you-go funding from the API settings, so an API-only developer does not need to buy a Starter subscription first.\[3]
API usage is priced by model and tokens. The published model table contains separate input and output rates, and Mammouth warns that the table may change; its Model Explorer is the current source for the full catalog. The service says the displayed rates are upper bounds and that the actual charge can be lower when provider availability permits.\[3]
In practice:
* buy a chat plan when people will work in the Mammouth interface;
* use PAYG when software, an automation, or an internal tool will make the calls;
* count the included API credit as a small developer allowance, not as the plan's main value;
* price an API workload from its actual model, input tokens, and output tokens rather than the subscription multiplier.
## What Mammouth API access includes
Mammouth documents an OpenAI-compatible chat-completions API at `https://api.mammouth.ai/v1`. Existing OpenAI client code can generally point to that base URL, use a Mammouth key, and keep the familiar message format.\[3]
The API is useful when a team wants one endpoint for several LLM providers, or needs to connect models to tools such as n8n, Make, Cline, or an internal application. It is not the same product experience as the Mammouth chat app. Projects, hosted conversations, voice controls, and other interface features do not automatically appear in an API integration.
Two budgeting details are easy to overlook:
1. A long system prompt and accumulated conversation history are input tokens on every request.
2. Model aliases can change their underlying route over time. Mammouth's `mammouth-recommended` alias is explicitly designed to follow its current price-performance pick, so pin a model ID when reproducibility matters.
## Document and file limits
Quota is only one limit. Mammouth's documentation also publishes hard boundaries for uploaded material:\[2]
| Limit | Published maximum |
| ---------------------------- | -------------------: |
| Total input length | 4,000,000 characters |
| Files in one conversation | 20 |
| Combined file size | 100 MB |
| Individual PDF | 100 MB |
| Individual non-PDF file | 20 MB |
| Image-only/scanned PDF | 50 pages and 20 MB |
| Standard extracted text | 30,000 characters |
| Large-context extracted text | 150,000 characters |
The four-million-character input ceiling includes the document content, prompt, and custom instructions. For long or multiple documents, Mammouth may extract the portions it considers most relevant instead of passing every character into the selected model. A file being accepted therefore does not guarantee that every page enters the model context.
For document-heavy work, test retrieval with questions whose answers sit near the beginning, middle, and end of the source. If the application misses late or obscure details, splitting the document by topic can be more effective than upgrading solely for a larger quota.
## Which Mammouth plan makes sense?
**Choose Starter** if use is mostly ordinary chat, search, and occasional creative work. It is also the lowest-cost way to evaluate the complete hosted app, though developers testing only the API can start with PAYG instead.
**Choose Standard** when Starter's premium-model threshold is reached regularly within the three-hour window. The threefold quota increase is a meaningful step without the large jump to Expert.
**Choose Expert** when expensive reasoning, large documents, or media generation are part of daily paid work. The tenfold quota is the strongest value per unit, but the top-tier monthly price can be difficult to justify when usage remains intermittent.
Do not upgrade to obtain API access alone. First estimate the API bill from the intended models and token volume, then compare PAYG with the included balance. Likewise, do not assume Expert removes every threshold: Mammouth still describes the service as quota-based.
## Where reAPI fits—and where it does not
Disclosure: we operate reAPI. reAPI is a developer API service, not a Mammouth-style consumer subscription, so it should not be presented as a replacement for Mammouth's hosted chat, projects, or three-hour app allowance.
It becomes relevant when the requirement shifts from “which Mammouth chat plan should I buy?” to “which API should my application call?”
| Decision point | Mammouth | reAPI |
| -------------- | -------------------------------------------------------- | ----------------------------------------------------------------------- |
| Main product | Hosted multi-model chat and creative app | Developer API for models inside your own product |
| API focus | OpenAI-compatible LLM access beside the subscription app | OpenAI-compatible chat plus dedicated image, video, and audio endpoints |
| Billing | App subscription plus optional API PAYG balance | PAYG credits without a required monthly subscription |
| User interface | Finished chat, projects, files, and creator tools | Model playground and API; your team builds the customer interface |
| Best fit | A person who wants one ready-made AI workspace | A SaaS or internal tool that needs programmable model access |
If you want a finished chat product, Mammouth is the closer fit. If you want to put models inside software you own, browse the [reAPI model catalog](/models) and run the first workload through the [API quickstart](/docs/api/quickstart). Compare live model availability, output quality, data handling, and total cost on the same representative requests.
Readers who have concluded that Mammouth itself is the wrong product can use the separate [Mammouth AI alternatives guide](/blog/mammouth-ai-alternatives). That article covers consumer and developer replacements in more depth.
## FAQ
### How much does Mammouth AI cost?
In the US-region view checked on July 30, 2026, Mammouth showed Starter at $12 per month, Standard at $24, and Expert at $72 before VAT. Other regions can show local currency and tax-inclusive prices. Standard includes three times the Starter quota; Expert includes ten times the Starter quota.\[1]\[2]
### Is Mammouth AI unlimited?
No. Usage quotas fully renew after three hours. When a model threshold is reached, Mammouth may switch the conversation to a lighter model until the allowance renews.\[2]
### Does a Mammouth subscription include API access?
Yes. Starter, Standard, and Expert include $2, $4, and $10 of monthly API credit respectively. The API balance is separate from the app's chat quota.\[3]
### Can I use the Mammouth API without a subscription?
Yes. Mammouth documents pay-as-you-go API funding without requiring a consumer chat subscription.\[3]
### When do Mammouth quotas reset?
Mammouth says quotas are fully renewed after three hours. Consumption varies with the selected model, conversation length, prompt and response length, uploaded documents, and tools used.\[2]
### Is VAT included in Mammouth's listed price?
It depends on location. The official page showed prices excluding VAT in a US-region view and VAT-inclusive euro prices in a European view when checked on July 30, 2026. Use the currency and tax treatment shown for your region at checkout.\[1]
### Is reAPI a Mammouth alternative?
It is an alternative to Mammouth's API layer, not to the finished Mammouth chat app. Choose Mammouth when people need its hosted workspace. Choose reAPI when developers need chat, image, video, or audio models behind an application they control.
## References
1. Mammouth AI. *Pricing.* Region-localized Starter, Standard, Expert, team, and enterprise plan information. Retrieved July 30, 2026 from mammouth.ai/pricing
2. Mammouth AI. *Quota System & Usage Limitations.* Three-hour renewal, plan multipliers, fallback behavior, and document limits. Retrieved July 30, 2026 from info.mammouth.ai/docs/quota-policy
3. Mammouth AI. *API Documentation.* Included monthly credits, PAYG access, OpenAI compatibility, model pricing, and aliases. Retrieved July 30, 2026 from info.mammouth.ai/docs/api-quick-start
---
# MiniMax H3 vs Seedance 2.5: Which AI Video Model Wins? (https://reapi.ai/blog/minimax-h3-vs-seedance-2-5)
**Choose MiniMax H3 when you need a documented video API today; wait for Seedance 2.5 when 30-second continuity and timestamp-directed editing matter more than immediate access.** H3 ships with 4–15-second 2K output, native stereo sound, mixed references, and published pricing. ByteDance has demonstrated a more ambitious 30-second Seedance workflow, but not its final API model ID, resolution, limits, or price.\[1]\[2]
That availability gap prevents a clean quality verdict. This comparison focuses on verified capability and production fit rather than ranking selected launch demos.
## TL;DR
* **MiniMax H3 is callable now:** 4–15 seconds, up to 2K, native stereo audio, and a public V2 API.\[1]
* **Seedance 2.5 promises longer continuity:** ByteDance officially demonstrates 30-second output, expanded references, timed segments, and video editing.\[2]
* H3 documents up to nine images, three videos, and three audio references. Seedance 2.5's final public reference limit is not published.
* MiniMax lists H3 at $0.13/s for 2K and $0.09/s for 768p; no official 2.5 price exists.\[3]
* H3 is the safer choice for a product launch this week. Seedance 2.5 is the more interesting migration target for long, tightly directed scenes.
* Developers can test H3 now on [reAPI's MiniMax H3 model page](/models/minimax-h3); [Seedance 2.0](/models/seedance-2-0) is the callable fallback while 2.5 remains unpriced.
## MiniMax H3 vs Seedance 2.5 specs

| Capability | MiniMax H3 | Seedance 2.5 |
| ---------------------- | ---------------------------------------------------------------------------------------------------- | -------------------------------------------------- |
| Status on August 2 | Public launch and documented API | Official demos; public API contract not found |
| Output duration | 4–15 seconds | 30 seconds demonstrated |
| Resolution | Up to 2K | Not stated on official 2.5 page |
| Audio | Native stereo audio | Multilingual speech/lip-sync demonstrated |
| Reference inputs | Up to 9 images, 3 videos, 3 audio clips | Expanded multimodal input; final limit unpublished |
| First/last frame | Documented image-to-video mode | Not yet documented as an API field |
| Timeline direction | Prompt-directed | Second-level timed segments demonstrated |
| Existing-video editing | Generalized reference/editing is part of the model design; API surface needs task-level verification | Selective edit demonstrated |
| Direct price | $0.13/s 2K; $0.09/s 768p | Unpublished |
| Weights | Announced for release; verify current license/download | No open-weight announcement |
The table deliberately leaves cells uncertain where the vendor has not published a contract. A promotional capability and a stable API parameter are different evidence levels.
## H3 wins current API availability
MiniMax launched H3 on July 31 and published an API guide covering text-to-video, first/last-frame image-to-video, and reference-to-video. The output duration accepts whole seconds from four through 15, and the same model generates native stereo sound.\[1]
Seedance 2.5 has an official product page, but ByteDance's live model catalog and price sheet still list the Seedance 2.0 family. A developer cannot safely design a request form around rumored 2.5 parameters.\[4]
For production planning, documented constraints beat a more exciting demo. H3 lets a team calculate cost, validate media inputs, handle errors, and ship.
## Seedance 2.5 wins the announced duration story
Thirty seconds in one pass reduces edit seams. A 60-second sequence needs at least four 15-second H3 generations, but only two 30-second Seedance generations in the ideal case. Every removed join lowers the chance of wardrobe, lighting, voice, and screen-direction drift.
Longer output also concentrates failure cost. If second 28 is unusable, a 30-second render may need a full reroll. Seedance's timed prompt segments could mitigate that by making the clip behave like a compact shot list, but the public release will determine how precisely actions land.
Choose duration based on continuity, not headline size. Short product beats and social shots fit H3's current range. Dialogue scenes and longer blocking are where Seedance 2.5 could matter more.
## H3 has the clearer native-audio contract
H3 generates stereo sound and video together. Prompts can direct dialogue, music, ambience, and effects, while audio files can also serve as references when paired with an image or video.\[1]
ByteDance demonstrates multilingual performance and synchronized presentation for Seedance 2.5, so audio is plainly part of the product direction. What is missing is the public parameter reference: supported audio formats, duration limits, pricing, and whether output audio can be disabled or separately controlled.
For an automated pipeline that must validate inputs today, H3 wins. For multilingual performance quality, neither launch reel replaces a fixed test set with the voices and languages you actually use.
## Reference control is a different kind of strength
H3 publishes exact mixed-reference ceilings: nine images, three videos, and three audio files, with total video and audio reference durations capped at 15 seconds each. It asks the creator to describe relationships in natural language—identity from one image, movement from a video, voice from audio.\[1]
Seedance 2.5 demonstrates individually addressed images and promises expanded multimodal references, but ByteDance has not published the final allocation. The more distinctive feature is timed direction: assigning events to parts of a 30-second clip and editing an existing video without regenerating everything.
H3 currently offers a known reference budget. Seedance 2.5 offers a potentially richer directing model. Do not turn “potentially” into an API schema.
## Price comparison
MiniMax's direct rate makes H3 easy to estimate:
| Example output | 768p direct | 2K direct |
| -------------- | ----------: | --------: |
| 6 seconds | $0.54 | $0.78 |
| 10 seconds | $0.90 | $1.30 |
| 15 seconds | $1.35 | $1.95 |
Input-video seconds and extra images can add charges under MiniMax's published rules.\[3] Routed providers have their own prices; for example, [reAPI's MiniMax H3 page](/models/minimax-h3) should be checked for the live rate rather than assuming MiniMax direct pricing.
There is no honest Seedance 2.5 cost table yet. Any article assigning it a per-second price is using an estimate, a provider waitlist, or a Seedance 2.0 rate. Budget comparisons should wait for the official row.
## Which model should you choose?
### Choose MiniMax H3 when
* the application must launch now;
* 2K delivery is useful;
* native stereo audio and mixed references are required;
* a 15-second shot ceiling is acceptable;
* cost predictability matters more than maximum duration.
For this route, start with [MiniMax H3 on reAPI](/models/minimax-h3). The live page exposes the request modes, current rate, and a playground, so the model decision can be tested with real inputs instead of launch examples.
### Wait for Seedance 2.5 when
* one-pass 30-second scenes could remove costly continuity joins;
* timestamped direction and selective video editing are central to the workflow;
* you can run a controlled evaluation after the API appears;
* the project does not depend on a current public price or model ID.
### Keep Seedance 2.0 when
* your existing pipeline is stable;
* the migration would not solve a measurable production problem;
* you need documented Seedance behavior while 2.5 remains unavailable.
## FAQ
### Is MiniMax H3 better than Seedance 2.5?
H3 is better for immediate API implementation because its contract is public. There is not enough reproducible evidence to declare it the higher-quality model overall.
### Which model makes longer videos?
Seedance 2.5 officially demonstrates 30-second continuous generation. MiniMax H3 documents a 4–15-second range.
### Does MiniMax H3 generate audio?
Yes. It generates native stereo sound and supports audio references alongside image or video references.
### Does Seedance 2.5 support 4K?
The official 2.5 promotion does not state an output resolution. Native 4K claims should remain unverified until ByteDance publishes the model specification.
### Can I call both through an API?
MiniMax H3 has a public API and is callable through [reAPI](/models/minimax-h3). Seedance 2.5 did not have a verified public model ID at the time of this comparison; [Seedance 2.0](/models/seedance-2-0) remains callable through the same reAPI account and task workflow.
## The practical verdict
MiniMax H3 wins the decision you can make today: it has a model ID, limits, resolution, native audio, and price. Seedance 2.5 wins the more ambitious preview, with twice the continuous duration and a stronger editing narrative. Ship with [H3](/models/minimax-h3) or [Seedance 2.0](/models/seedance-2-0) on reAPI now, then test 2.5 against the same prompts when its contract—not another demo—arrives.
## References
1. MiniMax. *MiniMax H3 launch and Video Generation Guide.* Published July 31, 2026. [minimax.io](https://www.minimax.io/blog/minimax-h3)
2. ModelArk. *Doubao Seedance 2.5 official promotion.* Retrieved August 2, 2026. [ark.volcengine.com](https://ark.volcengine.com/promotion?modelName=seedance-2-5)
3. MiniMax API. *Pay-as-you-go pricing.* Retrieved August 2, 2026. [platform.minimax.io](https://platform.minimax.io/docs/guides/pricing-paygo)
4. Volcano Engine. *Current Seedance model catalog.* Retrieved August 2, 2026. [volcengine.com](https://www.volcengine.com/docs/82379/1330310)
### Further reading
* reAPI. *Hailuo H3 API guide.* [reapi.ai/blog/hailuo-h3-minimax-h3-api-guide](/blog/hailuo-h3-minimax-h3-api-guide)
* reAPI. *Seedance 2.5 launch status.* [reapi.ai/blog/seedance-2-5-release-status](/blog/seedance-2-5-release-status)
* reAPI. *Best Seedance 2.5 alternatives.* [reapi.ai/blog/best-seedance-2-5-alternatives](/blog/best-seedance-2-5-alternatives)
---
# MiniMax M3 API: 1M Context, Pricing, and Coding Guide (2026) (https://reapi.ai/blog/minimax-m3-api-guide)
The **MiniMax M3 API** is built for coding agents, long-running tool workflows,
and multimodal analysis. It accepts text, images, and video, supports up to one
million tokens of context, and exposes thinking controls for trading latency
against deeper reasoning. The open weights are also available, but the hosted
API is the practical route for teams that do not want to operate an 800GB-class
checkpoint.
This guide separates MiniMax's direct API behavior from the current reAPI route.
That distinction matters because model IDs, pricing bands, and request surfaces
are not identical.
## TL;DR
* MiniMax released M3 on June 1, 2026, then published the model weights under
the MiniMax Community License.\[1]
* The hosted model supports up to **1M tokens**, with native text, image, and
video input.\[2]
* MiniMax prices direct API calls differently below and above 512K input tokens.
* reAPI currently uses one flat rate: **$0.60 input, $2.40 output, and $0.12
cache read per million tokens**.
* Use M3 for repository analysis, coding agents, document-heavy work, and
multimodal tasks. Do not assume that a 1M window makes retrieval, compaction,
or prompt organization unnecessary.
## MiniMax M3 API specifications
| Specification | MiniMax M3 |
| ------------------------- | ------------------------------------ |
| Release date | June 1, 2026 |
| Context window | Up to 1M tokens |
| Guaranteed hosted context | At least 512K tokens |
| Input modalities | Text, image, video |
| Output | Text |
| Reasoning modes | Enabled, adaptive, disabled |
| Recommended sampling | `temperature: 1.0`, `top_p: 0.95` |
| reAPI model ID | `minimax/minimax-m3` |
| reAPI endpoint | `/v1/chat/completions` |
| Weights | Published on Hugging Face and GitHub |
MiniMax describes M3 as a coding and agentic model powered by MiniMax Sparse
Attention. The architecture is designed to reduce the compute and memory cost
of million-token attention. That is a vendor claim about its architecture, not
a guarantee that every one-million-token request will be fast or equally
accurate across the entire sequence.
The published checkpoint is about 427 billion parameters and the main Hugging
Face repository is roughly 854GB before local serving overhead. Local deployment
is possible through frameworks such as SGLang, vLLM, Transformers, and
KTransformers, but it is not a casual single-GPU model.\[1]
## MiniMax M3 API pricing explained
MiniMax's direct API separates ordinary and long-context calls. Its current
promotional rate card lists the following standard-service prices:
| Route | Input / 1M | Output / 1M | Cache read / 1M |
| ----------------------------- | ---------: | ----------: | --------------: |
| MiniMax direct, input ≤512K | $0.30 | $1.20 | $0.06 |
| MiniMax direct, input 512K–1M | $0.60 | $2.40 | $0.12 |
| reAPI current rate | $0.60 | $2.40 | $0.12 |
The direct prices above are discounted rates shown by MiniMax on August 1,
2026\. The page also displays higher crossed-out list prices, so production
budgets should store the retrieval date rather than treating the discount as a
permanent property of the model.\[3]
reAPI uses a flat token rate instead of changing the price at the 512K boundary.
This is easier to estimate, although short-context calls can be cheaper through
MiniMax's direct promotional tier.

### A realistic cost example
Suppose a repository-analysis job sends 180,000 uncached input tokens and
returns 12,000 output tokens through reAPI:
```text
input = 180,000 / 1,000,000 × $0.60 = $0.1080
output = 12,000 / 1,000,000 × $2.40 = $0.0288
total = $0.1368
```
The same workflow becomes cheaper when repeated prefixes hit the cache. Put
stable tool definitions, repository instructions, and system context first;
keep changing task details near the end. MiniMax says automatic prompt caching
uses prefix matching and applies to inputs of at least 512 tokens.\[4]
## How to call MiniMax M3 through reAPI
The reAPI route uses the familiar OpenAI Chat Completions shape. The important
detail is the namespaced model ID.
```bash
curl https://api.reapi.ai/v1/chat/completions \
-H "Authorization: Bearer $REAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "minimax/minimax-m3",
"messages": [
{
"role": "system",
"content": "Review code conservatively. State assumptions before editing."
},
{
"role": "user",
"content": "Find the race condition in this queue worker and propose a minimal patch."
}
],
"temperature": 1,
"top_p": 0.95,
"max_tokens": 8192
}'
```
Use the exact ID returned by model discovery when integrating with a production
key. The [MiniMax M3 model page](/models/minimax-m3) and
[API documentation](/docs/minimax-m3) show the current reAPI surface.
MiniMax's direct endpoint uses `MiniMax-M3` rather than the reAPI namespaced ID.
Do not copy one provider's model string into another provider's request and
assume the gateway will normalize it.
## Thinking modes and agent behavior
The open model card documents three reasoning modes:
* `enabled`: always use explicit reasoning.
* `adaptive`: let M3 decide when additional reasoning is useful.
* `disabled`: minimize latency and maximize throughput.
For code review, debugging, multi-step planning, and tool-heavy agents, start
with adaptive or enabled reasoning. For autocomplete, classification, and short
transformations, disabled reasoning can reduce delay. MiniMax says thinking on
and off share the same direct API token rates; the cost difference comes from
how many tokens the request ultimately consumes, not a separate reasoning
surcharge.\[2]
Reasoning controls can differ across hosted routes. If an application depends
on a particular `thinking` field, test the accepted request schema before
shipping rather than assuming all OpenAI-compatible endpoints forward it.
## Where the one-million-token window helps
A long context window is most useful when the model needs several related
artifacts at once:
* a repository plus issue history and build logs;
* a long contract set with cross-document references;
* a video transcript paired with selected frames;
* multi-step agent traces that should remain available for later decisions.
It is less useful when the prompt is an unfiltered dump. Large inputs increase
latency and can bury the relevant evidence. A production agent should still
deduplicate files, retrieve likely sections first, summarize completed tool
traces, and preserve exact source locations for verification.
If your decision is mainly about open weights and managed reliability, compare
[Kimi K3 with Claude Opus 5](/blog/kimi-k3-vs-claude-opus-5). For another
lower-cost one-million-token coding model, see the
[GLM-5.2 API guide](/blog/glm-5-2-api-guide).
## MiniMax M3 limitations
The main constraint is operational scale. Publishing the weights gives teams
deployment control, but it does not make a 427B-parameter multimodal model easy
to serve. Quantization reduces memory, while long contexts add KV-cache and
throughput pressure back into the system.
Benchmark claims also need context. MiniMax reports 59.0% on SWE-Bench Pro and
66.0% on Terminal-Bench 2.1, but the release notes explain that different
models sometimes used different scaffolds and that several results came from
internal infrastructure.\[2] Treat those scores
as evidence of intended capability, not a forecast of success in your own agent.
Finally, one-million-token availability can depend on service capacity. MiniMax
describes 512K as guaranteed and the upper band as up to 1M. Test the exact
region, account, concurrency, and request size you plan to use.
## FAQ
### Is MiniMax M3 open weight?
Yes. MiniMax publishes M3 weights on Hugging Face and code on GitHub under the
MiniMax Community License. Review that license before commercial deployment.
### What is the MiniMax M3 API context window?
The hosted API supports up to one million tokens, with MiniMax describing 512K
as the guaranteed minimum. Calls above 512K use a higher direct pricing band.
### Does MiniMax M3 support images and video?
Yes. MiniMax describes M3 as natively multimodal with text, image, and video
input. Its output is text rather than generated images or video.
### What model ID does reAPI use?
Use `minimax/minimax-m3` with `https://api.reapi.ai/v1/chat/completions`.
### Is MiniMax M3 cheaper than GPT or Claude?
Its token rates are lower than many frontier managed models, but cost per token
is not cost per completed task. Compare tool reliability, retries, output length,
latency, and review effort on your own workload.
## Conclusion
The MiniMax M3 API is a strong fit for teams that need long context, coding
agents, multimodal input, and the option to self-host later. Use the hosted route
first when you want predictable operations; evaluate the open weights when data
control or customization justifies the infrastructure. Most importantly, budget
the correct context band and test the agent harness—not only the model name.
## References
1. MiniMax AI. *MiniMax M3 official model repository and local deployment guidance.* [huggingface.co/MiniMaxAI/MiniMax-M3](https://huggingface.co/MiniMaxAI/MiniMax-M3)
2. MiniMax. *MiniMax M3: Frontier Coding, 1M Context, Native Multimodality.* [minimax.io/blog/minimax-m3](https://www.minimax.io/blog/minimax-m3)
3. MiniMax API Platform. *MiniMax M3 pay-as-you-go pricing.* [platform.minimax.io](https://platform.minimax.io/subscribe/token-plan?tab=api-enterprise)
4. MiniMax API Docs. *Automatic prompt caching.* [platform.minimax.io/docs/api-reference/text-prompt-caching](https://platform.minimax.io/docs/api-reference/text-prompt-caching)
---
# Nano Banana API Free Tier: What Google Actually Charges (https://reapi.ai/blog/nano-banana-api-free-tier)
The short answer about a Nano Banana API free tier is that there isn't one. Google's published pricing lists the free tier for every Nano Banana image model as **Not available**\[1]. Not rate-limited, not capped at some quota. Not offered.
That catches people off guard because Nano Banana genuinely is free inside Google AI Studio, and the two things get conflated constantly. Clicking around the model in a browser and calling the Nano Banana API from your own code are different products under different commercial terms, and only one of them meters you from the first request.
## TL;DR
* **No Nano Banana API free tier exists on any model in the family.** Google lists free-tier input and output as "Not available" for Lite, 2, Pro, and the older 2.5 Flash Image\[1].
* **AI Studio is free; the API is not.** That gap is where the confusion comes from.
* **Billing is per token, not per picture.** A 1K image consumes 1,120 output tokens\[1].
* **Batch mode is exactly half price** and also has no free tier\[1].
* **Cheapest realtime route through Google is Lite at about $0.0336 per 1K image**\[1].
* **On reAPI the same Lite model runs $0.02 per image**, roughly 40% under Google's realtime rate.
* **Signup credit is $0.10**, about five Lite images. Enough to prove an integration works, not to run one.
## What Google actually publishes
The pricing page splits every model into a free-tier column and a paid-tier column. Across the Nano Banana API family, the free column is empty\[1].
| Model | API ID | Free tier | Paid, per image |
| ------------------ | ----------------------------- | ------------- | -------------------------------------------------------- |
| Nano Banana 2 Lite | `gemini-3.1-flash-lite-image` | Not available | \~$0.0336 at 1K |
| Nano Banana 2 | `gemini-3.1-flash-image` | Not available | $0.045 at 0.5K, $0.067 at 1K, $0.101 at 2K, $0.151 at 4K |
| Nano Banana Pro | `gemini-3-pro-image` | Not available | $0.134 at 1K and 2K, $0.24 at 4K |
| Nano Banana | `gemini-2.5-flash-image` | Not available | $0.039 |

Read the Lite row against the Pro row. The spread is about 4x at 1K. That gap is why tier selection moves your bill more than prompt tuning ever will.
## Why so many pages claim it is free
Three things drive the confusion, and each is partly true.
**Google AI Studio costs nothing.** You open the model in a browser, generate images, pay nothing. Real, and the origin of most "Nano Banana is free" claims. It is also not an API: no key, no programmatic calls, no terms that let you put it behind your own product.
**Consumer Gemini apps bundle image generation.** Also real, also not the Nano Banana API.
**New cloud accounts get promotional credit.** Promotional credit is not a free tier. It runs out, and the meter underneath is the paid rate.
None of that changes the answer. When a page quotes a specific "free tier limit" for the Nano Banana API, check it against Google's own table, because that table says the column is not available.
## The per-token detail that breaks naive estimates
Google bills image output by token rather than by picture. Lite output runs $30 per 1M tokens, Nano Banana 2 runs $60, and Pro runs $120\[1]. The per-image numbers above are Google's own conversions published beside those token rates.
The concrete conversion is worth internalizing: a 1K image at 1024×1024 consumes **1,120 output tokens**, which is how $30 per 1M becomes $0.0336 per image\[1].
Two consequences follow.
Resolution changes token count, so one model costs different amounts per image depending on output size. Nano Banana 2 ranges from $0.045 at 0.5K to $0.151 at 4K, a 3.4x spread on a single model\[1].
Input bills separately. Lite input is $0.25 per 1M tokens for text, image, or video; Nano Banana 2 input is $0.50\[1]. Editing pipelines that send a source image on every call carry input cost that a pure text-to-image estimate quietly ignores.
Budget from the token rates and your real resolution mix. Per-image shorthand is fine for rough comparison and misleading for planning.
## Batch mode halves the rate
Every tier of the Nano Banana API also has a batch rate, and it is consistently half of the realtime rate\[1].
| Model | Realtime per 1K image | Batch per 1K image |
| ------------------ | --------------------- | ------------------ |
| Nano Banana 2 Lite | $0.0336 | **$0.0168** |
| Nano Banana 2 | $0.067 | **$0.034** |
| Nano Banana Pro | $0.134 | **$0.067** |
Batch has no free tier either, so this is not a way around the answer above. It is a way to cut the bill in half when latency does not matter: overnight asset generation, bulk catalog work, backfills.
The tradeoff is that batch jobs are queued rather than answered immediately, which rules them out for anything a user is waiting on. If your workload is genuinely offline, batch through Google directly is the cheapest route in this entire comparison.
## What the Nano Banana API costs on reAPI
reAPI prices Nano Banana 2 Lite at **$0.02 per image**, flat, at 1K.
Against Google's realtime $0.0336 for the same model that is roughly 40% lower. Credits are the unit: 1 credit is $0.001, so a Lite image is 20 credits, and failed generations are refunded automatically, so the meter only moves on output you actually receive.
Set that honestly against batch. Google's batch rate of $0.0168 undercuts the $0.02 here. If your work is offline and you are willing to run the queue, that is the cheaper path. The $0.02 buys realtime response through the same unified endpoint as every other model on the platform, which is a different thing than being cheapest in absolute terms.
Two limits stated plainly. Lite is 1K only, so 2K and 4K need a different tier regardless of price. And there is no free tier here either. What exists is a **$0.10 signup credit**, roughly five Lite images. That covers confirming your integration end to end. It is not a trial you can build a product on, and describing it as one would be dishonest.
## Choosing a tier
Across the Nano Banana API the useful question is not which model is best but which failures you can afford to pay for.
**Lite** handles ordinary generation and editing at 1K. At $0.02 the cost of regenerating a bad result is lower than the engineering time spent preventing one.
**Nano Banana 2** buys resolution headroom to 4K and better handling of harder prompts.
**Pro** is the only tier positioned for reasoning-heavy work: dense compositions, many simultaneous constraints, in-image text that has to stay coherent. At $0.134 it is about 6.7x a Lite image, so it earns its slot only when the discard rate below it is high enough to close that gap.
A Nano Banana API routing rule that survives contact with production: draft at Lite, escalate only what fails. Most teams find the escalation rate is far below what they assumed.
## Calling it
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nano-banana-2-lite",
"prompt": "a ceramic mug on a walnut desk, morning light from the left"
}'
```
Media inputs are public http(s) URLs only, with no base64 on any model, so reference images need hosting before the call. Live rates sit on [reapi.ai/models/nano-banana-2-lite](/models/nano-banana-2-lite), and that table is canonical rather than any figure quoted here, because these numbers move.
## FAQ
### Is there a Nano Banana API free tier?
No. Google lists the free tier as "Not available" for every Nano Banana image model, Lite included\[1].
### Why do people say Nano Banana is free?
Because Google AI Studio is free in a browser. That is a playground, not API access: no key, no programmatic calls, no right to put it behind your own product.
### What is the cheapest way to call the Nano Banana API?
Batch mode on Lite, at $0.0168 per 1K image, if you can tolerate queueing\[1]. For realtime calls, Google publishes $0.0336 and reAPI runs the same model at $0.02.
### How much does a Nano Banana Pro image cost?
$0.134 at 1K and 2K, and $0.24 at 4K on Google's published realtime rates\[1].
### Is the Nano Banana API billed per image or per token?
Per token. A 1K image consumes 1,120 output tokens, and image output runs $30 to $120 per 1M tokens by tier\[1].
### Do I get anything free on reAPI?
A $0.10 signup credit, roughly five Lite images. It covers verifying an integration, not running one.
### Can Lite generate 2K or 4K images?
No. Lite is 1K only. Higher resolutions need Nano Banana 2 or Pro.
### Does editing an existing image cost more than generating one?
Usually. On the Nano Banana API the source image counts as input tokens billed separately from output, at $0.25 per 1M for Lite\[1].
## Budgeting without a free tier
The honest framing is that a Nano Banana API free tier does not exist and should not be planned around. What exists instead is a cheap floor. Batch on Lite at $0.0168 is the absolute minimum if latency is irrelevant, and $0.02 realtime on reAPI is low enough that per-generation cost stops being the variable worth optimizing at all.
Start at Lite, measure how often output actually misses your bar, and escalate only that slice. Then check the live tables before committing a budget, because both sides of this comparison changed in the weeks before it was written, and the Nano Banana API pricing will keep moving.
## References
1. Google. *Gemini API pricing — Nano Banana model family free tier, realtime and batch rates.* Retrieved July 2026 from [ai.google.dev/gemini-api/docs/pricing](https://ai.google.dev/gemini-api/docs/pricing)
### Further reading
* reAPI. *How to use Nano Banana 2 Lite.* [reapi.ai/blog/how-to-use-nano-banana-2-lite](/blog/how-to-use-nano-banana-2-lite)
* reAPI. *Nano Banana Pro vs Nano Banana 2.* [reapi.ai/blog/nano-banana-pro-vs-nano-banana-2](/blog/nano-banana-pro-vs-nano-banana-2)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# Nano Banana Pro vs Nano Banana 2: Which Tier to Use (https://reapi.ai/blog/nano-banana-pro-vs-nano-banana-2)
The Nano Banana Pro vs Nano Banana 2 question usually gets asked as if the two were one engine with different stickers. That assumption is wrong, and the difference is not subtle: they run on **different underlying model lines**.
Nano Banana 2 is built on Gemini 3.1 Flash Image. Nano Banana Pro is built on Gemini 3 Pro Image, a separate and more capable line aimed at reasoning-heavy generation\[1]. That is why the comparison is hierarchical rather than lateral, and why the routing rule between them is about what a frame has to survive, not about brand preference.
## TL;DR
* **Different model lines**, not two configurations of one model. Pro sits on Gemini 3 Pro Image; 2 and Lite sit on Gemini 3.1 Flash\[1].
* **Pro is the only tier rated high on reasoning**, which is what dense infographics and multi-constraint compositions actually need\[1].
* **Resolution is the hard divider.** Lite caps at 1K. Pro reaches 4K\[1].
* **On reAPI, Pro runs $0.042 per image at 1K and 2K**, and $0.045 at 4K, against Google's published $0.1072 and $0.192.
* **Lite runs $0.020 per image**, 1K only.
* **Base Nano Banana 2 is not on the reAPI gateway.** The available tiers are Pro and Lite.
## What each tier actually is
Google's own positioning table sorts the family on four axes\[1].
| Tier | Underlying model | Latency | Visual quality | Reasoning |
| ------------------- | --------------------------- | ------- | -------------- | --------- |
| Nano Banana 2 Lite | Gemini 3.1 Flash-Lite Image | Low | Medium | Low |
| Nano Banana 2 | Gemini 3.1 Flash Image | Medium | High | Medium |
| **Nano Banana Pro** | **Gemini 3 Pro Image** | High | High | **High** |
The row that decides most real routing decisions is the last one. Lite and 2 differ mainly in speed and fidelity, and both are rated below Pro on reasoning. Pro is the only tier Google rates high there, and reasoning is what a model needs when a composition has to be *correct* rather than merely attractive: a labeled diagram, a periodic table, a packaging mockup with a compliant nutrition panel, a multi-constraint scene where four things must all be true at once.
## The comparison that matters

| Dimension | Nano Banana Pro | Nano Banana 2 | Nano Banana 2 Lite |
| -------------------- | ----------------------------- | ---------------------- | --------------------------- |
| Model line | Gemini 3 Pro Image | Gemini 3.1 Flash Image | Gemini 3.1 Flash-Lite Image |
| Reasoning rating | **High** | Medium | Low |
| Max resolution | 4K | up to 4K | **1K ceiling** |
| Typical latency | High | Medium (\~20s at 1K) | **Low (\~4s at 1K)** |
| On reAPI | **Yes** | No | **Yes** |
| reAPI rate per image | $0.042 at 1K/2K, $0.045 at 4K | n/a | $0.020 |
Two things fall out of that table.
**The resolution ceiling is the cleanest disqualifier.** If the deliverable needs to survive a large crop or print-adjacent review, Lite is out by definition at 1K. That decision needs no quality judgment at all.
**The price gap is smaller than the tier language suggests.** Pro at $0.042 is about 2.1 times Lite's $0.020. That is a real multiple at volume, but it is not the order-of-magnitude gap the words "Pro" and "Lite" imply. For a hero frame, paying 2.2x for 4K plus high reasoning is rarely the wrong call.
## When each one is right
**Reach for Pro when the frame ships.** Packaging art, campaign stills, marketplace assets, anything at 4K, and anything where the composition has to be factually correct rather than plausible. The high reasoning rating is doing the work here, not the resolution number.
**Reach for Lite when you are iterating.** At roughly four seconds and $0.020 an image, generating a dozen directions costs about a quarter and takes under a minute. Draft there, then render the winner once on Pro.
**The mixed route is the actual answer for most pipelines.** Blind-and-discard variants go to Lite; finalists go to Pro. A team generating twenty candidates and shipping one pays $0.40 in drafts plus $0.045 for the final 4K frame, about $0.45 for the set. Running all twenty-one on Pro would cost $0.88. Same output, roughly half the spend, and the saving grows with the discard ratio.
That routing rule is worth stating as a policy rather than a preference: **the discard rate decides the tier, not the subject matter.**
## What Lite gives up
Worth being specific, because "good enough for volume" hides real constraints\[1].
* **1K ceiling.** No 2K, no 4K.
* **Weaker editing.** Lite's editing Elo trails base Nano Banana 2 by a real margin, so demanding edits belong higher up the family.
* **Low reasoning rating.** Complex multi-constraint scenes and dense factual infographics are exactly what it is not built for.
What it keeps is the part that matters for drafting: ten aspect ratios including 16:9 and 21:9, legible in-image text, character consistency, and multi-image composition.
## Calling both on reAPI
Both tiers run on the same async task endpoint, so switching is a model-string change:
```bash
# Draft: fast and cheap
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nano-banana-2-lite",
"prompt": "packaging concept for a matcha tea tin, soft studio light, 3:4",
"aspect_ratio": "3:4"
}'
```
```bash
# Final: 4K, high reasoning
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3-pro-image-preview",
"prompt": "packaging shot for a matcha tea tin with a compliant nutrition panel, studio light",
"resolution": "4k"
}'
```
Three platform notes. Media inputs are **public http(s) URLs only**, with no base64 on any model. Resolution is a priced dimension on Pro, so it is a cost choice rather than a quality toggle. And live rates are on [reapi.ai/models/gemini-3-pro-image-preview](/models/gemini-3-pro-image-preview) and [reapi.ai/models/nano-banana-2-lite](/models/nano-banana-2-lite), which are canonical rather than any number quoted here.
## FAQ
### What is the difference between Nano Banana Pro and Nano Banana 2?
They run on different model lines. Pro is Gemini 3 Pro Image, rated high on reasoning; Nano Banana 2 is Gemini 3.1 Flash Image, rated medium. Pro is the higher-fidelity, higher-resolution path\[1].
### Is Nano Banana Pro worth the extra cost?
For frames that ship at high resolution or need factual correctness in the composition, yes. On reAPI the gap is $0.042 against $0.020, roughly 2.1x, which is easy to justify for a final asset and hard to justify for drafts.
### Can Nano Banana 2 Lite produce 4K?
No. Lite caps at 1K. That is the cleanest reason to escalate\[1].
### Which tier is best for image editing?
Not Lite. Its editing quality trails base Nano Banana 2 by a real margin, so demanding edits belong on a higher tier\[1].
### Is base Nano Banana 2 available on reAPI?
No. The gateway carries Nano Banana Pro and Nano Banana 2 Lite. If your workflow specifically needs the middle tier, call it through Google directly.
### How should I route between them?
By discard rate. Iterate on Lite where most candidates get thrown away, render finalists on Pro. A twenty-draft, one-final cycle costs about half of running everything on Pro.
### Do they use the same API surface on reAPI?
Yes. Same async task endpoint, same auth, different model string and a resolution field on Pro.
## Routing by discard rate, not by prestige
The Nano Banana Pro vs Nano Banana 2 question is usually asked as a quality comparison, and answered better as an economics one. Pro is a genuinely different model line with the only high reasoning rating in the family and the resolution headroom to match. Lite is fast and cheap enough that throwing away most of what it makes is the intended workflow.
So the decision rule is not which model is better. It is how many candidates you expect to discard before something ships. High discard rates belong on the cheap tier; the frame that survives belongs on Pro. Set that up once as a router and the blended cost lands far below an all-Pro pipeline while the shipped work is indistinguishable. Read Nano Banana Pro vs Nano Banana 2 as a routing question and it answers itself.
## References
1. Google. *Gemini API — image generation models, tiers, and capabilities.* Retrieved July 2026 from [ai.google.dev/gemini-api/docs/image-generation](https://ai.google.dev/gemini-api/docs/image-generation)
2. Google. *Gemini Developer API pricing.* Retrieved July 2026 from [ai.google.dev/gemini-api/docs/pricing](https://ai.google.dev/gemini-api/docs/pricing)
### Further reading
* reAPI. *How to use Nano Banana 2 Lite.* [reapi.ai/blog/how-to-use-nano-banana-2-lite](/blog/how-to-use-nano-banana-2-lite)
* reAPI. *How to use GPT Image 2.* [reapi.ai/blog/how-to-use-gpt-image-2](/blog/how-to-use-gpt-image-2)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# Does Nano Banana Have a Watermark? SynthID Explained (https://reapi.ai/blog/nano-banana-watermark-synthid)
"Does Nano Banana have a watermark" gets asked as one question but describes two completely different things, and conflating them is why the answers online contradict each other.
One is a **visible badge** in a corner, added by a product tier. The other is **SynthID**, an imperceptible provenance marker embedded by Google across its generative AI products\[1]. They have different causes, and only one of them is a product-tier question.
## TL;DR
* **Visible badges come from the surface, not the model.** They are a consumer-tier feature, and API output does not carry them.
* **SynthID is different**: a digital watermark embedded directly into AI-generated images, audio, text, and video, imperceptible to humans\[1].
* **SynthID is embedded across Google's generative AI consumer products**\[1].
* **You can check for it** by uploading the file to Gemini and asking whether it was created or altered by Google AI\[1].
* **A SynthID Detector portal exists**, currently in testing with journalists and media professionals\[1].
* **Provenance marking is not a bug to route around.** This article covers detection, not removal.
## The two things people mean

**A visible watermark** is a logo or badge composited onto the output. It exists because a free or entry tier is subsidised, and the badge is part of what you are trading for the free generation. It is a product-packaging decision made by whichever surface you used.
**SynthID** is a watermarking tool designed specifically for AI-generated content. Google describes it as embedding digital watermarks directly into AI-generated images, audio, text, or video, imperceptible to humans but detectable by SynthID's technology\[1]. Its purpose is to let people identify AI-generated or altered content, which is transparency infrastructure rather than a limitation on your account.
Most people searching for watermark removal are looking at the first one and assuming it is the second. If a badge is visibly sitting in the corner of your image, that is a tier question.
## If it is a visible badge
The badge comes from the surface you generated on, so the fix is the surface, not the file.
**API output does not carry a consumer-tier badge.** When you generate through an image API you get the image the model produced. On reAPI that means Nano Banana 2 Lite at $0.020 per image or Nano Banana Pro at $0.042, with no badge and no plan attached\[2].
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nano-banana-2-lite",
"prompt": "product hero on a stone counter, soft window light",
"aspect_ratio": "16:9"
}'
```
That is the entire answer to the visible-badge version of the question: it is a property of the tier you generated on, and the paid API tier does not add one.
## If it is SynthID
This one is not a tier feature and should not be treated as one.
SynthID watermarks are embedded across Google's generative AI consumer products, and they are designed to survive as a provenance signal rather than to inconvenience you\[1]. The marker exists so that an image can later be identified as AI-generated or AI-altered, which is increasingly what platforms, publishers, and regulators expect to be able to check.
**How to check whether content carries it**\[1]:
* **Ask Gemini.** Upload the image, video, or audio clip to a chat and ask whether it was created or altered by Google AI. Gemini checks for a SynthID watermark and reports what it finds.
* **SynthID Detector**, a verification portal for uploading an image, video, or audio file. Google is currently collaborating with journalists and media professionals to test it, with an early-tester waitlist.
That detection path is worth knowing in both directions: for verifying content you received, and for understanding what is knowable about content you produced.
## What this means if you are building something
**Disclose rather than obscure.** If your product generates images for users, the useful posture is telling them the output is AI-generated. Provenance signals are moving toward being expected, and a product that labels its own output ages better than one that gets identified later.
**Do not build removal into a pipeline.** Beyond the practical problem that these markers are designed to be robust, stripping provenance from AI-generated media is the behavior that platform policies and emerging disclosure rules are specifically aimed at. It is not a technical corner worth optimizing.
**Know which marker you are dealing with.** A visible badge is a billing question. An invisible provenance marker is a compliance and transparency question. They have nothing to do with each other, and treating the second like the first is how teams end up with a policy problem they did not intend.
**Check what you receive.** If you accept user-uploaded images into a workflow where provenance matters, the Gemini check is a one-step way to find out whether something was AI-generated or edited\[1].
## FAQ
### Does Nano Banana add a watermark?
Two different things get called that. A visible badge is a consumer-tier feature and is not added to API output. SynthID is an imperceptible provenance marker Google embeds across its generative AI consumer products\[1].
### How do I get images without a visible badge?
Generate through an API tier rather than a free consumer tier. On reAPI that is $0.020 per image on Nano Banana 2 Lite and $0.042 on Nano Banana Pro\[2].
### What is SynthID?
Google DeepMind's watermarking tool for AI-generated content, embedding digital watermarks into images, audio, text, and video that are imperceptible to humans but detectable by SynthID's technology\[1].
### How do I check if an image has a SynthID watermark?
Upload it to Gemini and ask whether it was created or altered by Google AI. A separate SynthID Detector portal is in testing with journalists and media professionals\[1].
### Does SynthID apply to video and audio too?
Yes. Google describes SynthID as covering images, audio, text, and video\[1].
### Can SynthID be removed?
That is not something this guide covers. These markers are provenance infrastructure, and stripping them from AI-generated media is what disclosure rules and platform policies are aimed at.
### Is a visible badge the same as SynthID?
No. A badge is composited onto the picture by a product tier. SynthID is an imperceptible signal embedded in the content itself\[1].
### Should my product label AI-generated images?
Increasingly that is the expectation. Labelling your own output is cheaper than being identified later by someone else's detector.
## Two questions wearing one word
The reason "does Nano Banana have a watermark" produces contradictory answers is that half the people answering are talking about a corner badge and the other half are talking about provenance infrastructure.
If a logo is visibly on your image, that is a tier you generated on, and an API tier does not add one. If the question is whether the content carries an identifiable signal that it came from a generative model, the answer is that Google embeds SynthID across its generative AI products, it is imperceptible by design, and there is a detection path in Gemini for checking it. Those are different facts about different layers, and only the first is a watermark you should be trying to change.
## References
1. Google DeepMind. *SynthID — watermarking and identifying AI-generated content, and how to detect it.* Retrieved July 2026 from [deepmind.google/technologies/synthid](https://deepmind.google/technologies/synthid/)
2. reAPI. *Model catalog — per-image rates for the image models.* [reapi.ai/models](/models)
### Further reading
* reAPI. *Nano Banana Pro vs Nano Banana 2.* [reapi.ai/blog/nano-banana-pro-vs-nano-banana-2](/blog/nano-banana-pro-vs-nano-banana-2)
* reAPI. *How to use Nano Banana 2 Lite.* [reapi.ai/blog/how-to-use-nano-banana-2-lite](/blog/how-to-use-nano-banana-2-lite)
* reAPI. *Where to use Seedance 2.0.* [reapi.ai/blog/where-to-use-seedance-2-0](/blog/where-to-use-seedance-2-0)
---
# OpenAI Astra's 10 Math Breakthroughs: What They Mean (https://reapi.ai/blog/openai-astra-math-breakthroughs)
**OpenAI's next major model, Astra, has moved beyond solving hard exercises and into producing new mathematical results.** On August 1, 2026, OpenAI published ten advances generated by an internal version of the model, together with a 249-page manuscript, reasoning walkthroughs, and Lean formalizations. If the results survive independent scrutiny, this will matter more than another benchmark record. It still does not mean that the mathematics community has already accepted “ten conjectures solved by AI”\[1].
The important signal is not one lucky proof. It is the possibility that a general-purpose model can produce research-grade work repeatedly, across several fields, at a low marginal inference cost.
## TL;DR
* OpenAI released ten different kinds of results: proofs, counterexamples, and improved bounds. Calling all ten “solved conjectures” is convenient but imprecise.
* Astra is an internal, unreleased model. OpenAI has not announced public access, API pricing, or a release date.
* The paper and Lean files make the claims unusually testable, but formal verification does not replace expert review of the statement, assumptions, novelty, or significance.
* OpenAI estimates that the solution-finding tokens would cost about $2,000 at Sol API rates. That is an inference-cost estimate, not the total cost of the research program.
* If the results hold, the immediate shift is cheaper proof search and technical experimentation—not the disappearance of mathematicians.
## What did OpenAI Astra actually produce?
OpenAI's announcement lists ten advances spanning high-dimensional geometry, coding theory, group theory, operator algebras, arithmetic circuit complexity, quantum complexity, lattice problems, and extremal combinatorics\[1].
| Area | Reported Astra result |
| ------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| High-dimensional sphere packing | Determines the asymptotic strength of the Cohn–Elkies linear program and improves the general packing bound |
| Binary and spherical codes | Gives exponentially stronger bounds across the parameter range |
| Group theory | Constructs a non-sofic group, resolving a central open question |
| Operator algebras | Produces counterexamples to Connes's rigidity conjecture |
| Arithmetic circuit complexity | Improves lower bounds for circuits and formulas computing the permanent |
| Quantum complexity | Proves exponential parallel repetition for arbitrary finite two-player entangled games |
| Lattice problems | Establishes polynomial-factor hardness for approximating the Euclidean closest vector problem |
| Convex geometry | Proves the sharp Ehrhart volume bound in every dimension |
| Ramsey theory | Gives a superexponential lower bound for multicolor triangle Ramsey numbers |
| Extremal graph theory | Disproves compactness and degeneracy conjectures, resolving two Erdős problems |
The [full collection](https://cdn.openai.com/pdf/ten-proofs-oai.pdf) runs to 249 pages\[2]. Some chapters prove a conjecture, some disprove one by construction, and others improve a long-standing upper or lower bound. The results do not all have the same shape or scholarly weight.
The accurate description is therefore that OpenAI has submitted ten potentially major advances in mathematics and theoretical computer science. Whether every chapter is correct, new, and as important as the announcement suggests must be judged by specialists in the relevant fields.
## Why this matters more than an AI math-olympiad medal
An olympiad problem can be extremely difficult, but it is still a closed problem. A human has selected the question, the conclusion is known, and a solution of manageable length is expected to exist.
Research mathematics removes those supports. A researcher may not know whether a statement should be proved or disproved. The expensive work includes choosing a direction, connecting distant bodies of theory, searching for counterexamples, recognizing a useful construction, and abandoning an attractive approach before it consumes months.
OpenAI had already shown one example of this behavior in May 2026. An internal general-purpose model used ideas from algebraic number theory to disprove the long-standing Erdős unit-distance conjecture. External mathematicians checked the proof, and Tim Gowers described it as a milestone in AI mathematics\[3].
One result can still be treated as an exceptional hit. Ten results across unrelated domains are harder to dismiss that way. If outside review confirms them, the story is no longer simply that a model can occasionally help with a proof. It is that a model may have acquired a repeatable research capability.
That distinction is the real source of the excitement.
## What the reported $2,000 token cost does—and does not—mean
OpenAI says the tokens used to find the ten solutions would cost roughly $2,000 at Sol API rates\[1].
That figure does not mean the research cost $2,000. It excludes training Astra, running the underlying infrastructure, selecting problems, paying researchers, organizing the manuscripts, formalizing the arguments, and investigating failed directions. The announcement also does not provide the denominator: how many open problems were tested to obtain these ten results.
What the number does reveal is a possible collapse in the marginal cost of exploration.
A human researcher may spend weeks establishing that a promising route cannot work. A model can try many routes in parallel: test special cases, search for counterexamples, combine lemmas from separate literatures, run computations, and retain only the branches that survive initial checking.
If that process becomes reliable, research budgets start to buy a search space rather than a single line of attack. AI may not replace the person who asks the important question, but it can make mathematical trial and error far cheaper.
Astra itself is not publicly available. OpenAI has not announced an API, a price, or a release date. To compare models that are actually available now, use the [reAPI model catalog](/models); an Astra listing should not be inferred from these research results.
## From Astra output to accepted mathematics

The public workflow can be separated into four stages:
1. **Model discovery.** Astra searches for a construction, counterexample, or proof and produces a candidate argument.
2. **Human preparation.** Researchers use the model to turn that argument into a readable manuscript.
3. **Formal checking.** The argument is encoded in Lean so its logical steps can be checked by a small trusted kernel.
4. **Community review.** Independent mathematicians verify that the formal statement matches the original problem and assess novelty, context, and significance.
OpenAI has published Lean 4 files for all ten results, along with build instructions and a path for independent proof checking\[4]. That is a much stronger form of evidence than a polished model transcript. Reviewers can inspect definitions, rebuild the project, and challenge specific assumptions.
The four stages should not be compressed into the sentence “the machine proved it.” They answer different questions, and none of the later stages is decorative.
## What does Lean verification guarantee?
Lean checks whether a formalized theorem follows from the definitions, axioms, and earlier results supplied to the system. It does not accept “obviously” as a proof step, and it is not persuaded by fluent prose. This makes it well suited to checking long AI-generated arguments.
But a successful Lean build is not an all-purpose certificate of mathematical importance or even a complete guarantee that the original informal claim was captured correctly.
Experts still need to ask:
* Does the formal theorem faithfully represent the open problem?
* Are the definitions standard for that field?
* Did the formalization introduce an assumption stronger than intended?
* Was a difficult claim moved into an imported lemma or premise?
* Is the result genuinely new?
* Does it have the significance claimed in the announcement?
Lean can establish that a derivation is valid inside a formal environment. It cannot independently decide that the environment represents the intended mathematics, nor can it search the entire scholarly record for priority.
Formal verification and peer review are complements, not substitutes.
## Why independent review still matters
OpenAI is the model developer, the publisher of the paper, and the party taking responsibility for the formalizations. Releasing the manuscript and code is a strong transparency measure, but a claim of this size still needs review by experts without a stake in the model launch.
OpenAI's own First Proof experiment shows why. In February 2026, the company released ten research-level proof attempts from an internal model. One attempt that initially appeared likely to be correct was later judged incorrect after feedback from the problem authors and the mathematical community\[5].
Frontier proofs can fail at a single subtle step buried inside dozens of pages. As models become better at producing confident, professional mathematical prose, surface quality becomes an even less reliable guide to correctness.
The Astra release offers considerably stronger evidence than a raw proof attempt because it includes full manuscripts and formal files. As of August 2, 2026, the responsible conclusion is still narrower than the viral headline: OpenAI has released a highly consequential, unusually well-documented set of candidate advances that now require broad independent scrutiny.
## Will Astra replace mathematicians?
Not immediately. It is more likely to rearrange what mathematicians spend their time doing.
Research includes problem selection, conjecture formation, literature mapping, proof search, technical derivation, verification, explanation, and generalization. Astra appears most likely to compress the cost of proof search and technical experimentation—the parts that can consume large amounts of time without producing a publishable result.
If candidate proofs become abundant, other skills become more scarce:
* deciding which questions are worth solving;
* identifying the actual idea inside a long machine-generated argument;
* turning a technically correct proof into a concept humans can understand;
* recognizing where a new method transfers;
* designing reliable human–AI research workflows;
* auditing a growing volume of machine-produced mathematics.
The mathematician of the near future may act more like a research director, problem designer, and proof auditor. Technical work will not vanish, but manually producing every intermediate step may no longer sit at the center of the profession.
That change also creates institutional questions. How should an AI-generated proof be credited? What level of model assistance must a paper disclose? Can researchers without access to private frontier models compete? If public papers and formal libraries help train commercial systems, what obligations do developers have to the mathematical community?
OpenAI says it would be misleading to claim human authorship for an argument generated entirely by its system\[1]. The [Leiden Declaration on Artificial Intelligence and Mathematics](https://leidendeclaration.ai/) goes further, calling for standards around attribution, training data, research equity, corporate power, and the preservation of human mathematical understanding\[6].
These are not side issues. The more capable Astra is, the faster the research community will need workable answers.
## Does Astra prove that AGI has arrived?
No.
Mathematics has explicit rules, a large digitized literature, and outcomes that can often be checked automatically. That makes it unusually suitable for a training and verification loop. Progress in formal reasoning does not automatically transfer at the same rate to fields that depend on physical experiments, noisy observations, tacit laboratory knowledge, or ambiguous goals.
The public record also does not show Astra independently completing the entire research cycle. Humans still selected or curated the problems, prepared the manuscripts, released the work, and assumed responsibility for it. Without data on failed attempts, outsiders cannot calculate the model's success rate on an arbitrary open problem.
The evidence supports a more specific claim: a general-purpose AI system may now be able to produce original, end-to-end arguments on some research-level mathematical problems, at a quality high enough to merit formal and expert review.
That is already a major result. It does not need an AGI label to be consequential.
## The real turning point in AI mathematics
Computers have long helped mathematicians run calculations, search literature, and verify finite cases. Astra is attempting something different: choosing proof directions, connecting concepts, constructing counterexamples, and sustaining a novel argument across many pages.
If the ten results survive review, mathematical discovery will have a visible new production model. Candidate proofs can be searched in parallel, checked with formal systems, and scaled with additional inference resources.
Mathematics, however, is not only the act of moving a statement from “unknown” to “proved.” A proof is also valuable because it explains why a result is true and exposes structures that can be used elsewhere.
The future bottleneck may therefore be less about whether machines can generate more theorems and more about whether humans can understand, organize, and retain agency over them. Astra has not ended the role of the mathematician. It may have opened the era of automated mathematical research.
## FAQ
### Is OpenAI Astra publicly available?
No. OpenAI describes Astra as its next major model and says these results came from an internal version. As of August 2, 2026, the announcement includes no public release date, API access, or pricing\[1].
### Did Astra solve ten mathematical conjectures?
Not in one uniform sense. The collection includes proofs, counterexamples, and improved bounds. The manuscripts and Lean files are public, but specialists still need to review each result independently.
### Does a Lean proof guarantee that the result is correct?
Lean can guarantee that a formal theorem follows from its stated assumptions inside the formal system. It cannot by itself guarantee that the theorem exactly matches the original informal problem, that the assumptions are appropriate, or that the result is new and significant.
### Did the ten results cost only $2,000?
No. The $2,000 figure is OpenAI's estimate for the solution-finding tokens at Sol API rates. It excludes model training, infrastructure, staff, failed experiments, problem selection, manuscript preparation, and formalization.
### Will Astra make mathematicians obsolete?
The nearer-term effect is task reallocation. Proof search and technical trial and error may become heavily automated, while problem selection, interpretation, verification, conceptual compression, and research governance become more important.
## Further reading
* [Browse currently available AI models](/models) — capabilities and access for released models, not unreleased Astra research systems.
* [What is reAPI?](/blog/what-is-reapi) — how one API provides access to models from multiple providers.
## References
1. OpenAI. *Ten advances in mathematics and theoretical computer science.* August 1, 2026. [openai.com/index/ten-advances-in-mathematics](https://openai.com/index/ten-advances-in-mathematics/)
2. OpenAI. *Ten Advances in Mathematics and Theoretical Computer Science.* August 2026. [cdn.openai.com/pdf/ten-proofs-oai.pdf](https://cdn.openai.com/pdf/ten-proofs-oai.pdf)
3. OpenAI. *An OpenAI model has disproved a central conjecture in discrete geometry.* May 20, 2026. [openai.com/index/model-disproves-discrete-geometry-conjecture](https://openai.com/index/model-disproves-discrete-geometry-conjecture/)
4. OpenAI. *Ten Advances in Mathematics and Theoretical Computer Science — Lean 4 formalizations.* Retrieved August 2, 2026. [github.com/openai/ten-proofs](https://github.com/openai/ten-proofs)
5. OpenAI. *Our First Proof submissions.* February 20, 2026. [openai.com/index/first-proof-submissions](https://openai.com/index/first-proof-submissions/)
6. Leiden Declaration Working Group. *Leiden Declaration on Artificial Intelligence and Mathematics.* May 2026. [leidendeclaration.ai](https://leidendeclaration.ai/)
---
# Qwen Image 2 API: Pricing, Text Rendering, and Editing (2026) (https://reapi.ai/blog/qwen-image-2-api-guide)
The **Qwen Image 2 API** combines text-to-image and image editing in one model.
Send only a prompt to generate a new image; add one source image to edit it.
QwenCloud lists the accelerated model at $0.035 per image, while reAPI's current
route charges $0.031 for one output.
The model is especially relevant when prompts contain exact copy, structured
layouts, or detailed scene constraints. It is not a universal replacement for
every image model, but its flat price and shared generation/editing request shape
make it easy to budget.
## TL;DR
* Qwen Image 2 handles both image generation and image editing.
* QwenCloud advertises improved text rendering and prompts up to 1,000 tokens.\[1]
* Official QwenCloud pricing is **$0.035 per image**; reAPI currently charges
**$0.031 per image**.
* The reAPI route returns one image per request and accepts one public source
image URL for editing.
* `nsfw_checker` controls an additional reAPI output check. It is not a promise
that upstream model or platform policies disappear.
## Qwen Image 2 API specifications
| Specification | reAPI Qwen Image 2 route |
| ------------------------- | ----------------------------------- |
| Model ID | `qwen-image-2` |
| Tasks | Text-to-image, single-image editing |
| Output count | One image per request |
| Output formats | PNG, JPEG |
| Text-to-image ratios | `1:1`, `3:4`, `4:3`, `9:16`, `16:9` |
| Additional editing ratios | `2:3`, `3:2`, `21:9` |
| Source image | One public HTTP(S) URL |
| Deterministic control | Optional `seed` |
| Current reAPI price | $0.031 per image |
Alibaba also offers Qwen Image 2 through Model Studio. That official surface
supports up to six outputs per call and resolutions up to 2048×2048, while the
QwenCloud marketplace describes a 120 RPM limit.\[2]
Those details do not automatically carry across providers. A gateway can expose
a narrower, simpler schema even when the underlying family supports more modes.
## Qwen Image 2 API pricing
| Provider surface | Price | Billing unit |
| -------------------- | -----: | ------------------- |
| QwenCloud list price | $0.035 | One generated image |
| reAPI current price | $0.031 | One generated image |
Both generation and editing use the same reAPI price. Aspect ratio does not
change the charge, and the current route has no `n` field. To produce four
variants, submit four requests and track their seeds rather than pretending one
request contains a batch.
At 10,000 outputs per month, the simple headline comparison is:
```text
QwenCloud list price: 10,000 × $0.035 = $350
reAPI current price: 10,000 × $0.031 = $310
```
That estimate excludes failed requests, application storage, retries, and any
post-processing pipeline. Save the rate-card date in production budgets because
gateway and vendor prices can change independently.

## How to generate an image with Qwen Image 2
The generation request needs a prompt and model ID. The remaining fields are
optional.
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer $REAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen-image-2",
"prompt": "Editorial poster for a typography workshop. Warm ivory paper, black serif headline reading TYPOGRAPHY LAB, cobalt italic annotation, one orange registration stamp.",
"image_size": "16:9",
"output_format": "png",
"seed": 4281,
"nsfw_checker": true
}'
```
Use a public, durable location for the returned image after generation. Image
APIs often return task assets from temporary storage, and an accessible URL is
not the same as permanent object storage.
The [Qwen Image 2 model page](/models/qwen-image-2) exposes the current controls,
while the [Qwen Image 2 API docs](/docs/qwen-image-2) should be treated as the
wire-format source of truth.
## How to edit an image
Editing uses the same endpoint and model. Add `image_url`; the model infers the
task from the request shape.
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer $REAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen-image-2",
"prompt": "Keep the product and camera angle unchanged. Replace the gray background with warm ivory paper and add a soft cobalt shadow.",
"image_url": "https://example.com/product-source.png",
"image_size": "3:2",
"output_format": "png",
"nsfw_checker": true
}'
```
Editing prompts work better when they separate the change from the invariants.
State what must change, then name what must remain fixed: subject identity,
camera angle, product proportions, text, lighting direction, or background.
That structure is more reliable than asking the model to “improve everything.”
## Writing prompts for text rendering
QwenCloud positions text rendering as a core strength, but legibility still
depends on prompt complexity and layout density. Use these practical rules:
1. Quote the exact visible copy.
2. Specify where the text sits and which line is largest.
3. Keep the first attempt to one or two short phrases.
4. Name the typography class—serif, sans-serif, condensed, monospaced—rather
than a copyrighted typeface.
5. Generate the visual, then verify every character before publishing.
For example:
```text
Create a 16:9 conference poster.
Exact headline: “BUILD WITH IMAGES”.
Place the headline in two large black serif lines on the left.
Add a small cobalt italic note: “generation + editing”.
No other text, logos, badges, or interface elements.
```
Long prompts are useful for composition, but adding more words does not always
improve exact typography. Treat the official 1,000-token support as an input
capacity, not a recommendation to fill the full allowance.
## Safety controls and `nsfw_checker`
The reAPI playground keeps `nsfw_checker` enabled. Direct callers can set it to
`false`, which skips an additional output-classification layer on the reAPI
route. It does not override Alibaba policies, model alignment, upstream request
inspection, or applicable law.
This distinction matters for benign edge cases such as swimwear, medical
illustration, anatomy education, and some fine-art prompts. If a legitimate
workflow is blocked, log the stage and error code before changing a control.
Repeatedly rephrasing prompts without knowing whether the rejection happened at
the gateway, provider, or model layer makes debugging slower.
For the broader policy model, see
[how uncensored image API claims actually work](/blog/uncensored-ai-image-api-content-filters).
## When to choose Qwen Image 2
Choose Qwen Image 2 when you need flat per-image billing, a shared generation
and edit endpoint, aspect-ratio control, and prompts that contain structured
copy. It is a good fit for posters, product variations, editorial graphics, and
single-source edits.
Choose another route when you need multi-reference editing, native 4K output,
or multiple images returned by one gateway request. The
[GPT Image 2 guide](/blog/how-to-use-gpt-image-2) covers a different high-control
editing workflow, while [Nano Banana Pro vs Nano Banana 2](/blog/nano-banana-pro-vs-nano-banana-2)
explains Google's professional and value tiers.
## Common Qwen Image 2 API mistakes
### Passing multiple source images
The current reAPI schema accepts one `image_url`, not an array. Use a model that
explicitly supports multi-reference input when several sources are essential.
### Using edit-only ratios for text-to-image
`2:3`, `3:2`, and `21:9` are available on the edit surface but rejected for
plain text-to-image through the current route. Pick one of the five generation
ratios for a new image.
### Treating a seed as a byte-for-byte guarantee
A seed improves reproducibility, but infrastructure, model revisions, and
backend changes can still alter the result. Store the prompt, source asset,
model ID, seed, and generation date together.
## FAQ
### How much does the Qwen Image 2 API cost?
QwenCloud lists $0.035 per image. reAPI currently charges $0.031 per output for
its `qwen-image-2` route.
### Can Qwen Image 2 edit an existing image?
Yes. Add one public `image_url` and an edit instruction to the same generation
endpoint.
### Does Qwen Image 2 generate readable text?
Text rendering is an advertised capability and often works best with short,
explicit copy. Always verify spelling before production use.
### Can I generate several images in one reAPI request?
No. The current route returns one image per request. Submit separate jobs for
variants.
### Is Qwen Image 2 uncensored?
No API should be described as unconditionally uncensored. reAPI exposes an
optional additional NSFW checker, but upstream rules and safety systems remain.
## Conclusion
The Qwen Image 2 API offers a compact workflow: one model for generation and
editing, flat pricing, useful aspect-ratio control, and a strong emphasis on
text-heavy compositions. Keep provider-specific parameters separate, verify
rendered copy, and treat optional gateway moderation as one layer rather than
the entire safety system.
## References
1. QwenCloud. *Qwen-Image-2.0 model overview, pricing, and limits.* [qwencloud.com/models/qwen-image-2.0](https://www.qwencloud.com/models/qwen-image-2.0)
2. Alibaba Cloud Model Studio. *Image generation and editing model comparison.* [alibabacloud.com/help/en/model-studio/image-model](https://www.alibabacloud.com/help/en/model-studio/image-model)
3. Qwen Team. *Qwen-Image-2.0 Technical Report.* [arxiv.org/abs/2605.10730](https://arxiv.org/abs/2605.10730)
---
# Run a 70B LLM on a 4GB GPU with AirLLM: The Honest Guide (https://reapi.ai/blog/run-70b-llm-on-4gb-gpu-airllm)
**Yes, a 70B LLM can execute while using about 4GB of GPU memory with
AirLLM—but the model does not fit inside a 4GB graphics card.** AirLLM keeps
the checkpoint on disk, loads one transformer layer into VRAM, computes that
layer, releases it, and repeats. The technique trades memory for storage I/O
and latency.\[1]\[2]
That distinction is the entire story. A 70B model still needs roughly 130GB
of weight storage at the precision used in the original demonstration. The
4GB figure describes peak VRAM during a narrowly configured inference run,
not total machine memory, download size, or interactive performance.
## TL;DR
* **AirLLM makes extreme offloading possible.** It stores layer-wise shards
on disk and moves only the active layer to the GPU.
* **The original result measured under 4GB on a 16GB Nvidia T4.** It did not
demonstrate a fast chatbot running entirely from a physical 4GB card.\[1]
* **Disk capacity and speed still matter.** The first run downloads and
reshards the checkpoint, and every generated token repeatedly streams
weights through the compute device.
* **Short context is part of the trick.** The original example used an input
length of 100 tokens; a larger KV cache and runtime buffers consume more
memory.\[1]
* **Expect research or batch speeds, not responsive chat.** AirLLM's author
explicitly positions low-end hardware for offline work rather than
interactive applications.\[1]
* **Use an API when output speed matters.** Extreme local offloading is useful
for learning and occasional private jobs; hosted inference is usually the
simpler production path.
## What “run a 70B LLM on a 4GB GPU” actually means
Four different resources are compressed into one headline. Separating them
prevents most bad hardware decisions.
| Resource | What AirLLM changes | What it does not change |
| ---------- | ------------------------------------------------------------------------- | ----------------------------------------------------- |
| GPU VRAM | Keeps roughly one layer plus runtime state resident | The full checkpoint is not in VRAM |
| System RAM | Uses lazy loading and a meta device to avoid materializing the full model | Python, tokenizer, buffers, and OS memory still exist |
| Disk | Stores the complete model as layer-oriented shards | The download does not become 4GB |
| Time | Prefetch can overlap part of loading and compute | Storage traffic remains the central bottleneck |
The original 2023 walkthrough used a Llama 2-based 70B checkpoint with 80
transformer layers. One layer was estimated at about 1.6GB, while the KV cache
for its 100-token example was about 30MB. The measured process stayed below
4GB of GPU memory on an Nvidia T4.\[1]
AirLLM's current repository extends the same idea to Llama 3.x, Qwen,
DeepSeek, Mixtral, Phi, Gemma, and other families. Its current reference table
still lists a full-precision Llama 3.x 70B run at approximately 4GB of
VRAM.\[2] Treat that as a project claim and
a memory target—not a throughput benchmark for every 4GB card.
## How AirLLM layer-wise inference works

A transformer runs its blocks in sequence. Layer 12 consumes the hidden state
from layer 11; layer 13 waits for layer 12. AirLLM exploits that order with a
five-part pipeline.
1. **Create an empty model shell.** Hugging Face Accelerate's meta device
initializes the architecture without allocating real storage for every
parameter.\[3]
2. **Reshard the checkpoint by layer.** Safetensors files are rearranged so
loading one layer does not require reading an unrelated multi-gigabyte
shard.
3. **Load one layer to the compute device.** Only that layer and the required
runtime tensors occupy the GPU at that moment.
4. **Compute, release, and continue.** The hidden state moves forward while
the layer weights leave VRAM.
5. **Repeat for every generated token.** Prefetching overlaps some storage I/O
with computation, but it cannot remove the repeated data movement.
FlashAttention reduces the temporary memory used by attention through tiled,
I/O-aware computation.\[4] It helps the active
layer fit, while layer streaming solves the separate problem of where inactive
weights wait.
## Hardware and storage you still need
The GPU is only one component. Before downloading a 70B checkpoint, verify the
rest of the machine.
* **A compatible compute path.** The headline demonstrations target Nvidia
CUDA. AirLLM also documents Apple-silicon and CPU paths, but their memory and
performance characteristics are different.\[2]
* **Enough disk for the checkpoint and conversion.** The project warns that
first-run layer splitting is disk-intensive. Its `delete_original` option
can remove the original checkpoint after conversion when storage is tight.
* **Fast local storage.** NVMe does not make layer streaming free, but a slow
hard drive makes an already I/O-bound loop substantially worse.
* **A short initial context and output.** Start with a tiny prompt and 20–40
new tokens. Longer context increases the KV cache, while longer output
repeats the full layer traversal more times.
* **Model access.** Gated Meta checkpoints require a Hugging Face token and
acceptance of the model license.
Do not begin with a 70B download just to test whether the environment works.
Run an 8B or smaller supported model first, verify CUDA and storage paths, then
scale up.
## How to try AirLLM with a 70B model
The current project quickstart uses `AutoModel`, which chooses the appropriate
implementation from a Hugging Face repository ID.\[2]
Install a PyTorch build compatible with your CUDA driver first, then install
AirLLM.
```bash
python -m venv .venv
source .venv/bin/activate
pip install airllm
```
Keep secrets in environment variables rather than source code:
```bash
export HF_TOKEN="your_hugging_face_token"
```
Then run a deliberately small generation:
```python
import os
from airllm import AutoModel
MODEL_ID = "meta-llama/Llama-3.3-70B-Instruct"
MAX_LENGTH = 128
model = AutoModel.from_pretrained(
MODEL_ID,
hf_token=os.environ["HF_TOKEN"],
layer_shards_saving_path="/data/airllm-shards",
)
prompt = ["Explain layer-wise inference in three short sentences."]
tokens = model.tokenizer(
prompt,
return_tensors="pt",
return_attention_mask=False,
truncation=True,
max_length=MAX_LENGTH,
padding=False,
)
result = model.generate(
tokens["input_ids"].cuda(),
max_new_tokens=32,
use_cache=True,
return_dict_in_generate=True,
)
print(model.tokenizer.decode(result.sequences[0]))
```
This is a minimal adaptation of the repository quickstart, not a universal
environment lockfile. AirLLM, Transformers, PyTorch, CUDA, and a model's remote
code can have version-specific constraints, so check the current repository
issues before installing it on a production machine.
## What happens on the first run
The first launch is not representative of later launches. AirLLM must download
the model, inspect its architecture, split the checkpoint into layer shards,
and write those shards to the configured path. Interrupting that conversion or
running out of disk can leave an incomplete safetensors header; the project's
FAQ recommends clearing the incomplete cache and rerunning after making space.
\[2]
Monitor four signals separately:
```bash
nvidia-smi -l 1 # GPU memory and utilization
free -h # system memory
df -h /data # free disk space
iostat -xz 1 # storage saturation, if sysstat is installed
```
A low VRAM number is not success by itself. Record time to first token, seconds
per output token, disk read volume, and whether repeated runs reuse completed
shards.
## Why AirLLM is slow even when it fits
Normal GPU inference loads weights once and reuses them for many tokens and
requests. Extreme layer offloading reverses that advantage. Each new token must
pass through the model's full stack while weights move from storage into the
GPU in small pieces.
AirLLM added prefetching to overlap loading with computation and offers 4-bit
or 8-bit block-wise weight compression to reduce disk traffic. The repository
reports up to a threefold improvement from compression, but actual performance
depends on the model, storage, GPU, context, and software versions.\[2]
This makes the method more credible for:
* one-off evaluation of a model that otherwise cannot load;
* offline document classification or extraction;
* low-volume private batch processing where latency is secondary;
* studying memory scheduling and model architecture.
It is a poor default for live chat, agent loops, high concurrency, or any API
with a latency target.
## AirLLM vs quantization vs an API
| Approach | Local weights | Typical goal | Main trade-off |
| ---------------------- | :-----------: | ------------------------------------------------------ | ----------------------------------------------------------- |
| AirLLM layer streaming | Yes | Make an oversized model execute | Very low throughput and heavy disk I/O |
| 4-bit quantization | Yes | Make a model smaller and faster | A dense 70B model still needs far more than 4GB for weights |
| CPU/GPU offload | Yes | Split a moderately oversized model across RAM and VRAM | Requires substantial system RAM |
| Hosted API | No | Get interactive or production inference | Remote execution, usage cost, provider trust |
Choose AirLLM when the experiment is the point. Choose a smaller quantized
model when local interactivity is the point. Choose an API when the 70B-class
model and usable response time are both requirements.
The same distinction applies to much larger claims. Our [Kimi K3 on a 4GB GPU
analysis](/blog/unbelievable-run-kimi-k3-2-8-trillion-parameters-on-a-single-4gb-gpu)
explains why sparse experts change the streaming unit but do not erase the
checkpoint. For a current long-context API example, see the [MiniMax M3 API
guide](/blog/minimax-m3-api-guide), or browse the live [model catalog](/models).
## A practical decision checklist
Before attempting a 70B LLM on a 4GB GPU, answer these questions:
1. Is the objective to prove execution, or to build a responsive product?
2. Can the disk hold the original model and its layer-sharded copy during
conversion?
3. Is the checkpoint architecture explicitly supported by the current AirLLM
release?
4. Can the workload tolerate long time-to-first-token and low throughput?
5. Does the model license permit the intended use?
6. Have you tested the same software stack with a small checkpoint first?
If the answer to questions two through four is no, the 4GB headline is not a
useful deployment plan.
## FAQ
### Can a 70B LLM really run on a 4GB GPU?
Yes, through extreme layer-wise streaming. Only a small part of the model is
resident in VRAM at once; the full checkpoint remains on disk. That is not the
same as loading a 70B model into 4GB.
### Did the original AirLLM test use an actual 4GB graphics card?
The 2023 article says the team tested on a 16GB Nvidia T4 and measured less
than 4GB of GPU memory usage.\[1] The current
repository separately lists Llama 3.x 70B at approximately 4GB VRAM.
### How much disk space does a 70B model need?
It depends on the checkpoint precision and format. The original walkthrough
described roughly 130GB of parameters, and layer conversion may temporarily
require both the original and converted copies. Check the repository files
before downloading and leave room for interrupted or partial conversions.
### Is AirLLM fast enough for a chatbot?
Usually not on low-end hardware. The original author warns that a T4 setup is
slow and better suited to offline work.\[1]
### Does AirLLM train a 70B model in 4GB?
No. Training must retain or recompute activations and gradients for
backpropagation. AirLLM's layer-wise technique addresses inference, not full
training.\[1]
### Is a 4-bit 70B model small enough for 4GB VRAM?
No. Seventy billion parameters at four bits require a theoretical 35GB just
for raw weights, before quantization metadata and runtime memory. Quantization
helps, but it does not close that gap.
## The honest verdict on 70B inference with 4GB VRAM
AirLLM turns a hard memory ceiling into a scheduling problem. That is a real
technical result: a 70B LLM can execute with roughly 4GB of VRAM when the
runtime streams layer shards from much larger storage. The price is repeated
I/O, slow generation, a large checkpoint, and a fragile software stack.
Use it to study extreme inference or finish low-volume offline jobs. For an
interactive application, use a smaller local model or call a hosted model
through the [reAPI quickstart](/docs/api/quickstart). The useful lesson is not
that 70B has become a 4GB model. It is that VRAM no longer has to hold every
weight at the same time.
## References
1. Gavin Li. *Unbelievable! Run 70B LLM Inference on a Single 4GB GPU with This New Technique.* November 30, 2023. [huggingface.co/blog/lyogavin/airllm](https://huggingface.co/blog/lyogavin/airllm)
2. AirLLM. *AirLLM repository, current quickstart, supported models, configuration, and FAQ.* Retrieved August 2, 2026. [github.com/lyogavin/airllm](https://github.com/lyogavin/airllm)
3. Hugging Face Accelerate. *Big Model Inference and the meta device.* [huggingface.co/docs/accelerate/usage\_guides/big\_modeling](https://huggingface.co/docs/accelerate/usage_guides/big_modeling)
4. Dao et al. *FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.* NeurIPS 2022. [arxiv.org/abs/2205.14135](https://arxiv.org/abs/2205.14135)
---
# Runway Alternatives in 2026: A Head-to-Head With reAPI (https://reapi.ai/blog/runway-alternatives)
Runway is very good at what it is. Gen-4.5, Aleph 2.0 and Act-Two are frontier models the company trained itself, the browser suite around them is the most complete creative environment anyone ships, and the developer side is honest pay-as-you-go with a published price for every model.\[1] Nothing below argues otherwise, and if Runway's own models are what you came for, no alternative substitutes for them.
The reason people search for Runway alternatives is narrower than "Runway is bad." It is that a large share of what Runway sells through its API is not Runway's model. Seedance, Kling, Veo, Hailuo, Grok Imagine and Nano Banana all run on Runway's platform at Runway's own published rates, and if those are the models your product actually calls, the question is simply whether that rate is competitive. This is a head-to-head on that specific overlap, priced from both vendors' published numbers on 9 August 2026.
## TL;DR
* **Runway wins on its own models.** Gen-4 Turbo at $0.05 a second and Gen-4.5 at $0.12 have no equivalent anywhere, because nobody else has the weights.\[2]
* **reAPI wins on the shared catalogue.** Seedance 2.5 at 720p is $0.30 a second on Runway's API against $0.266824 on reAPI, and at 480p it is $0.20 against $0.118589.\[2]\[3]
* **The gap widens on reference work.** A 10-second reference producing a 5-second 720p clip is $3.00 on Runway and $2.40 on reAPI, a 20% difference.\[2]\[3]
* **Credit granularity differs 10x.** Runway sells credits at $0.01; reAPI at $0.001.\[2]\[3]
* **Honest limits:** reAPI has no Gen-4.5, no asset library, no zero-data retention or IP indemnification, and a free tier worth $0.10. Those sections are below, not buried.
## Where the money actually goes
Anyone evaluating Runway alternatives should start here, because Runway runs two pricing systems and they behave nothing alike.
**Creative plans** are subscriptions with a credit allowance. Runway publishes its own conversion, which makes them unusually easy to audit: Standard is $12 a month annually for 625 credits, which they translate as 52 seconds of Gen-4.5. Pro is $28 for 187 seconds. Max is $76 for 791 seconds.\[1]
| Creative plan | Annual price | Stated Gen-4.5 output | Implied rate |
| ------------- | ------------ | --------------------- | ------------ |
| Standard | $12/mo | 52 s | $0.231 /s |
| Pro | $28/mo | 187 s | $0.150 /s |
| Max | $76/mo | 791 s | $0.096 /s |
The entry tier charges 2.4 times more per second than the top tier for identical output. That is not a criticism unique to Runway, it is how every credit-allowance plan works, but it is worth seeing before you pick a tier.
**Runway Dev** is the API, and it is straightforward metered pricing: credits at $0.01 each, a published rate per model, "from $0.05 / sec" for video.\[1]\[2] That $0.05 floor is Gen-4 Turbo, Runway's own cheapest model. It is not the rate for the third-party models most people are actually calling.
## Seedance 2.5, priced from both sides
Runway's API bills Seedance 2.5 at 30 credits per second of output at 720p and 20 at 480p, plus 15 and 10 credits per second of input and reference video, with reference images and audio free and an 80-credit minimum per generation.\[2] At $0.01 a credit that converts cleanly:
| Job | Runway API | reAPI | Difference |
| ---------------------------------- | ---------- | --------- | ---------- |
| 720p, per second | $0.30 | $0.266824 | 11% lower |
| 480p, per second | $0.20 | $0.118589 | 41% lower |
| 720p, 5 s clip | $1.50 | $1.335 | 11% lower |
| 720p, 30 s clip | $9.00 | $8.005 | 11% lower |
| 480p, 5 s clip | $1.00 | $0.593 | 41% lower |
| 720p, 10 s reference to 5 s output | $3.00 | $2.401 | 20% lower |
Sources: Runway's published API pricing\[2] and reAPI's model page.\[3]
The 480p row is the one worth looking at twice. Runway prices 480p at two-thirds of 720p; reAPI prices it at 44%. If your pipeline runs previews or drafts at 480p before committing to a final render, that gap compounds across every iteration you throw away.
One structural difference favours Runway on a narrow case: reference images and audio are free on their side, where reAPI folds all reference material into the same billing. If your workflow is many still references and no source video, run your own numbers rather than trusting the table above.
## Feature by feature
| | Runway | reAPI |
| ------------------------------------------------------ | ---------------------------------------------------------------------------------------- | --------------------------------------------------------------------- |
| Own frontier video models | Gen-4.5, Aleph 2.0, Act-Two\[1] | None |
| Cheapest first-party rate | Gen-4 Turbo, $0.05 /s\[2] | Not applicable |
| Browser creative suite, Characters, Recipes, Workflows | Yes\[1] | No |
| Asset storage | 5 GB free, 500 GB on Pro\[1] | None |
| Zero-data retention, IP indemnification, uptime SLA | Enterprise tier\[1] | Not offered |
| Model router that picks a model for you | Yes\[2] | No |
| Seedance 2.5, 720p per second | $0.30\[2] | $0.266824\[3] |
| Seedance 2.5, 480p per second | $0.20\[2] | $0.118589\[3] |
| Credit unit | $0.01\[2] | $0.001\[3] |
| Per-generation minimum on Seedance 2.5 | 80 credits\[2] | None |
| Credit expiry | Monthly on Creative plans, one month rollover on Max\[1] | No expiry |
| Subscription required | For Creative tiers; Dev is pay-as-you-go\[1] | Never |
| Free start | Dev free with no card; 125 credits on Creative\[1] | 100 credits, worth $0.10 |
| Relaxed-moderation tier | No published equivalent | `nsfw_checker` on selected models\[4] |
| Catalogue size | Larger, plus first-party models | 62 models\[5] |
Six of those fifteen rows go to Runway, and the first one goes to them permanently. No list of Runway alternatives can hand you Gen-4.5, because nobody else has the weights.
## Use Runway directly if you
Skip the rest of this article in any of these cases.
**You need Gen-4.5, Aleph 2.0 or Act-Two.** These are Runway's own weights. No aggregator carries them and none will.
**You want a creative environment, not an API.** The browser suite, asset library, Characters and Recipes are a product, not a wrapper, and reAPI has no equivalent to any of it.
**You need enterprise guarantees.** Zero-data retention, uncapped IP indemnification, uptime SLAs and formal governance are on Runway's Enterprise tier.\[1] reAPI offers none of those, and pretending otherwise would be the fastest way to lose your trust.
**Gen-4 Turbo fits the job.** At $0.05 a second it is cheaper than anything in the shared catalogue, on either platform.
## Honest limitations on the reAPI side
**The free tier is $0.10.** 100 credits at 1 credit = $0.001. That is roughly three 1K images or a fraction of one video. It exists to let you verify the API responds, not to let you evaluate quality for free. Runway's Dev tier starting free with no credit card is genuinely more generous at the trial stage.
**No self-hosting, no custom weights, no LoRA.** If you fine-tune, this is the wrong platform.
**No compliance surface.** No SOC 2, no HIPAA, no VPC deployment, no indemnification.
**Smaller catalogue.** 62 models on the live catalogue\[5], and none of them are first-party.
**No creative tooling.** No editor, no storage, no project management. reAPI returns a URL.
## Using both, which is what most teams end up doing
The two products sit at different layers, and the sensible arrangement is usually not a migration at all.
Keep Runway for the work that needs Runway: Gen-4.5 shots, Aleph edits, anything your team is doing by hand in the browser, anything covered by an enterprise agreement. Route the programmatic third-party generation, the Seedance and Kling and Veo calls your product makes on a schedule, through the cheaper metered path. The 11% to 41% you save on the shared catalogue funds the Gen-4.5 work.
That split also isolates a real operational risk. If one provider has a bad day, the workload that has a second path keeps running.
## Moving the shared-catalogue calls, about an hour
**Step 1: find out what you actually spend, 20 minutes.** Pull last month's Runway usage and split it into first-party models and everything else. If the third-party share is under about 20%, the rest of this is not worth doing.
**Step 2: map the model ids, 15 minutes.** Runway's `seedance2_5` is `doubao-seedance-2.5-face` on reAPI, with the same parameter surface. Full reference at [reapi.ai/docs/seedance-2-5](/docs/seedance-2-5).
**Step 3: change the base URL and key, 5 minutes.**
```bash
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer rk_live_your_key" \
-H "Content-Type: application/json" \
-d '{
"model": "doubao-seedance-2.5-face",
"prompt": "Your prompt",
"resolution": "720p",
"duration": 5
}'
```
**Step 4: run both for a week, 20 minutes of setup.** Send a slice of production traffic to each and compare output and invoices on the same prompts. Do not take either vendor's word for it, including this article's.
## FAQ
### Which Runway alternatives are actually cheaper?
On the models both platforms carry, yes, by 11% to 41% on Seedance 2.5 at the specs measured here.\[2]\[3] On Runway's own models the question does not apply, because reAPI does not have them.
### Can I get Gen-4.5 anywhere else?
No. Runway trained it and does not license it to aggregators.
### Does Runway's API require a subscription?
No. Runway Dev is pay-as-you-go and starts free without a credit card.\[1] The subscription tiers are the Creative product.
### What about Runway's unlimited offers?
Runway advertised 7 days of unlimited Seedance 2.5 on Max plans, on a model described as launching later, in a promotion expiring 14 August 2026.\[1] Unlimited tiers are worth pricing carefully; there is a four-step method for that in [reapi.ai/blog/unlimited-ai-video-generator-cost-test](/blog/unlimited-ai-video-generator-cost-test).
### Which has more models?
Runway, once you count its first-party models. reAPI's live catalogue shows 62.\[5] If catalogue size is your deciding factor, neither of these is the largest option on the market.
### Is there a downtime risk running only one provider?
Yes, on either platform, which is the argument for using both rather than switching entirely.
### Do reAPI credits expire?
No. Runway's Creative-plan credits are monthly, with one month of rollover on Max.\[1]
### What happens to my money if a generation fails?
On reAPI a failed task refunds the reserve in full.\[3] Check the equivalent policy wherever you buy, because at video rates a silent charge on a rejected prompt is a real line item.
## Deciding by layer, not by brand
The useful question is not which platform is better. Runway and reAPI are not competing for the same slot. Runway sells frontier models it trained plus the environment to use them by hand. reAPI sells metered access to a shared catalogue for code you are writing.
Split your workload along that line and the answer falls out. First-party models and manual creative work stay on Runway. Scheduled, programmatic, high-volume calls against models both platforms carry belong wherever the per-second rate is lower, and on Seedance 2.5 that is reAPI by 11% at 720p and 41% at 480p. Most teams evaluating Runway alternatives discover they wanted a second path rather than a replacement.
Live rates are on the [Seedance 2.5 model page](/models/seedance-2-5), and the full catalogue is at [reapi.ai/models](/models).
## References
1. Runway. *Pricing — Creative plans, Runway Dev pay-as-you-go, credit conversions and the Seedance 2.5 offer.* Retrieved 9 August 2026 from [runwayml.com/pricing](https://runwayml.com/pricing)
2. Runway. *API Pricing & Costs — per-model credit rates and the $0.01 credit price.* Retrieved 9 August 2026 from [docs.dev.runwayml.com/guides/pricing](https://docs.dev.runwayml.com/guides/pricing/)
3. reAPI. *Seedance 2.5 — published per-second rate band and refund behaviour.* Retrieved 9 August 2026 from [reapi.ai/models/seedance-2-5](/models/seedance-2-5)
4. reAPI. *doubao-seedance-2.5-face — parameter reference including `nsfw_checker`.* Retrieved August 2026 from [reapi.ai/docs/seedance-2-5](/docs/seedance-2-5)
5. reAPI. *Model catalogue.* Retrieved 9 August 2026 from [reapi.ai/models](/models)
### Further reading
* reAPI. *Unlimited AI Video Generator Plans: Run the Math First.* [reapi.ai/blog/unlimited-ai-video-generator-cost-test](/blog/unlimited-ai-video-generator-cost-test)
* reAPI. *Seedance 2.5 API Pricing: fal vs kie vs WaveSpeed vs reAPI.* [reapi.ai/blog/seedance-2-5-api-pricing-compared](/blog/seedance-2-5-api-pricing-compared)
---
# Seedance 2.0 Character Consistency: References, Voice, Shots (https://reapi.ai/blog/seedance-2-0-character-consistency-guide)
Character consistency is the reason Seedance 2.0 exists in its current shape: the model takes up to 9 images, 3 video clips, and 3 audio tracks as references in a single generation\[1], and ByteDance's own prompt guide documents a syntax for pinning a subject to specific reference images\[2]. Most of the frustration I see in community threads ("same character came back different," "it ignored my second reference") traces to rules that are actually written down and almost never read.
This guide is those rules, assembled from ByteDance's official prompt documentation, the API reference, and Dreamina's own tutorials, plus the honest boundaries: what has no official mechanism (cross-request character locks, seeds) and what belongs to which platform (the lipsync question).
## TL;DR
* **Seedance 2.0's input budget is 9 images + 3 videos + 3 audio clips**, with videos and audio each capped at 15 seconds total; audio cannot be sent without other content\[1].
* **The official binding syntax exists**: define your subject from a numbered image ("subject\@image 1"), and put the most important references first, because order carries weight\[2].
* **ByteDance recommends 4–5 references, not the maximum**, one headshot plus one full-body shot per character, and advises against multi-angle sets of the same face\[2].
* **There is no seed parameter and no persistent character ID.** Cross-video consistency comes from reusing the identical reference set, and from `return_last_frame` continuation\[1]\[3].
* **Voice guidance is real, lipsync wording is platform-specific**: the API docs describe audio references steering the performance\[1]; the explicit "lip-sync" promise appears in Dreamina's documentation, not the API's\[4].
* **Multi-shot has an official pattern**: number your shots ("Shot 1… Shot 2… Shot 3…") and treat exact per-shot timings as unreliable by ByteDance's own warning\[2].
## Seedance 2.0's 12-slot reference budget, precisely
Seedance 2.0 accepts references through a content array where each item declares a role. The documented limits: up to 9 reference images (JPEG/PNG/WebP among others, aspect ratio between 0.4 and 2.5, 300 to 6,000 pixels a side, under 30MB each), up to 3 reference videos totaling no more than 15 seconds, and up to 3 audio references also totaling 15 seconds\[1]. Audio is a modifier, not a subject: the API rejects requests that send audio with nothing else\[1].
One structural rule catches nearly everyone: first-frame mode and multimodal reference mode are mutually exclusive request shapes\[1]. You either hand the model an exact opening frame to animate, or a pile of references to compose from; mixing the two mental models in one request is the fastest route to "it ignored my image."
And the rule that generates the most confusing rejections: real human faces cannot be uploaded directly as references. That restriction and its sanctioned consent-based routes are their own topic, covered in our [not-eligible guide](/blog/seedance-not-eligible-explained).
## The official syntax for locking a character
ByteDance's prompt guide is unusually concrete about binding. The canonical pattern defines each subject from a numbered image, with a shorthand the docs render as "subject\@image 1"\[2]; the launch blog uses the same @-style reference in its own example prompts\[5], and Dreamina's tutorial mirrors it as @AssetName inside its editor\[6]. The point of the syntax is disambiguation: with several references in play, the prompt says explicitly which image owns the face, which owns the outfit, which owns the location.
Around that syntax, the guide's four working rules\[2]:
1. **Use 4–5 references, not 12.** The docs recommend against maxing the slots; every extra asset dilutes attention.
2. **Per character: one headshot plus one full-body image.** That pair anchors identity better than a stack of angles.
3. **Skip multi-view sets of the same person.** ByteDance explicitly advises against them; they invite identity drift rather than preventing it.
4. **Order by importance.** Earlier in the prompt means more influence. Put the character before the location, the location before the mood board.
That is the whole trick most "consistency hack" videos re-sell: define subjects explicitly, feed fewer and better references, order them deliberately.
## The same character across many videos
Here is the boundary the documentation draws and no tutorial should blur: Seedance 2.0 has no persistent character ID, no seed input, and no cross-request memory. Neither ByteDance's parameter reference nor reAPI's API surface exposes a seed at all\[1]\[3], so the "does Seedance use seed numbers" question has a clean answer: no, determinism control is reference-driven, not seed-driven.
What works instead, in production order:
**Freeze the reference set.** Same character images, same order, same binding phrases, request after request. Seedance 2.0's consistency comes from identical inputs, so store the set alongside your prompts and treat any change as a new character version.
**Chain with `return_last_frame`.** The API can hand back the final frame of a generation\[1]; feed it forward as the first frame of the next clip and you get scene-to-scene continuity with zero identity re-rolling at the joins.
**Let generated faces re-enter legally.** Content the platform generated for your account within the past 30 days is trusted as input even when it contains faces\[7], which is what makes iterative character work possible at all under the face rules.
For scale, ByteDance's technical report ranks Seedance 2.0 first on subject-consistency evaluation among peers\[5]; the machinery above is how that capability actually gets exercised across a series rather than a single clip.
## Voice, music, and the lipsync question
Seedance 2.0's audio references steer three things: the sound of the output, the timing of the performance, and the voice character. The API accepts up to three clips totaling 15 seconds in the reference\_audio role\[1], and the practical workflow for "make my character perform this song" is exactly what it sounds like: attach the track (a public URL on API platforms), bind your character from images, and prompt the performance.
On lipsync specifically, the sources split and it is worth being precise. ByteDance's API documentation describes joint audio-video generation and audio-guided output; the words "lip-sync" do not appear. Dreamina's product documentation is the surface that promises it, stating that a voice sample guides voice character and "aligns lip-sync, pacing, expressions"\[4]. In practice the same model family powers both, but if your pipeline contractually needs mouth-accurate sync, test on your own material rather than citing a docs page, because the API-side wording deliberately promises less.
Voice consistency across clips follows the same logic as visual consistency: reuse the same voice sample in the same slot every time. There is no voice ID to pin, so the sample is the ID.
## Multi-shot prompts that survive generation
The official pattern for multiple shots in one clip is numbered storyboard prose: "Shot 1: … Shot 2: … Shot 3: …"\[2]. Two warnings straight from the guide: exact per-shot second timings are not reliably honored, so write sequence and emphasis rather than timecodes\[2]; and shot count multiplies complexity, so the fewer-references rule matters double in multi-shot prompts.
For anything longer than one generation can hold (Seedance 2.0 caps at 15 seconds\[1]), the chain is storyboard → per-clip prompts with the frozen reference set → `return_last_frame` joins. That is also the workflow that Seedance 2.5's announced 30-second single-pass generation and segment-level prompt control are designed to collapse; our [Seedance 2.5 pre-launch guide](/blog/seedance-2-5-what-we-know-2026) tracks what is confirmed there.
If you want to run all of this over an API, the whole reference surface (images, videos, audio, first frame, last-frame return) is exposed on [reAPI's Seedance 2.0](/models/seedance-2-0) with per-second billing, and reference-mode requests bill lower than pure text-to-video\[3].
## FAQ
### How many reference images can Seedance 2.0 take?
Up to 9 images, plus 3 videos (15s total) and 3 audio clips (15s total) in one request\[1]. ByteDance's own guidance is to use 4–5 well-chosen assets rather than the maximum\[2].
### How do I keep the same person across multiple Seedance 2.0 videos?
Freeze one reference set (headshot + full body), bind it with the subject\@image syntax, reuse it identically in every request, and chain scenes with `return_last_frame`\[1]\[2]. There is no character-lock parameter; the reference set is the lock.
### Does Seedance 2.0 use seed number inputs?
No. No seed parameter exists in ByteDance's documented request schema or on reAPI's surface\[1]\[3]. Repeatability comes from identical references and prompts, and it is soft repeatability, not bitwise.
### Can Seedance 2.0 lipsync to a Suno song?
Attach the track as an audio reference and prompt the vocal performance; the API documents audio-guided generation\[1], while the explicit lip-sync alignment claim comes from Dreamina's docs\[4]. For release-quality sync, validate on your own footage.
### How many audio clips can Seedance 2.0 take?
Three, totaling no more than 15 seconds, and never alone; audio must accompany other content in the request\[1].
### How do I prompt multiple shots in one Seedance 2.0 clip?
Numbered storyboard style, "Shot 1 / Shot 2 / Shot 3," per the official guide, which also warns that exact per-shot durations are approximate\[2].
### Why did my character's face get rejected?
Real-person faces cannot be uploaded directly as references; that is a model-level rule with documented consent-based exceptions\[7]. Full breakdown in our [not-eligible guide](/blog/seedance-not-eligible-explained).
## The consistency playbook on one line
Bind subjects explicitly, feed 4–5 deliberate references with the character first, freeze that set for the whole series, chain clips through the last frame, and reuse the same voice sample when sound matters. Everything above comes from ByteDance's own documentation rather than folklore, and every piece of it runs over one endpoint on [reAPI](/models/seedance-2-0). That is Seedance 2.0 character consistency without the mystery: fewer, better references, used the same way every time.
## References
1. Volcano Engine / BytePlus (ByteDance). *Seedance 2.0 API reference — content roles, reference limits, first-frame exclusivity, return\_last\_frame.* Retrieved July 2026 from [docs.byteplus.com/en/docs/ModelArk/1520757](https://docs.byteplus.com/en/docs/ModelArk/1520757)
2. Volcano Engine / BytePlus (ByteDance). *Seedance 2.0 official prompt guide — subject binding syntax, reference count and ordering, multi-shot pattern.* Retrieved July 2026 from [docs.byteplus.com/en/docs/ModelArk/2222480](https://docs.byteplus.com/en/docs/ModelArk/2222480)
3. reAPI. *Seedance 2.0 — API documentation and model page.* Retrieved July 2026 from [reapi.ai/docs/seedance-2-0](/docs/seedance-2-0)
4. Dreamina (CapCut). *Seedance 2.0 tool page — voice sample and lip-sync alignment.* Retrieved July 2026 from [dreamina.capcut.com/tools/seedance-2-0](https://dreamina.capcut.com/tools/seedance-2-0)
5. ByteDance Seed. *Seedance 2.0 launch blog and technical report (arXiv:2604.14148).* Retrieved July 2026 from [seed.bytedance.com/en/blog/seedance-2-0-official-launch](https://seed.bytedance.com/en/blog/seedance-2-0-official-launch)
6. Dreamina (CapCut). *Seedance 2.0 prompt tutorial — @AssetName references and Multiframes.* Retrieved July 2026 from [dreamina.capcut.com/resource/seedance-2-0-prompt](https://dreamina.capcut.com/resource/seedance-2-0-prompt)
7. Volcano Engine / BytePlus (ByteDance). *Trusted-input exemptions for generated content and consent-verified assets.* Retrieved July 2026 from [docs.byteplus.com/en/docs/ModelArk/2291680](https://docs.byteplus.com/en/docs/ModelArk/2291680)
### Further reading
* reAPI. *Seedance 2.0 "Not Eligible": Why It Happens, What Works.* [reapi.ai/blog/seedance-not-eligible-explained](/blog/seedance-not-eligible-explained)
* reAPI. *What Is Seedance 2.0 and How to Use It (2026 Guide).* [reapi.ai/blog/what-is-seedance-2-0-and-how-to-use-it](/blog/what-is-seedance-2-0-and-how-to-use-it)
* reAPI. *Seedance 2.5: What We Know Before the Public Launch.* [reapi.ai/blog/seedance-2-5-what-we-know-2026](/blog/seedance-2-5-what-we-know-2026)
---
# Seedance 2.0 Cost Per Second: The Real Billing Model (https://reapi.ai/blog/seedance-2-0-cost-per-second)
Seedance 2.0 bills by output duration. You pay for the seconds of finished video the model returns, not for how long a GPU ran, how long your prompt was, or how large your input files were. On reAPI the Seedance 2.0 cost per second runs from **$0.045 to $1.04**, and which end of that range you land on is decided by three choices you make in the request.
Two of those choices are obvious: speed tier and resolution. The third one is not, and it is where most of the money leaks. There is a cheaper rate that applies only when you upload a source **video**, and passing images or first/last frames does not qualify, even though it feels like it should.
## TL;DR
* **Billing is per second of output**, with duration clamped to a 4 to 15 second range.
* **Three dimensions set the rate**: speed tier (standard or Fast), resolution (480p through 4K), and whether the request carries an uploaded source video.
* **The uploaded-video rate is roughly 39% cheaper** at 720p, $0.125/s against $0.205/s.
* **It only triggers on `video_urls`.** Image inputs and first/last frame references stay on the higher rate.
* **A 5-second 720p clip costs $1.03** from a text prompt, or **$0.63** if it is driven by an uploaded video.
* **Fast at 480p is the cheapest path** at $0.078/s, or $0.045/s with an uploaded video.
## The full rate card
Rates below are per second of output video\[1].
| Resolution | Standard, from prompt | Standard, with uploaded video | Fast, from prompt | Fast, with uploaded video |
| ---------- | --------------------- | ----------------------------- | ----------------- | ------------------------- |
| 480p | $0.095 | **$0.058** | $0.078 | **$0.045** |
| 720p | $0.205 | **$0.125** | $0.165 | **$0.100** |
| 1080p | $0.510 | **$0.310** | n/a | n/a |
| 4K | $1.040 | **$0.640** | n/a | n/a |
Two structural facts fall out of that table.
**Resolution is the steepest lever.** Going from 480p to 4K on the standard tier multiplies the rate by about 11. Going from standard to Fast at the same resolution saves roughly 18 to 20%. If a clip does not need to be 4K, resolution is the first thing to cut, not the tier.
**Fast stops at 720p.** There is no Fast path to 1080p or 4K, so any high-resolution work is on the standard tier by definition.
## The rate that catches people out

The cheaper column is labeled "& uploaded videos," and the label is precise in a way that is easy to misread.
That rate applies **only when the request carries an uploaded source video**. In API terms, only when `video_urls` is populated. Every other kind of reference material, an image, a first frame, a last frame, audio, leaves you on the base rate.
This matters because image-to-video and first/last-frame workflows *feel* like reference-driven generation. You are supplying source material, the model is transforming it rather than inventing from nothing, and the intuition is that it should cost less. It does not.
At 720p the gap is 39%: $0.205 per second from a prompt or an image, against $0.125 per second when a source video is attached. On a 10-second clip that is $2.05 against $1.25.
The practical consequence is a forecasting error rather than a bug. Teams model their costs on the cheaper column because their workflow "uses references," then the invoice arrives priced on the base column. If you are budgeting an image-to-video pipeline, budget it at the from-prompt rate.
## Worked examples
Multiply the rate by output seconds. That is the whole model.
**A 5-second 720p clip from a text prompt:** 5 × $0.205 = **$1.03**
**The same clip driven by an uploaded video:** 5 × $0.125 = **$0.63**
**A 10-second 1080p clip from a prompt:** 10 × $0.510 = **$5.10**
**A 5-second 480p draft on Fast:** 5 × $0.078 = **$0.39**
**A 15-second 4K hero shot:** 15 × $1.040 = **$15.60**
At volume, the shape of the spend matters more than the rate. A thousand 5-second 720p clips a month from prompts runs about **$1,025**. The same thousand clips at 480p on Fast runs about **$390**. The same thousand at 1080p runs about **$2,550**.
The formula for a mixed workload:
```text
monthly cost ≈ Σ (output seconds × rate for that resolution, tier, and input mode)
```
Duration is clamped to between 4 and 15 seconds, so a request for 3 seconds bills as 4, and there is no single-call path past 15.
## Picking a configuration
**Draft on Fast at 480p.** At $0.078 per second, a 5-second iteration costs 39 cents. Iterate there, then render the approved shot once at the resolution you actually ship.
**Only pay for 1080p and 4K on final output.** The 480p-to-4K multiplier is roughly 11x. A 4K draft is the most expensive mistake available in this pricing model.
**If your pipeline can be video-driven, make it video-driven.** The 39% discount at 720p is real, but only for `video_urls`. If you are already producing a source clip, feeding it in rather than describing it saves meaningfully.
**Do not assume image-to-video is discounted.** Budget it at the from-prompt rate.
## Calling it and knowing the bill
The endpoint is async: submit returns a `task_id`, then poll.
```bash
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "doubao-seedance-2.0",
"prompt": "a slow dolly through a neon-lit alley after rain, cinematic",
"resolution": "720p",
"duration": 5
}'
```
That request bills at the from-prompt rate. Adding `video_urls` moves it to the cheaper column:
```bash
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "doubao-seedance-2.0",
"prompt": "restyle this shot as neon-lit and rain-soaked, keep the camera move",
"video_urls": ["https://example.com/source.mp4"],
"resolution": "720p",
"duration": 5
}'
```
Two platform rules worth knowing. Media inputs are **public http(s) URLs only**, with no base64 accepted on any model. And rates move, so the live table on [reapi.ai/models/seedance-2-0](/models/seedance-2-0) is canonical rather than any figure quoted in an article, including this one. Request and response shapes are in the [reapi.ai/docs/seedance-2-0](/docs/seedance-2-0) reference.
## FAQ
### How much does Seedance 2.0 cost per second?
From $0.045 to $1.04 per second of output, depending on speed tier, resolution, and whether the request carries an uploaded source video. A 720p clip from a text prompt is $0.205 per second\[1].
### Does prompt length affect the price?
No. Billing is by output duration only. Prompt length, input file size, and server-side queue time do not change the charge.
### Why is my image-to-video generation billed at the higher rate?
Because the cheaper rate applies only when an uploaded source **video** is present. Images, first frames, and last frames stay on the base rate even though they are reference material.
### What does a 10-second 720p video cost?
$2.05 from a text prompt, or $1.25 when driven by an uploaded video.
### Is Fast available at 1080p?
No. Fast covers 480p and 720p. Anything above that runs on the standard tier.
### What is the minimum and maximum duration?
Duration is clamped between 4 and 15 seconds, so shorter requests bill as 4 seconds and there is no single-call path beyond 15.
### How do I estimate a monthly bill?
Sum your expected output seconds per configuration and multiply each by its rate. A thousand 5-second 720p prompt-driven clips is roughly $1,025 a month.
### Where is the authoritative price?
The live table on the [Seedance 2.0 model page](/models/seedance-2-0). Rates change, so treat any number in an article as a planning figure.
## Forecasting the bill instead of discovering it
The Seedance 2.0 cost per second is simple arithmetic once you know which of the four columns you are actually in. The mistake that costs real money is not misreading the rate card, it is assuming that supplying reference material puts you in the cheaper column when only an uploaded video does that.
So the useful habit is small: before you model unit economics, check whether the pipeline sends `video_urls` or an image. Draft at 480p on Fast, reserve 1080p and 4K for shots that ship, and multiply expected output seconds by the live rate rather than the one you remember.
## References
1. reAPI. *Seedance 2.0 model page — live per-second rates by resolution, tier, and input mode.* Retrieved July 2026 from [reapi.ai/models/seedance-2-0](/models/seedance-2-0)
### Further reading
* reAPI. *Seedance 2.5 features.* [reapi.ai/blog/seedance-2-5-features](/blog/seedance-2-5-features)
* reAPI. *Seedance 2.0 endpoint reference.* [reapi.ai/docs/seedance-2-0](/docs/seedance-2-0)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# Is Seedance 2.0 Uncensored? Safety Filters and API Control (2026) (https://reapi.ai/blog/seedance-2-0-uncensored-safety-filter-api)
**Seedance 2.0 is not literally uncensored. On reAPI, direct API callers can
set `nsfw_checker: false` to request a more permissive Flexible route on
compatible Standard models, but that switch does not remove every content,
model, rights, or platform rule.** The safer description is **less restricted
Seedance 2.0**, not a rule-free video model.\[1]
That distinction matters because “Seedance 2.0 uncensored” is used online to
mean several different things: no extra prompt filter, fewer output blocks,
support for real-person references, or no moderation at all. Those are not the
same product. This guide defines the terms, shows exactly what reAPI's
`nsfw_checker` flag controls, and explains which safety boundaries remain.
## TL;DR
* **The default is moderated.** `nsfw_checker` defaults to `true`.
* **The API offers an explicit control.** Compatible Standard model requests
can send `"nsfw_checker": false`; the normal browser playground keeps the
control locked on.\[1]
* **“False” requests a route, not immunity.** reAPI uses a compatible Flexible
channel when available. If it cannot, the request may stay on a
safety-checked Standard route.
* **Official is different.** The Official model IDs do not expose this switch,
and the direct official channel does not accept real-person reference
uploads.\[1]\[3]
* **Some rejections still happen.** Model-level or output-policy failures can
remain terminal even after an optional checker is disabled.
* **Rights still apply.** Consent, copyright, acceptable-use rules, and your
own application's moderation obligations do not disappear.
## What “Seedance 2.0 uncensored” can mean
There is no industry-standard definition of “uncensored AI video.” Treat it as
a marketing label until a provider states which layer is actually changed.
| Phrase people use | What it may actually mean | What it does not prove |
| ---------------------- | --------------------------------------------------------- | --------------------------------------------------- |
| No extra prompt filter | The host does not add a pre-generation keyword gate | The model will accept every prompt |
| NSFW checker off | One host-side check or route choice is disabled | All provider and output checks are gone |
| Real-person support | A designated model/channel accepts authorized face inputs | Consent is optional |
| Relaxed moderation | Fewer borderline, non-explicit prompts are rejected | Prohibited content is allowed |
| Fully uncensored | Supposedly no checks or rules at any layer | A credible commercial API rarely makes this promise |
For reAPI, the precise claim is simple: **the Seedance 2.0 API exposes an
optional `nsfw_checker` field on compatible Standard requests.** It is on by
default. Setting it to false asks the router to use the Flexible path when the
request and channel support it.\[1]
That is useful control for legitimate production teams dealing with false
positives in editorial fashion, stylized performance, fictional characters,
historical drama, or other lawful material that a broad safety classifier may
misread. It is not a prompt-obfuscation feature and should not be used to evade
law, consent, copyright, or the service's acceptable-use terms.
## Seedance 2.0 censorship is a stack, not one switch

A generated video can be stopped at multiple points. Turning off one optional
checker does not remove the others.
### 1. Client or product-layer checks
A consumer app may screen the prompt before it ever reaches Seedance. It can
also block uploads, certain words, or categories based on its own audience and
app-store obligations. Two sites using the same model can therefore feel very
different.
This is why a prompt rejected in one interface may be accepted through an API
without any change to the underlying model. The difference can be the host's
extra filter rather than the model itself.
### 2. Route-level moderation
An API host can send a request through a Standard, Flexible, face-enabled, or
official route. On reAPI, `nsfw_checker` participates in that routing decision.
The default `true` keeps the standard safety behavior; `false` requests the
more permissive compatible path.\[1]
This is the layer most people mean when they search for an “uncensored
Seedance 2.0 API.” More accurately, they are looking for an API without an
additional conservative host check.
### 3. Model and provider policy
The generation model still has learned and enforced boundaries. The upstream
provider can reject an input, refuse to render a scene, or return a policy
error. A routing flag cannot promise that every requested frame will be
generated.
BytePlus publishes separate ModelArk acceptable-use terms for its generative
AI services, and its official Seedance documentation describes input
restrictions for real-person reference media.\[3]\[5]
### 4. Output review
A prompt can pass and the rendered result can still fail a final output check.
This is not contradictory: generative models are probabilistic, so a harmless
brief can sometimes produce an image that crosses a policy threshold.
On reAPI, a task that ends with error `80006` is a content-policy failure. It
is different from a provider submission or infrastructure failure, and should
be treated as terminal rather than retried in a loop.\[7]
## What `nsfw_checker: false` actually does on reAPI
The `nsfw_checker` name is easy to overread. It is a Boolean API control, not a
promise that Seedance has no safety system.
| Behavior | `true` or omitted | `false` |
| ----------------------------- | -------------------------- | ---------------------------------------- |
| Default | Yes | No; must be explicit |
| Intended route | Standard safety behavior | Compatible Flexible route when available |
| Browser playground | Locked on for normal users | Direct API only |
| Official model IDs | No comparable toggle | Not supported |
| Guaranteed generation | No | No |
| Rights and policy obligations | Still apply | Still apply |
The important implementation detail is graceful compatibility. If no
compatible Flexible channel is available for the selected model and request,
reAPI can use the Standard route instead, where safety checking remains on.
Therefore `nsfw_checker: false` should not be treated as a deterministic “all
content accepted” mode.\[1]
It also does not convert the Official channel into a face-enabled route. The
official IDs—`doubao-seedance-2.0-official` and
`doubao-seedance-2.0-fast-official`—are a separate channel with their own input
contract. If a lawful workflow needs real-person references, use a designated
face variant and retain the subject's consent and asset rights.
## How to call the less-restricted Seedance 2.0 API
Use the normal asynchronous video endpoint and add the explicit Boolean. This
example stays intentionally benign: the flag is demonstrated as an API
setting, not as a way to encode prohibited content.
```bash
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer $REAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "doubao-seedance-2.0-face",
"prompt": "A cinematic editorial performance by a fictional adult character, original blue costume, dramatic red stage lighting, slow dolly-in, no logos",
"resolution": "720p",
"size": "16:9",
"duration": 5,
"nsfw_checker": false
}'
```
The POST returns an asynchronous task. Poll the shared task endpoint until it
completes:
```bash
curl https://reapi.ai/api/v1/tasks/TASK_ID \
-H "Authorization: Bearer $REAPI_API_KEY"
```
A successful task returns the generated file in `output.video_urls`. If the
task fails, inspect the structured error code before deciding whether to retry.
Infrastructure failures may be retryable; a content-policy decision should be
treated as a final result for that request.\[7]
Three practical rules make this safer in production:
1. **Keep the default on unless the use case needs the Flexible route.** This
gives ordinary requests the standard safety posture.
2. **Expose the setting only to trusted users or internal workflows.** A public
app still needs abuse controls, traceability, and a takedown process.
3. **Moderate the final asset for its destination.** An API's acceptance does
not mean the result is suitable for an app store, ad network, marketplace,
workplace, or public feed.
## Which Seedance 2.0 model should you use?
The route name and the model name answer different questions. Choose the model
for capability first, then choose the moderation behavior supported by that
surface.
| Need | Recommended surface | Why |
| ------------------------------------------ | ------------------------------- | -------------------------------------------------------------------------------- |
| General text or reference video generation | `doubao-seedance-2.0` | Full-resolution Standard family |
| Authorized real-person reference media | `doubao-seedance-2.0-face` | Face-enabled Standard variant |
| Faster iterations at 480p/720p | `doubao-seedance-2.0-fast-face` | Lower-latency face-enabled workflow |
| Lowest-friction official direct route | `doubao-seedance-2.0-official` | Separate Official contract, but no real-person uploads or `nsfw_checker` control |
The full model family supports text-to-video, image-to-video, first/last-frame
transitions, and multimodal reference workflows. Resolution and input limits
vary by variant, so use the [live Seedance 2.0 model page](/models/seedance-2-0)
and [API reference](/docs/seedance-2-0) as the contract rather than copying a
third-party “uncensored” label.\[1]\[2]
## Why safe Seedance prompts still get censored
A rejection does not automatically mean the prompt was malicious. Broad
classifiers trade precision for coverage, and that creates false positives.
Common causes include:
* a fictional or synthetic face that looks like a real person;
* theatrical makeup, skin-toned materials, or close crops that lose context;
* an original design that resembles protected character or brand imagery;
* a neutral prompt paired with a reference asset carrying metadata or visual
details the classifier flags;
* a valid prompt whose probabilistic output, rather than input, crosses a
threshold.
Do not respond by hiding keywords or repeatedly resubmitting the same rejected
request. Instead, identify which layer rejected it.
For a lawful false positive, a good diagnostic sequence is:
1. Confirm the prompt and all references are original, licensed, or used with
documented consent.
2. Run the same benign request once with `nsfw_checker: true` and once with
`false` on a compatible Standard model.
3. Record whether the failure occurred before submission, during provider
generation, or at final output review.
4. Check the task's structured error code and avoid automatic retry on a
content-policy result.
5. If the issue is a real-person input, move to a face-enabled workflow rather
than trying to disguise the subject.
This A/B test tells you whether the optional host check was the cause. It does
not attempt to defeat a model-level decision.
## What a less-restricted API still requires from SaaS teams
If you are integrating Seedance into a product, moderation becomes a product
design responsibility rather than a box to remove.
BytePlus's rules for platform customers call for end-user identity controls,
traceability from generated content to user accounts, warnings, content
takedowns, account restrictions, audit logs, incident response, and rights
records for trusted or real-person assets.\[6]
Your exact obligations depend on your service and jurisdiction, but the
engineering pattern is broadly useful:
* keep a durable task-to-user mapping;
* log the model ID, route choice, input asset IDs, and policy outcome;
* rate-limit repeated failures and suspicious bursts;
* require a rights attestation for real-person or branded references;
* provide reporting, takedown, and account-enforcement tools;
* separate internal creative testing from public user-generated output;
* review the destination platform's rules before publishing.
This is the practical meaning of responsible API control. A more permissive
route can reduce false positives for trusted creators while your application
keeps the safeguards appropriate to its users and distribution channel.
## FAQ
### Is Seedance 2.0 uncensored?
No—not in the literal sense of having no rules. reAPI supports a less-restricted
Flexible path for compatible Standard requests when direct API callers set
`nsfw_checker: false`. Model, provider, output, rights, and platform policies
can still reject or prohibit content.\[1]
### Does Seedance 2.0 support NSFW content?
The field name is `nsfw_checker`, but the Boolean should not be interpreted as
blanket approval for any content category. It controls an optional checking
and routing layer. The acceptable-use policy and model/output decisions remain
in force.\[5]
### How do I turn off the Seedance safety filter?
On compatible reAPI Standard model requests, send
`"nsfw_checker": false` through the direct API. The normal browser playground
keeps the setting locked on. This changes the eligible route; it does not
disable every safety layer or authorize policy violations.
### Can the Official Seedance 2.0 API disable moderation?
The Official model IDs documented by reAPI do not expose `nsfw_checker`.
BytePlus's official video API has its own input and policy contract, including
restrictions around direct real-person reference uploads.\[3]\[4]
### Does `nsfw_checker: false` guarantee my video will generate?
No. A compatible Flexible route may still return a model or output-policy
failure. If a compatible route is unavailable, the task may use Standard
safety behavior instead. Treat the flag as a route preference, not a success
guarantee.
### Can I use a real person's face?
Use a face-enabled model only when you have the subject's authorization and
the necessary rights. The Official direct channel does not accept real-person
reference uploads, while reAPI's Standard face variants are designed for
authorized real-person workflows.\[1]
### Will a failed moderation check consume credits?
reAPI refunds failed generations automatically. Check the task response to
distinguish a refunded failure from a completed generation and to identify the
error code.\[1]\[7]
## The accurate label is “less restricted,” not “no rules”
The useful part of a Seedance 2.0 uncensored API is not the word
“uncensored.” It is explicit, documented control over one moderation layer.
On reAPI, the safe default stays on, trusted direct API workflows can request a
compatible Flexible route, and the Official channel remains a separate
contract.
That clarity lets a SaaS team reduce avoidable false positives without
pretending the rest of the safety stack has vanished. Start with the default,
use `nsfw_checker: false` only where the workflow needs it, retain rights and
consent records, and handle final policy decisions as final. The current
[Seedance 2.0 API documentation](/docs/seedance-2-0) lists the supported model
IDs and request fields, while the [model page](/models/seedance-2-0) carries
live pricing and capability details.
**Disclosure:** reAPI publishes this article and operates the routed Seedance
2.0 endpoint described above. reAPI behavior comes from our public API
reference; official-channel restrictions and policy obligations come from
BytePlus documentation. “Uncensored” is discussed as a search and marketing
term, not as a promise that prohibited content is accepted.
## References
1. reAPI. *Seedance 2.0 API reference — channels, `nsfw_checker`, billing, inputs, and real-person variants.* Retrieved August 1, 2026. [reapi.ai/docs/seedance-2-0](/docs/seedance-2-0)
2. reAPI. *Seedance 2.0 model page and live pricing.* Retrieved August 1, 2026. [reapi.ai/models/seedance-2-0](/models/seedance-2-0)
3. BytePlus ModelArk. *Dreamina Seedance 2.0 series tutorial.* Updated July 31, 2026. [docs.byteplus.com/api/docs/ModelArk/2291680](https://docs.byteplus.com/api/docs/ModelArk/2291680)
4. BytePlus ModelArk. *Create a video generation task — official API reference.* Updated July 31, 2026. [docs.byteplus.com/en/docs/modelark/1520757](https://docs.byteplus.com/en/docs/modelark/1520757)
5. BytePlus. *GenAI Acceptable Use Policy.* Updated July 11, 2026. [docs.byteplus.com/en/docs/legal/acceptable\_use\_policy\_byteplus\_genai](https://docs.byteplus.com/en/docs/legal/acceptable_use_policy_byteplus_genai)
6. BytePlus ModelArk. *Platform Customer Code of Conduct and Default Handling Rules.* Updated June 16, 2026. [docs.byteplus.com/en/docs/ModelArk/2353368](https://docs.byteplus.com/en/docs/ModelArk/2353368)
7. reAPI. *API error codes and task failure responses.* Retrieved August 1, 2026. [reapi.ai/docs/api/errors](/docs/api/errors)
### Further reading
* reAPI. *What Is Seedance 2.0 and How to Use It.* [reapi.ai/blog/what-is-seedance-2-0-and-how-to-use-it](/blog/what-is-seedance-2-0-and-how-to-use-it)
* reAPI. *Seedance 2.0 Character Consistency Guide.* [reapi.ai/blog/seedance-2-0-character-consistency-guide](/blog/seedance-2-0-character-consistency-guide)
* reAPI. *How Long Can Seedance Videos Be?* [reapi.ai/blog/how-long-can-seedance-videos-be](/blog/how-long-can-seedance-videos-be)
---
# Seedance 2.0 vs Happyhorse 1.0: Picking a Video Model 2026 (https://reapi.ai/blog/seedance-2-0-vs-happyhorse-1-0-2026)
For two months in 2026, ByteDance's Seedance 2.0 sat at the top of the Artificial Analysis Video Arena. Then on April 7, an anonymous model named Happyhorse 1.0 appeared on the leaderboard and took both the text-to-video and image-to-video crowns. Three days later, CNBC and Bloomberg confirmed it: Happyhorse is Alibaba's ATH division's first public entry. The model went from stealth to #1 to identified inside a single news week\[1].
If you're picking Seedance 2.0 vs Happyhorse 1.0 in 2026, the leaderboard isn't the whole story. They're built for different workflows. Seedance 2.0 generates multi-shot sequences with phoneme-level lip-sync in 8+ languages. Happyhorse 1.0 uniquely lets you rewrite an existing video, has 50+ baked-in style presets, and currently leads the headline leaderboard. Different tools, different jobs.
## TL;DR
* **Origins.** Seedance 2.0 from ByteDance Seed, released February 12, 2026\[2]. Happyhorse 1.0 from Alibaba's ATH Innovation Division (Tongyi Lab + others), stealth-debuted on Artificial Analysis April 7, 2026; identity confirmed April 10, 2026\[1].
* **Leaderboard.** Happyhorse 1.0 currently leads both text-to-video and image-to-video Video Arena rankings on Artificial Analysis as of late April 2026\[3]. Seedance 2.0 held the top spot before Happyhorse displaced it.
* **Architecture.** Happyhorse is a 15B-parameter unified 40-layer Transformer with native audio in 7 languages\[1]. Seedance 2.0 is a unified multimodal audio-video Transformer with phoneme-level lip-sync in 8+ languages\[2].
* **Unique to Happyhorse.** EDIT mode: send a `video_url` and the model rewrites or restyles the source clip while preserving motion. 50+ style presets baked in.
* **Unique to Seedance 2.0.** Multi-shot single-generation (multiple cuts inside one 15-second clip), 9 image + 3 video + 3 audio multi-modal references in one request, dedicated face-aware variants, 21:9 cinematic aspect.
* **Pricing.** Happyhorse 1.0 on reAPI: $0.1625/s at 720P, $0.2875/s at 1080P (0% markup, list-price passthrough)\[4]. Seedance 2.0 on reAPI: $0.0865–$0.4048/s depending on tier and reference mode\[5].
* **The split.** Happyhorse for video editing / restyling and the latest leaderboard quality. Seedance 2.0 for multi-shot storyboards, lip-synced dialogue, reference-heavy compositions.
## Where each model comes from
Seedance 2.0 launched on February 12, 2026 from ByteDance's Seed research group\[2]. ByteDance describes it as a "next-generation video creation model" with a "unified multimodal audio-video joint generation architecture" supporting text, image, audio, and video inputs. The model went viral in China for photorealistic clips of named celebrities, and Disney sent ByteDance a cease-and-desist letter on February 13, 2026 over training-data concerns\[6]. Seedance 2.0 ships with C2PA watermarking by default.
Happyhorse 1.0 had a stranger debut. The model appeared on the Artificial Analysis Video Arena leaderboard on April 7, 2026 with no listed creator, climbed to #1 in both text-to-video and image-to-video blind tests, and stayed there for three days before Alibaba publicly claimed authorship on April 10, 2026\[7]. The team is Alibaba's ATH Innovation Division: Tongyi Lab plus Alibaba Platform Technology and Taotian Tech\[1]. Bailian enterprise API testing went live April 27, 2026; full commercial launch in May 2026\[1].
Both models are accessible through reAPI on the same OpenAI-compatible endpoint (`POST /api/v1/videos/generations`). Switching between them is a one-field change.
## What each model can actually do
| Capability | Seedance 2.0 | Happyhorse 1.0 |
| ------------------------------ | -------------------------------------------------------------------------------------- | --------------------------------------------------------------------- |
| Text-to-video | yes | yes |
| Image-to-video (single ref) | yes | yes (`first_frame_image`) |
| Image-to-video (multi-ref) | up to 9 images | up to 9 images (R2V mode) |
| First/last-frame interpolation | yes (`image_with_roles`) | first-frame only (no last-frame anchoring) |
| **Video editing / restyling** | **no** | **yes (EDIT mode)** |
| Reference video for style | REF mode, ≤3 clips, generates new video | EDIT mode, source clip itself |
| Reference audio | up to 3 clips (must accompany visual) | source clip's audio (EDIT, `audio_setting: "origin"`) |
| Audio synthesis | native joint, phoneme-level lip-sync, 8+ languages\[2] | native, synchronized, 7 languages\[1] |
| Multi-shot in single output | yes, multiple cuts in one generation\[2] | no (single continuous shot) |
| Resolution ceiling | 1080p | 1080P |
| Duration | 4–15s | 3–15s (T2V/I2V/R2V); EDIT = source length capped at 15s |
| Aspect ratios | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, adaptive | 16:9, 9:16, 1:1, 4:3, 3:4 |
| Face-aware variants | yes (`-face`, `-fast-face`) | no separate variant |
| Style presets | none built in | 50+ baked in\[1] |
| C2PA watermark | default on | optional (`watermark` flag, default off) |
The biggest single architectural divide:
**Happyhorse 1.0 has an EDIT mode that Seedance 2.0 doesn't.** Send a `video_url` to Happyhorse and it rewrites the clip. Restyle the character into 3D cartoon, keep the original motion, optionally preserve the original audio track. Seedance 2.0's REF mode uses video as a style reference but generates a new clip from scratch; it doesn't modify the source. For content remix, animation restyling, or "take this rough footage and make it look like X" workflows, Happyhorse is the only choice between the two.
**Seedance 2.0's multi-shot single-output is the inverse advantage.** One prompt, one 15-second video, multiple cuts and transitions inside it\[2]. Happyhorse outputs single continuous shots. For storyboard-driven content where you want an edited sequence in one API call, Seedance saves you a stitching pass.
## The leaderboard
As of April 2026, Happyhorse 1.0 holds the #1 position on the Artificial Analysis Video Arena for both text-to-video and image-to-video\[3]. The benchmark runs blind A/B tests; humans pick between two anonymized model outputs and Elo scores update from those preferences. CNBC and Bloomberg both flagged Happyhorse's debut as displacing Seedance 2.0 from the top spot it had held since February\[7]\[8].
Three caveats stay relevant. The Video Arena measures aggregate human preference on prompt-level outputs and doesn't capture multi-shot quality, lip-sync fidelity, or workflow ergonomics. Independent comparisons report Seedance 2.0 edging Happyhorse in audio-synced output by a small margin while Happyhorse leads by a wider margin in silent video\[9]. And the leaderboard delta might compress as more diverse prompts hit Happyhorse since it only debuted publicly on April 7.
For projects where leaderboard standing matters, Happyhorse 1.0 is the model to cite as of May 2026. For projects where workflow fit matters more than the headline number, the leaderboard is informational at best.
## EDIT mode vs multi-shot output
The two unique capabilities map to different shipping pipelines.
**Happyhorse EDIT mode** is for workflows that already have raw video and need to transform it: restyle existing footage (live action to anime, 2D to 3D, day to night), apply a brand visual identity to UGC clips, repaint a character while preserving original motion. You send `prompt` + `video_url` (+ optionally up to 5 `image_urls` as style references), and Happyhorse outputs the source clip rewritten to match the prompt. EDIT mode bills by the source clip's actual length (server-probed via ffmpeg), capped at 15 seconds. The `duration` parameter is ignored in EDIT mode.
**Seedance 2.0 multi-shot output** is for workflows that need a complete edited sequence from a single prompt: 15-second product spots with multiple cuts, storyboarded explainer videos, brand vignettes with B-roll plus hero shot plus transition built in, lip-synced dialogue scenes that flow across shot boundaries. You send `prompt` (+ optionally `image_urls` / `image_with_roles` / `video_urls` / `audio_urls` for reference modalities), and Seedance produces a multi-cut clip up to 15 seconds with natural shot transitions inside the single output\[2]. No stitching pass needed.
If your pipeline starts from raw video, pick Happyhorse. If it starts from a prompt and needs an edited result, pick Seedance.
## Audio
Both generate native synchronized audio. Details differ.
**Seedance 2.0** outputs joint audio-video in a single forward pass with phoneme-level lip-sync across 8+ languages\[2]. The model accepts up to 3 reference audio clips (`audio_urls`) that must accompany visual references; reference audio gives the model a soundtrack to align with rather than synthesize freely. `generate_audio: true` triggers fresh joint synthesis; `generate_audio: false` outputs silent video.
**Happyhorse 1.0** generates native synchronized audio in 7 languages\[1]. The audio knob is `audio_setting`, but it only applies in EDIT mode: `"auto"` generates fresh audio, `"origin"` keeps the source video's original soundtrack. In T2V / I2V / R2V modes, audio is generated automatically without a toggle.
For non-English lip-synced dialogue, Seedance's 8+ languages and phoneme-level alignment is a real edge. For preserving a source audio track during video restyling, Happyhorse's `audio_setting: "origin"` is the only path between the two.
## Price math
Cheapest 1080p 5-second clip across providers, May 2026:
| Provider | Model | Tier / mode | 5s 1080p |
| -------- | -------------- | -------------------- | --------------------------------------------- |
| reAPI | Happyhorse 1.0 | any mode | **$1.44**\[4] |
| fal.ai | Happyhorse 1.0 | any mode | $1.40\[10] |
| reAPI | Seedance 2.0 | standard, ref mode | **$1.23**\[5] |
| reAPI | Seedance 2.0 | standard, text mode | $2.02\[5] |
| fal.ai | Seedance 2.0 | standard, with audio | $2.00\[9] |
| fal.ai | Seedance 2.0 | fast, no audio | $0.75\[9] |
At 720P 5 seconds:
| Provider | Model | Tier / mode | 5s 720P |
| -------- | ----------------- | ----------- | --------------------------------------------- |
| reAPI | Happyhorse 1.0 | any mode | **$0.81**\[4] |
| fal.ai | Happyhorse 1.0 | any mode | $0.70\[10] |
| reAPI | Seedance 2.0 fast | text mode | $0.72\[5] |
| reAPI | Seedance 2.0 fast | ref mode | **$0.43**\[5] |
Two pricing observations worth knowing.
**Happyhorse 1.0 has flat per-resolution pricing**: $0.1625/s at 720P and $0.2875/s at 1080P, regardless of mode (T2V/I2V/R2V/EDIT all cost the same). reAPI passes the price through with 0% markup\[4]; what you see on the model page is what you pay.
**Seedance 2.0 has tier × mode × resolution pricing**: text mode costs more than reference mode at every cell, so a workflow that always feeds at least one reference image gets the cheaper rate automatically\[5]. The Fast variant at 720P in reference mode is the cheapest verifiable Seedance 2.0 cell.
For predictable per-second budgeting at 1080P, Happyhorse is simpler. For workflows that exploit Seedance's reference-mode discount, Seedance can come out cheaper.
## Picking Seedance 2.0 vs Happyhorse 1.0 in practice
**Pick Happyhorse 1.0 when:**
* Your pipeline transforms existing video (EDIT mode is unique to Happyhorse)
* You want one of the 50+ baked-in style presets
* You need predictable per-second budgeting (one rate per resolution)
* The output gets evaluated against current leaderboard standings
* A single continuous shot is what your downstream editor wants
**Pick Seedance 2.0 when:**
* A single 15-second multi-shot clip replaces what would otherwise be 4 stitched outputs
* Your dialogue needs phoneme-level lip-sync in non-English languages
* Your pipeline feeds the model multiple reference modalities (images + reference video + audio bed)
* 21:9 cinematic ultrawide or `adaptive` (input-matching) aspect ratios are required
* Real-person uploads are part of the workflow (use the `-face` variants)
Most production pipelines that ship at scale will run both. Seedance 2.0 for the prompt-to-finished-clip path, Happyhorse 1.0 for the source-video-to-restyled-clip path. They're complements more than substitutes.
## FAQ
### Is Happyhorse 1.0 free to use?
Not at the API level. Happyhorse 1.0 is paid-tier on every provider that exposes it (fal.ai, reAPI, Alibaba Cloud Bailian for enterprise). Alibaba's stealth launch on Artificial Analysis was free during the benchmark window, but the public API has been paid since the Bailian commercial launch in May 2026.
### Can Happyhorse 1.0 do multi-shot output like Seedance 2.0?
No. Happyhorse 1.0 generates single continuous shots. To produce multi-cut sequences, generate clips separately and edit them in post, or stay with Seedance 2.0 where multi-shot is built into a single generation call\[2].
### Does Seedance 2.0 have a video editing mode like Happyhorse?
Not directly. Seedance 2.0's REF mode accepts a reference video, but it generates new content using the video as a style reference rather than rewriting the source clip while preserving motion. Happyhorse 1.0's EDIT mode is the closest thing between the two to a true video-to-video transformation.
### Why did Happyhorse top the leaderboard so quickly?
The Artificial Analysis Video Arena ranks models by blind human preference. Happyhorse appeared with no public identity and won enough head-to-head comparisons against Seedance 2.0, Veo 3.1, and Sora to climb to #1 in three days\[7]. The leaderboard captures aggregate visual quality on text and image conditioning, not multi-shot, lip-sync, or workflow features.
### What languages does each model support for audio?
Seedance 2.0: 8+ languages with phoneme-level lip-sync\[2]. Happyhorse 1.0: 7 languages with synchronized audio\[1]. Specific language lists aren't fully published by either vendor; assume the major Asian and European languages are covered, with quality variance expected.
### Is Happyhorse 1.0 open source?
Not yet. Alibaba has stated open-source weight release as a future intent, but as of late April 2026, no weights have been published\[9].
### Can I switch between them with one code change?
On reAPI, yes. Both run on `POST /api/v1/videos/generations`. Switching from Seedance 2.0 to Happyhorse 1.0 means changing `"model": "doubao-seedance-2.0"` to `"model": "happyhorse-1.0"` and adapting reference fields (`image_urls` carries across; Seedance's `image_with_roles` becomes Happyhorse's `first_frame_image`; Seedance's `video_urls` style reference becomes Happyhorse's `video_url` for EDIT mode; `size` and `resolution` work the same way).
### What about regional availability?
Happyhorse 1.0 is available globally via fal.ai (April 27 onward) and via Alibaba Cloud Bailian for enterprise customers\[1]. Seedance 2.0 is widely available through fal.ai and ByteDance's consumer surfaces (Dreamina, CapCut), with some restrictions in specific markets reported by industry observers\[9]. reAPI exposes both globally on the same endpoint.
## So which video model wins for your workflow
The short answer for Seedance 2.0 vs Happyhorse 1.0 in 2026: it depends on whether your input is video or text.
If your pipeline starts from existing footage and needs to transform it, restyle, animate-from-still, swap visual identity while preserving motion, Happyhorse 1.0's EDIT mode is the only choice between the two and it currently leads the public quality leaderboard. If your pipeline starts from a prompt and needs an edited multi-cut output in a single API call, Seedance 2.0's multi-shot single-generation is the better fit.
The Seedance 2.0 vs Happyhorse 1.0 decision isn't really about which model wins. Both are competent at the basics, both currently sit in the top tier of available options, and both will likely live behind one OpenAI-compatible endpoint in any production pipeline that ships at meaningful scale. Pick the model whose unique capability matches the unique constraint in your workflow. The rest is leaderboard noise.
## References
1. Apiyi.com. *HappyHorse API is now live on Alibaba Cloud Bailian.* Retrieved May 2026 from [help.apiyi.com/en/happyhorse-api-bailian-launch-apiyi-en.html](https://help.apiyi.com/en/happyhorse-api-bailian-launch-apiyi-en.html)
2. ByteDance Seed. *Official Launch of Seedance 2.0.* February 12, 2026. [seed.bytedance.com/en/blog/official-launch-of-seedance-2-0](https://seed.bytedance.com/en/blog/official-launch-of-seedance-2-0)
3. Artificial Analysis. *Happyhorse — Quality, Generation Time & Price Analysis.* Retrieved May 2026 from [artificialanalysis.ai/video/model-families/happyhorse](https://artificialanalysis.ai/video/model-families/happyhorse)
4. reAPI. *Happyhorse 1.0 — Model page (live pricing).* Retrieved May 2026 from [reapi.ai/models/happyhorse-1-0](/models/happyhorse-1-0)
5. reAPI. *Seedance 2.0 — Model page (live pricing).* Retrieved May 2026 from [reapi.ai/models/seedance-2-0](/models/seedance-2-0)
6. Wikipedia contributors. *Seedance 2.0.* Retrieved May 2026 from [en.wikipedia.org/wiki/Seedance\_2.0](https://en.wikipedia.org/wiki/Seedance_2.0)
7. CNBC. *Alibaba revealed as creator of AI video generation model 'HappyHorse-1.0'.* April 10, 2026. [cnbc.com/2026/04/10/alibaba-happyhorse-ai-video-model-benchmark-reveal.html](https://www.cnbc.com/2026/04/10/alibaba-happyhorse-ai-video-model-benchmark-reveal.html)
8. Bloomberg. *Video AI Model Developed by Alibaba Tops Global Ranking on Debut.* April 10, 2026. [bloomberg.com/news/articles/2026-04-10/stealth-alibaba-video-ai-model-tops-global-ranking-on-debut](https://www.bloomberg.com/news/articles/2026-04-10/stealth-alibaba-video-ai-model-tops-global-ranking-on-debut)
9. BuildFastWithAI. *Happy Horse vs Seedance 2.0: Which AI Video Model Wins? (2026).* Retrieved May 2026 from [buildfastwithai.com/blogs/happy-horse-vs-seedance-2-0-2026](https://www.buildfastwithai.com/blogs/happy-horse-vs-seedance-2-0-2026)
10. fal.ai. *HappyHorse-1.0 — Official API Partner.* Retrieved May 2026 from fal.ai/happyhorse-1.0
### Further reading
* ByteDance Seed. *Seedance 2.0 product page.* [seed.bytedance.com/en/seedance2\_0](https://seed.bytedance.com/en/seedance2_0)
* Alibaba Cloud. *Compare and Select Video Generation Models.* [alibabacloud.com/help/en/model-studio/use-video-generation](https://www.alibabacloud.com/help/en/model-studio/use-video-generation)
* reAPI. *Happyhorse 1.0 — API reference.* [reapi.ai/docs/happyhorse-1-0](/docs/happyhorse-1-0)
* reAPI. *Seedance 2.0 — API reference.* [reapi.ai/docs/seedance-2-0](/docs/seedance-2-0)
* reAPI. *Veo 3.1 vs Seedance 2.0: Picking a Video Model in 2026.* [reapi.ai/blog/veo-3-1-vs-seedance-2-0-2026](/blog/veo-3-1-vs-seedance-2-0-2026)
---
# Seedance 2.0 vs Kling 3.0: Benchmarks, Prices, Verdict (https://reapi.ai/blog/seedance-2-0-vs-kling-3-0-2026)
Seedance 2.0 vs Kling 3.0 is the closest thing AI video has to a heavyweight title fight in 2026: ByteDance's blind-test champion against Kuaishou's feature-packed challenger with native 4K and the deepest audio stack in the category. I pulled the specs from both vendors' own documentation, the Elo numbers from the two arenas that matter, and the per-second prices from every route you can actually buy, all verified on July 3, 2026.
The short version holds few surprises for anyone watching the leaderboards, but the details decide real projects: the quality gap is bigger than most people assume, and so is the feature gap running the other way.
## TL;DR
* **Blind tests favor Seedance 2.0, consistently and by a lot**: #1 on Artificial Analysis text-to-video at Elo 1,222 versus Kling 3.0's best variant at #5 with 1,106\[1], and #1 on LMArena image-to-video at 1,474 versus 1,360\[2].
* **Kling 3.0 wins the spec sheet**: native 4K generation\[3], dialogue in five languages with multi-character voice assignment\[4], first-and-last-frame control, and up to 7 reference images on its Omni variant\[4].
* **Prices are closer than reputations suggest**: official API rates run $0.084 to $0.42/s for Kling 3.0\[5]; Seedance 2.0 spans $0.0703 to $0.7776/s on official token billing\[6].
* **On reAPI both undercut their list rates**: Kling 3.0 from $0.077/s and its 4K at $0.3685/s; Seedance 2.0 from $0.0400/s with references, Mini from $0.03/s\[7]\[8].
* **Watch the reseller spread on Kling**: fal.ai matches official pricing exactly, while Replicate charges up to double for the same tiers\[9].
* Both models run on the same reAPI endpoint, so the honest answer to "which one" is a $2 A/B test on your own prompts.
## Head-to-head on paper
| | Seedance 2.0 | Kling 3.0 |
| ---------------- | -------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| Maker, launch | ByteDance, Feb 12, 2026\[10] | Kuaishou, Feb 5, 2026\[11] |
| Clip length | 4–15s\[6] | 3–15s\[4] |
| Resolutions | 480p–4K (4K in 10-bit H.265)\[6] | 720p, 1080p, native 4K since April 2026\[3] |
| Reference inputs | 9 images + 3 videos + 3 audio in one request\[6] | Up to 7 images (Omni), 4-image element fusion, video input editing\[4] |
| Frame control | First-frame based modes | First AND last frame\[4] |
| Audio | Generated jointly with video\[6] | Native dialogue + SFX, 5 languages, multi-character voice assignment, voice cloning control\[4] |
| Multi-shot | Prompt-driven | Dedicated auto/custom Multi-Shot sequencing\[4] |
Read that table one way: Kling 3.0 is built like a production tool, Seedance 2.0 like a reference-fusion engine. Kling gives you storyboard-grade controls, endpoints for dialogue-driven scenes, and the only native 4K pipeline of the two consumer-facing lineups. Seedance takes more of your source material in a single request, including audio and video references, which is why character-consistent series work gravitates to it.
## What the blind tests say
Elo from paired blind votes is the least gameable public signal we have, and it is one-sided here.
On Artificial Analysis, Seedance 2.0 (720p) holds #1 in text-to-video at Elo 1,222; Kling 3.0's best entry, the 1080p Pro tier, sits fifth at 1,106, a 115-point gap, with three more Kling variants clustered at 1,094 to 1,099\[1]. Image-to-video is wider: 1,195 against 1,072, a 123-point spread\[1]. On LMArena's image-to-video board, built from over 115,000 votes, Seedance 2.0 ranks #1 at 1,474 while kling-v3-pro lands thirteenth at 1,360\[2]. Kling 3.0 has no entry at all on LMArena's text-to-video board as of July 3.
A 115-point Elo gap converts to the higher-rated model winning roughly two out of three blind matchups. That is not a rounding error; it is a visible quality difference in photorealism and motion coherence, and it is why Seedance 2.0 stays the default recommendation for social-length realism. Where the votes do not reach: none of these boards measures 4K output, dialogue quality across languages, or multi-shot narrative control, which is precisely the territory Kling 3.0 stakes out.
## Where Kling 3.0 actually wins
**Native 4K.** Kling shipped what it calls the first native 4K video model in April 2026, rendering at 3840×2160 rather than upscaling\[3]. Seedance 2.0 also offers a 4K tier through its official API\[6], but on most platforms its lineup is served up to 1080p, so if your deliverable is contractually 4K, Kling is frequently the practical route.
**The audio stack.** Kling 3.0 generates dialogue and effects natively, assigns distinct voices across three or more characters, supports five languages plus dialect accents, and can lock a specific voice via its control feature\[4]. Seedance's joint audio generation is strong for ambience and sync, but multi-character scripted dialogue is Kling's home turf.
**Directorial control.** First-and-last-frame interpolation, element references that fuse up to four images into one subject, and Multi-Shot sequencing with automatic or custom shot breakdowns\[4] make Kling the more scriptable camera. Seedance answers with raw reference bandwidth: nine images plus video plus audio in one request\[6].
If your work is "make this exact character do exact things across a storyboard with spoken lines," the spec sheet argues Kling. If it is "make the most convincing 15 seconds possible from this pile of references," the votes argue Seedance.
## Price per second, every real route
Official list rates first. Kling 3.0's API bills flat per second: $0.084/s at 720p and $0.112/s at 1080p without audio, $0.126 and $0.168 with audio, $0.42/s for 4K\[5]. Seedance 2.0's official billing is token-based and converts to $0.0703 (480p), $0.1512 (720p), $0.3742 (1080p), and $0.7776/s (4K)\[6].
Two reseller notes from checking the usual suspects: fal.ai lists Kling 3.0 at exactly the official rates, while Replicate charges up to double ($0.168/s for standard without audio, $0.336/s for pro with audio)\[9]. Same model, 2x spread; always check the per-second number, not the brand.
On reAPI, both families run below list, on one endpoint, as of July 2026:
| Tier | reAPI rate |
| --------------------- | ------------------------------------------------------------------------------------------ |
| Kling 3.0 Standard | $0.077/s (no audio) · $0.11/s (audio)\[8] |
| Kling 3.0 Pro | $0.099/s · $0.1485/s\[8] |
| Kling 3.0 4K | $0.3685/s\[8] |
| Seedance 2.0 Standard | $0.0834–$0.4048/s by resolution, references discounted\[7] |
| Seedance 2.0 Fast | from $0.0400/s with references\[7] |
| Seedance 2.0 Mini | from $0.03/s\[7] |
The cost verdict depends on the job. Dialogue-free 720p volume: Seedance Fast and Mini own the floor. 1080p with audio: Kling Pro at $0.1485/s versus Seedance Standard-with-references at $0.2456/s makes Kling the cheaper premium tier. 4K: Kling at $0.3685/s is roughly half Seedance's official 4K rate.
## Which one for which job
**Choose Seedance 2.0** for photoreal social content, character consistency from reference stacks, image-to-video where fidelity to the source matters, and any volume pipeline where Fast and Mini tiers cut costs 40 to 70 percent. The blind-test margin is real and your audience votes the same way the arenas do.
**Choose Kling 3.0** for 4K deliverables, dialogue scenes in multiple languages, storyboard-driven multi-shot narratives, and first-to-last-frame precision work. The features are not gimmicks; they are the parts of production the leaderboards do not score.
**Or stop choosing.** Both run on [reAPI's videos endpoint](/models/seedance-2-0) with per-second billing; switching is the `model` string. Run your three toughest prompts through [Seedance 2.0](/models/seedance-2-0) and [Kling 3.0](/models/kling-3-0), watch them side by side, and let the outputs settle it for a couple of dollars.
## FAQ
### Is Seedance 2.0 better than Kling 3.0?
On measured quality, yes: it leads Kling 3.0 by 115 to 123 Elo across the Artificial Analysis and LMArena boards\[1]\[2]. On production features (4K, dialogue, shot control), Kling 3.0 is ahead. "Better" depends on which of those your project bills for.
### Which is cheaper, Seedance 2.0 or Kling 3.0?
At the floor, Seedance: its Fast and Mini tiers run $0.03 to $0.0672/s on reAPI, below anything in Kling's lineup\[7]. At the premium end it flips: Kling's 4K costs $0.3685/s versus Seedance's official 4K at $0.7776/s\[6]\[8].
### Does Kling 3.0 have better audio than Seedance 2.0?
For scripted dialogue, yes: five languages, multi-character voice assignment, and voice locking are documented Kling 3.0 features\[4]. Seedance generates audio jointly with video and takes audio as a reference input, which suits ambience and music-driven clips\[6].
### Is Kling 3.0's 4K really native?
Kuaishou markets it as the first native 4K generation model, rendering at 3840×2160 rather than upscaling\[3]. Independent leaderboards do not yet score 4K output, so treat "native" as documented and "better at 4K" as your own eyeball test.
### Why do Kling 3.0 prices differ so much between platforms?
Reseller margin. fal.ai matches Kling's official per-second rates; Replicate lists the same tiers at up to 2x\[9]; reAPI runs below list\[8]. The model is identical, so the per-second number is the whole comparison.
### Can I A/B test both models on one API?
Yes. Both are live on reAPI's async videos endpoint with per-second billing and signup credits; the only change between requests is the `model` field\[7]\[8].
## The verdict that survives contact with projects
Treat Seedance 2.0 vs Kling 3.0 as a division of labor, not a duel. The arena numbers say Seedance for anything where perceived realism decides, and the spec sheets say Kling for 4K, dialogue, and shot-by-shot control. Prices overlap enough that neither wins on cost alone, and both sit one model-string apart on the same endpoint. Run the A/B on your own material; in our experience the winner of Seedance 2.0 vs Kling 3.0 changes by use case far more often than by taste.
## References
1. Artificial Analysis. *Text to Video and Image to Video Leaderboards.* Retrieved July 2026 from [artificialanalysis.ai/video/leaderboard/text-to-video](https://artificialanalysis.ai/video/leaderboard/text-to-video)
2. LMArena. *Image to Video Leaderboard.* Retrieved July 2026 from [arena.ai/leaderboard/image-to-video](https://arena.ai/leaderboard/image-to-video)
3. Kling AI (Kuaishou). *Kling AI introduces native 4K video model.* Retrieved July 2026 from [kling.ai/blog/kling-ai-introduces-native-4k-video-model](https://kling.ai/blog/kling-ai-introduces-native-4k-video-model)
4. Kling AI (Kuaishou). *Kling VIDEO 3.0 model user guide.* Retrieved July 2026 from [kling.ai/quickstart/klingai-video-3-model-user-guide](https://kling.ai/quickstart/klingai-video-3-model-user-guide)
5. Kling AI (Kuaishou). *Developer API pricing.* Retrieved July 2026 from [kling.ai/dev/pricing](https://kling.ai/dev/pricing)
6. BytePlus (ByteDance). *ModelArk — Seedance 2.0 specifications and pricing.* Retrieved July 2026 from [docs.byteplus.com/en/docs/ModelArk/1544106](https://docs.byteplus.com/en/docs/ModelArk/1544106)
7. reAPI. *Seedance 2.0 — model page and live pricing.* Retrieved July 2026 from [reapi.ai/models/seedance-2-0](/models/seedance-2-0)
8. reAPI. *Kling 3.0 — model page and live pricing.* Retrieved July 2026 from [reapi.ai/models/kling-3-0](/models/kling-3-0)
9. Replicate. *kwaivgi/kling-v3-video — pricing.* Retrieved July 2026 from replicate.com/kwaivgi/kling-v3-video
10. ByteDance Seed. *Seedance 2.0 — official model page.* Retrieved July 2026 from [seed.bytedance.com/en/seedance2\_0](https://seed.bytedance.com/en/seedance2_0)
11. Kuaishou Investor Relations. *Kling AI launches 3.0 model.* Retrieved July 2026 from [ir.kuaishou.com/news-releases](https://ir.kuaishou.com/news-releases/news-release-details/kling-ai-launches-30-model-ushering-era-where-everyone-can-be)
### Further reading
* reAPI. *Veo 3.1 vs Seedance 2.0.* [reapi.ai/blog/veo-3-1-vs-seedance-2-0-2026](/blog/veo-3-1-vs-seedance-2-0-2026)
* reAPI. *Cheapest Seedance 2.0 in 2026: Real Prices, Compared.* [reapi.ai/blog/cheapest-seedance-2-0-2026](/blog/cheapest-seedance-2-0-2026)
* reAPI. *Kling 3.0 API documentation.* [reapi.ai/docs/kling-3-0](/docs/kling-3-0)
---
# Seedance 2.1 and Seedance 2.0 Mini: What's Actually Coming (https://reapi.ai/blog/seedance-2-1-and-seedance-2-0-mini-preview)
Two new ByteDance video models are reportedly in the pipeline. Seedance 2.1 is said to carry a 20% generation-quality improvement over Seedance 2.0, per a Pandaily report citing unnamed sources\[1]. Seedance 2.0 Mini is a lighter tier that, per an early report from inside the API-platform world, will undercut the current Fast variant on price while beating it on output. There's a wrinkle most coverage skipped: the same day the Seedance 2.1 report ran, ByteDance told Chinese financial media the rumor was untrue\[2].
So this is a pre-announcement story with a denial attached. Worth covering anyway. The sourcing pattern looks a lot like the runup to Seedance 2.0's own launch, and if either model ships, the pricing math changes for anyone running video generation at volume. Here's what's sourced, what's grapevine, what the numbers look like against today's live Seedance 2.0 rates, and how to be ready on day one.
## TL;DR
* **The Seedance 2.1 report.** Pandaily (May 20, 2026): Seedance 2.1 is in preparation with "a reported 20% improvement in generation quality," attributed mainly to temporal consistency and physics simulation\[1].
* **The denial.** Same day, AASTOCKS reported ByteDance "clarified that the relevant rumors are untrue"\[2].
* **The Mini report.** Seedance 2.0 Mini surfaced through a platform that serves the Seedance family, citing its own internal channels: a lighter variant priced well below Seedance 2.0 Fast, performing above it. Well-placed, not yet official.
* **Where 2.0 stands.** Seedance 2.0 currently ranks #1 on Artificial Analysis's text-to-video (with audio) leaderboard at Elo 1,215 and #1 on image-to-video at Elo 1,194\[3]\[4].
* **Today's price floor.** On reAPI, Seedance 2.0 runs $0.0400–$0.4048 per second depending on variant, resolution, and reference mode\[5]. Mini's claimed positioning would push under that floor.
* Nothing here is officially announced. Treat every date and number below the leaderboard scores as soft until ByteDance publishes a model card.
## The report, the denial, and what's still standing
The Pandaily piece is specific on the claim and vague on the source. The exact line: Seedance 2.1 comes "with a reported 20% improvement in generation quality over the current 2.0 version, according to sources familiar with the matter"\[1]. The improvement is attributed "primarily" to "advances in temporal consistency — the model's ability to maintain visual coherence across frames — and improved physics simulation for generated scenes"\[1].
ByteDance's response came fast. AASTOCKS, the Hong Kong financial wire, reported the same day that "ByteDance clarified that the relevant rumors are untrue"\[2]. No Tier-1 outlet (Reuters, Bloomberg, The Information) has touched Seedance 2.1 at all.
I'd read the denial narrowly. "The rumors are untrue" most plausibly targets the *imminent release* framing rather than the model's existence; companies rarely deny that a successor to their flagship is in training, because one always is. Seedance 2.0 itself followed this arc — leaked details, official silence, then a launch on February 12, 2026\[6]. But a narrow reading is still a reading. The honest version: a Seedance 2.1 is presumably in development, the 20% number is unverified, and the timeline is officially disputed.
Seedance 2.0 Mini sits in a different sourcing category. The claim surfaced through an API platform that serves the Seedance family, attributed to its own internal channels: a second variant priced well below Seedance 2.0 Fast while performing above it. That's a vantage point worth taking seriously, because platforms hosting a model family tend to hear about new variants before the public does. It still isn't official: no mention exists on ByteDance Seed's site or in Volcano Engine's docs. A Medium post has since circulated a "$0.073 per second" figure for Mini; it cites no source, so I'm not repeating it as fact.
## Where Seedance pricing sits in June 2026
To see what Mini would actually disrupt, here are the live per-second rates for the Seedance 2.0 family on reAPI as of June 2026 (text mode · reference mode)\[5]:
| Variant | 480p | 720p | 1080p |
| ----------------- | ----------------- | ----------------- | ----------------- |
| Seedance 2.0 | $0.0834 · $0.0506 | $0.1796 · $0.1086 | $0.4048 · $0.2456 |
| Seedance 2.0 Fast | $0.0672 · $0.0400 | $0.1444 · $0.0865 | — |
Reference mode (any image, video, or audio reference attached) bills lower than pure text-to-video at every cell. Face-aware variants for real-person inputs sit roughly 30–40% above their base counterparts. The current floor is Fast at 480p with a reference: $0.0400 per second, or about $0.20 for a 5-second clip.
Mini's claimed slot is well below Fast. If that holds upstream, the floor drops below two cents per second at the cheap end once it propagates through to API platforms. That would be the cheapest rate ever attached to the Seedance name, for a model claimed to outperform the tier above it.
That combination is what makes the rumor worth tracking despite the sourcing. The standard release pattern is a new flagship at a higher price while last year's model becomes the budget tier. Mini inverts it: the budget tier would be the *new* model. Pareto improvements at the bottom of a price ladder are rare, which is exactly why the claim deserves skepticism until there's a price sheet.
## What a 20% bump would do to Seedance 2.1's leaderboard math
The 20% claim has no stated metric, and that matters less than it seems, because Seedance 2.0's current position makes almost any version of it newsworthy. As of June 11, 2026, Dreamina Seedance 2.0 720p holds #1 on Artificial Analysis's text-to-video leaderboard (with audio) at Elo 1,215, with HappyHorse-1.0 second at 1,122 and SkyReels V4 third at 1,106\[3]. It also holds #1 on image-to-video at Elo 1,194\[4]. The one chart it doesn't lead is no-audio text-to-video, where HappyHorse-1.0 sits ahead, 1,293 to 1,274\[3].
So Seedance 2.1 wouldn't be a comeback story; it would be a lead-extension story. If the temporal-consistency and physics framing from the Pandaily report is accurate\[1], the lift would land precisely on the axes where blind-test voters punish video models hardest: objects that morph between frames, motion that ignores momentum, hands.
The commercial stakes are lopsided too. The same AASTOCKS piece that carried the denial notes the Seedance series "has reportedly captured 80% of the AI video generation market share, while Kling accounted for 14%"\[2]. Take the precision of that figure with salt, but the direction is clear: ByteDance is defending a lead, not chasing one. Incremental quality releases are how leads get defended.
For builders, my advice is the opposite of exciting: if Seedance 2.0 already clears your quality bar, a 20% bump on Seedance 2.1 doesn't change your stack. Migration costs are real: the QA set gets re-run, prompts get re-validated, content-policy behavior gets re-checked. Upgrade when your failure modes are quality-shaped, not because a leaderboard number moved.
## Why Seedance 2.0 Mini is the half worth watching
Mini changes economics rather than ceilings, and for most production workloads economics bind first. Three places a sub-Fast price tier with Fast-or-better quality would matter immediately:
**Prompt iteration.** Prompt tuning for video is brutally expensive compared to images because every test render costs real money. At Fast's $0.0672/s text rate, a hundred 5-second 480p test clips run about $34. Cut the per-second rate meaningfully below that and the iteration loop tightens for the same budget.
**Short-form volume.** For social clips, drafts, and B-roll, Fast-level output already passes. The binding constraint is unit cost per clip, and that's the exact dial Mini reportedly turns.
**Eval pipelines.** If you maintain a reference set of generated clips to score new models against (and after this year's release pace, you should), regenerating it on every model update is a recurring bill. Cheaper reference generation compounds.
The catch stays the same: well-placed, but not official. If Mini never ships, Fast at $0.0400/s in reference mode remains the family's value play. Nothing about that math requires waiting.
## A launch-day playbook for Seedance 2.1 and Mini
Three things to do before touching production routing, whenever either model goes live:
1. **Re-run your own eval set, not the vendor's.** Same prompts that drove your last Seedance 2.0 evaluation, same seeds where supported. Marketing benchmarks measure what the marketer chose to measure.
2. **Test 2.0's known weak spots first.** Text rendered inside video frames, identity drift across multi-shot cuts, physics-heavy motion. The Pandaily report says temporal consistency and physics are where Seedance 2.1's 20% lives\[1]; that's a falsifiable claim, so falsify it.
3. **Benchmark Mini against Fast, not Standard.** "Better than Fast, cheaper than Fast" is the actual claim on the table. If Mini merely ties Fast at a lower price, switching already pays; if it beats Fast, the migration decides itself.
## Running Seedance 2.0 on reAPI while you wait
The current lineup is live today: `doubao-seedance-2.0`, `doubao-seedance-2.0-fast`, plus face-aware versions of both for real-person inputs. One endpoint (`POST /api/v1/videos/generations`), per-second billing computed at submit time, automatic refund on failure\[5]. Multi-shot sequences, up to 9 image + 3 video + 3 audio references per request, 4–15 second outputs at 480p/720p/1080p.
When new family members ship, they land on the same request schema, so an upgrade is a one-field model swap. That's also the cheapest way to act on this whole story: build your eval harness against Seedance 2.0 now, and pointing it at Seedance 2.1 or Mini later becomes an afternoon, not a sprint. For how 2.0 stacks against the competition in the meantime, we've published head-to-heads against [Veo 3.1](/blog/veo-3-1-vs-seedance-2-0-2026), [HappyHorse 1.0](/blog/seedance-2-0-vs-happyhorse-1-0-2026), and [Gemini Omni](/blog/gemini-omni-vs-seedance-2-0-2026).
## FAQ
### Is Seedance 2.1 officially announced?
No. The only Seedance 2.1 report is Pandaily's May 20, 2026 piece citing anonymous sources\[1], and ByteDance denied the rumors the same day through AASTOCKS\[2]. There is no model card, system card, or official blog post.
### What is Seedance 2.1's 20% improvement supposed to cover?
Per the report: generation quality overall, driven "primarily" by temporal consistency (frame-to-frame coherence) and physics simulation\[1]. No benchmark or metric was named, which is why launch-day testing on your own prompts matters more than the headline number.
### What is Seedance 2.0 Mini?
A reported lighter variant priced below Seedance 2.0 Fast while outperforming it. The claim comes from a platform-side report citing internal channels; no official documentation mentions it yet.
### When did Seedance 2.0 actually launch?
ByteDance Seed officially launched it on February 12, 2026\[6]. CapCut began rolling it out globally through Dreamina on March 26, and Volcano Engine opened API access (globally via BytePlus) on April 15, 2026\[7]. Some posts compress this into a single "April launch," which understates how long the model has been in the wild.
### How good is Seedance 2.0 right now?
It leads two of the three Artificial Analysis video leaderboards as of June 2026: #1 in text-to-video with audio (Elo 1,215) and #1 in image-to-video (Elo 1,194), with a #2 slot behind HappyHorse-1.0 in no-audio text-to-video\[3]\[4].
### What does Seedance 2.0 cost on reAPI today?
From $0.0400/s (Fast, 480p, reference mode) up to $0.4048/s (Standard, 1080p, text mode). The [live pricing table](/models/seedance-2-0#pricing) always reflects current rates\[5].
### Will Seedance 2.1 cost more than Seedance 2.0?
Nobody has published Seedance 2.1 pricing, including the leak. Flagship refreshes usually land at or above the outgoing Standard rate, so that's the safe assumption, but it is an assumption, not a report.
### Should I wait for Mini before building?
No. Mini is unofficial and undated. Build on Seedance 2.0 Fast now ($0.0400–$0.1444/s), keep your prompts and eval set portable, and switching later is a model-string change.
## How to position before anything ships
The asymmetry is the whole reason to care. If the reports are wrong, you've lost nothing by having a portable eval set and a price-aware routing setup. If they're right, you can move the day Seedance 2.1 or Seedance 2.0 Mini appears, while everyone else is still re-reading the launch post. The leaderboard lead is verified\[3], today's prices are live, and the rest is two reports and a denial. Run [Seedance 2.0 on reAPI](/models/seedance-2-0) now, keep the harness warm, and let Seedance 2.1 prove the 20% on your own prompts when it arrives.
## References
1. Pandaily. *ByteDance to Launch Seedance 2.1 Video Generation Model with 20% Quality Boost.* Retrieved June 2026 from [pandaily.com/bytedance-seedance-2-1-video-generation-2026](https://pandaily.com/bytedance-seedance-2-1-video-generation-2026)
2. AASTOCKS. *ByteDance Negates Imminent Release of AI Video Generation Model Seedance 2.1.* Retrieved June 2026 from [aastocks.com/en/mobile/news.aspx?newsid=NOW.1525415](https://www.aastocks.com/en/mobile/news.aspx?newsid=NOW.1525415\&newstype=61\&newssource=AAFN)
3. Artificial Analysis. *Text to Video Leaderboard.* Retrieved June 2026 from [artificialanalysis.ai/video/leaderboard/text-to-video](https://artificialanalysis.ai/video/leaderboard/text-to-video)
4. Artificial Analysis. *Image to Video Leaderboard.* Retrieved June 2026 from [artificialanalysis.ai/video/leaderboard/image-to-video](https://artificialanalysis.ai/video/leaderboard/image-to-video)
5. reAPI. *Seedance 2.0 — model page and live pricing.* Retrieved June 2026 from [reapi.ai/models/seedance-2-0](/models/seedance-2-0)
6. ByteDance Seed. *Seedance 2.0 Official Launch.* Retrieved June 2026 from [seed.bytedance.com/en/blog/official-launch-of-seedance-2-0](https://seed.bytedance.com/en/blog/official-launch-of-seedance-2-0)
7. Pandaily. *ByteDance Opens Seedance 2.0 Video Generation API.* Retrieved June 2026 from [pandaily.com/byte-dance-opens-seedance-2-0-video-generation-api](https://pandaily.com/byte-dance-opens-seedance-2-0-video-generation-api)
### Further reading
* TechCrunch. *ByteDance reportedly pauses global launch of its Seedance 2.0 video generator.* [techcrunch.com/2026/03/15/bytedance-reportedly-pauses-global-launch](https://techcrunch.com/2026/03/15/bytedance-reportedly-pauses-global-launch-of-its-seedance-2-0-video-generator/)
* CapCut Newsroom. *Dreamina Seedance 2.0 global rollout.* [capcut.com/newsroom/dreamina-seedance-2](https://www.capcut.com/newsroom/dreamina-seedance-2)
---
# Seedance 2.5 API Pricing: fal vs kie vs WaveSpeed vs reAPI (https://reapi.ai/blog/seedance-2-5-api-pricing-compared)
Someone opened a thread on r/Seedance\_AI on 3 August asking two questions at once: which host has Seedance 2.5 uncensored, and which one has it cheapest.\[6] The censorship half got argued about for a day. The price half got one estimate and then nothing. In a separate thread a week earlier, somebody asked flat out how many credits a 720p 30-second clip costs and nobody answered at all.\[7]
Nobody answers because Seedance 2.5 API pricing is not one number per platform. It is four numbers, and two of them are measured on a different clock than the others. Every host publishes a per-second rate that looks comparable, then bills reference-video jobs on a formula that is not. Below are the published rates on fal.ai, kie.ai, WaveSpeed and reAPI, pulled on 9 August 2026 and converted to a unit you can divide.
## TL;DR
* **reAPI is the lowest of these four on every tier.** $0.118589/s at 480p and $0.266824/s at 720p without a reference clip, $0.071153/s and $0.160094/s with one.\[1]
* **fal.ai is the highest on every tier.** $0.4730/s at 720p without video input, 1.77x reAPI's rate for output from the same weights.\[3]
* **All four discount reference-video jobs, then bill you for the source footage too.** kie.ai states it plainly: "No video = Price × Output; With video = Price × (Input + Output)."\[4]
* **A 720p job with a 10-second reference and a 5-second output** costs $2.40 on reAPI, $2.85 on kie.ai, $3.30 on WaveSpeed and $4.26 on fal.ai.\[1]\[3]\[4]\[5]
* **One caveat against reAPI:** a minimum-billing floor makes very short references against long outputs cost more than kie.ai's simpler formula. The threshold is worked out below.
## What the four platforms actually charge
All four sell the same model with the same surface. fal's published schema lists 480p and 720p, durations of 4 to 30 seconds plus auto, up to 30 reference images and 10 reference videos\[3], the same envelope ByteDance described at launch.\[2] So this is a rate comparison, not a capability comparison.
Seedance 2.5 API pricing starts with the simple case, where you send a prompt and maybe some reference images but no source clip:
| Platform | 480p | 720p |
| --------- | ------------ | ------------ |
| reAPI | $0.118589 /s | $0.266824 /s |
| kie.ai | $0.140 /s | $0.315 /s |
| WaveSpeed | $0.18 /s | $0.36 /s |
| fal.ai | $0.2205 /s | $0.4730 /s |
Sources: reAPI's model page\[1], kie.ai's rate of 28 and 63 credits per second at $0.005 per credit\[4], WaveSpeed's published table\[5], and fal's own pricing note\[3].
Spread from cheapest to dearest is 1.77x. For a 30-second 720p clip that is $8.01 against $14.19, for output from the same weights.
## Reference-video jobs bill on a different clock
Here is the part that makes single-number comparisons useless. Seedance 2.5's headline capability is reference-driven generation, and every platform discounts the per-second rate when you supply a source clip. None of them stops charging for the source.
kie.ai states the rule directly: "No video = Price × Output; With video = Price × (Input + Output)."\[4] WaveSpeed says the same thing in its own words, pricing reference jobs on "the combined duration of reference video duration + output duration."\[5] fal's token formula multiplies by "(input video duration + output video duration)."\[3] reAPI bills the same way.
| Platform | 480p reference rate | 720p reference rate | Billed over |
| --------- | ------------------- | ------------------- | -------------- |
| reAPI | $0.071153 /s | $0.160094 /s | input + output |
| kie.ai | $0.085 /s | $0.190 /s | input + output |
| WaveSpeed | $0.11 /s | $0.22 /s | input + output |
| fal.ai | \~$0.1323 /s | \~$0.2838 /s | input + output |
The practical consequence is that a long reference is expensive everywhere. Twenty seconds of source footage producing ten seconds of output bills thirty seconds on all four, not ten. Budget from counted seconds, not from output length.
## The floor that works against us
One rule on reAPI deserves stating because it is the only place our rate stops being the lowest. Billable seconds on a reference job are the greater of input-plus-output and ceil(5 × output ÷ 3), capped at 60. That floor mirrors the minimum-token table ByteDance publishes for the Seedance 2.5 series, so the underlying cost exists on every host reselling the model. reAPI's billing exposes it. The other three published formulas do not mention it either way.\[3]\[4]\[5]
In practice: a 2-second reference against a 30-second output bills as 50 seconds on reAPI, not 32. At that shape kie.ai comes out ahead, $6.08 against $8.01.
Checked across output lengths from 4 to 30 seconds, the threshold is clean. **If your reference footage runs at least half as long as your target output, reAPI is the cheapest of these four.** Below that ratio the floor dominates and kie.ai wins. Above it, reAPI wins at every length I tested.
Most reference work sits comfortably above that line, because the point of supplying twenty seconds of footage is rarely to get four seconds back. But if your workflow is a two-second clip driving a thirty-second generation, price it on kie.ai first.
## Seedance 2.5 API pricing for three real jobs
Rates are abstract. Here are three shapes people actually run, all 720p, priced on each platform's published formula.
**A five-second clip from a prompt, no reference material.**
| Platform | Cost |
| --------- | ------ |
| reAPI | $1.335 |
| kie.ai | $1.575 |
| WaveSpeed | $1.80 |
| fal.ai | $2.365 |
**A ten-second reference clip restyled into a five-second output.** Fifteen counted seconds on all four.
| Platform | Cost |
| --------- | ------ |
| reAPI | $2.401 |
| kie.ai | $2.850 |
| WaveSpeed | $3.30 |
| fal.ai | $4.257 |
**A twenty-second reference cut down to ten seconds.** Thirty counted seconds on all four.
| Platform | Cost |
| --------- | ------ |
| reAPI | $4.803 |
| kie.ai | $5.700 |
| WaveSpeed | $6.60 |
| fal.ai | $8.514 |
The gap between cheapest and dearest holds near 1.8x across all three shapes. On a hundred clips a month at the second shape, that is $240 against $426.
## One caveat on credit packs
kie.ai's pricing note says high-tier top-ups carry a 10% bonus, which drops its effective rate roughly 10% below the listed price.\[4] Applied to 720p without video input that is about $0.284/s against reAPI's $0.267/s, so the ranking holds and the gap narrows to roughly 6%. Anyone buying credits in bulk on kie.ai is paying less than its listed rates suggest, which is worth knowing before treating a list price as final. Neither WaveSpeed nor fal publishes a comparable volume bonus.
## FAQ
### Which platform has the cheapest Seedance 2.5 API pricing?
Of the four compared here, reAPI on every tier: $0.266824/s at 720p without a reference clip and $0.160094/s with one.\[1] The one exception is very short references against long outputs, where kie.ai's simpler formula wins.
### Why is fal.ai nearly double reAPI for the same model?
fal publishes $0.4730/s for 720p against reAPI's $0.266824/s.\[3]\[1] Both resell the same ByteDance weights with the same parameter surface, so the difference is margin and packaging, not capability.
### How many credits is a 720p 30-second Seedance 2.5 clip on reAPI?
8,005 credits, which is $8.005 at 1 credit = $0.001, with no reference video.\[1] At 480p the same clip is 3,558 credits.
### Does turning off the safety checker change the price?
Not on reAPI. `nsfw_checker` selects where a task runs, not a rate tier, so the same request bills the same either way.\[9]
### Why does my reference job cost more than duration times rate?
Because all four platforms count input seconds as billable. A 20-second reference producing 10 seconds bills 30 seconds everywhere.\[4]\[5]\[3]
### Is 480p worth it to save money?
At 720p you pay roughly 2.2x the 480p rate on every platform here. ByteDance's parameter table lists only 480p and 720p for Seedance 2.5,\[9] so on most routes 720p is the ceiling and 480p is a real saving rather than a downgrade from something better.
### Do these prices include audio?
Yes on all four. Seedance 2.5 generates a synced audio track as part of the same job, and none of the four publishes a separate audio charge.\[2]
### What about failed generations?
reAPI refunds the reserve in full when a task fails, including moderation failures.\[1] Refund behaviour on a failed job is worth checking wherever you buy, because at these rates a silent charge on a rejected prompt is a real line item.
## Picking a platform by workload shape
For anything reference-driven, which is the workflow ByteDance built Seedance 2.5 around, reAPI is the cheapest of these four at both resolutions and stays cheapest on every job shape where the reference runs at least half the output length. For plain text-to-video and image-to-video it is also the lowest of the four, by 12% against kie.ai and 44% against fal.ai.
The one shape to price elsewhere is a very short reference driving a long generation, where reAPI's minimum-billing floor hands the win to kie.ai. Everything else lands the same way, and fal.ai does not win any shape measured here. Its published Seedance 2.5 API pricing is the highest on every tier in this comparison, at both resolutions, with and without reference video.
Full parameter reference and runnable examples are at [reapi.ai/docs/seedance-2-5](/docs/seedance-2-5), and the live rate band is on [reapi.ai/models/seedance-2-5](/models/seedance-2-5).
## References
1. reAPI. *Seedance 2.5 — model page, published per-second rate band and refund behaviour.* Retrieved 9 August 2026 from [reapi.ai/models/seedance-2-5](/models/seedance-2-5)
2. ByteDance Seed. *One-take Creation, Flexible Referencing: Introducing Seedance 2.5.* Published 31 July 2026, retrieved August 2026 from [seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5](https://seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5)
3. fal.ai. *Seedance 2.5 Reference to Video — pricing note and queue schema.* Retrieved 9 August 2026 from [fal.ai/models/bytedance/seedance-2.5/reference-to-video](https://fal.ai/models/bytedance/seedance-2.5/reference-to-video)
4. kie.ai. *Seedance 2.5 — credits-per-second pricing and video-input formula.* Retrieved 9 August 2026 from [kie.ai/seedance-2-5](https://kie.ai/seedance-2-5)
5. WaveSpeed. *bytedance/seedance-2.5/text-to-video — pricing tables with and without reference videos.* Retrieved 9 August 2026 from [wavespeed.ai/models/bytedance/seedance-2.5/text-to-video](https://wavespeed.ai/models/bytedance/seedance-2.5/text-to-video)
6. r/Seedance\_AI. *Monday Seedance 2.5 checkup.* Posted 3 August 2026, retrieved August 2026 from [reddit.com/r/Seedance\_AI/comments/1veptk2](https://www.reddit.com/r/Seedance_AI/comments/1veptk2)
7. r/Seedance\_AI. *Which site is the best for Seedance 2.5.* Posted 31 July 2026, retrieved August 2026 from [reddit.com/r/Seedance\_AI/comments/1vbxwcj](https://www.reddit.com/r/Seedance_AI/comments/1vbxwcj)
8. r/Seedance\_AI. *Seedance 2.0 Providers Keep Making Big Claims, Who Actually Proves It?* Posted 16 July 2026, retrieved August 2026 from [reddit.com/r/Seedance\_AI/comments/1uxvyno](https://www.reddit.com/r/Seedance_AI/comments/1uxvyno)
9. reAPI. *doubao-seedance-2.5-face — parameters, resolutions and billing dimensions.* Retrieved August 2026 from [reapi.ai/docs/seedance-2-5](/docs/seedance-2-5)
### Further reading
* reAPI. *Seedance 2.5 Unlimited Plans: The Math That Can't Work.* [reapi.ai/blog/seedance-2-5-unlimited-plans-math](/blog/seedance-2-5-unlimited-plans-math)
* reAPI. *Seedance 2.5 Uncensored: A Provider Test You Can Run.* [reapi.ai/blog/seedance-2-5-uncensored-provider-test](/blog/seedance-2-5-uncensored-provider-test)
---
# Seedance 2.5 Camera Control: Write Numbers, Not Adjectives (https://reapi.ai/blog/seedance-2-5-camera-control)
Generate a thirty-second shot with "smooth cinematic pan" in the prompt and the virtual lens hurls itself across the room like a runaway drone. The subject gets left behind in a blur. This is the single most common frustration with Seedance 2.5 camera control, and it is not a bug.
Adjectives are not velocities. The model maps descriptive words to default motion vectors, and words like *smooth*, *epic*, and *dramatic* map to fast ones. Fixing it means writing the prompt like a script supervisor rather than a mood board.
## TL;DR
* **Replace adjectives with numbers.** "Slow lateral dolly at 2 seconds per meter" beats "fast pan across room."
* **Name the rig.** Dolly, crane, locked tripod, tracking shot. Mechanical vocabulary constrains the motion model in a way mood words do not.
* **Split the shot into timestamped beats** so the engine knows when a move starts, holds, and stops.
* **Anchor space with reference material** rather than describing it, which is what stops coordinate drift over a long take.
* **Longer generations drift more.** The 30-second window makes explicit constraints more necessary, not less.
## Why the default is too fast
The model works in a compressed latent space, and text cues get mapped to velocity vectors before anything is rendered. Generic descriptors carry no physical weight, so the mapping resolves toward sweeping, high-acceleration trajectories.
Three failure modes follow from that, and they compound over a long take.
**Semantic gap.** "Smooth" and "dramatic" are not velocity metrics. Neural weights read them as permission for maximum acceleration rather than instructions for deliberate movement.
**Latent drift.** Without hard constraints, spatial coordinates wander across an extended timeline. A wall that was on the left at second three is somewhere else at second twenty.
**Pacing breakdown.** Fast default translations destroy viewer orientation within the first few seconds, which is the part of the clip that decides whether anyone keeps watching.
## Replace adjectives with numbers
This is the whole technique, and the table below is the working version of it.
| Failure mode | Avoid | Use instead |
| ---------------- | ------------------------------------------- | ----------------------------------------------------------------------------------- |
| Hyper-fast pan | "Smooth, epic camera pan across the room" | "Lateral tracking shot at 0.5 m/s, fixed focal length, 4-second linear progression" |
| Runaway zoom | "Dramatic cinematic zoom into the subject" | "Gradual focal push-in from wide (24mm) to medium (50mm) over a 10-second window" |
| Coordinate drift | "Fast drone shot flying over the landscape" | "Low-altitude stable aerial glide, altitude locked at 3 meters, vector heading 090" |
Three rules generalize from those rows.
**Define numeric velocity.** Substitute metric or time-based pacing for every descriptive word. Meters per second, or seconds per meter, both work. What matters is that a number appears.
**Lock the mechanical rig.** Say dolly, crane, locked tripod, or tracking shot explicitly. Naming real equipment forces the motion model toward physically plausible paths, because those words carry spatial constraints that "cinematic" does not.
**Add time triggers.** Break a long shot into stages so the engine knows when to move and when to hold.
## Timestamped beats

The 30-second window is where prompt structure stops being optional. A single paragraph describing an entire shot gives the model no reason to hold still at any point, so it moves continuously.
Structure it as beats instead:
```text
0-6s: [Locked tripod] Medium shot, subject centered, no camera movement.
Ambient motion only — steam rising, fabric shifting.
6-14s: [Slow dolly in] Push from medium to close at 0.3 m/s, constant speed,
no easing. Focal length fixed at 50mm.
14-22s: [Hold] Camera static at close range. Subject turns toward the lens.
22-30s: [Lateral track left] 0.4 m/s, subject remains centered in frame,
background parallax reveals the window.
```
Two things that structure buys you. Each beat has an explicit velocity, so no segment inherits a default. And the holds are declared rather than implied, which is what prevents the model from filling silence with movement.
## Anchor the space instead of describing it
Reference material is the strongest control surface available, and it is underused because it feels like extra work.
Describing a room in text gives the model a probability distribution over rooms. Supplying a reference gives it geometry. For camera work specifically, that difference decides whether the wall stays put across a twenty-second move.
The practical hierarchy, strongest first: a reference **video** whose camera move you want echoed, then still frames establishing the space from the angles the shot will pass through, then text describing what neither of those covers. Text is the fallback, not the foundation.
This also interacts with billing on the 2.0 generation, where supplying an uploaded video moves the request onto a cheaper rate. Worth knowing when a pipeline is video-driven anyway.
## Troubleshooting
**Motion blur on every move.** Velocity is too high for the frame rate to resolve. Cut the numeric speed rather than adding "less blur" to the prompt, which the model reads as a style note rather than a physics constraint.
**The camera drifts off the subject over a long take.** Add a framing constraint to each beat: "subject remains centered," "subject occupies the left third." Framing is a per-beat property, and stating it once at the top does not survive to second twenty-five.
**Movement starts before it should.** A beat without an explicit hold gets interpreted as an invitation to move. Declare the static segments.
**Every generation looks different despite identical prompts.** Space is being re-invented each run. That is a reference-material problem, not a prompt problem.
## Calling it
```bash
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2-5",
"prompt": "0-6s: [Locked tripod] medium shot, no camera movement. 6-14s: [Slow dolly in] push to close at 0.3 m/s, focal length fixed 50mm. 14-22s: [Hold] static, subject turns to lens. 22-30s: [Lateral track left] 0.4 m/s, subject centered.",
"resolution": "1080p",
"duration": 30
}'
```
Adding reference material anchors the geometry:
```bash
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2-5",
"prompt": "Match the camera move in the reference. Subject remains centered throughout.",
"video_urls": ["https://example.com/camera-reference.mp4"],
"image_urls": ["https://example.com/set-wide.jpg", "https://example.com/set-close.jpg"],
"resolution": "1080p",
"duration": 30
}'
```
Media inputs are **public http(s) URLs only**, with no base64 accepted. Current rates and the full request surface are on [reapi.ai/models/seedance-2-5](/models/seedance-2-5).
## FAQ
### Why does Seedance move the camera so fast by default?
Because descriptive words map to default velocity vectors, and generic descriptors resolve toward high acceleration. Adjectives are not velocity metrics.
### How do I specify camera speed precisely?
Use a number and a unit. "0.5 m/s" or "2 seconds per meter" both constrain the model; "slow" does not.
### What is the timestamp prompt format?
Split the shot into time ranges, each with a bracketed rig type, an explicit velocity, and a framing constraint. Declare the holds as well as the moves.
### How do I stop the camera drifting off my subject?
Put a framing constraint in every beat rather than once at the top. Framing does not persist across a long take by itself.
### Does reference material help with camera control?
Yes, more than prompt wording does. A reference video whose move you want echoed is the strongest anchor, followed by stills establishing the space.
### Why do identical prompts give different camera paths?
Because the space is being re-derived on each run. Supply reference geometry rather than describing the room in text.
### What causes motion blur in generated shots?
Velocity too high for the frame rate to resolve. Lower the numeric speed rather than asking for less blur.
## Directing instead of describing
The mental shift that fixes Seedance 2.5 camera control is small and specific. Stop writing what the shot should feel like and start writing what the rig does: which piece of equipment, moving how fast, for how long, holding when, framed how.
Every adjective in a camera prompt is a decision handed to a model that will resolve it toward motion. Every number is a decision you kept. Longer generations make that trade sharper, because thirty seconds of unconstrained interpretation drifts further than five. Write the beats, name the rig, give it geometry to hold onto, and the output stops being a lottery.
## References
1. Volcano Engine. *Seedance and the Doubao video model family.* Retrieved July 2026 from [volcengine.com](https://www.volcengine.com/)
### Further reading
* reAPI. *Seedance 2.5 features.* [reapi.ai/blog/seedance-2-5-features](/blog/seedance-2-5-features)
* reAPI. *Seedance 2.0 cost per second.* [reapi.ai/blog/seedance-2-0-cost-per-second](/blog/seedance-2-0-cost-per-second)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# Seedance 2.5 for E-commerce Video: A Real Workflow (https://reapi.ai/blog/seedance-2-5-ecommerce-video)
The failure mode in AI product video is not ugly output. It is a logo that warps halfway through a camera move, lighting that shifts mid-scene, and text the model invented and pasted across your packaging. Those are the things that make a clip unusable for an ad, and none of them are fixed by a better prompt alone.
Building a Seedance 2.5 e-commerce video workflow that survives contact with a real campaign comes down to how you organize the inputs, not how you phrase the request.
## TL;DR
* **Assign roles to references rather than dumping files.** Product shots are identity anchors, storyboard frames are structure, motion clips are camera guidance.
* **A white-model reference fixes shape drift.** Feed a plain 3D form alongside the product photo and the model respects the object's real dimensions through a camera move.
* **Use localized editing** to fix one region instead of regenerating a 30-second clip over a lighting glitch.
* **Write negative constraints explicitly.** No invented text, no added logos, no music.
* **Log the winning setup.** Prompt, reference hierarchy, and edit steps, per campaign.
* **Match format to funnel stage.** Lifestyle demo, technical explainer, and unboxing proof are different briefs, not one video re-cut.
## Organize the asset library first

High-converting output fails on messy inputs more often than on weak prompts. Sort source material into three buckets before writing anything.
**Identity assets.** High-resolution pack shots and brand style frames. These carry the things that must not drift: logo geometry, packaging typography, exact color.
**Motion references.** Short clips demonstrating the camera path or the human movement you want echoed. A reference video is a far stronger control than any sentence describing a dolly move.
**Script and tone cues.** Whatever locks the register of the spot, so the 30-second version does not wander tonally from the 6-second cutdown.
The 50-reference capacity is leverage only if each file has a job. Dumping fifty images gives the model fifty competing signals. Naming the product photo as the identity anchor and the storyboard as the structural guide gives it a hierarchy to resolve against, which is what keeps it from inventing a brand style of its own.
## The white-model trick for shape drift
The hardest defect to prompt away is motion drift: the product subtly changes shape as the camera pans, so the bottle is slightly wrong at second twelve in a way a viewer feels without naming.
The fix is structural rather than textual. Upload a **white model**, a plain untextured 3D form of the product, alongside the real product photo. The white model gives the engine physical dimensions and spatial orientation to respect, and the photo supplies the texture layered on top.
This works because it separates two problems the model otherwise solves together and badly: what shape is this object, and what does its surface look like. Answer the first with geometry and the second with a photograph, and the pan stops deforming the packaging.
## Fix regions, not clips
Regenerating a full 30-second take because of one harsh highlight is the biggest efficiency loss in this pipeline.
Localized editing isolates a region of the frame and corrects it while preserving the surrounding composition and timing. A character's expression, a specular hit on the product, a background element that reads wrong. The rest of the shot, including the camera move you already approved, stays intact.
The workflow discipline this enables is worth stating plainly: **approve the motion first, then fix the details.** Once the camera move is right, every subsequent problem is a local edit rather than a new roll of the dice.
## Write the constraints, including the negative ones
Product video has a specific failure the model volunteers unprompted: text. Invented taglines, garbled label copy, a logo that is almost yours. State the prohibition rather than hoping.
A production brief that works looks closer to this than to a description:
```text
Product: [name], reference image 1 is the identity anchor.
Geometry: reference image 2 is a white model — respect these proportions exactly.
Camera: 0-4s locked medium. 4-10s slow orbit left at 0.3 m/s. 10-15s hold.
Lighting: soft key from camera left, no hard specular on the label.
Constraints: do not add text, do not add logos, do not alter package typography,
no music, no on-screen captions.
Framing: product occupies the center third throughout.
```
Every line there is a decision removed from the model. The negative constraints are not politeness, they are the difference between a clip you can ship and one legal sends back.
## Match the format to the funnel
One hero video is not a campaign. Creative burns out fast enough that the useful posture is producing variations and measuring which survive.
| Format | Funnel stage | Objective |
| ------------------- | ------------ | ------------------------------------- |
| Lifestyle demo | Top | Emotional connection, brand awareness |
| Technical explainer | Middle | Educate on a specific benefit |
| Unboxing or proof | Bottom | Reduce purchase hesitation |
Those are three different briefs with different camera logic, not one asset re-cut three ways. The lifestyle demo can afford an ambient opening; the bottom-funnel proof clip cannot, because it has about two seconds to establish that the product is real.
## Log what worked
Most teams lose their winning formula because it lived in someone's session history.
Keep a production README per campaign holding the exact prompt, the reference hierarchy including which file played which role, and the local edit steps applied after the first generation. That file is what lets the next product launch reuse an aesthetic that performed, instead of reverse-engineering it from the output.
## The request
```bash
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2-5",
"prompt": "0-4s: [locked medium] product centered, soft key from camera left. 4-10s: [slow orbit left] 0.3 m/s, product remains centered. 10-15s: [hold] static. Do not add text, logos, or captions. Preserve package typography exactly.",
"image_urls": [
"https://example.com/product-packshot.jpg",
"https://example.com/product-whitemodel.jpg",
"https://example.com/storyboard-frame.jpg"
],
"resolution": "1080p",
"duration": 15
}'
```
Media inputs are **public http(s) URLs only**, with no base64 on any model, so assets need to be hosted before the call. Rates and the full request surface are on [reapi.ai/models/seedance-2-5](/models/seedance-2-5).
## Troubleshooting
**The logo warps during camera moves.** Add a white-model reference. This is a geometry problem, not a prompt problem.
**The model invents text on the packaging.** Add an explicit negative constraint and re-run. Text suppression has to be stated.
**Lighting shifts mid-clip.** Declare the key light position in the prompt and hold it across beats rather than mentioning it once.
**Every variation looks different.** The reference hierarchy is not stable. Fix which file is the identity anchor and keep it first across the whole batch.
**Output is technically fine but does not convert.** That is a format problem, not a generation problem. Check whether the clip matches its funnel stage.
## FAQ
### How do I stop the product changing shape during a camera move?
Upload a white model, a plain untextured 3D form, alongside the product photo. Geometry from one, texture from the other.
### How many reference images should I use?
Fewer with clear roles beats many without. Assign an identity anchor, a structural guide, and motion references rather than uploading everything available.
### How do I fix one bad detail without regenerating the clip?
Localized editing corrects a region while preserving the rest of the composition and the camera move you already approved.
### How do I stop the model adding text to my packaging?
State it as a negative constraint in the prompt. Invented text is a default behavior, not an accident.
### What length should an e-commerce product video be?
Match it to funnel stage rather than to a fixed number. Top-of-funnel lifestyle work can breathe; bottom-funnel proof clips have about two seconds to establish credibility.
### Should I generate one hero video or many variants?
Many. Creative burns out quickly, and the value of a fast pipeline is producing variations to measure rather than perfecting a single asset.
### What should I record after a successful generation?
The prompt verbatim, the reference hierarchy with each file's role, and any local edits applied afterward.
## Treating the pipeline as the product
The instinct with a capable video model is to keep rewriting the prompt until something good comes out. That produces one good clip and no repeatable process.
The Seedance 2.5 e-commerce video workflow that scales looks different: a curated asset library with roles assigned, geometry supplied rather than described, negative constraints stated, motion approved before details are polished, and the winning configuration written down. The prompt is the smallest part of it. What makes the next campaign fast is that the last one left behind a README rather than a folder of renders.
## References
1. Volcano Engine. *Seedance and the Doubao video model family.* Retrieved July 2026 from [volcengine.com](https://www.volcengine.com/)
### Further reading
* reAPI. *Seedance 2.5 camera control.* [reapi.ai/blog/seedance-2-5-camera-control](/blog/seedance-2-5-camera-control)
* reAPI. *Seedance 2.5 features.* [reapi.ai/blog/seedance-2-5-features](/blog/seedance-2-5-features)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# Seedance 2.5 Features: 30-Second Clips and 50 References (https://reapi.ai/blog/seedance-2-5-features)
At the Volcano Engine FORCE conference in Beijing on June 23, 2026, ByteDance previewed Seedance 2.5, the next version of its Doubao video model. The headline Seedance 2.5 features are a single continuous 30-second clip generated directly, support for up to 50 multimodal reference materials in one generation, and region-level editing\[1].
One framing before anything else, because it changes how you should read every number below: this was a **preview, not a release**. ByteDance demonstrated the model on stage, described it as being in enterprise beta, and set public launch for early July 2026\[1]. That date has passed without a public API. We track that separately in [Seedance 2.5 release status](/blog/seedance-2-5-release-status).
So everything here is a vendor claim measured against stage demos, and every independent benchmark that exists still describes the shipping predecessor, Seedance 2.0. This guide separates the two.
## TL;DR
* **Three announced pillars**: 30-second single clips (up from 15), up to 50 multimodal references (up from 12), and region-level editing plus 3D previz\[1].
* **Native 4K is not one of them.** The native-4K pipeline belongs to the Seedance 2.0 line announced at the same event. Several outlets merged the two\[1].
* **No independent benchmark for 2.5 exists**, because it has not shipped publicly. Any 2.5 arena score circulating now is unverified.
* **The predecessor already leads.** Seedance 2.0 sits at #1 on Artificial Analysis's Text-to-Video Arena with an Elo of 1,219 among models with audio\[2].
* **Pricing is unpublished.** For context, Seedance 2.0 normalizes to roughly $9 per minute of 1080p, against about $24 for Veo 3.1 and $20 for Kling 3.0 Pro\[2].
## What ByteDance actually announced

| Capability | Seedance 2.0 | Seedance 2.5 (claimed) |
| -------------------- | ------------ | ------------------------------------ |
| Single-clip duration | 15 seconds | **30 seconds, generated directly** |
| Reference materials | up to 12 | **up to 50, full multimodal** |
| Editing control | limited | **region-level editing + 3D previz** |
According to the keynote, the 30-second clip is produced as one native generation rather than several short clips joined at the seams, and the 50-reference ceiling combining images, video, and text in a single joint generation is the highest ByteDance is aware of in a commercial video model\[1].
## Why 30 seconds is harder than it sounds
Doubling maximum clip length from 15 to 30 seconds sounds incremental. In video generation it is one of the harder problems.
Most models hold quality for a few seconds and then drift. Characters subtly change appearance, lighting shifts, motion loses physical plausibility. The common workaround is generating several short clips and stitching them, which introduces visible seams and continuity errors at every join\[1].
Generating a coherent 30-second take directly, if it holds outside curated demos, removes a real production tax. For the formats ByteDance named on stage, film and TV pre-visualization, advertising, and short-form animated drama, a continuous half-minute is often the difference between a usable shot and a clip that needs manual repair.
The honest caveat: a single 30-second clip is a *ceiling*. Sustained quality across that full window is exactly what independent testing has to confirm.
## Fifty references is the consistency play

The jump from 12 to 50 reference inputs is the upgrade most relevant to professional work, and it is the one likeliest to be undersold. A reference can be an image, a video, or text, and all 50 feed a single joint generation\[1].
In practice this is a bid for consistency. Feed the model your character turnarounds, product shots, brand colors, and style frames, and it has far more grounding to hold the same face, the same packaging, and the same look across a sequence.
For an advertiser, character and product fidelity across shots is the entire job. A model that drifts on the logo is unusable regardless of how cinematic the motion is.
## Region editing and 3D previz
The third pillar is control. ByteDance demonstrated region-level editing: replacing a subject, background, or product inside an existing shot without changing the original motion, camera move, or lighting\[1].
That matters for iteration. Instead of regenerating an entire clip and hoping the rest survives, you swap one element and keep everything already approved. For localizing a product into a different market or swapping a model, that targeted edit is often the whole job.
The keynote also showed 3D white-model previsualization, letting creators block out shots and camera moves in a rough 3D scene before committing to a full generation\[1]. That brings storyboarding and camera blocking into the generation step rather than leaving them to trial and error.
## The 4K confusion worth clearing up
Native 4K, including 4K 10-bit, was discussed at FORCE, but it belongs to the **Seedance 2.0 line**, not to the 2.5 headline slide, which leads with duration, references, and editing control\[1].
Several outlets folded 4K into the 2.5 spec sheet. We keep them separate because that is what the official materials show, and because a spec that gets attributed to the wrong model is exactly the kind of error that survives into buying decisions. Our earlier piece on [what we know about Seedance 2.5](/blog/seedance-2-5-what-we-know-2026) flagged the same press-claimed 4K spec.
## Where the shipping version stands
There is no independent benchmark for Seedance 2.5. What can be reported is where Seedance 2.0 stands on a neutral blind human-preference leaderboard.

| Model | Text-to-Video Arena Elo |
| ----------------------- | ----------------------- |
| **Seedance 2.0 (720p)** | **1,219** |
| HappyHorse-1.0 | 1,124 |
| Kling 3.0 1080p Pro | 1,105 |
| Google Veo 3.1 | 1,094 |
Seedance 2.0 currently sits at #1 among models with audio, and leads the Image-to-Video board as well\[2]. A first-place arena standing reflects aggregate human preference on sampled prompts, not a guaranteed win on every brief, and emphatically not a 2.5 number.
What it does tell you is that the baseline 2.5 builds on is already at the front of the field, which raises the bar for what the upgrade has to prove.
## Pricing context
ByteDance has not published Seedance 2.5 pricing. For context, Artificial Analysis normalizes the shipping Seedance 2.0 to roughly $9 per minute of 1080p video, against about $24 per minute for Google Veo 3.1 and roughly $20 for Kling 3.0 Pro\[2].
If 2.5 lands anywhere near that band, the pitch is not just quality but quality per dollar, which was the posture across the whole FORCE lineup. Seedream 5.0 for image, Seed-Audio 1.0 for audio, and Doubao 2.1 Pro for text shipped as previews at the same event, making it a full-stack multimodal release aimed at being a single cheaper provider for an entire generation pipeline\[1].
## Who the upgrades are aimed at
These are not hobbyist features. They target people for whom consistency and iteration are the job.
**Advertisers and brand teams** are the clearest fit. The 50-reference ceiling exists so a campaign can hold a product, a logo, and a spokesperson on-model across every shot, and region editing turns localizing an ad into a targeted edit rather than a reshoot.
**Pre-visualization teams** get 3D white-model blocking and longer takes, which move storyboarding into the generation step.
**Short-form animation studios** need exactly the multi-shot narrative coherence a continuous 30-second clip enables.
**Performance marketers producing volume** benefit if the cost posture holds, because the economics of generating many on-brand variants shift decisively.
Consider the concrete case. A consumer-electronics brand needs the same 20-second hero spot in four markets. With a drift-prone short-clip model that is four near-from-scratch generations plus manual continuity fixes. With a reference stack holding the product fixed and region editing swapping only the talent and packaging copy, versions two through four become edits of the first. Same motion, same lighting, different market.
That workflow is why the boring upgrades, references and editing, may matter more than the headline 30 seconds.
## What to test when it ships
Three concrete checks, because demos are curated and the model was still in enterprise beta at preview:
1. **Does the 30-second clip stay coherent end to end**, or does drift appear at second 18 the way it does on shorter models pushed past their comfort zone?
2. **Do 50 references actually hold** a character and a brand across shots, or does the marginal reference stop contributing well before the ceiling?
3. **Does region editing leave the rest of the frame untouched**, including lighting and camera motion?
Until those have answers from someone other than the vendor, treat the numbers as ByteDance's rather than the field's.
## FAQ
### Is Seedance 2.5 available now?
Not publicly. It was previewed on June 23, 2026 at FORCE and described as being in enterprise beta, with public launch set for early July 2026\[1]. That date has passed without a public API. The version you can actually use is Seedance 2.0.
### Does Seedance 2.5 really generate 30-second videos?
ByteDance claims a single continuous 30-second clip generated directly, double the 15-second ceiling of Seedance 2.0 and without stitching. That was demonstrated on stage; independent confirmation waits on public access\[1].
### Does Seedance 2.5 support native 4K?
Native 4K was discussed at FORCE but belongs to the Seedance 2.0 line, not the 2.5 headline slide. Some coverage merged them\[1].
### How does it compare to Veo, Kling, and Sora?
No Seedance 2.5 benchmark exists. Shipping Seedance 2.0 leads the Artificial Analysis Text-to-Video and Image-to-Video arenas on blind human preference at a lower normalized price than Veo 3.1 or Kling 3.0 Pro\[2].
### How many reference materials can Seedance 2.5 take?
Up to 50, combining images, video, and text in one joint generation, against 12 on Seedance 2.0\[1].
### What else did ByteDance announce at FORCE?
Seedream 5.0 for image, Seed-Audio 1.0 for audio, and Doubao 2.1 Pro for text, alongside platform figures including roughly 180 trillion tokens processed per day\[1].
### Can I use Seedance 2.5 through reAPI today?
No. Seedance 2.5 has no public API to route to. Seedance 2.0 is available and is the model behind the arena standing above.
## Reading a preview as a preview
Seedance 2.5 is a confident set of claims from the team whose current model already tops the independent video arenas. The upgrades it leads with are the right ones for professional production, where consistency and iteration matter more than one cinematic clip.
But the gap between a polished stage demo and a model holding quality across 30 seconds on *your* prompts is exactly what public access will reveal, and that access is later than announced. Until then the useful posture is the one this guide takes throughout: treat the Seedance 2.5 features as announced capabilities with a vendor's name on them, use Seedance 2.0 for work that ships this week, and keep the test list ready for the day the API opens.
## References
1. Volcano Engine. *Seedance and the Doubao model family, announced at the FORCE conference.* Retrieved July 2026 from [volcengine.com](https://www.volcengine.com/)
2. Artificial Analysis. *Text-to-Video Arena leaderboard.* Retrieved July 2026 from [artificialanalysis.ai/text-to-video/arena](https://artificialanalysis.ai/text-to-video/arena)
### Further reading
* reAPI. *Seedance 2.5 release status.* [reapi.ai/blog/seedance-2-5-release-status](/blog/seedance-2-5-release-status)
* reAPI. *Seedance 2.5: what we know before the public launch.* [reapi.ai/blog/seedance-2-5-what-we-know-2026](/blog/seedance-2-5-what-we-know-2026)
* reAPI. *Seedream 5.0 Pro and Seedance 2.5 workflow.* [reapi.ai/blog/seedream-5-0-pro-seedance-2-5-workflow](/blog/seedream-5-0-pro-seedance-2-5-workflow)
---
# Seedance 2.5 on Higgsfield? Availability & Alternatives (https://reapi.ai/blog/seedance-2-5-on-higgsfield)
**Seedance 2.5 is not callable on Higgsfield as of August 2, 2026.** Higgsfield now has a Seedance 2.5 coming-soon notification page, but the page says access has not arrived. Its working Seedance routes remain 2.0 and 2.0 Fast, including a time-limited Unlimited offer built on Enhanced Seedance 2.0 Fast.\[1]\[2]
That answer is easy to blur because “Seedance 2.5 Higgsfield” combines a model with a platform. This guide separates them, explains what a Higgsfield subscription buys today, and shows the closest working alternatives while 2.5 remains in rollout.
## TL;DR
* Higgsfield has a **Seedance 2.5 coming-soon page**, but it does not provide a callable model, price, or production limits.\[1]
* Higgsfield Unlimited uses **Enhanced Seedance 2.0 Fast** at 480p or 720p, not an unlimited 2.5 or 4K plan.\[2]
* No verified Higgsfield price exists for Seedance 2.5 because its coming-soon page publishes no generation price.
* Dreamina has an official 2.5 page but labels the model “coming soon” and tells users to check account-level rollout.\[3]
* Atlas Cloud has a 2.5 early-access page, while its callable API remains Seedance 2.0.\[4]
* If you need a production API today, use [Seedance 2.0](/models/seedance-2-0) or [MiniMax H3](/models/minimax-h3) on reAPI and keep the model ID configurable for a later 2.5 test.
## Is Seedance 2.5 available on Higgsfield?
No public Higgsfield page checked for this article confirms live Seedance 2.5 generation. Higgsfield's dedicated 2.5 page is a notification surface, while the working video interface and plan material identify Seedance 2.0 and Seedance 2.0 Fast.\[1]\[5]
This is not a minor naming difference. Seedance 2.5 is the announced successor with 30-second continuous generation, expanded multimodal reference handling, second-level direction, and editing demonstrations. Seedance 2.0 is the shipping model with a documented 4–15-second generation range and a real API schema.\[6]
The fastest verification method is the model picker inside your signed-in Higgsfield account. Product access can vary by region and plan, but a promotional blog or third-party directory is not proof that the generator is callable. Look for an explicit `Seedance 2.5` option, its credit cost, duration controls, and output settings before buying a plan for that model.

## What Higgsfield offers Seedance users today
Higgsfield is a creator studio, not merely a model endpoint. Its value comes from putting generation beside reusable characters, camera-oriented tools, advertising workflows, project organization, and other video models.
Higgsfield currently presents Seedance across several product surfaces:
| Higgsfield route | What it provides | Important limit |
| ------------------------------ | ----------------------------------------------------- | -------------------------------------------------------------------------------------------- |
| Standard plan access | Seedance 2.0 and 2.0 Fast through the web studio | Uses plan credits unless a specific promotion applies |
| Seedance Unlimited promotion | 30 days of credit-free Enhanced Seedance 2.0 Fast | Standard queue, 480p/720p, up to 8 or 15 seconds by plan\[2] |
| Seedance 2.5 notification page | Product description and an availability alert | Coming soon; no callable model or published price\[1] |
| Priority access language | Higher plans advertise earlier access to new releases | This is not confirmation that 2.5 is currently available\[5] |
Higgsfield's published plan guide starts full-model Plus/Pro access at $29 per month and Ultra/Max at $79, while credits and offers can vary by region.\[5] Those are platform prices, not a Seedance 2.5 price. Do not divide the monthly fee by a guessed number of 2.5 clips.
## Why the Seedance 2.5 pricing question has no answer yet
ByteDance's official 2.5 promotion describes capabilities but does not publish a model ID, parameter reference, or billing row. Its current catalogs still document the Seedance 2.0 family.\[6]
Three numbers are often confused:
1. The price of a Higgsfield subscription.
2. The credits charged for a shipping Seedance 2.0 generation.
3. The eventual upstream cost of Seedance 2.5.
Only the first two can be checked today. The third determines whether providers will bundle 2.5 into existing plans, put it behind a higher tier, or charge per generation. Until that happens, “Seedance 2.5 pricing on Higgsfield” is a status question, not a calculation.
## Working alternatives to Seedance 2.5 on Higgsfield
### Use Seedance 2.0 on reAPI now
[Seedance 2.0 on reAPI](/models/seedance-2-0) is the direct route for developers who searched for Higgsfield but actually need an API beneath their own product. The live model page exposes the supported variants, duration and resolution rules, reference inputs, and current per-second charges. Media jobs follow one submit-and-poll task flow, and a job that finishes in failure is automatically refunded.
This does not reproduce Higgsfield's editing studio. It solves a different problem: programmatic Seedance generation with a documented contract today, while the model ID remains configurable for a later 2.5 migration. Start with the [API quickstart](/docs/api/quickstart) if the application does not yet have a task queue.
### Use MiniMax H3 on reAPI when the deadline cannot wait
MiniMax H3 is not Seedance, but it is the closest newly released alternative for mixed-reference video. It supports 4–15-second output at up to 2K with native stereo audio and a documented API.\[7] You can inspect the live schema and price on the [MiniMax H3 model page](/models/minimax-h3), then compare it with Seedance 2.0 using the same prompts.
### Use Seedance 2.0 on Higgsfield
Choose this when you already value Higgsfield's editor and need Seedance output now. It is the least disruptive creator route, but it retains the current model's duration and plan-credit constraints.
### Check Dreamina for official account rollout
Dreamina is ByteDance's creator-facing surface. Its official Seedance 2.5 page says “coming soon,” while its access guide tells users to sign in and check whether the model appears for their account.\[3] This is the most relevant consumer route to monitor, not proof of universal access.
### Monitor Atlas Cloud early access
Atlas Cloud is preparing day-one API access and has a 2.5 cohort page. Its live Seedance endpoint is still 2.0, with an asynchronous submit-and-poll API.\[4] Treat that page as a rollout signal, not a reason to make a production schedule depend on an unpublished endpoint.
## Which route should you choose?
| Your priority | Best current route |
| --------------------------------------- | ---------------------------------------------------------------------------------- |
| Shipping Seedance through an API now | [Seedance 2.0 on reAPI](/models/seedance-2-0) |
| Testing a new mixed-reference model now | [MiniMax H3 on reAPI](/models/minimax-h3) |
| Higgsfield editing and creator tools | Seedance 2.0 inside Higgsfield |
| Official ByteDance creator access | Monitor Dreamina account rollout |
| Monitoring third-party 2.5 rollout | Atlas Cloud early-access list and reAPI's [coming-soon page](/models/seedance-2-5) |
Do not pay for an annual plan solely on an assumption about future model access. Buy the workflow that exists, then treat 2.5 as an upgrade when its button, price, and limits are visible.
## FAQ
### Does Higgsfield have Seedance 2.5?
Higgsfield has a public Seedance 2.5 coming-soon page, but it does not yet provide live generation, a model price, or production limits. Its callable creator workflow remains on Seedance 2.0 and 2.0 Fast.
### Is Higgsfield Seedance Unlimited actually Seedance 2.5?
No. Higgsfield says the promotion uses Enhanced Seedance 2.0 Fast at 480p or 720p.\[2]
### How much will Seedance 2.5 cost on Higgsfield?
No price has been published. Existing subscription prices and Seedance 2.0 credit costs should not be presented as 2.5 pricing.
### Can I use Seedance 2.5 on Atlas Cloud?
Atlas Cloud has an early-access page, but its publicly documented callable Seedance API is 2.0. Verify the model identifier in the live catalog before sending production traffic.
### What is the best alternative while waiting?
Use [Seedance 2.0 on reAPI](/models/seedance-2-0) if workflow continuity matters. Use [MiniMax H3](/models/minimax-h3) if you need a newly released model with 2K output, native audio, and a public API today. Use Higgsfield's Seedance 2.0 workflow instead when a browser studio matters more than programmatic access.
## Wait for a model selector, not a headline
The useful answer to “Seedance 2.5 Higgsfield” is currently no: Higgsfield has a coming-soon page, not live 2.5 generation. If you need a creator studio, keep producing with Seedance 2.0 on Higgsfield. If you need an API beneath your own product, start with [Seedance 2.0 on reAPI](/models/seedance-2-0), keep the model ID configurable, and test 2.5 only when a real schema and price arrive.
## References
1. Higgsfield. *Seedance 2.5 coming-soon notification page and current video model surfaces.* Retrieved August 2, 2026. higgsfield.ai
2. Higgsfield. *What Creators Really Get with Seedance Unlimited on Higgsfield.* Updated July 9, 2026. geo.higgsfield.ai
3. Dreamina. *Official Seedance 2.5 AI Video Generator and access guide.* Retrieved August 2, 2026. [dreamina.capcut.com](https://dreamina.capcut.com/seedance/seedance-2-5)
4. Atlas Cloud. *Seedance 2.0 API and Seedance 2.5 early access.* Retrieved August 2, 2026. atlascloud.ai
5. Higgsfield. *AI pricing and plans.* Retrieved August 2, 2026. geo.higgsfield.ai
6. Volcano Engine. *Doubao Seedance 2.5 promotion and current model catalog.* Retrieved August 2, 2026. [ark.volcengine.com](https://ark.volcengine.com/promotion?modelName=seedance-2-5)
7. MiniMax. *MiniMax H3 launch and video API guide.* Published July 31, 2026. [minimax.io](https://www.minimax.io/blog/minimax-h3)
### Further reading
* reAPI. *Seedance 2.5 launch status.* [reapi.ai/blog/seedance-2-5-release-status](/blog/seedance-2-5-release-status)
* reAPI. *Best Seedance 2.5 alternatives.* [reapi.ai/blog/best-seedance-2-5-alternatives](/blog/best-seedance-2-5-alternatives)
* reAPI. *Atlas Cloud vs Higgsfield.* [reapi.ai/blog/atlas-cloud-vs-higgsfield](/blog/atlas-cloud-vs-higgsfield)
---
# Seedance 2.5 Pricing: 53% More Per Token Than Seedance 2.0 (https://reapi.ai/blog/seedance-2-5-pricing-per-token)
Seedance 2.5 pricing went live on ByteDance's docs ahead of the model's August 7 API launch, and it lands at $10.70 per million tokens without video input, against $7.00 for Seedance 2.0\[1]. That is 52.9% more expensive per token for the same two resolutions.
The per-token framing matters more than it sounds, because AI video is quoted per second and metered per token. The gap between those two units is where most cost surprises live, and Seedance 2.5 widens it in a way that breaks any conversion factor you built for 2.0.
## TL;DR
* Seedance 2.5 costs **$10.70 per million tokens** without video input and **$6.40 with it**, against $7.00 and $4.30 for Seedance 2.0 — 52.9% and 48.8% increases\[1].
* Only **480p and 720p** are published for 2.5. No 1080p or 4K rate exists yet, and offline inference is marked "not supported yet"\[1].
* Billing runs on `(input_video_seconds + output_seconds) × width × height × fps / 1024`, with fps fixed at 24\[1]. Your reference clip is metered exactly like generated frames.
* **480p changed frames between versions.** Back-solving ByteDance's own worked examples puts 2.5 at \~854×480 and 2.0 at a taller frame, which is why 480p rises 46% per second while the token rate rises 53%.
* 2.5 accepts up to **30 seconds** of reference video, double 2.0's 15. A five-second 720p job runs $1.244 with a short reference and $4.838 with a 30-second one\[1].
* Seedance 2.0 on reAPI runs $0.023 to $0.780 per billable second depending on tier\[2].
## Seedance 2.5 pricing, and the rate that looks wrong
ByteDance publishes token rates per model and input mode\[1]:
| Model | No video input | With video input |
| ------------------------- | -------------- | ---------------- |
| Seedance 2.5 (480p, 720p) | 10.70 | 6.40 |
| Seedance 2.0 (480p, 720p) | 7.00 | 4.30 |
| Seedance 2.0 (1080p) | 7.70 | 4.70 |
| Seedance 2.0 (4K) | 4.00 | 2.40 |
Seedance 2.5 occupies a single row because it ships with a single rate. Read the 4K row underneath it again, though: it is the cheapest tier per token by a wide margin, 43% below 480p.
It is also, by a wide margin, the most expensive output you can buy. A 3840×2160 frame carries 19.4 times the pixels of what 480p actually renders. The rate falls 43% while the token count rises 1940%. Anyone comparing providers by scanning the rate column will conclude 4K is a bargain, and will be wrong by a factor of eleven.
This is the recurring hazard with token-metered video. The number on the price sheet is only half of the multiplication.
## Where the tokens come from
The formula ByteDance documents is\[1]:
```
tokens = (input_video_seconds + output_seconds) × width × height × fps / 1024
```
Frame rate is fixed at 24. Pixel dimensions per label:
| Label | Pixels | Tokens per second |
| ------------------- | ----------- | ----------------- |
| 480p (Seedance 2.5) | \~854 × 480 | 9,607 |
| 480p (Seedance 2.0) | 864 × 496 | 10,044 |
| 720p | 1280 × 720 | 21,600 |
| 1080p | 1920 × 1080 | 48,600 |
| 4K | 3840 × 2160 | 194,400 |
Checking the formula against a published example: five seconds of 720p on 2.0 with no reference is 5 × 21,600 = 108,000 tokens, times $7.00 per million, or $0.756. ByteDance's own worked example for that exact configuration reads $0.756\[1]. The formula is exact, not an approximation.
## The 480p frame changed and nobody announced it
This one is not in the release notes. It falls out of dividing ByteDance's worked examples by their own token rates.
Their published five-second, 16:9, no-reference examples\[1]:
| Model | 480p | 720p |
| ------------ | ----------------- | ----------------- |
| Seedance 2.5 | $0.514 ($0.103/s) | $1.156 ($0.231/s) |
| Seedance 2.0 | $0.352 ($0.070/s) | $0.756 ($0.151/s) |
Divide each price by its token rate to recover the token count, then by 24/1024 to recover pixels. Seedance 2.5 lands at 9,607 tokens per second of 480p, roughly 854×480. Seedance 2.0 lands at about 10,057, a slightly taller frame.
720p resolves to 21,600 tokens per second on both, which is exactly 1280×720. That the same arithmetic produces a clean, verifiable number at 720p is what makes the 480p result trustworthy rather than a rounding artifact.
The consequence: at 720p, where the frame is unchanged, per-second cost rises by the same 53% as the token rate. At 480p, where the frame shrank about 4.5%, per-second cost rises only 46%. If you carry a per-second conversion factor from 2.0 to 2.5, it is wrong at 480p by roughly five percent, in your customer's favor if you are lucky and yours if you are not.
## Your reference clip bills like generated video
`input_video_seconds` sits inside the same parenthesis as the output. A reference-to-video job pays for the source you uploaded at the identical rate as the frames the model invented.
Providers present this as a discounted per-second rate for reference mode, which reads as a bargain right up until you total it. Seedance 2.5's own published range makes the point without help: a five-second 720p generation costs $1.244 with a short reference and $4.838 with a 30-second one\[1]. Same output, 3.9 times the invoice.
Seedance 2.5 also widened the input window to 30 seconds, where 2.0 caps at 15\[1]. The ceiling on a single job roughly doubled alongside it.
There is a floor as well, and it is unpublished as a number. ByteDance's examples price two-second and four-second inputs identically, which implies a minimum around four seconds. The docs point at a spreadsheet calculator instead of stating it\[1]. If you resell this, customers sending two-second clips cost you more than your arithmetic predicts.
## What Seedance 2.0 costs per second today
Until 2.5 opens, 2.0 is what you can actually call. Per second of billable duration on reAPI, in USD\[2]\[3]:
| Tier | Text or image input | With reference video |
| ---------------------- | ------------------- | -------------------- |
| Seedance 2.0 Mini 480p | $0.036 | $0.023 |
| Mini 720p | $0.077 | $0.047 |
| Fast 480p | $0.059 | $0.034 |
| Fast 720p | $0.124 | $0.075 |
| Seedance 2.0 480p | $0.072 | $0.044 |
| Seedance 2.0 720p | $0.154 | $0.094 |
| Seedance 2.0 1080p | $0.383 | $0.233 |
| Seedance 2.0 4K | $0.780 | $0.480 |
Multiplied to a minute, that spans $1.38 to $46.80, a 34x range. Worth internalizing before you pick a default resolution for a product, because that decision compounds across every generation your users trigger.
Per-minute figures are a unit conversion rather than a purchasable thing, incidentally. A single generation caps at 15 seconds, so a minute of finished video is four clips minimum and realistically more once you discard the bad takes.
The practical move is to stop paying flagship rates while you are still hunting for the prompt. Mini at 480p is 18 cents for five seconds. The same five seconds at 4K is $3.90. Parameters are identical across the family, so iterating cheap and rendering expensive costs one model-id change. Twenty iterations at Mini plus one 4K render is $7.50; twenty iterations at 4K is $78.
## FAQ
### When does Seedance 2.5 launch?
The API opens August 7, 2026. Pricing was published ahead of it\[1].
### How much does Seedance 2.5 cost per second?
At the published rates, roughly $0.103 per second at 480p and $0.231 at 720p for a five-second clip with no reference video\[1]. Reference video adds your input duration to the billed total.
### Does Seedance 2.5 support 1080p or 4K?
No published rate exists for either as of August 5, 2026. Only 480p and 720p appear on the pricing page\[1].
### Why is 4K cheaper per token than 480p?
Unexplained by the vendor, and it does not apply to Seedance 2.5 anyway, which publishes one rate covering both its resolutions. On 2.0, what matters is that the token count rises far faster than the rate falls, so 4K remains the most expensive output despite the lowest rate.
### Is Seedance 2.0 being retired?
Nothing announced. Its rates and capabilities stand, and 1080p and 4K currently exist only there.
### How do I estimate a job before running it?
Compute `(input_seconds + output_seconds) × tokens_per_second` from the pixel table, multiply by the rate, divide by a million. Then bill downstream on the returned `usage.completion_tokens`, since the formula is an estimate until the job finishes\[1].
### Does a failed generation still cost money?
On reAPI, no. Failed jobs refund automatically\[3].
## Budgeting for the switch
The honest summary is that Seedance 2.5 pricing represents a real increase rather than a repackaging: 53% more per token, 46 to 53% more per second depending on resolution, and no 1080p or 4K option at launch. Whether that is worth paying depends on output quality nobody outside ByteDance has measured yet.
What you can do before August 7 is fix your arithmetic. Rebuild your per-second conversion from the 2.5 pixel dimensions rather than inheriting the 2.0 factor, budget reference jobs on input plus output, and meter customers on returned token counts instead of estimates. Those three corrections cover most of the gap between a quote and an invoice, and they apply whatever Seedance 2.5 pricing settles at once the model is actually live.
## References
1. BytePlus. *ModelArk — model pricing (token rates, calculation formula, and per-video examples).* Retrieved August 2026 from [docs.byteplus.com/en/docs/ModelArk/1544106](https://docs.byteplus.com/en/docs/ModelArk/1544106)
2. reAPI. *Seedance 2.0 — model page and live pricing.* Retrieved August 2026 from [reapi.ai/models/seedance-2-0](/models/seedance-2-0)
3. reAPI. *Seedance 2.0 Mini — model page and live pricing.* Retrieved August 2026 from [reapi.ai/models/seedance-2-0-mini](/models/seedance-2-0-mini)
4. BytePlus. *ModelArk — Seedance 2.0 API reference (resolution, duration and pixel dimensions).* Retrieved August 2026 from [docs.byteplus.com/en/docs/ModelArk/1520757](https://docs.byteplus.com/en/docs/ModelArk/1520757)
---
# Seedance 2.5 Launch Status: 30 Seconds Confirmed, 4K Unverified (https://reapi.ai/blog/seedance-2-5-release-status)
**Seedance 2.5 has an official ByteDance product page and an impressive set of demos, but it does not yet have a public API model ID, price, or final parameter table.** As of August 2, 2026, ByteDance has confirmed 30-second continuous output, expanded multimodal references, second-level control, controllable video editing, and broader multilingual presentation. It has not confirmed native 4K for version 2.5, an exact 30-image/10-video/10-audio allowance, or named Maya and Blender integrations.\[1]\[2]\[3]
That distinction matters because the most shared Seedance 2.5 summaries combine three different things: capabilities shown by ByteDance, figures repeated from launch coverage, and features that already belong to Seedance 2.0. This guide separates them, explains what the confirmed upgrade means for production, and compares the preview with the now-documented MiniMax H3 release.
## TL;DR
* **Confirmed by ByteDance:** one-pass 30-second video, more multimodal reference input, second-level scene control, deeper video editing, and multilingual output demonstrations.\[1]
* **Demonstrated, not fully specified:** the official page uses eight image references in one example and timestamped instructions in another, but publishes no final input-limit table.\[1]
* **Not first-party verified for Seedance 2.5:** native 4K, 30 images plus 10 videos plus 10 audio clips, 50 total references, 10-bit output, and direct Maya or Blender integration.
* **Not a public API launch yet:** the current official model catalog and pricing page list the Seedance 2.0 family, not a callable 2.5 model ID.\[2]\[3]
* **The practical move:** keep shipping on [Seedance 2.0](/models/seedance-2-0), prepare a reference-role test set, and treat 2.5 as a migration target until its real API contract appears.
## Seedance 2.5 specs: confirmed versus unverified

| Claim | Status on August 2, 2026 | What the evidence says |
| ------------------------------------- | --------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 30-second continuous generation | **Officially confirmed** | ByteDance's ModelArk promotion says “30s continuous output” and shows a segmented 30-second example.\[1] |
| Expanded multimodal references | **Officially confirmed, limit unknown** | The page promises more reference input and demonstrates eight images; no final image/video/audio ceiling is published.\[1] |
| Timestamp or segment control | **Officially demonstrated** | Prompts assign actions to time ranges inside a continuous clip.\[1] |
| Video editing | **Officially demonstrated** | One example removes background people while preserving the main subject.\[1] |
| Multilingual speech and lip sync | **Officially demonstrated** | The product page includes a multilingual performance example.\[1] |
| Native 4K in Seedance 2.5 | **Unverified** | The 2.5 promotion does not state a resolution. 4K is already documented for Seedance 2.0.\[1]\[2] |
| 30 images, 10 videos, 10 audio inputs | **Unverified** | This exact split appears in third-party summaries, not in the official 2.5 page or API documentation. |
| 50 total references | **Unverified as an API limit** | ByteDance says “more,” while its public demo visibly addresses eight images.\[1] |
| Maya and Blender integration | **Unverified** | No named DCC plug-in, import format, or integration guide appears in the first-party material checked. |
| Public API availability | **Not found** | No 2.5 model ID or price row appears in the current official catalog.\[2]\[3] |
The table does not imply that the unverified features are impossible. It means they should not be used in purchasing decisions, API schemas, or production estimates until ByteDance publishes a parameter table.
## What 30-second generation changes in practice
Moving from Seedance 2.0's documented 15-second ceiling to a 30-second single pass is more than doubling a number.\[4] It reduces the number of seams in any sequence longer than one shot.
A 60-second product film built from 15-second generations needs at least four clips and three joins. At 30 seconds, the theoretical minimum falls to two clips and one join. Fewer joins mean fewer opportunities for wardrobe, lighting, geometry, voice, and screen direction to drift.
The tradeoff is failure cost. A bad detail near second 27 can make a longer generation expensive to discard. The best workflow will probably remain modular:
1. Block the story into five- to ten-second beats.
2. Describe those beats with explicit time ranges.
3. Generate a 30-second take when continuity is more valuable than independent rerolls.
4. Keep difficult dialogue, product text, or transformation shots separate when precision matters more.
Until public pricing appears, “30 seconds” should be treated as a creative ceiling, not the default duration for every request.
## The multimodal reference system is the real upgrade
Longer clips attract the headline, but the reference system is more important for repeatable production. The official Seedance 2.5 example addresses eight images individually and gives each one a role inside the scene.\[1] That suggests a workflow where references are not merely blended into a style average; they can be assigned to characters, objects, environments, and beats.
The useful way to prepare is to build a reference manifest now:
| Reference role | Recommended asset | What it should control |
| ------------------ | --------------------------------------- | -------------------------------------------------- |
| Character identity | Clean front and three-quarter portraits | Face, hair, age, and wardrobe anchors |
| Product identity | Neutral product packshot | Shape, materials, color, and logo placement |
| Environment | Wide establishing frame | Architecture, palette, weather, and time of day |
| Motion | Short source clip | Gesture, blocking, camera path, or physical timing |
| Audio | Clean speech or sound reference | Voice, rhythm, ambience, or lip-sync target |
Do not design around the rumored 30/10/10 split yet. Design around roles. When the API arrives, the manifest can be compressed to the actual limits without rewriting the creative plan.
## Timestamp control turns a prompt into a compact edit timeline
ByteDance's demonstration divides a continuous generation into timed segments. That makes a Seedance prompt closer to a shot list:
```text
0–6s: Wide establishing shot. The actor enters from frame left.
6–14s: Track backward at walking speed. Keep the product in the right hand.
14–22s: Medium close-up. The actor stops and delivers one short line.
22–30s: Slow orbit to the product. Hold the final composition for two seconds.
```
This is not proof of frame-exact editing. “Second-level control” describes the interface demonstrated by ByteDance; it does not guarantee that every action starts on the requested frame.\[1] Prompts should still leave physical transitions enough time to complete.
The editing demo is also notable. Removing everyone except one protagonist from an existing clip is a different task from text-to-video. If the public API exposes this capability, teams could use one model for generation, cleanup, and selective revision instead of rerendering the entire shot. The missing information is the contract: masks, reference-video limits, output resolution, and pricing remain undocumented.
## Native 4K is the wrong claim to lead with
Several summaries describe Seedance 2.5 as a native 4K launch. The official 2.5 promotion does not mention 4K. Meanwhile, ByteDance's current model catalog already documents 4K output for Seedance 2.0.\[1]\[2]
There are three plausible explanations: 2.5 inherits the existing 4K path, 4K is available only on a product surface not yet documented, or coverage attached a 2.0 capability to the 2.5 announcement. None supports writing “Seedance 2.5 launches with native 4K” as a settled fact.
For production, resolution is only one part of quality. Temporal stability, facial continuity, typography, motion coherence, and sound synchronization usually determine whether a shot survives review. Evaluate those at the delivery resolution you actually need rather than treating 4K as a proxy for reliability.
## Is there a Maya or Blender integration?
Not in the official material currently available. Some coverage interprets 3D-oriented demonstrations as Maya or Blender integration. A model accepting a render, viewport capture, depth-like image, or turntable video is not the same as a native plug-in that reads scene geometry, cameras, materials, or animation curves.
A real DCC integration announcement should identify at least one of these:
* a supported Maya or Blender plug-in;
* a scene or interchange format such as USD, Alembic, or FBX;
* preserved camera and object metadata;
* an official installation guide and version requirements.
Until that exists, the safe description is **3D-reference-friendly video generation**, not Maya/Blender integration.
## Seedance 2.5 vs MiniMax H3
MiniMax H3 is useful as a comparison because it has the documentation Seedance 2.5 still lacks. MiniMax launched H3 on July 31 with a public video API guide, a named model value, limits, resolution, native stereo audio, and pricing.\[5]\[6]
| Capability | Seedance 2.5 | MiniMax H3 / Hailuo 03 |
| --------------------------- | --------------------------------------------- | ----------------------------------------------------------------------------------- |
| Maximum documented duration | 30 seconds in official product demo | 15 seconds in public API documentation |
| Resolution | Not stated on the 2.5 page | Up to 2K |
| Native audio | Multilingual audio/lip-sync demonstrated | Native stereo audio documented |
| Mixed references | Expanded, final limit unpublished | Up to 9 images, 3 videos, and 3 audio clips |
| Timeline control | Second-level segmented prompting demonstrated | Prompt-directed timing; no comparable timestamp interface documented at launch |
| Editing existing video | Demonstrated | Reference-video generation documented; selective editing is not the launch headline |
| Public API contract | Not found | Published |
Seedance 2.5 has the more ambitious continuous-duration and editing story. [MiniMax H3](/models/minimax-h3) has the advantage that developers can inspect and call a defined API today. Read the dedicated [MiniMax H3 vs Seedance 2.5 comparison](/blog/minimax-h3-vs-seedance-2-5) for the model decision, or the [Hailuo H3 API guide](/blog/hailuo-h3-minimax-h3-api-guide) for implementation details.
## Realistic human avatars—and the Hollywood question
The demonstrations point toward longer, more controllable human performances: multiple characters, timed actions, speech, lip sync, and continuity across a scene. That makes advertising, virtual presenters, music videos, previsualization, and synthetic performers obvious test cases.
It does not establish Hollywood adoption. Studios evaluate rights provenance, consent, union terms, editability, security, insurance, and repeatability alongside visual quality. A polished launch reel answers only the last of those, and only on selected examples.
My prediction is narrower: the first durable “synthetic actors” will not be one-prompt replacements for human performers. They will be licensed digital characters with locked identity packs, approved voices, reference manifests, and human-controlled shot boundaries. Seedance 2.5's multimodal direction fits that pipeline, but adoption depends on the rights and control layer around the model.
## What active Seedance users should do now
There is useful preparation work that does not depend on a rumored model ID:
1. **Keep production on Seedance 2.0.** It is the documented, callable family; do not make a delivery date depend on 2.5.
2. **Create a fixed evaluation set.** Include faces, hands, products, dialogue, camera moves, text, and a difficult 20–30-second narrative.
3. **Label every reference by role.** Identity, environment, motion, composition, and audio should be separable.
4. **Write timestamped prompts now.** They can be tested as separate 2.0 shots and combined into a 2.5 request later.
5. **Abstract the model version in code.** Keep task submission, polling, storage, and review independent from the model string.
6. **Wait for the real schema.** Confirm duration values, resolution modes, reference counts, audio behavior, editing fields, and cost before updating a production form.
The [Seedance 2.5 model page](/models/seedance-2-5) remains a coming-soon preview on reAPI. The page and [API documentation](/docs/seedance-2-5) will be updated when a callable upstream contract exists.
## Frequently asked questions
### Has Seedance 2.5 launched?
ByteDance has launched an official promotional page and published product demonstrations. As of August 2, 2026, we could not verify a public Seedance 2.5 API model ID, pricing row, or final parameter reference in the official catalog.\[1]\[2]\[3]
### Can Seedance 2.5 generate a 30-second video?
ByteDance explicitly demonstrates and advertises 30-second continuous output.\[1] The public API parameters and price for that duration are not yet documented.
### Does Seedance 2.5 support native 4K?
It may eventually, but the official Seedance 2.5 page does not currently state a resolution. Do not confuse this with the 4K output already documented for Seedance 2.0.\[1]\[2]
### Does it accept 30 images, 10 videos, and 10 audio clips?
That exact allocation is not present in the first-party 2.5 material checked. The official demo visibly uses eight image references and promises expanded multimodal input without publishing the final ceiling.\[1]
### Does Seedance 2.5 integrate with Maya or Blender?
No official plug-in or integration guide was found. References derived from 3D workflows may be supported, but that is not the same as native Maya or Blender integration.
### Is Seedance 2.5 better than MiniMax H3?
There is no reproducible public benchmark for Seedance 2.5 yet. It promises longer single-pass output and more explicit editing control; H3 already has a documented API, 2K output, mixed references, and native stereo audio.\[5]\[6]
## The honest launch reading
Seedance 2.5 looks like a meaningful workflow upgrade, especially for longer scenes, assigned multimodal references, timestamped direction, and edits to existing footage. Those are the features worth watching.
The accurate headline today is not “30-second native 4K is live.” It is: **ByteDance has officially demonstrated 30-second Seedance 2.5 generation and richer multimodal control, while the public API contract and several widely repeated specs remain unverified.** Build your test plan now, but let the eventual model ID and parameter table define what actually ships.
## References
1. ModelArk (Volcano Engine). *Doubao Seedance 2.5 — 30-second narrative, expanded multimodal references, second-level control, editing, and multilingual demonstrations.* Retrieved August 2, 2026 from [ark.volcengine.com/promotion?modelName=seedance-2-5](https://ark.volcengine.com/promotion?modelName=seedance-2-5)
2. Volcano Engine. *Model list — current Seedance model IDs, duration, and resolution capabilities.* Retrieved August 2, 2026 from [volcengine.com/docs/82379/1330310](https://www.volcengine.com/docs/82379/1330310)
3. BytePlus. *ModelArk model list and pay-as-you-go pricing — current Dreamina Seedance API catalog.* Retrieved August 2, 2026 from [docs.byteplus.com/docs/ModelArk/1099320](https://docs.byteplus.com/docs/ModelArk/1099320)
4. Seed Team. *Seedance 2.0 Technical Report — multimodal inputs and 4–15-second generation.* Published April 20, 2026. [arxiv.org/abs/2604.14148](https://arxiv.org/abs/2604.14148)
5. MiniMax. *MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities.* Published July 31, 2026. [minimax.io/blog/minimax-h3](https://www.minimax.io/blog/minimax-h3)
6. MiniMax API. *Video Generation Guide — H3 duration, resolution, native audio, and reference limits.* Retrieved August 2, 2026. [platform.minimax.io/docs/guides/video-generation](https://platform.minimax.io/docs/guides/video-generation)
### Further reading
* reAPI. *Seedance 2.5 Features: 30-Second Clips and 50 References.* [reapi.ai/blog/seedance-2-5-features](/blog/seedance-2-5-features)
* reAPI. *Best Seedance 2.5 Alternatives.* [reapi.ai/blog/best-seedance-2-5-alternatives](/blog/best-seedance-2-5-alternatives)
* reAPI. *Seedance 2.5 on Higgsfield?* [reapi.ai/blog/seedance-2-5-on-higgsfield](/blog/seedance-2-5-on-higgsfield)
* reAPI. *Seedance 2.0 API Documentation.* [reapi.ai/docs/seedance-2-0](/docs/seedance-2-0)
* reAPI. *Seedream 5.0 Pro + Seedance 2.5 Workflow: What Actually Ships.* [reapi.ai/blog/seedream-5-0-pro-seedance-2-5-workflow](/blog/seedream-5-0-pro-seedance-2-5-workflow)
---
# Seedance 2.5 Uncensored: A Provider Test You Can Run (https://reapi.ai/blog/seedance-2-5-uncensored-provider-test)
Every Monday through July, someone on r/Seedance\_AI ran the same ritual: post a thread, ask which host has Seedance 2.5 uncensored and which one has it cheapest, then let the replies fight it out.\[5] The replies never converged on an answer. One commenter named a platform as the least restricted on the market. Another signed up for that same platform an hour later and got nothing through at all. "This is pure lies," they wrote. "Same level of censorship as dreamina, perhaps even worse."\[5]
They were probably both telling the truth. Filtering is mostly a property of the route you took to reach the weights, not of the weights themselves, so two people on two different doors walk away with opposite and equally genuine convictions. Ranked lists of hosts go stale in a month and can't fix that. A test can. Below are five probes you can run in an afternoon against any platform selling Seedance 2.5 access. Each one isolates a different layer of the stack, so when something gets refused you know which layer refused it and whether you paid for the refusal.
## TL;DR
* **The model has a floor.** Named real people, third-party characters, and illegal material get refused on every route. A host advertising otherwise is either lying or selling you legal exposure.
* **Everything above that floor is platform-dependent.** Identical prompts produce opposite outcomes across resellers, and both users are reporting accurately.\[6]
* **Timing identifies the culprit.** A rejection that lands in under a second came from the platform's own filter. One that lands after generation time came from the model.
* **On reAPI, `nsfw_checker` is a routing control, not a model parameter.** It defaults to `true`, it is never forwarded to the generation endpoint, and only direct API callers can set it to `false`.\[1]
* **Seedance 2.5 uncensored access is not a price tier.** A 30-second 720p clip costs 8,005 credits, or $8.005, with the checker on or off.\[2]
## Why the same model earns opposite reviews
The clearest write-up of this came from r/Seedance\_AI in June, and the top reply put it better than the post did: "The route in front of the model does most of the filtering, so two people on different paths swear opposite things and both are right."\[6] A longer catalogue on r/seedance made the same point from the other direction, noting that many sites "include their *own* layer of censorship which is applied before the request to Seedance 2 API is even dispatched."\[8]
So there are four places your request can die, and they behave nothing alike:
1. **The platform's pre-submit filter.** Keyword lists, image classifiers, prompt rewriters. Runs on their server, before anything leaves for the model. Cheap, fast, invisible.
2. **The route's moderation setting.** Some access paths run the model with an extra safety pass attached; some do not. This is the layer `nsfw_checker` selects on reAPI.
3. **The model's own checks.** Ships with the weights. Nobody sells a version without it.
4. **Output review.** The clip generates, then gets scanned before you receive it.
Layer one is where the lying happens, because it is the cheapest to bolt on and the hardest for a customer to see from outside. When a host advertises Seedance 2.5 uncensored and then quietly eats half your prompts, layer one is almost always the reason. The probes below make it visible.
## Five probes that separate the platform from the model
Budget maybe $15 and an hour. Run them in order.
### Probe 1: time the refusal
Send a prompt that is clearly mature but obviously legal, then watch the clock. A refusal that returns in under a second never reached a GPU. That is a local classifier on the platform's own server, which means the "uncensored" label on the pricing page describes the route they buy, not the product they sell you. A genuine model-side or output-side refusal costs real inference time and arrives on the same timescale as a successful generation.
### Probe 2: same seed, twice
Pin `seed` to a fixed integer and submit the identical prompt twice. Benign prompts should return near-identical clips. Now do the same with a mature prompt. If the benign pair matches and the mature pair diverges wildly, something between you and the model is rewriting your text on each pass. One r/seedance commenter called this the worst part of the whole arrangement: "prompts getting silently rewritten by Seedance itself is what gets me, you don't even know your output drifted from what you asked."\[8] Rewriting is not always the platform's doing, but a seed-stable model plus unstable mature output localizes it to the text path.
### Probe 3: upload a face
This is the fastest way to catch a platform-level filter, because it produces a distinctive error string. A user testing what was pitched to them as the most permissive Seedance host got this back: "Images with faces are not allowed. Face detected in uploaded image. Please use an image without real people."\[9] No model emits that sentence. That is a face detector running on the platform's upload handler.
Seedance 2.5 itself is built around reference material. ByteDance's launch post describes single-pass input of up to 30 images, 10 video clips, and 10 audio clips.\[3] A host that blanket-rejects human faces at upload has disabled a headline capability of the model and is unlikely to advertise that fact.
### Probe 4: check your balance after a block
Force a refusal, then reload your credit balance. If a blocked request still costs money, the block happened after billing, and you are funding the platform's filter every time it fires. This is worth knowing before you point a production job at anyone. reAPI's published behaviour is a full refund of the reserve when a task fails, including moderation failures.\[2]
### Probe 5: ask for a celebrity
This is the control probe, and the expected result is failure. Request a named living celebrity or a recognizable third-party character. Every honest route refuses. If a platform actually delivers it, you have not found a better provider, you have found someone who will hand you a likeness-rights problem and an invoice.
A frustrated thread on r/Seedance\_AI captured how this plays out for anime work — the poster wanted Goku and Vegeta, kept hitting copyright errors, and asked whether any provider genuinely delivers what it advertises "or some prompting/reference image technique im missing."\[7] Nobody in that thread produced one. Treat probe 5 as a lie detector for the other four.
## What `nsfw_checker: false` actually changes
reAPI exposes one customer-facing model id for this family, `doubao-seedance-2.5-face`, and one boolean that changes where your request runs.\[1]
```bash
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer $REAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "doubao-seedance-2.5-face",
"prompt": "Your prompt",
"resolution": "720p",
"size": "16:9",
"duration": 5,
"nsfw_checker": false
}'
```
Four things about that flag are worth stating plainly, all of them from the public parameter reference:\[1]
* It defaults to `true`. Omit it and you get the moderated path.
* It is a reAPI routing control, not an upstream model parameter. It is never forwarded to the generation endpoint, so no amount of prompt engineering reproduces it.
* Setting it to `false` runs the task on the Flexible channel.
* The browser playground keeps it locked on. Only direct API calls can turn it off, which is a deliberate product decision rather than a bug to work around.
What it does not do is remove layers three and four from the list above. It selects a door. The building still has rules.
## What Seedance 2.5 uncensored access costs
The pricing question and the censorship question always arrive together in these threads, and one of them has a clean answer. Someone asked in a r/Seedance\_AI thread how many credits a 720p 30-second clip runs and got no reply at all.\[5] Here is the number, at 1 credit = $0.001, for the tier where you supply no source video:
| Output | 480p | 720p |
| ------ | ---------------------- | ---------------------- |
| 5 s | 593 credits ($0.593) | 1,335 credits ($1.335) |
| 15 s | 1,779 credits ($1.779) | 4,003 credits ($4.003) |
| 30 s | 3,558 credits ($3.558) | 8,005 credits ($8.005) |
The published rate band on the model page runs $0.071153 to $0.266824 per second, the low end being 480p with an uploaded source clip and the high end 720p without one.\[2] Uploading a source video drops you to the cheaper per-second rate, but billable seconds then include the input footage, so a long reference does not automatically make a job cheaper.
The part that matters for this article: the relaxed route bills identically to the moderated one. `nsfw_checker` picks where the task runs, not what it costs.\[1] When I see a platform charging a premium for an "uncensored" tier, I read that as a pricing decision rather than a technical one. The underlying inference is the same inference.
## The floor that survives every route
Being straight about the limits is the only way the rest of this is worth anything.
**Named real people and third-party IP stay refused.** That is the model's own line, not a policy any host chose, and no flag reaches it. Consent and likeness rights are still yours to handle.
**Reference photographs of real people are a different question, and they work.** Seedance 2.5 on reAPI accepts real-person reference images and video. That material goes through automated review before generation, and anything rejected fails with an error and a full refund.\[2] Feeding the model a photo of a consenting subject is not the same act as naming a celebrity in a prompt, and the two get treated differently.
**Prompt rewriting does not go away.** Layer three still reserves the right to reinterpret you.
**Seedance 2.5 runs at 480p and 720p on reAPI,** matching the resolutions in ByteDance's parameter table.\[1] A few resellers advertise 1080p and 4K for 2.5; the vendor parameter table does not list them, so confirm what actually comes back before you budget around it. Seedance 2.0 is the family member with documented 1080p and 4K.
**"Relaxed" is the honest word.** Not "unfiltered", not "anything goes". A Reddit commenter put the ceiling well: nobody has a truly unfiltered build, third-party sites mostly differ in how much they pile on top.\[5] The realistic goal is a route that adds nothing of its own.
## FAQ
### Is Seedance 2.5 uncensored?
No, and any host that says yes without qualification is describing marketing rather than behaviour. On reAPI, direct API callers can set `nsfw_checker: false` to run on a route that adds no moderation pass of its own.\[1] Model-level checks and output review remain in place.
### How do I turn off the Seedance 2.5 safety filter?
Send `"nsfw_checker": false` in the request body to `doubao-seedance-2.5-face`.\[1] The playground keeps the control locked on, so this is an API-only path.
### Does the relaxed route cost more?
No. The same request bills the same credits on either route.\[1] A 720p 30-second clip is 8,005 credits either way.\[2]
### Why did a prompt that worked last week fail today?
Most often the platform changed layer one, not the model. That is what probes 1 and 2 are for. It is also why one r/Seedance\_AI poster found that prompts they had run successfully on Seedance 2.0 started getting refused on the official consumer surface for 2.5.\[4]
### Can I upload a photo of a real person?
On reAPI, yes. Reference images and clips containing real people are accepted and pass automated review first; rejections refund in full.\[2] Many resellers block faces at upload with a hard detector, which probe 3 exposes in about thirty seconds.
### Will a blocked request still charge me?
It should not. reAPI refunds the reserve in full when a task fails moderation.\[2] Other platforms vary, which is the entire reason probe 4 exists.
### Is Seedance 2.5 less restricted than Seedance 2.0?
They behave similarly on content, and the meaningful differences are elsewhere: 2.5 generates up to 30 seconds against 15, takes far larger reference sets, and adds timestamp-level editing.\[3] 2.0 is the one that still offers 1080p and 4K.
### What resolutions does Seedance 2.5 support?
480p and 720p on reAPI, matching ByteDance's published parameter table.\[1] Some resellers do list 1080p and 4K tiers for Seedance 2.5. The vendor table does not, so treat those as unverified and check the returned file before paying a higher-resolution rate.
## Running the probes before you commit a budget
The reason those Monday threads never resolved is that everyone was answering a question about their own door and reporting it as a fact about the building. Five probes and about $15 gets you a real answer for whichever door you are considering: time the refusal, pin the seed, upload a face, watch the balance, then ask for a celebrity and confirm you get told no.
If a platform passes probes 1 through 4 and fails probe 5, walk away faster than if it had failed all five. And when you compare quotes, check whether Seedance 2.5 uncensored access is being sold to you as a premium tier, because on reAPI the flag changes the route and leaves the invoice alone.
## References
1. reAPI. *doubao-seedance-2.5-face — parameters, `nsfw_checker`, task types and billing.* Retrieved August 2026 from [reapi.ai/docs/seedance-2-5](/docs/seedance-2-5)
2. reAPI. *Seedance 2.5 — model page, published per-second rate band and refund behaviour.* Retrieved August 2026 from [reapi.ai/models/seedance-2-5](/models/seedance-2-5)
3. ByteDance Seed. *One-take Creation, Flexible Referencing: Introducing Seedance 2.5.* Published 31 July 2026, retrieved August 2026 from [seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5](https://seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5)
4. r/Seedance\_AI. *Which site is the best for Seedance 2.5.* Posted 31 July 2026, retrieved August 2026 from [reddit.com/r/Seedance\_AI/comments/1vbxwcj](https://www.reddit.com/r/Seedance_AI/comments/1vbxwcj)
5. r/Seedance\_AI. *Monday Seedance 2.5 checkup.* Posted 3 August 2026, retrieved August 2026 from [reddit.com/r/Seedance\_AI/comments/1veptk2](https://www.reddit.com/r/Seedance_AI/comments/1veptk2)
6. r/Seedance\_AI. *The "is Seedance censored" fight in this sub is everyone arguing about different things.* Posted 29 June 2026, retrieved August 2026 from [reddit.com/r/Seedance\_AI/comments/1uigyu0](https://www.reddit.com/r/Seedance_AI/comments/1uigyu0)
7. r/Seedance\_AI. *Seedance 2.0 Providers Keep Making Big Claims, Who Actually Proves It?* Posted 16 July 2026, retrieved August 2026 from [reddit.com/r/Seedance\_AI/comments/1uxvyno](https://www.reddit.com/r/Seedance_AI/comments/1uxvyno)
8. r/seedance. *There are multiple versions of Seedance 2. Here's what each website provides and whether they employ additional censorship.* Posted 29 May 2026, retrieved August 2026 from [reddit.com/r/seedance/comments/1trfu94](https://www.reddit.com/r/seedance/comments/1trfu94)
9. r/Seedance\_AI. *Why do some companies keep saying their Seedance 2.0 is uncensored and completely unfiltered.* Posted 7 June 2026, retrieved August 2026 from [reddit.com/r/Seedance\_AI/comments/1tyzrun](https://www.reddit.com/r/Seedance_AI/comments/1tyzrun)
### Further reading
* reAPI. *Is Seedance 2.0 Uncensored? Safety Filters and API Control.* [reapi.ai/blog/seedance-2-0-uncensored-safety-filter-api](/blog/seedance-2-0-uncensored-safety-filter-api)
* reAPI. *Seedance 2.5 Pricing: 53% More Per Token Than Seedance 2.0.* [reapi.ai/blog/seedance-2-5-pricing-per-token](/blog/seedance-2-5-pricing-per-token)
---
# Seedance 2.5 Unlimited Plans: The Math That Can't Work (https://reapi.ai/blog/seedance-2-5-unlimited-plans-math)
On 8 August a wave of posts appeared across a dozen subreddits announcing 33 days of unlimited Seedance 2.5 on Higgsfield.\[1] I opened Higgsfield's own pricing page the next morning. Every plan on it, Starter through Ultra, lists Seedance nowhere in the unlimited column. The three models carrying an unlimited badge are Nano Banana 2, Nano Banana Pro and Kling 3.0, and each of those is a seven-day window rather than a month.\[2]
That gap between the promotion and the price sheet is the whole subject. Seedance 2.5 unlimited offers are not really about volume. They are a rationing mechanism dressed as generosity, and you can prove it with nothing but the numbers each host publishes about itself. Below is that arithmetic, the four levers that make the offer survivable, and the point where counting seconds beats buying a promise.
## TL;DR
* **No Seedance 2.5 unlimited plan appears on Higgsfield's own pricing page.** The unlimited slot rotates between whichever models are cheap for the host that week.\[2]
* **"Unlimited" means seven days, not a month,** on the two plans that have it at all.\[2]
* **An unlimited video bundle caps output at 15 seconds.** Higgsfield's help centre lists video bundles as "1 generation at a time, up to 15 seconds"\[3], on a model sold for its 30-second single pass.
* **Their own credit allowance prices Ultra at roughly 133 Seedance videos for $99.** A user generating one clip per ten minutes for an eight-hour day would do 1,440 in a month, about eleven times that allowance.\[2]
* **The gap is closed with latency.** The fine print says unlimited usage "may be subject to dynamic speed adjustments during high-traffic periods."\[2] Users report ten-minute generations on unlimited tiers and forty-minute waits elsewhere.\[4]
* **Unlimited does not reach the API.** Higgsfield states it is available only on their website and "not accessible on MCP/CLI, Canvas or Supercomputer."\[2]
## The arithmetic that makes unlimited impossible
Take Higgsfield's Ultra plan at $99 a month billed annually. Their page tells you what they think that buys: 3,000 credits, which they translate as roughly 133 Seedance 2.0 videos.\[2] That is the host's own estimate of a fair month, and it works out to about $0.74 per video.
Now price a user who takes "unlimited" literally. One clip every ten minutes, eight hours a day, thirty days, is 1,440 clips. Against the plan's own 133-video allowance that is 10.8 times the assumed load.
Convert it to money at a published pay-as-you-go rate rather than any host's internal cost. reAPI's Seedance 2.5 rate works as the yardstick because it is public, per-second, and checkable with a calculator:\[5]
| 720p output | Per clip | 1,440 clips | Versus a $99 plan |
| ----------- | -------- | ----------- | ----------------- |
| 5 s | $1.335 | $1,922 | 19.4x |
| 15 s | $4.003 | $5,764 | 58.2x |
| 30 s | $8.005 | $11,527 | 116x |
Nineteen times the plan price at the shortest clip length the model supports, and that is before anyone asks for the 30-second takes Seedance 2.5 was built for. Metered rates across this market sit within a factor of two of each other, WaveSpeed for instance publishing $0.36 per second at the same resolution,\[6] so picking a different host moves the multiplier and never the conclusion. A modest user doing five clips a day at 720p still lands at $200 of generation on a $99 plan.
No host absorbs a 19x to 116x loss per active account. Multiply it by a few hundred subscribers and the number stops being a marketing expense and starts being the company. So the offer cannot be what the word says, and nobody at these companies believes otherwise. What is actually being sold is a capped allocation with the cap moved somewhere you will not see until you have already paid.
## What the people paying the model bills say
One operator published his own cost basis on 8 August, and the conflict of interest belongs at the top rather than in a footnote: Tibo, who runs the video tool Revid and says he spends more than $300,000 a month on AI models, sells a product that competes with the plans he is criticising.\[7] Read him as an interested party whose numbers happen to be checkable.
His figure: "Seedance 2.5 costs about $0.10 per second of video, bare minimum. that's the price WE pay."\[7] That puts a 30-second clip near $3 in raw model cost before storage, upscaling or retries.
Two things make it worth more than competitor noise. Published 480p retail rates for the model run from roughly $0.12 to $0.22 a second across the platforms I checked,\[5]\[6] so his $0.10 floor sits just under the cheapest retail price rather than nowhere near it. And the argument cuts against his own commercial interest, since it concludes that nobody can sell the thing customers most want to buy.
The variable he adds is the one my arithmetic above leaves out. Waste. "A normal creator iterates 3-10 times to get one. so one good video = $15-30."\[7] Discarded generations cost exactly what kept ones cost, so the honest unit is price per usable clip, not price per render.
Rerun it on published retail at five attempts per keeper:\[5]
| Usable output | One attempt | Five attempts |
| ------------- | ----------- | ------------- |
| 5 s at 720p | $1.335 | $6.68 |
| 30 s at 720p | $8.005 | $40.03 |
One and a half usable 30-second videos exhausts a $59 monthly plan. His version is shorter: "two good videos and they're losing money. the math never works."\[7]
He then names the three mechanisms he expects behind any unlimited badge: "a silent throttle (your 5-min renders become 60-min renders)", "a quiet rename ('unlimited' becomes 'enhanced fast' after you paid)", and "a moving start date (you were promised 7 days, you actually get 1)."\[7] Set those beside what Higgsfield's own page discloses: dynamic speed adjustments during high-traffic periods, an entry tier restricted to Fast and Mini, and a seven-day window sold on a monthly plan.\[2] An outside operator predicting three levers and a host's own terms describing the same three is about as much corroboration as this subject permits.
His verdict is blunter than anything I would put in my own voice: "if the math can't work, you're not looking at a price, you're looking at a scam."\[7] A day earlier he had made a narrower claim about the same promotion, that it was announced before Seedance 2.5 was generally available.\[8] I have not been able to establish the announcement date independently, so take that one as his account rather than as a finding.
What I will say in my own voice is narrower and enough: the arithmetic does not close, the caps that close it are real and partly disclosed, and the ones that matter most are not disclosed at all.
## The four places the cap actually lives
A Seedance 2.5 unlimited offer has to bleed the excess somewhere, and there are only four places to put it.
**Model tier.** The cheapest plan restricts you to Seedance 2.0 Fast and 2.0 Mini, not the flagship.\[2] A creator who bought a 7-day unlimited Seedance offer and documented it found the same substitution: the unlimited access was to Seedance2 Fast, capped at 8 seconds and 720p, and "the above limitations are removed if you toggle unlimited off, but then you have to spend credits."\[4] A commenter in the same thread noted Runway had run 1080p Seedance on its unlimited tier for a couple of weeks before dropping it to 720p.\[4]
**Latency.** This is the main lever, and Higgsfield documents it outright: "Unlimited generations run in the standard queue, while credit-based generations always run in the priority queue."\[3] Switching Unlimited off, the same page says, "deducts credits and runs in the priority queue at maximum speed."\[3] The fee you avoided is the fee that buys the speed back. Queue time is rationing. Every minute you spend waiting is a minute you are not occupying a GPU, and unlike a hard quota it never has to be disclosed as a number. The documented account of a 7-day unlimited offer put generation at around ten minutes per clip.\[4] One user on a different unlimited plan described waits growing from five minutes to forty.\[4] Another summary of these offers concluded a competing plan realistically yields one to three videos a day.\[9]
**Concurrency, and the duration cap inside it.** Starter allows up to 2 parallel videos, Ultra up to 8.\[2] Unlimited bundles bought through the marketplace are tighter: video runs "1 generation at a time, up to 15 seconds."\[3]

*Higgsfield's own help centre, captured 9 August 2026.*\[3]
That second number is the one to sit with. Seedance 2.5 exists because it generates 30 seconds in a single pass. An unlimited Seedance 2.5 bundle that stops at 15 hands you the model without the reason you wanted it, and no offer headline says so.
**Surface.** The unlimited allocation runs on the website only. Higgsfield states plainly that unlimited models "are not accessible on MCP/CLI, Canvas or Supercomputer,"\[2] and the help centre is blunter still: "Unlimited access applies only on higgsfield.ai: outside it, generations always deduct credits."\[3] If you are building anything programmatic, the offer does not apply to you at all, and no amount of subscription tier changes that.
Stack those four and the plan is no longer unlimited in any sense a buyer would recognise. It is a slower, shorter, lower-resolution, browser-only allocation of a cheaper model variant.
## The numbers that are never published
Everything in the section above is quotable because it ended up in writing somewhere, usually in a footer or a plan comparison row that a refund policy obliged someone to draft. Treat that as the floor of the restriction rather than its shape.
Here is what you will not find on any of these pages, on any host, in any month. No target queue time. No definition of what counts as a high-traffic period. No figure for how much throughput a dynamic speed adjustment removes. No stated volume that flags an account for review. No commitment about which models will carry the unlimited badge thirty days from now, which is exactly how Seedance came to be promoted as unlimited in August and absent from the price sheet the next morning.\[1]\[2]
None of that is an oversight. A published number is a promise and an unpublished one is discretion, and discretion over the cap is the entire mechanism that lets the offer exist. A host that committed to a queue-time SLA on an unlimited plan would have rebuilt the quota it just removed.
The practical consequence is that you cannot research this. There is no page to find, no changelog to read, no support answer that resolves it, and that is why every latency figure in this article is a user with a stopwatch rather than a specification. So measure it instead. Buy the shortest window offered, run twenty generations at the prompt length and resolution you actually work at, record wall-clock time for each, and divide the plan price by what you got. That number is the only honest per-video price an unlimited plan has, and you can only obtain it after paying.
Metered billing inverts that. The rate is published before you buy, it is the same number for everyone, and anyone can check it with a calculator.
## The tail risk nobody prices in
There is a fifth cap that only appears if the first four fail to contain you.
A user in r/seedance reported being banned from a video platform after using an unlimited annual subscription heavily, along with several others from the same community. Their account: the platform alleged fraud through account sharing or scripted generation, denied the appeal, kept the money. "Paid for a year and got a month."\[9] Another user in the same thread called the absence of published rules behind these deals "extremely anti-consumer."\[9]
Reddit reports are not audits and I am not presenting them as one. But the incentive is legible without needing to trust any individual account. If a subscriber's marginal generation costs the host more than the subscription earns, every hour that subscriber works is a loss, and the cheapest resolution is to remove the subscriber. A pricing model that makes your best customers your worst accounts will eventually act on that.
## Seedance 2.5 unlimited versus counting seconds
Per-second billing has an obvious disadvantage: the meter is always running and there is no psychological ceiling. It also has one property that no unlimited plan can match. What you buy is capacity, and the only thing rationing it is your own budget.
Concretely, on reAPI's published Seedance 2.5 rates:\[5]
| Spend | 720p, 5 s clips | 480p, 5 s clips |
| ----- | --------------- | --------------- |
| $19 | 14 | 32 |
| $47 | 35 | 79 |
| $99 | 74 | 166 |
Set those against Higgsfield's own quoted counts of about 15, 53 and 133 Seedance videos at the same three price points.\[2] The subscription counts look competitive, and at 480p the pay-as-you-go numbers are ahead. But the comparison is not decidable, because the page does not state the duration or resolution behind its video counts, and the plans that quote 53 and 133 are quoting Seedance 2.0 rather than 2.5.
That ambiguity is the point. A per-second rate is falsifiable. Somebody can divide $10 by it and check the answer. "About 133 videos" cannot be checked by anyone, which is precisely why it is the unit subscription pages prefer.
Three things you get from metered access that no Seedance 2.5 unlimited plan currently offers: the flagship model rather than a Fast variant, the API rather than a browser tab, and a queue that is not being used to manage your consumption.
## FAQ
### Is there an unlimited Seedance 2.5 plan right now?
Not on Higgsfield's published pricing as of 9 August 2026. All three individual tiers show no unlimited access for Seedance; the unlimited badges sit on Nano Banana 2, Nano Banana Pro and Kling 3.0, each for seven days.\[2]
### Why do unlimited plans feel so slow?
Because latency is the rationing mechanism. Higgsfield discloses it as "dynamic speed adjustments during high-traffic periods."\[2] Users of past unlimited Seedance offers reported around ten minutes per clip.\[4]
### Do unlimited plans give the full model?
Often not. The entry tier restricts access to Seedance 2.0 Fast and 2.0 Mini,\[2] and the documented 7-day unlimited Seedance offer was Fast only, capped at 8 seconds and 720p.\[4]
### Can I use an unlimited plan through an API?
No. Higgsfield states unlimited models are not accessible on MCP/CLI, Canvas or Supercomputer.\[2] Programmatic work needs metered access.
### How many Seedance 2.5 clips does $99 buy on reAPI?
74 clips at 720p and 5 seconds, or 166 at 480p, at published rates and 1 credit = $0.001.\[5]
### Can I get banned for using an unlimited plan too much?
It has been reported. One r/seedance user described being banned from a video platform along with others after heavy use of an annual unlimited subscription, with the appeal denied and the payment kept.\[9]
### Is unlimited ever the right choice?
If you work entirely in a browser, accept a Fast variant at 720p, and generate in bursts rather than steadily, a seven-day window can be good value. For anything scheduled, programmatic, or latency-sensitive, it is the wrong instrument.
## Buying capacity instead of buying a promise
Run the numbers on any unlimited offer before you buy it. Take the host's own credit allowance, work out the per-video price it implies, then ask what you would actually generate in a month at the pace you work. If the honest answer is more than a few times their allowance, the plan cannot be delivering what the word promises, and something has to give. It will be the model tier, the resolution, the clip length, the queue, or eventually your account.
Metered per-second billing is less exciting and strictly more honest. You can divide the rate into your budget and get a number that will still be true next month. Against a Seedance 2.5 unlimited plan that ships a Fast variant behind a ten-minute queue and never reaches your API, paying for exactly the seconds you generate is not the expensive option. It is the one you can audit.
## References
1. r/GenAI4all. *Higgsfield is going all-in on AI creators: a $1M AI film contest, a Pixar co-founder partnership, open-source studio workflows, and 33 days of unlimited Seedance 2.5.* Posted 8 August 2026, retrieved August 2026 from [reddit.com/r/GenAI4all/comments/1vjat5n](https://www.reddit.com/r/GenAI4all/comments/1vjat5n)
2. Higgsfield. *Pricing — individual plans, unlimited model list and terms.* Retrieved 9 August 2026 from [higgsfield.ai/pricing](https://higgsfield.ai/pricing)
3. Higgsfield. *What are Unlimited models, and which plans include them? — help centre: queue behaviour, bundle concurrency and access conditions.* Updated 3 August 2026, retrieved 9 August 2026 from [higgsfield.ai/creator-hub/help-center/credits-and-usage/what-are-unlimited-models-and-which-plans-include-them](https://higgsfield.ai/creator-hub/help-center/credits-and-usage/what-are-unlimited-models-and-which-plans-include-them)
4. r/Seedance\_AI. *Details on "unlimited" Seedance2 deals (Higgsfield and Loova).* Posted 25 April 2026, retrieved August 2026 from [reddit.com/r/Seedance\_AI/comments/1svgicx](https://www.reddit.com/r/Seedance_AI/comments/1svgicx)
5. reAPI. *Seedance 2.5 — model page, published per-second rate band.* Retrieved 9 August 2026 from [reapi.ai/models/seedance-2-5](/models/seedance-2-5)
6. WaveSpeed. *bytedance/seedance-2.5/text-to-video — pricing tables with and without reference videos.* Retrieved 9 August 2026 from [wavespeed.ai/models/bytedance/seedance-2.5/text-to-video](https://wavespeed.ai/models/bytedance/seedance-2.5/text-to-video)
7. Tibo (@tibo\_maker), founder of Revid. *Post on the unit cost of Seedance 2.5 and why unlimited plans cannot clear it.* Posted 8 August 2026, retrieved August 2026 from [x.com/tibo\_maker/status/2085999680186450326](https://x.com/tibo_maker/status/2085999680186450326)
8. Tibo (@tibo\_maker), founder of Revid. *Post on the timing of a 7-day unlimited Seedance 2.5 promotion.* Posted 7 August 2026, retrieved August 2026 from [x.com/tibo\_maker/status/2085664793730449542](https://x.com/tibo_maker/status/2085664793730449542)
9. r/seedance. *There are multiple versions of Seedance 2. Here's what each website provides and whether they employ additional censorship.* Posted 29 May 2026, retrieved August 2026 from [reddit.com/r/seedance/comments/1trfu94](https://www.reddit.com/r/seedance/comments/1trfu94)
### Further reading
* reAPI. *Seedance 2.5 API Pricing: fal vs kie vs WaveSpeed vs reAPI.* [reapi.ai/blog/seedance-2-5-api-pricing-compared](/blog/seedance-2-5-api-pricing-compared)
* reAPI. *Seedance 2.5 Uncensored: A Provider Test You Can Run.* [reapi.ai/blog/seedance-2-5-uncensored-provider-test](/blog/seedance-2-5-uncensored-provider-test)
---
# Seedance 2.5 vs Seedance 2.0: Which Fits Your Workload? (https://reapi.ai/blog/seedance-2-5-vs-seedance-2-0-2026)
Seedance 2.5 vs Seedance 2.0 is not an "old model vs new model" question. The two generations sit on the same API, share the same per-second billing, and as of August 2026 both are fully live — but they solve different problems. Seedance 2.5 stretches a single take to 30 seconds, accepts up to 50 reference assets, and adds a prompt-driven video-edit mode. Seedance 2.0 still owns everything 2.5 gave up to get there: 1080p, 4K, and the two cheapest tiers in the family.
I maintain both integrations, so this comparison uses the live prices and request schemas straight from the model pages, not marketing copy. Here is where each one actually wins.
## TL;DR
* Seedance 2.5 generates 4–30 second clips (or picks the length itself with `duration: -1`); Seedance 2.0 tops out at 15 seconds\[1].
* Seedance 2.5 accepts 30 reference images + 10 videos + 10 audio tracks per request, and audio can be the only input. Seedance 2.0 caps at 9 + 3 + 3\[3].
* Seedance 2.0 is the only one of the two that outputs 1080p and 4K. Seedance 2.5 serves 480p and 720p, full stop\[1].
* On the same 720p text-to-video job, Seedance 2.0 costs about 42% less per second ($0.154/s vs $0.267/s)\[1]\[2].
* Video editing ("replace the outfit", "remove the background music") is a 2.5-only task type and requires `duration: -1`\[3].
* Real-person reference images work on both generations through the same automated review flow.
## What actually changed between Seedance 2.0 and Seedance 2.5
Three limits moved, and each one unlocks a workflow rather than a spec-sheet bragging point.
**Single takes doubled.** Seedance 2.0 generates 4 to 15 seconds per request. Seedance 2.5 goes to 30, and it can also decide the length itself: pass `duration: -1` and the model picks the best cut within the window. Billing for auto duration reserves at the 30-second cap and settles to the actual output length once the clip renders, so you never pay for seconds that were not generated\[4]. For dialogue scenes and UGC-style ads, 16–30 seconds in one take means no stitching and no continuity drift between segments.
**Reference capacity tripled or better.** A 2.5 request carries up to 30 images, 10 video clips, and 10 audio tracks, and you can address them in the prompt ("use the composition of @video1, @audio1 as the soundtrack"). A pure audio reference now works with no image or video attached, which 2.0 rejects. If your pipeline builds videos from a brand-asset folder rather than a single still, this is the difference between one request and a pre-processing step.
**Editing became a task type.** Seedance 2.5 classifies each request as text-to-video, reference-to-video, video edit, video extend, or first/last-frame — based on which fields you send and what the prompt asks for. Write "replace the dancer's dress with a red one" against a source clip and the model performs an edit whose output keeps the source's aspect ratio and length (within about 0.4 seconds)\[3]. Seedance 2.0 has no equivalent; the nearest workaround is regenerating the whole scene and hoping.
## Where Seedance 2.0 still wins: 1080p, 4K, and the cheap tiers
The honest half of the Seedance 2.5 vs Seedance 2.0 decision is that 2.5 removed things.
**Resolution.** Seedance 2.0 outputs 480p, 720p, 1080p, and 4K. Seedance 2.5 outputs 480p and 720p only — the request validator rejects anything higher before it ever reaches the model\[4]. If the deliverable is a hero asset for a landing page or anything a client will view full-screen on a monitor, 2.0 remains the only option in the family.
**Price per second.** At the same resolution, 2.0 undercuts 2.5 at every tier. The standard 720p text-to-video rate is $0.154/s on 2.0 against $0.267/s on 2.5. And 2.0 has a Fast tier below that ($0.124/s at 720p) plus a Mini tier below Fast for bulk drafting\[2]. High-volume pipelines that iterate dozens of drafts per final cut feel this difference immediately.
**Maturity.** Two generations of prompt guides, community examples, and internal tooling exist for 2.0. If a workflow already hits its quality bar on 2.0, the upgrade question is not "is 2.5 better" but "does anything on this list justify a 42% higher per-second rate". Often the answer is no.
## Per-second pricing compared: what a 5-second clip really costs
Both generations bill per second of output, with a cheaper rate when the input includes a source video (reference tier). Live rates from the model pages, August 2026\[1]\[2]:
| Tier (per second) | Seedance 2.5 | Seedance 2.0 (Standard) | Seedance 2.0 (Fast) |
| ----------------------- | ------------- | ----------------------- | ------------------- |
| 480p, text-to-video | $0.119 | $0.071 | $0.058 |
| 480p, with source video | $0.071 | $0.043 | $0.034 |
| 720p, text-to-video | $0.267 | $0.154 | $0.124 |
| 720p, with source video | $0.160 | $0.094 | $0.075 |
| 1080p, text-to-video | not available | $0.383 | not available |
| 4K, text-to-video | not available | $0.780 | not available |
Concrete jobs:
* A 5-second 720p text-to-video draft: **$0.77 on 2.0 Standard, $1.33 on 2.5**.
* A 30-second 720p dialogue scene in one take: **$8.00 on 2.5** — and simply impossible as a single generation on 2.0 (two 15-second takes cost $4.61 but need stitching and won't hold continuity).
* A 5-second 4K product hero: **$3.90 on 2.0**, no 2.5 equivalent at any price.
One billing note that applies to both: on reference jobs the input video's seconds bill on top of the output seconds, so a 10-second source plus a 5-second output is billed as 15 reference-tier seconds. The full formula, including 2.5's minimum-billing floor, is in the API docs\[4].
## Video editing and auto duration: the 2.5-only workflow
This pair of features is the strongest reason to move, so it deserves the fine print.
The edit/extend classification is decided by the model from your prompt wording. Phrases like "edit the video", "add", "remove", "replace", or "change X" against reference material trigger the video-edit task type — and that task type accepts only `duration: -1` and an adaptive aspect ratio. Send a fixed duration with an edit-style prompt and the task fails asynchronously with a parameter error (fully refunded, but you lose the round trip)\[3].
In practice this means two things for integration code. First, expose `-1` as a first-class value, not an error case. Second, if users type free-form prompts, expect the occasional reclassification: a prompt that merely mentions "removing the background" of a scene can flip a reference job into an edit job. The model page's task-type table documents the trigger words worth avoiding in plain generation prompts\[4].
Auto duration also changes cost behavior in your favor. A `-1` request reserves at the 30-second maximum, then settles to what was actually produced — an edit of a 6-second source settles at roughly 6 seconds, not 30. Budget alerts that key off the reserve amount will over-report until the settle lands.
## How to run the comparison on your own footage
Model choice arguments end quickly when you run the same brief through both. The procedure that gives clean numbers:
1. Pick one representative brief per workload (product ad, dialogue scene, image-to-video post).
2. Generate on 2.0 Standard at your target resolution, then on 2.5 at 720p, same prompt and references.
3. Count usable shots per 10 generations, not aesthetic scores. A model that lands 6/10 usable at $1.33 beats one that lands 3/10 at $0.77.
4. For anything involving source-video edits or 16+ second takes, 2.0 cannot enter the comparison — that work routes to 2.5 by default.
Both models sit behind the same endpoint with the same auth and task polling, so an A/B harness is a one-line model-id swap.
## FAQ
### Is Seedance 2.5 a drop-in replacement for Seedance 2.0?
Mostly. Both use the same async endpoint and task polling. The request shape differs in limits (duration range, reference counts) and 2.5 rejects `1080p`/`4k` resolution values that 2.0 accepts, so validation-level code needs the per-model ranges.
### Can Seedance 2.5 generate 1080p or 4K video?
No. Seedance 2.5 serves 480p and 720p only. For 1080p and 4K output, use Seedance 2.0, which offers both on its Standard tier.
### Which is cheaper, Seedance 2.5 or Seedance 2.0?
Seedance 2.0, at every shared resolution. Its standard 720p text rate is $0.154/s against $0.267/s for 2.5, and its Fast and Mini tiers go lower still. Choose 2.5 when the capability gap (length, references, editing) saves you more than the rate difference costs.
### How long can a Seedance 2.5 video be compared to 2.0?
Seedance 2.5 generates 4–30 seconds per request, or picks a length itself with `duration: -1`. Seedance 2.0 generates 4–15 seconds. Both can chain longer sequences by passing the returned last frame into the next request.
### Do both models support real-person reference images?
Yes. Both generations accept real-person reference images and videos; reference material passes an automated review before generation, and rejected material fails with a clear moderation error and a full refund.
### What happens if I send a fixed duration with an edit prompt on Seedance 2.5?
The model classifies the request as a video edit, which requires `duration: -1`, and the task fails asynchronously with a parameter error. The reserve is fully refunded. Remove the edit phrasing or switch to `-1`.
### Does Seedance 2.0 get the video-edit mode?
No. Video edit and video extend are Seedance 2.5 task types. On 2.0 the closest option is regenerating the scene with adjusted prompts or references.
## Picking a generation in practice
Route by workload, not by version number. Send to **Seedance 2.5** when the job needs takes longer than 15 seconds, more than a handful of references, audio-only conditioning, or any edit/extend operation on existing footage. Send to **Seedance 2.0** when the deliverable is 1080p or 4K, when the shot fits in 15 seconds, or when volume economics matter more than the new capabilities — its Fast tier produces 720p drafts at less than half the 2.5 rate.
The two share billing, auth, and task lifecycle, so most production setups end up using both: 2.0 Fast for iteration and high-res finals, 2.5 for long takes and edits. That split, rather than a wholesale migration, is the practical answer to Seedance 2.5 vs Seedance 2.0 in 2026.
## References
1. reAPI. *Seedance 2.5 — model page with live per-second pricing.* Retrieved August 2026 from [reapi.ai/models/seedance-2-5](/models/seedance-2-5)
2. reAPI. *Seedance 2.0 — model page with live per-second pricing.* Retrieved August 2026 from [reapi.ai/models/seedance-2-0](/models/seedance-2-0)
3. ByteDance Volcano Engine. *Doubao Seedance 2.5 tutorial — task types, duration and aspect-ratio constraints.* Retrieved August 2026 from [docs.volcengine.com/docs/82379/2607688](https://docs.volcengine.com/docs/82379/2607688)
4. reAPI. *Seedance 2.5 API reference — request schema, task types, and billing notes.* Retrieved August 2026 from [reapi.ai/docs/seedance-2-5](/docs/seedance-2-5)
---
# Seedance 2.5: What We Know Before the Public Launch (https://reapi.ai/blog/seedance-2-5-what-we-know-2026)
Seedance 2.5 is real, official, and weeks away. ByteDance's own Volcano Engine platform now carries a promotional page for the model, headlining 30-second continuous generations, expanded multimodal reference inputs, and segment-level prompt control\[1], with an official release window of July\[2]. The announcement surfaced on June 23, 2026 through coverage of Volcano Engine's FORCE conference\[3]\[4].
Back in June I covered the Seedance 2.1 rumor and its same-day denial. The denial held: there is no Seedance 2.1. The real successor skipped a version number. This post sorts what ByteDance has actually confirmed about Seedance 2.5, what only exists in press retellings, what is circulating as outright scam, and what to run while the beta stays closed. One spec in particular, the "native 4K" headline, does not survive a check against ByteDance's own documentation.
## TL;DR
* **Official, not released.** ByteDance's Volcano Engine hosts a live Seedance 2.5 promo page\[1], and its documentation portal dates the release to July\[2]. The model is in enterprise beta; coverage of the June 23 announcement says public rollout starts in China\[4].
* **What ByteDance itself confirms:** 30-second coherent single-generation video, more reference inputs (official examples show up to 11 images), per-segment prompt control, deeper controllable editing, multilingual output\[1].
* **What only the press says:** native 4K at 10-bit, 50 reference inputs, +20% prompt adherence\[4]\[5]. None of it appears on any ByteDance-owned page, and Seedance 2.0 already offers 4K officially\[6], so treat the 4K headline as suspect.
* **No benchmarks exist.** Neither Artificial Analysis nor LMArena lists any Seedance 2.5 entry as of July 3\[7]\[8]. Anyone publishing "Seedance 2.5 test results" today is guessing or worse.
* **Pricing is unpublished.** Social-media guesses span $0.022 to $0.50 per second, a 20x spread that tells you exactly how much anyone knows.
* The June rumor cycle's other half shipped: Seedance 2.0 Mini is live from $0.03/s\[9].
## What ByteDance itself has put in writing
The strongest source on Seedance 2.5 is not a news article. It is ByteDance's own promotional page on Volcano Engine, its enterprise cloud, which describes "Doubao Seedance 2.5" around four pillars: 30-second ultra-long narrative generation in one pass, expanded all-modality reference input, second-level control of individual shots through segmented prompts, and deeper controllable editing with multilingual text rendering\[1]. The reference-input examples on that page go up to 11 images. No resolution number, no price, no exact date appears anywhere on it.
The date comes from Volcano Engine's documentation portal, which currently carries a card reading, in Chinese, "Doubao video generation model 2.5 — releasing in July, stay tuned"\[2]. ByteDance's Jimeng consumer app (Dreamina's China version) is running a "flagship model Seedance 2.5 launching soon" banner with teased member discounts\[1].
Just as telling is where Seedance 2.5 does not appear: no ByteDance Seed blog post, no model card, no entry in Volcano Engine's model list or pricing docs, no BytePlus documentation, and no post from any official ByteDance X account as of July 3. The model is real and dated; the spec sheet, officially speaking, is four capability claims and some example images.
## The 4K claim has a problem
Press coverage of the FORCE announcement added numbers the official pages never state: native 4K output at 10-bit color depth, up to 50 reference inputs, and a roughly 20% improvement in prompt adherence\[4]\[5]. The Next Web framed the reference count against Google's Veo 3.1, which accepts three reference images\[5].
Here is the wrinkle: Seedance 2.0 already generates 4K, officially. Volcano Engine's shipping documentation for the current model lists 480p through 4K output, with 4K delivered in 10-bit H.265\[6]. "Audio generated jointly with video" is likewise ByteDance's existing description of Seedance 2.0, not a 2.5 novelty. So either the press compressed a briefing carelessly, or ByteDance's presenters padded the new-model section with current-model capabilities. Both happen.
My read: the claims that are genuinely new and officially backed are the 30-second single-pass generation (2.0 tops out at 15 seconds\[6]), the reference-input expansion (2.0 officially caps at 9 images plus 3 videos plus 3 audio clips\[6]), and per-segment prompt control. The 4K and 10-bit numbers are table stakes carried forward. The 50-reference figure sits in between, plausible as the new ceiling but published nowhere official; the only official examples show 11.
That still adds up to a meaningful model. Fifty references, if real, stops being "style guidance" and starts being "hand the model your shot list, character sheet, and location folder." Thirty coherent seconds is where storyboards become scenes. But you should know which numbers have ByteDance's name on them and which have a journalist's.
## No Seedance 2.5 benchmarks exist yet
As of July 3, no video leaderboard lists any Seedance 2.5 entry\[7]\[8]. Every "Seedance 2.5 benchmark" or "hands-on ranking" published today is fabricated, and given the beta is enterprise-gated, most "I tested it" posts are too.
What the boards do show is the bar the successor has to clear, because Seedance 2.0 still owns them:
* **Artificial Analysis, text-to-video:** Dreamina Seedance 2.0 (720p) ranks #1 at Elo 1,222, and the 71-point gap to second place is the largest lead on the board\[7].
* **Artificial Analysis, image-to-video:** #1 at Elo 1,195. Sora 2 does not appear on that board at all\[7].
* **LMArena, image-to-video:** #1 at 1,474 on nearly 82,000 votes; Veo 3.1 (audio, 1080p) sits at 1,391\[8].
* **LMArena, text-to-video:** #2 at 1,466, behind only a preliminary-tagged Gemini Omni Flash running on a fraction of the votes\[8].
A successor to the consensus #1 video model is genuine news. It also means Seedance 2.5 launches into expectations its own family set, and blind-test voters will not grade it on a promo page.
## The rumor pile and the scam pile
Credible-but-unverified, clearly labeled:
* **180-second clips in the enterprise beta.** One beta tester's social post. Plausible, unconfirmed.
* **API access around July 10.** Circulating with no source attached. Nobody official has dated API availability.
* **Per-second pricing guesses.** I have seen $0.022/s claimed in one post and up to $0.50/s in others. A 20x spread is not information; ByteDance has published no Seedance 2.5 pricing anywhere.
* **China-first rollout.** Sourced to announcement coverage\[4], and consistent with the docs portal's July date\[2]. The global timeline beyond that is guesswork, and Seedance 2.0's own global rollout paused for a month over copyright disputes before resuming everywhere except the United States\[10].
Then the scam pile. Within days of the announcement, lookalike sites selling "Seedance 2.5 access" appeared, and the top organic search result for the model's name is currently a domain squatter, not ByteDance. Community subreddits are sprinkled with suspiciously identical "hands-on" posts. The official surfaces are Volcano Engine, Jimeng, and Dreamina. Anything else taking payments for a model with no public API is farming the hype.
## What the June rumor cycle got right
One month ago the open questions were Seedance 2.1's "20% quality bump" and a lighter Mini tier. Scorecard: the Mini half shipped. [Seedance 2.0 Mini is live on reAPI](/models/seedance-2-0-mini) at $0.048/s for 480p and $0.103/s for 720p in text mode (from $0.03/s with a video reference), exactly the below-Fast slot the platform-side report described\[9]. The 2.1 half resolved differently: the denial held, and the real successor arrived as Seedance 2.5 with a bigger swing than any incremental quality bump.
Two lessons carry forward. Platform-side sourcing about the Seedance family has a decent hit rate. And ByteDance denials are worth reading narrowly; they deny the story, not the roadmap. Both now apply to the July rumors about Seedance 2.5's API timing.
## Running Seedance today, switching on launch day
The practical question is not whether Seedance 2.5 is exciting but what to build on this month. My answer is the boring one: build on the released family, keep the harness portable.
The full [Seedance 2.0 lineup runs on reAPI](/models/seedance-2-0) right now, Standard and Fast tiers from $0.0400/s to $0.4048/s depending on resolution and reference mode, plus the Mini tier below both\[9]\[11]. One async endpoint, per-second billing, one model string to change later.
For the new model, reAPI already has a [Seedance 2.5 page](/models/seedance-2-5) live with the announced capability set; API access lands there when ByteDance opens the model beyond its beta. Day-one preparation is the same as I recommended in June: keep an eval set of your own prompts, price your workload per output second rather than per clip, and treat the eventual switch as a one-line change.
## FAQ
### Is Seedance 2.5 released?
No. It is officially announced, with a live promo page on Volcano Engine\[1] and an official July release window\[2], but access today is enterprise beta only. Coverage of the announcement says public rollout starts in China\[4].
### What is officially new in Seedance 2.5 versus Seedance 2.0?
Per ByteDance's own page: 30-second coherent generation in one pass (2.0 caps at 15 seconds), expanded reference input, per-segment prompt control, deeper editing, and multilingual text rendering\[1]\[6]. The widely reported 4K/10-bit spec is already shipping in Seedance 2.0\[6], and the 50-reference figure appears only in press coverage, not on any official page.
### When will the Seedance 2.5 API be available?
No official date. Volcano Engine's docs say the model releases in July\[2]; a "July 10 API" figure circulates on social media with nothing behind it. Seedance 2.0's pattern was consumer surfaces first, then API access roughly two months later\[10]. 2.5 may move faster, but that is inference, not information.
### How much will Seedance 2.5 cost?
Unpublished. Guesses in circulation run from $0.022 to $0.50 per second, which is a 20x spread. For scale, Seedance 2.0 currently runs $0.0400 to $0.4048 per second on reAPI depending on tier, resolution, and reference mode\[11].
### Can Seedance 2.5 really generate 30-second videos?
The 30-second single-generation claim is ByteDance's own, stated on the official promo page\[1]. What no one outside the beta has verified is quality at that length; long generations are historically where video models lose temporal consistency. The "30 seconds of 4K" framing in headlines merges an official claim with a press-added one.
### What are the expanded reference inputs for?
Consistency at production scale: characters, locations, props, and style references attached to one generation. Seedance 2.0 officially accepts up to 9 images, 3 videos, and 3 audio clips\[6]; the official 2.5 examples show 11 images\[1], and press coverage claims the ceiling is 50\[4].
### Is seedance2.ai or any "Seedance 2.5 early access" site official?
No. The official surfaces are Volcano Engine, Jimeng, and Dreamina (via CapCut). Lookalike domains appeared within days of the announcement, some selling access to a model that has no public API. Do not pay them.
### Will Seedance 2.5 be available in the US?
Unknown. Seedance 2.0's global rollout paused in March 2026 over copyright disputes and resumed everywhere except the United States\[10]. Nothing announced so far says whether Seedance 2.5 changes that.
## How to be ready for launch day
The honest summary: an official promo page with four capability claims, an official July window, a press spec sheet that partly recycles the current model, zero benchmarks, no pricing, and a scam ecosystem that moved faster than the model did. Seedance 2.0 has spent five months as the model to beat, and its successor now has to clear a bar it set itself. Keep your prompts portable, watch the [Seedance 2.5 page](/models/seedance-2-5) for the moment API access opens, and let Seedance 2.5 prove the 30 seconds and the reference ceiling on your own footage before you believe any headline.
## References
1. Volcano Engine (ByteDance). *Doubao Seedance 2.5 — official promotional page.* Retrieved July 2026 from [ark.volcengine.com/promotion?modelName=seedance-2-5](https://ark.volcengine.com/promotion?modelName=seedance-2-5)
2. Volcano Engine (ByteDance). *Documentation portal — "Doubao video generation model 2.5, releasing in July."* Retrieved July 2026 from [volcengine.com/docs/search?q=Seedance 2.5](https://www.volcengine.com/docs/search?q=Seedance%202.5)
3. The Information. *ByteDance Unveils Seedance 2.5 Video Model.* Retrieved July 2026 from [bytedance-unveils-seedance-2-5-video-model](https://seedance25ai.im)
4. CNET. *ByteDance Introduces New Seedance 2.5 Video Model.* Retrieved July 2026 from [cnet.com/tech/services-and-software/bytedance-introduces-new-seedance-2-5-video-model](https://www.cnet.com/tech/services-and-software/bytedance-introduces-new-seedance-2-5-video-model/)
5. The Next Web. *ByteDance's Seedance 2.5 pushes AI video to 4K and 30 seconds.* Retrieved July 2026 from [thenextweb.com/news/bytedance-seedance-2-5-ai-video-4k-30-seconds](https://thenextweb.com/news/bytedance-seedance-2-5-ai-video-4k-30-seconds)
6. Volcano Engine (ByteDance). *Doubao Seedance 2.0 — model specifications (resolutions, durations, reference limits).* Retrieved July 2026 from [volcengine.com/docs/82379/1330310](https://www.volcengine.com/docs/82379/1330310)
7. Artificial Analysis. *Text to Video and Image to Video Leaderboards.* Retrieved July 2026 from [artificialanalysis.ai/video/leaderboard/text-to-video](https://artificialanalysis.ai/video/leaderboard/text-to-video)
8. LMArena. *Text to Video and Image to Video Leaderboards.* Retrieved July 2026 from [arena.ai/leaderboard/text-to-video](https://arena.ai/leaderboard/text-to-video)
9. reAPI. *Seedance 2.0 Mini — model page and live pricing.* Retrieved July 2026 from [reapi.ai/models/seedance-2-0-mini](/models/seedance-2-0-mini)
10. Reuters. *ByteDance suspends launch of video AI model after copyright disputes.* Retrieved July 2026 from [reuters.com/technology/bytedance-suspends-launch-video-ai-model](https://www.reuters.com/technology/bytedance-suspends-launch-video-ai-model-after-copyright-disputes-information-2026-03-14/)
11. reAPI. *Seedance 2.0 — model page and live pricing.* Retrieved July 2026 from [reapi.ai/models/seedance-2-0](/models/seedance-2-0)
### Further reading
* The Verge. *Netflix gives ByteDance three days to stop Seedance AI theft.* [theverge.com/ai-artificial-intelligence/880542](https://www.theverge.com/ai-artificial-intelligence/880542/netflix-gives-bytedance-three-days-to-stop-seedance-ai-theft)
* CNBC. *Senators call on ByteDance to shut down Seedance.* [cnbc.com/2026/03/17/bytedance-seedance-shut-down](https://www.cnbc.com/2026/03/17/bytedance-seedance-shut-down-tiktok-marsha-blackburn-peter-welch.html)
* reAPI. *Seedance 2.1 and Seedance 2.0 Mini: What's Actually Coming.* [reapi.ai/blog/seedance-2-1-and-seedance-2-0-mini-preview](/blog/seedance-2-1-and-seedance-2-0-mini-preview)
---
# Seedance Credits and Quotas: Why the Balance Runs Out (https://reapi.ai/blog/seedance-credits-and-quotas)
"How many clips does my plan get me" is a question with no stable answer, and that is the actual problem. Credits convert to clips at a rate that changes with resolution, duration, and settings you may not be watching, so the number of videos left in a subscription is never a number you were told.
Per-second billing answers the same question with arithmetic instead.
## TL;DR
* **Credits hide the exchange rate.** The same balance buys a different number of clips depending on resolution, length, and input mode.
* **Per-second billing makes it multiplication**: rate × seconds, and nothing else moves.
* **A 5-second 720p clip is $1.03** from a prompt, or **$0.63** driven by an uploaded video\[1].
* **Resolution is the steepest lever**, roughly 11x from 480p to 4K.
* **Duration is clamped to 4–15 seconds**, so short requests bill as 4.
* **There is no "unlimited"** on any per-second model. What there is, is a rate you can multiply.
## Why credit balances never answer the question

A credit is a unit the platform defines, and a clip consumes a variable number of them. Three things move the conversion, usually at once:
**Resolution.** A 1080p clip does not cost twice a 480p one; on the underlying per-second rates it is closer to five times.
**Duration.** Billing is per second of output, so a 12-second clip is three times a 4-second one.
**Input mode.** On Seedance, supplying an uploaded source video moves the request onto a cheaper rate. Supplying an image does not.
Multiply those together and "how many videos do I have left" has no single answer, which is why the question keeps getting asked and never gets a satisfying reply.
## What per-second billing looks like
Rates per **second of output**\[1]:
| Resolution | From a text prompt | With an uploaded video |
| ---------- | ------------------ | ---------------------- |
| 480p | $0.095 | $0.058 |
| 720p | $0.205 | $0.125 |
| 1080p | $0.510 | $0.310 |
| 4K | $1.040 | $0.640 |
The fast tier runs lower still at 480p and 720p: $0.078 and $0.165 from a prompt, $0.045 and $0.100 with an uploaded video.
Now the same question, answered:
| What you want | Cost |
| ---------------------------------------------- | --------- |
| One 5-second 720p clip | **$1.03** |
| One 10-second 1080p clip | **$5.10** |
| 100 five-second 720p clips | **$103** |
| 1,000 five-second 480p drafts on the fast tier | **$390** |
No balance to track, no conversion rate, no wall that arrives mid-project.
## The formula
```text
cost = per-second rate × output seconds
```
For a mixed workload, sum the seconds per configuration and multiply each by its own rate. That is the whole model.
Two constraints attached to it\[1]:
**Duration is clamped between 4 and 15 seconds.** A 3-second request bills as 4. Longer pieces are assembled from multiple generations, each billed separately.
**The cheaper rate needs an uploaded video specifically.** Image inputs and first/last-frame references stay on the base rate, which is the single most common budgeting mistake on this model. Details in [Seedance 2.0 cost per second](/blog/seedance-2-0-cost-per-second).
## About "unlimited"
There is no unlimited tier on a per-second model, and it is worth being direct about why rather than pretending otherwise.
Generating video costs compute that scales with output length and resolution. Any plan advertising unlimited generation is either rate-limiting you somewhere you cannot see, degrading quality or queue priority under load, or reserving the right to change the terms. The cost does not disappear; it moves somewhere less visible than a line item.
Per-second pricing is the opposite trade. You give up the comfort of a flat monthly number and get an invoice you can predict to the cent before you run anything.
## Controlling spend without a quota
**Draft cheap, render once.** 480p on the fast tier is $0.078 per second, so a 5-second iteration costs 39 cents. Iterate there, then render the approved shot at delivery resolution. The 480p-to-4K spread is roughly 11x.
**Feed video when the pipeline already has it.** The uploaded-video rate is about 39% lower at 720p. If a source clip exists, passing it in rather than describing it saves real money.
**Cap duration deliberately.** Every second bills. A 15-second clip where 8 would do is nearly double the cost for the same shot.
**Forecast before you build.** Multiply expected seconds by rate for each configuration and you have a monthly number before writing the integration.
```bash
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "doubao-seedance-2.0",
"prompt": "product hero on a stone counter, soft window light, slow orbit",
"resolution": "480p",
"duration": 5
}'
```
That request costs $0.475. The same prompt at 1080p costs $2.55. Live rates are on [reapi.ai/models/seedance-2-0](/models/seedance-2-0).
## Image generation works the same way
Per-image rather than per-second, same principle of a knowable unit\[2]:
| Model | Per image |
| ----------------------- | --------- |
| Nano Banana 2 Lite | $0.020 |
| Seedream 5.0 Pro (1K) | $0.032 |
| Nano Banana Pro (1K/2K) | $0.042 |
Twenty drafts on Lite plus one final on Pro is about $0.44 for the set. There is no plan underneath that number either.
## FAQ
### How many videos can I generate before credits run out?
On a per-second model the question does not apply: you pay per second of output, so a $100 budget is 100 divided by your rate times your clip length. At 720p from a prompt that is roughly 97 five-second clips\[1].
### How much does one Seedance 2.0 video cost?
A 5-second 720p clip is about $1.03 from a text prompt, or $0.63 when driven by an uploaded video\[1].
### Is there an unlimited Seedance plan?
Not on per-second billing. Any unlimited offer is limiting somewhere less visible, because output length and resolution consume compute that scales.
### Why did my costs jump when I switched to 1080p?
Because resolution is the steepest lever. 1080p is roughly five times 480p per second, and 4K is roughly eleven times\[1].
### Does a failed or short generation still cost?
Duration is clamped to a 4-second minimum, so a 3-second request bills as 4 seconds.
### How do I cut cost without cutting quality?
Draft at 480p on the fast tier and render only the approved shot at delivery resolution. Feed an uploaded video where your pipeline already produces one.
### Do image models bill the same way?
Per image rather than per second, but with the same property: one knowable unit, no conversion rate\[2].
### How do I forecast a monthly bill?
Sum expected output seconds per configuration and multiply each by its rate. That number is the bill.
## Trading a comfortable number for a knowable one
Credit systems feel simpler because they produce one figure a month. They are harder to plan with, because the figure that matters, cost per finished clip, is derived from a conversion rate that moves with three variables at once.
Per-second billing inverts that. There is no monthly comfort number, and there is no wall arriving in week three either. What there is, is a rate table and multiplication, which is enough to answer "what will this campaign cost" before anyone approves it. For anyone who arrived here because a credit balance ran out mid-project, that is the actual difference.
## References
1. reAPI. *Seedance 2.0 model page — live per-second rates by resolution, tier, and input mode.* [reapi.ai/models/seedance-2-0](/models/seedance-2-0)
2. reAPI. *Model catalog — per-image rates for the image models.* [reapi.ai/models](/models)
### Further reading
* reAPI. *Seedance 2.0 cost per second.* [reapi.ai/blog/seedance-2-0-cost-per-second](/blog/seedance-2-0-cost-per-second)
* reAPI. *Where to use Seedance 2.0.* [reapi.ai/blog/where-to-use-seedance-2-0](/blog/where-to-use-seedance-2-0)
* reAPI. *Nano Banana Pro vs Nano Banana 2.* [reapi.ai/blog/nano-banana-pro-vs-nano-banana-2](/blog/nano-banana-pro-vs-nano-banana-2)
---
# Seedance Face Detected: Real Person Errors, Explained (https://reapi.ai/blog/seedance-face-detected-real-person-error)
A filmmaker posted to r/Seedance\_AI in April with a problem that had nothing to do with deepfakes. He generated a video, pulled the end frame, and fed that frame back in as the reference for the next shot. Seedance face detected a person in his own model's output and blocked it.\[4] His summary: "no consistent characters, no narrative workflows, no filmmaking use."
That thread ran to 118 comments and it is still the most common complaint in the Seedance ecosystem. The errors arrive in three different wordings depending on which host you are on, they fire on faces that were never real, and the advice you find online is almost entirely about tricking the classifier. This piece does the other thing: what each error actually means, which layer emits it, why synthetic faces trip it, and where real-person references are a supported input instead of a thing to smuggle past a filter.
## TL;DR
* **Three error strings, same family.** "Input image may contain real person", "Images with faces are not allowed. Face detected in uploaded image", and "The images or videos provided may contain likenesses of real people or other private information that cannot be processed."\[5]\[6]\[7]
* **They fire on fully synthetic faces.** The classifier keys on photorealism, not on provenance, so a character you generated yourself gets flagged for looking too good.\[6]
* **It is not one filter.** Some hosts run their own face detector on upload; some pass the reference through and let the upstream review decide. Same model, different door, different verdict.
* **Seedance 2.5 makes references the product.** ByteDance's launch describes up to 30 images, 10 video clips and 10 audio clips per generation.\[3] A blanket upload ban on faces removes most of that.
* **reAPI accepts real-person reference images and video.** Material passes automated review before generation, and anything rejected fails with an error and a full refund.\[1]
## The three errors and where each one comes from
The wording tells you roughly how far your request travelled before it died.
**"Images with faces are not allowed. Face detected in uploaded image. Please use an image without real people."** This one never left the platform. It is a face detector running on the host's own upload handler, and it makes no distinction between a photograph of your neighbour and a rendered character. A user hit it on a service that had been sold to them as the most permissive Seedance host available.\[5]
**"Input image may contain real person."** Reported against Seedance 2.0 on BytePlus by someone whose characters came out of ChatGPT. Their description of the failure mode is the clearest one I have read: "The face is just too photorealistic, so the filter thinks it's a real human. Same images work fine on Kling."\[6]
**"The images or videos provided may contain likenesses of real people or other private information that cannot be processed."** The longest of the three, reported in July by someone whose avatars had been working for weeks with no changes on their end. "Today literally every single one of them is getting rejected."\[7] A commenter in the same thread reported the identical break on fal.ai the same day.\[7]
That last detail matters more than it looks. When the same error appears across unrelated hosts on the same day, the change happened upstream, not at your provider. When it appears on one host and not another with identical inputs, the host added it.
## Why Seedance face detected fires on a face you made yourself
The detector is a classifier, not a provenance check. It has no access to C2PA metadata, no record of which model produced the pixels, and no way to know that the woman in your reference photo does not exist. It scores photorealism and facial structure and blocks above a threshold.
That produces a perverse gradient: the better your image generator, the more likely its output is rejected. One commenter put the absurdity plainly. "We create human ai faces in NB but cant use it on SD 2. SD2 should be able to distinguish between ai generated human faces from real photo faces. Not sure what the use case is if we cant use ai faces."\[4]
There is a second, quieter failure. Thresholds move. Nothing on your side changes, and a workflow that ran for six weeks starts failing every request, which is exactly what the July thread describes.\[7] If your production pipeline depends on a classifier's mood, you have no pipeline.
## What the bypass threads are doing, and why this is not one
Search Seedance face detected and most of what surfaces is circumvention. Grid overlays, scenery composites, selective blurring, sketch conversion, prompt-level laundering, several "DM me for the tool" replies. Some of them work. One post claimed a 90% success rate for its method.\[8]
I am not publishing those, for two reasons that have nothing to do with squeamishness.
The first is that the control exists for consent. A likeness check is the one piece of this stack that protects somebody who is not in the conversation. Defeating it with a blur filter defeats it for photographs of real people who never agreed to anything, not just for your synthetic actor. That is worth keeping intact even when it is annoying.
The second is that the techniques degrade what you paid for. Every one of them works by damaging the reference until the classifier stops recognising a face. The model then also gets a damaged reference, which is the input you were relying on for character consistency. You buy a pass at the door by throwing away the thing you came in for.
The useful question is not how to defeat the detector. It is which routes treat a real-person reference as a supported input with a review step, rather than as something to block on sight.
## Reference material as a first-class input
Seedance 2.5's design assumes reference material. ByteDance's launch post leads with it: up to 30 images, 10 video clips and 10 audio clips in a single pass, with reference capabilities spanning motion, style and multi-subject scenes.\[3] A host that rejects any image containing a face has disabled most of that surface.
On reAPI the model id itself reflects the position. There is one customer-facing Seedance 2.5 id, `doubao-seedance-2.5-face`, with no faceless variant to fall back to.\[2] The published behaviour is that real-person reference images and videos are accepted, that reference material passes an automated review before generation runs, and that anything rejected fails with a clear error and a full refund of the reserve.\[1]
The refund is the part that changes how you work. A false positive on a synthetic character costs you a retry and nothing else, so you can test a reference set instead of rationing attempts against a filter that might be having a bad day.
Two obligations do not move. Named living people and recognisable third-party characters stay refused regardless of route, and that is the model's own line rather than a host's policy setting. And consent for anyone appearing in your reference material remains yours to obtain. An API that accepts the upload is not an API that acquired permission for you.
## What reAPI accepts, exactly
Worth having in one place, because half the reference failures people report are format problems misread as moderation.\[2]
| Field | Limit | Formats |
| ------------------ | ----------------------------------------------------- | ------------------------------------------- |
| `image_urls` | up to 30, each under 30 MB | jpeg, png, webp, bmp, tiff, gif, heic, heif |
| `image_with_roles` | first frame, or first + last (max 2) | same as above |
| `video_urls` | up to 10, each 2–30 s and under 200 MB, combined 30 s | mp4, mov, 480p to 4K |
| `audio_urls` | up to 10, each 2–30 s and under 15 MB | wav, mp3 |
One platform rule catches people out: every reference must be a public HTTP(S) URL. Base64 and `data:` payloads are rejected across the whole reAPI surface, not just this model.\[2] A r/VeniceAI thread in August shows what that looks like when a host changes it without warning, with several users hitting the same "needs an accessible URL" wall in one afternoon.\[9]
Two more constraints are enforced before submit rather than minutes later: `image_urls` and `image_with_roles` are mutually exclusive, and a first/last-frame job only accepts `size: adaptive`.\[2] Both return immediately instead of failing after the task starts.
## FAQ
### What does the Seedance face detected error actually mean?
That a classifier scored your uploaded image as containing a real human face. It is not a claim that the person is real, only that the image looks photographic enough to cross a threshold.\[6]
### Why does it reject my AI-generated character?
Because photorealism is the signal. A stylised or obviously rendered face passes; a face good enough to be mistaken for a photograph does not. Several users report their best synthetic characters being the ones that fail.\[4]\[6]
### Can I upload a real person's photo to Seedance 2.5 on reAPI?
Yes. Real-person reference images and video are accepted and go through automated review first, with a full refund if the review rejects them.\[1] Consent and likeness rights remain your responsibility.
### Why did the same images stop working overnight?
Upstream thresholds get retuned. The July r/generativeAI thread documents exactly this, with the same failure appearing on more than one host on the same day.\[7]
### Does turning off the safety checker fix face rejections?
No. `nsfw_checker` selects which route runs your task, and likeness checks are not the thing it removes.\[2] A reference rejection and a content rejection are different failures.
### Will a rejected reference still charge me?
Not on reAPI. A task that fails review refunds the reserve in full.\[1]
### Can I use a named celebrity as a reference?
No, on any route. That limit sits with the model and no parameter reaches it.
### Are the bypass methods on Reddit safe to use?
They work by degrading the reference until the classifier stops seeing a face, which also degrades the input the model uses for character consistency. Applied to photographs of real people they defeat a consent control, which is a different problem from an inconvenience.
## Getting a character through without fighting a classifier
If you are building a narrative workflow, the classifier is not a puzzle to solve, it is a routing question. Hosts that run their own face detector on upload will block your synthetic lead however many overlays you stack on it, and the overlays cost you the consistency you were buying. Hosts that treat references as a reviewed input give you a verdict, a reason, and your credits back when the verdict is wrong.
Check which kind you are on before you build a production pipeline on top of it. Upload a single frame of your main character and read what comes back. If the answer is an instant Seedance face detected refusal with no generation time attached, you are talking to a doorman, not to the model.
## References
1. reAPI. *Seedance 2.5 — model page, real-person reference support and refund behaviour.* Retrieved August 2026 from [reapi.ai/models/seedance-2-5](/models/seedance-2-5)
2. reAPI. *doubao-seedance-2.5-face — reference field limits, formats and submit-time constraints.* Retrieved August 2026 from [reapi.ai/docs/seedance-2-5](/docs/seedance-2-5)
3. ByteDance Seed. *One-take Creation, Flexible Referencing: Introducing Seedance 2.5.* Published 31 July 2026, retrieved August 2026 from [seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5](https://seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5)
4. r/Seedance\_AI. *Seedance 2.0 is becoming unusable for filmmaking – face detection blocks even AI-generated content.* Posted 8 April 2026, retrieved August 2026 from [reddit.com/r/Seedance\_AI/comments/1sfp3ag](https://www.reddit.com/r/Seedance_AI/comments/1sfp3ag)
5. r/Seedance\_AI. *Why do some companies keep saying their Seedance 2.0 is uncensored and completely unfiltered.* Posted 7 June 2026, retrieved August 2026 from [reddit.com/r/Seedance\_AI/comments/1tyzrun](https://www.reddit.com/r/Seedance_AI/comments/1tyzrun)
6. r/generativeAI. *Seedance keeps rejecting my AI images as "real person"; any fix?* Posted 22 May 2026, retrieved August 2026 from [reddit.com/r/generativeAI/comments/1tkkxjq](https://www.reddit.com/r/generativeAI/comments/1tkkxjq)
7. r/generativeAI. *Anyone else getting hit with "likenesses of real people" errors on Seedance 2.0 all of a sudden??* Posted 7 July 2026, retrieved August 2026 from [reddit.com/r/generativeAI/comments/1upwj55](https://www.reddit.com/r/generativeAI/comments/1upwj55)
8. r/seedance. *90% bypass rate on Seedance 2.0 face detection.* Posted 17 April 2026, retrieved August 2026 from [reddit.com/r/seedance/comments/1sog2st](https://www.reddit.com/r/seedance/comments/1sog2st)
9. r/VeniceAI. *Reference videos now need accessible URL? What is this?* Posted 1 August 2026, retrieved August 2026 from [reddit.com/r/VeniceAI/comments/1vcbvpr](https://www.reddit.com/r/VeniceAI/comments/1vcbvpr)
### Further reading
* reAPI. *Seedance 2.5 Uncensored: A Provider Test You Can Run.* [reapi.ai/blog/seedance-2-5-uncensored-provider-test](/blog/seedance-2-5-uncensored-provider-test)
* reAPI. *Seedance 2.0 Character Consistency Guide.* [reapi.ai/blog/seedance-2-0-character-consistency-guide](/blog/seedance-2-0-character-consistency-guide)
---
# Seedance 2.0 "Not Eligible": Why It Happens, What Works (https://reapi.ai/blog/seedance-not-eligible-explained)
If Seedance 2.0 keeps stamping "Not eligible" on your reference image, you are not being throttled and you did not break a hidden rate limit. You are hitting a content check that ByteDance enforces at the model level, and its two triggers are documented in the company's own API rules: real human faces and protected intellectual property\[1]. No platform can wave it off, and most platforms do not even document it, which is why the error feels random.
This post assembles what is actually written down: ByteDance's own eligibility rules and error codes, what the platforms' interfaces are really checking, why the same face passes one day and fails the next, and the routes that legitimately work when your project involves real people.
## TL;DR
* **"Not eligible" is face and IP detection.** Higgsfield's own interface explains the failed state as "This image contains faces or IP and cannot be used," and refunds the credits\[2]\[3].
* **The rule comes from ByteDance, not the platform.** The official API documentation states Seedance 2.0 models "do not support direct upload of reference images or videos containing real human faces"\[1], and Higgsfield's team has said publicly the detection "is on Seedance 2.0's side"\[4].
* **There is no moderation off-switch.** The official Seedance 2.0 API exposes no safety parameter; input and output are both screened, always\[1]\[5].
* **"Sometimes it works" has boring explanations**: older Seedance 1.5 Pro is not under the same face restriction on Dreamina\[6], platforms stack their own filters on top, and detection is threshold-based, so borderline images flip between runs.
* **Legitimate routes exist**: ByteDance documents three exemption paths, including a consent-verified upload channel\[7], and [reAPI's face-enabled Seedance tier](/models/seedance-2-0) is built for exactly that workflow.
* Bypass tricks circulating on Reddit violate the acceptable-use policy you agreed to, and the model re-screens outputs anyway\[5].
## What "not eligible" actually means
Higgsfield is where most people meet this error, so start there. The platform runs every reference through a content check with three states, and the interface strings are unambiguous: "Checking content…", "Eligible", "Not eligible." When the check fails, the interface says why: "This image contains faces or IP and cannot be used," or on generations, "This generation may involve protected IP or likeness rights"\[2]. The pipeline even has a dedicated IP-detection stage, and flagged requests are not charged; Higgsfield's docs confirm credits come back automatically\[3].
So the error is not about image quality, format, or your account standing. It is a verdict on the content: the system believes your reference contains a real person's face or someone's protected property.
And it is not Higgsfield's verdict. When users complained on Reddit, the Higgsfield team answered directly: "The face detection you're running into is on Seedance 2.0's side, not something Higgsfield controls"\[4]. Every platform serving Seedance 2.0 inherits this.
## The rule underneath: no real faces, no protected IP
The primary source is ByteDance's own API documentation, which states it plainly: Seedance 2.0 series models "do not support direct upload of reference images or videos containing real human faces"\[1]. The error-code catalog is just as direct. `InputImageSensitiveContentDetected.PrivacyInformation` means the input "may contain real person"; the `.PolicyViolation` variant flags material that "may be related to copyright restrictions"\[5].
The consumer side matches. When CapCut rolled Seedance 2.0 out globally on Dreamina, its newsroom post said the launch was "restricting certain capabilities… including the ability to make videos from images or videos that contain real faces," alongside technology "designed to block the unauthorized generation of intellectual property" and invisible watermarking of outputs\[6]. Dreamina's model picker labels Seedance 2.0 with a flat "Real human faces are not supported."
The why is not mysterious either: Seedance 2.0's launch collided with studio copyright disputes serious enough that ByteDance paused the global rollout for a month\[8]. Deepfake liability plus IP litigation equals a hard filter at the model boundary.
Worth knowing: there is no toggle. The official API's parameter set contains no moderation switch of any kind; screening runs on inputs and outputs both, and the acceptable-use policy you sign prohibits depicting "a person (living or dead)'s voice or likeness without appropriate consent" and any attempt to bypass safety filters\[9].
## Why the same image passes one day and fails the next
The inconsistency has three documented-or-observable causes, and none of them is luck.
**Different models, different rules.** On Dreamina, the "real faces not supported" label sits on Seedance 2.0 and 2.0 Fast, but not on the older Seedance 1.5 Pro\[6]. If a workflow "used to work," it may have been running a different model, and some platforms switch backend versions without announcing it.
**Platforms stack their own filters.** Higgsfield's trust page notes that moderation "is applied at the model level" and that "content policies may vary depending on the model being used"\[2], and platforms layer additional checks with their own thresholds on top of ByteDance's. Same image, different platform, different verdict.
**Detection is probabilistic.** Community testing shows the frustrating edge: stylized and even fully AI-generated faces get flagged as "real" people regularly, and borderline images flip verdicts between attempts\[10]. A threshold classifier near its threshold is noisy. Celebrity likenesses appear to face an additional recognition layer that composition tricks do not move, which is consistent with the "protected likeness" framing in Higgsfield's interface strings\[2].
One more source of confusion worth killing: "visual restriction" has no official definition anywhere in ByteDance's or CapCut's published documentation. It circulates in third-party SEO posts, and its practical meaning is simply the face-and-IP restriction described above. Dreamina's own "not eligible" string, separately, is an age-gate message unrelated to image checks.
## The routes that legitimately work
Start with the split that matters: is your subject a real person, or is the detector wrong about a synthetic one?
**If your character is synthetic and still gets flagged**, you are fighting a false positive. The legitimate fixes are compositional: avoid tight, passport-style close-ups as the reference; prefer fuller scenes where the face is smaller in frame; and if the platform offers a first-frame mode alongside a universal reference mode, feed the character through the first frame, which is treated as scene context rather than a face reference. These adjust what the classifier sees; they do not defeat a correct detection.
**If your subject is a real person**, the answer is consent infrastructure, not tricks. ByteDance's own docs describe three exemption channels: content the platform itself generated for your account within the last 30 days can be trusted as input; a library of preset virtual avatars; and a verified route where a real person's identity and authorization are confirmed, after which their material is referenced through a dedicated asset ID\[7]. That last one is the official answer to "how do I make videos of myself."
On reAPI, the [Seedance 2.0 face tier](/models/seedance-2-0) is built for real-person reference workflows where you hold the subject's consent and the necessary rights, priced above the standard tier, with the same per-second billing and the same one-line model switch. Failed generations refund automatically, so an unexpected rejection costs nothing but time.
**What not to do:** the grid-overlay and face-obscuring hacks that circulate on Reddit. They violate the acceptable-use terms you agreed to\[9], output screening re-checks the result anyway\[7], and platforms refund flagged attempts precisely because they expect the filter to catch things. Build on the sanctioned routes and the error stops being part of your workflow.
## FAQ
### What does "not eligible" mean on Seedance 2.0?
The reference you uploaded failed a content check for real human faces or protected intellectual property. Higgsfield's interface states it directly: "This image contains faces or IP and cannot be used"\[2]. The check originates with ByteDance's model rules, not the platform\[1]\[4].
### Why does Seedance 2.0 block my AI-generated character?
False positive. The face detector estimates "is this a real person," and photorealistic synthetic faces sit near its threshold, so they get flagged often and inconsistently\[10]. Looser framing and first-frame placement usually resolve it for genuinely synthetic subjects.
### What is "visual restriction" on Seedance 2.0?
An unofficial phrase with no definition in any ByteDance or CapCut documentation. In practice it refers to the same documented restriction: no real faces, no protected IP in reference material\[1]\[6].
### Can I turn off Seedance 2.0's content moderation?
No. The official API has no moderation parameter, screening applies to both inputs and outputs, and the acceptable-use policy forbids circumventing it\[5]\[9].
### Which platform has the fewest Seedance restrictions?
The face and IP rules travel with the model, so no platform escapes them\[4]; platforms differ only in the extra filters they stack on top and in how clearly they surface the error. The real lever is not platform-shopping, it is using the consent-based channels for real-person work\[7].
### Do failed "not eligible" generations cost money?
Generally no. Higgsfield auto-refunds flagged requests\[3], and on reAPI failed tasks refund automatically as well.
### Why are celebrities blocked even in stylized images?
Likeness protection appears to run as an additional recognition layer beyond generic face detection, matching the "protected IP or likeness rights" language in platform interfaces\[2]. Composition changes do not move it, by design.
### How do I make Seedance 2.0 videos with my own face?
Through the verified-consent route: ByteDance's documentation describes identity verification plus authorization, after which your material is referenced via a dedicated asset channel\[7]. reAPI's face-enabled tier supports real-person reference workflows for API users who hold the subject's consent.
## Designing around the rules, not against them
The pattern behind every "not eligible" story is the same: a model-level filter, documented by ByteDance and inherited by every platform, doing exactly what its owner intends after a copyright firestorm. Fighting it wastes credits and violates terms; understanding it turns the error into a routing decision. Synthetic subjects get compositional fixes, real people get the consent channels, and the [face-enabled Seedance 2.0 tier on reAPI](/models/seedance-2-0) turns the sanctioned path into an API call. That is the entire playbook for making Seedance 2.0 not eligible errors disappear from your pipeline.
## References
1. BytePlus (ByteDance). *ModelArk — Seedance 2.0 input rules ("do not support direct upload of reference images or videos containing real human faces").* Retrieved July 2026 from [docs.byteplus.com/en/docs/ModelArk/1520757](https://docs.byteplus.com/en/docs/ModelArk/1520757)
2. Higgsfield. *Trust & safety page and production interface strings ("Not eligible", "contains faces or IP").* Retrieved July 2026 from higgsfield.ai/trust
3. Higgsfield. *API documentation FAQ — flagged requests are not charged.* Retrieved July 2026 from docs.higgsfield.ai/docs/help/faq
4. Higgsfield team (official account). *Reddit reply: "The face detection… is on Seedance 2.0's side."* Retrieved July 2026 from [reddit.com/r/HiggsfieldAI/comments/1sq1hms](https://www.reddit.com/r/HiggsfieldAI/comments/1sq1hms/)
5. BytePlus (ByteDance). *ModelArk — content moderation error codes (PrivacyInformation, PolicyViolation).* Retrieved July 2026 from [docs.byteplus.com/en/docs/ModelArk/1299023](https://docs.byteplus.com/en/docs/ModelArk/1299023)
6. CapCut Newsroom. *Dreamina Seedance 2.0 global rollout — face and IP restrictions, invisible watermarking.* Retrieved July 2026 from [capcut.com/newsroom/dreamina-seedance-2](https://www.capcut.com/newsroom/dreamina-seedance-2)
7. BytePlus (ByteDance). *ModelArk — trusted-input exemptions and consent-verified asset channel.* Retrieved July 2026 from [docs.byteplus.com/en/docs/ModelArk/2291680](https://docs.byteplus.com/en/docs/ModelArk/2291680)
8. Reuters. *ByteDance suspends launch of video AI model after copyright disputes.* Retrieved July 2026 from [reuters.com/technology/bytedance-suspends-launch-video-ai-model](https://www.reuters.com/technology/bytedance-suspends-launch-video-ai-model-after-copyright-disputes-information-2026-03-14/)
9. BytePlus. *Generative AI Acceptable Use Policy.* Retrieved July 2026 from [docs.byteplus.com/en/docs/legal/acceptable\_use\_policy\_byteplus\_genai](https://docs.byteplus.com/en/docs/legal/acceptable_use_policy_byteplus_genai)
10. r/Seedance\_AI (community). *Face detection blocks even AI-generated content — discussion thread.* Retrieved July 2026 from [reddit.com/r/Seedance\_AI/comments/1sfp3ag](https://www.reddit.com/r/Seedance_AI/comments/1sfp3ag/)
### Further reading
* reAPI. *What Is Seedance 2.0 and How to Use It (2026 Guide).* [reapi.ai/blog/what-is-seedance-2-0-and-how-to-use-it](/blog/what-is-seedance-2-0-and-how-to-use-it)
* reAPI. *Seedance 2.0 API documentation.* [reapi.ai/docs/seedance-2-0](/docs/seedance-2-0)
* CapCut. *Dreamina Community Guidelines.* [capcut.com/clause/dreamina-community-guidelines](https://www.capcut.com/clause/dreamina-community-guidelines?region=US\&lang=en)
---
# Seedream 5.0 Lite vs Pro: Price, Editing, and Quality (2026) (https://reapi.ai/blog/seedream-5-0-lite-vs-pro)
**Seedream 5.0 Lite is the volume model; Seedream 5.0 Pro is the production
model.** Lite currently costs $0.031 per image on reAPI, supports 2K and 3K
output, and can return up to four images in one request. Pro starts at $0.045,
returns one image per call, and is aimed at dense typography, annotation-guided
editing, multilingual layouts, and higher-fidelity production work.
That does not mean Pro is automatically the right choice. The useful decision
is whether the job needs one carefully controlled asset or many inexpensive
variations.
## TL;DR
* Choose **Seedream 5.0 Lite** for ideation, batches, general editing, style
transfer, and lower-cost application volume.
* Choose **Seedream 5.0 Pro** for text-rich graphics, precise guided edits,
multilingual layouts, and final production assets.
* reAPI currently charges Lite at **$0.031 per output** for both 2K and 3K.
* Pro currently starts at **$0.045 for the 1K pixel band** and **$0.089 for the
2K band**, before additional reference-image surcharges.
* The models use different request schemas. Do not swap only the model name and
expect the same payload to work.
## Seedream 5.0 Lite vs Pro specifications
| Feature | Seedream 5.0 Lite | Seedream 5.0 Pro |
| ------------------------- | ------------------------------- | ---------------------------------------------- |
| reAPI model ID | `doubao-seedream-5-0-lite` | `doubao-seedream-5-0-pro` |
| Current base price | $0.031/image | $0.045–$0.089/image |
| Output per request | 1–4 | Exactly 1 |
| Output control | 2K or 3K | Basic/high quality pixel bands |
| New image generation | Yes | Yes |
| Image editing | Yes | Yes |
| Reference input | Multiple image URLs | Up to 10 image URLs |
| Sequential generation | Yes | No `n` or batch field |
| Output format | JPEG or PNG | Provider-selected image output |
| Watermark control | Yes | Not exposed on current reAPI route |
| Additional output checker | No public field | `nsfw_checker` available to direct API callers |
| Best fit | Volume, ideation, general edits | Final assets, dense text, guided editing |
ByteDance introduced Lite in February 2026 as a reasoning-oriented image model
with real-time search, stronger world knowledge, and improved editing. Its own
release notes also call it a relatively small model and acknowledge room for
better structural stability, realism, and aesthetics.\[1]
The Pro page emphasizes high-density infographics, spatial annotations,
sketch-guided editing, layer separation, photographic rendering, and native
multilingual generation.\[2] Those are vendor
positioning claims. They explain intended use, but they do not replace a
same-prompt test on your own product category.
## Seedream 5.0 Lite vs Pro pricing
| Model and band | Current reAPI price | What changes the bill |
| -------------- | ------------------: | -------------------------------- |
| Lite 2K | $0.031/image | Output count `n` |
| Lite 3K | $0.031/image | Output count `n` |
| Pro 1K band | $0.045/image | Reference images after the first |
| Pro 2K band | $0.089/image | Reference images after the first |
Lite's 2K and 3K controls currently route to the same upstream price cell. Four
outputs therefore cost approximately $0.124 before any future rate change.
reAPI rounds the request in aggregate rather than rounding every image
separately.
Pro uses two pixel bands. Outputs at or below 2.36 million pixels use the lower
band; larger outputs use the higher band. The first reference image is included,
then each additional reference adds RMB 0.02—about $0.00294 at the fixed
conversion used by the current rate card.
The practical cost comparison depends on selection rate. If ten Lite attempts
are needed to find one usable asset, the nominally cheaper model can cost more
than one well-controlled Pro output. Track accepted assets, not only generated
assets.

## Image quality and instruction following
Lite is designed to understand intent before drawing. ByteDance highlights
world knowledge, visual reasoning, information graphics, style transfer, and
editing from short instructions. It also reports stronger internal Elo results
than Seedream 4.5, based on a vendor-run evaluation.\[1]
The methodology is useful context, but it is not an independent guarantee for
your prompts.
Pro raises the ceiling for complex production assets. Its official showcase
focuses on dense educational graphics, storyboards, interface-like layouts,
annotation-based transformations, and multilingual copy. Choose Pro when small
typographic or structural errors create expensive manual rework.
For open-ended visual ideation, the quality difference may matter less. Lite's
ability to produce several outputs per request gives art directors more options
per dollar, especially when the final result will be retouched anyway.
## Editing and reference-image control
Both models edit images, but the control surfaces differ.
Lite accepts `image_urls` and supports general image-to-image fusion. It is
useful for style references, subject consistency, and multi-image composition.
The current public schema does not impose a simple total reference count, so
applications should follow the live API docs and validate payload size rather
than assuming an unlimited array.
Pro accepts up to ten reference images and is positioned for annotation-guided
and spatial editing. The official showcase includes boxed regions, handwritten
instructions, sketch-to-production transformations, and translation that keeps
an existing layout.\[2]
Use clear reference roles in the prompt:
```text
Image 1 is the product whose shape must remain unchanged.
Image 2 supplies only the color palette.
Image 3 supplies only the paper texture.
Place the product in a warm editorial scene. Preserve the label spelling,
camera angle, and package proportions from Image 1.
```
The prompt prevents the model from treating every input as an equal visual
blend. It also gives a reviewer specific invariants to verify.
## API request differences
### Seedream 5.0 Lite request
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer $REAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "doubao-seedream-5-0-lite",
"prompt": "Four editorial concepts for a sustainable coffee package, ivory paper, cobalt type, one orange seal, no logo mockups.",
"size": "4:3",
"resolution": "3k",
"n": 4,
"output_format": "png",
"sequential_image_generation": "auto",
"max_images": 4,
"watermark": false
}'
```
When `n` is greater than one, the route uses automatic sequential generation.
The final bill scales with the number of images returned.
### Seedream 5.0 Pro request
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer $REAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "doubao-seedream-5-0-pro",
"prompt": "Create a publishable 16:9 product launch graphic. Preserve the package in Image 1. Use the material palette from Image 2. Exact headline: LESS PACKAGING, MORE COFFEE.",
"aspect_ratio": "16:9",
"image_urls": [
"https://example.com/package.png",
"https://example.com/material-board.png"
],
"quality": "high",
"nsfw_checker": true
}'
```
The [Seedream 5.0 Lite model page](/models/seedream-5-0-lite) and
[Seedream 5.0 Pro model page](/models/seedream-5-0-pro) show the current
parameters. Pro does not accept Lite's `n`, `resolution`, or
`sequential_image_generation` fields.
## Which Seedream model should you choose?
### Choose Lite for high-volume application workflows
Lite is better suited to social variants, mood boards, thumbnail concepts,
style exploration, and tools where users expect several options. Its lower
price and 1–4 output count reduce orchestration work.
### Choose Pro for final graphics and controlled edits
Pro is the safer default for infographics, e-commerce hero assets, layout-aware
translation, annotated revisions, and work where exact text or composition
matters. The higher per-image price can be offset by fewer discarded outputs.
### Use both in a two-stage pipeline
A practical production system uses Lite for exploration and Pro for the chosen
direction. Generate several compositions cheaply, select one, then send the
approved asset and an explicit edit brief to Pro. This keeps expensive runs
focused on decisions that have already been made.
For deeper Pro billing details, read
[Seedream 5.0 Pro pricing](/blog/seedream-5-0-pro-price). The
[Seedream 5.0 Pro bugs and fixes guide](/blog/seedream-5-0-pro-bugs-and-fixes)
covers retry-safe integration and inconsistent outputs.
## Safety and `nsfw_checker`
The Pro route exposes `nsfw_checker` to direct API callers, while the public
playground keeps it enabled. Turning off this additional output check does not
remove upstream rules or guarantee that a request will run. It also does not
change a developer's obligation to moderate user input and output.
The dedicated guide to
[Seedream 5 Pro's NSFW checker](/blog/seedream-5-pro-nsfw-checker) explains the
layers and billing implications. Lite and Pro should not be labeled
“uncensored” based on one optional gateway field.
## FAQ
### Is Seedream 5.0 Pro always better than Lite?
No. Pro targets controlled, production-grade work. Lite is often more suitable
for high-volume ideation and applications that need several choices.
### Which Seedream 5.0 model is cheaper?
Lite currently costs $0.031 per output on reAPI. Pro starts at $0.045 and moves
to $0.089 in the larger output-pixel band, before extra reference charges.
### Can Seedream 5.0 Lite generate several images?
Yes. The current reAPI route accepts `n` from one to four and uses sequential
generation when more than one image is requested.
### How many references does Seedream 5.0 Pro accept?
The current reAPI schema accepts up to ten reference image URLs.
### Does Pro support better text rendering?
ByteDance positions Pro for high-density infographics and multilingual text.
That makes it the stronger candidate, but spelling and layout still require
human verification.
## Conclusion
Seedream 5.0 Lite vs Pro is a workflow decision more than a simple quality
ranking. Use Lite when selection volume and cost matter; use Pro when a single
asset must survive close review. For many teams, the best system is Lite for
exploration followed by Pro for the final controlled edit.
## References
1. ByteDance Seed. *Introducing Seedream 5.0 Lite.* [seed.bytedance.com](https://seed.bytedance.com/en/blog/deeper-thinking-more-accurate-generation-introducing-seedream-5-0-lite)
2. ByteDance Seed. *Seedream 5.0 Pro model overview and showcase.* [seed.bytedance.com/en/seedream5\_0\_pro](https://seed.bytedance.com/en/seedream5_0_pro)
3. BytePlus ModelArk. *Seedream 5.0 Lite API availability updates.* [docs.byteplus.com](https://docs.byteplus.com/en/docs/modelark/1159177)
---
# Seedream 5.0 Pro Bugs: What's Documented, What's Random (https://reapi.ai/blog/seedream-5-0-pro-bugs-and-fixes)
Most complaints filed as Seedream 5.0 Pro bugs are not bugs. They are documented behavior nobody read the documentation for, or hard limits that reject the request before a pixel gets generated. Exactly one item on the list is genuinely random, and separating it from the rest is the whole point of this page.
So this is written as a lookup rather than an essay. Find your symptom in the table, read that section, apply the fix. The last section is the retest loop you need for the one failure that has no fix.
## TL;DR
* **The long-prompt failure is documented by ByteDance with a threshold.** The `prompt` field reference recommends staying under 300 Chinese characters or 600 English words and states the consequence: too many words scatter the information, so the model ignores details and the image comes back missing elements\[1].
* **Validation passing is not the model coping.** reAPI's hard cap is 4,000 characters\[6], well past 600 English words.
* **Anatomy errors are real and reported at volume.** A tester past 1,000 generations reported oversized heads and characters that "end up with three hands"\[5]. No parameter fixes this.
* **Many "failures" are enumerable rejections**: 512×512 is below the pixel floor, references cap at 10, each reference must be under 30 MB and 36 million pixels\[1]\[2].
* **The watermark is a default, not a defect.** On ByteDance's Ark API `watermark` is `true` unless you pass `false`\[1]. reAPI ships it off\[6].
* **Retesting costs cents**, which is what makes the stochastic failure manageable: ten runs of a suspect prompt is 32 cents\[7].
## Find your symptom
| What you're seeing | What's actually happening | Where to go |
| ---------------------------------------------------- | --------------------------------------------------- | ------------------------------------------------------------------ |
| Some things you asked for are missing from the image | Documented behavior past a documented prompt length | [Missing elements](#symptom-elements-from-your-prompt-are-missing) |
| Three hands, three legs, a head that's too big | Genuinely stochastic, clusters on occluded joints | [Anatomy](#symptom-extra-limbs-and-warped-proportions) |
| The API returns an error before generating anything | One of a finite list of documented limits | [Rejections](#symptom-the-request-fails-before-anything-generates) |
| An "AI generated" mark in the corner | The upstream default is watermark on | [Watermark](#symptom-a-watermark-you-didnt-ask-for) |
| You asked for four variations and got one image | Group output is unsupported on this model | [Rejections](#symptom-the-request-fails-before-anything-generates) |
| It worked last month and behaves differently now | Version-stamped snapshot changed | [Retest loop](#two-ways-to-run-a-retest-suite) |
## Symptom: elements from your prompt are missing
You write twelve requirements, you get nine, and nothing in the response tells you which three vanished. This is the most expensive failure in production and the best documented.
ByteDance's `prompt` parameter reference recommends no more than 300 Chinese characters or 600 English words, then states the reason plainly: when the word count runs too high the information gets scattered, so the model may ignore details and attend only to the main points, which causes the image to be missing some elements\[1]. That is the vendor describing the failure mode of its own product, with a number attached.
**Why word count is the wrong metric.** The disease is how many discrete things you asked to be individually correct. A 150-word paragraph about one subject in one room lands almost every time. A 150-word paragraph containing nine objects, two people, per-person wardrobe and three background details starts shedding whatever you named last or named vaguely. The useful budget is roughly eight discrete elements per generation.
**Why negations make it worse.** Seedream 5.0 Pro has no negative-prompt parameter, on ByteDance's API or on reAPI\[1]\[6]. A tail like "no retouching, no perfect symmetry, no plastic skin, no commercial smile, no artificial eyelashes, no flawless white teeth, no glamour lighting" spends thirty-odd words of the same attention budget on things you do not want.
There is a public example of this failing exactly as predicted. In a mid-July comparison thread, a roughly 290-word portrait prompt opened by specifying "a Southeast Asian woman in her mid-20s with warm medium-brown skin" and closed with a long negation list. A commenter's verdict: both models "outright ignored the first thing you said about the subject", and neither result "looks remotely SE Asian"\[5]. The most important attribute was in the first clause and it still lost.
**The fix.**
* Non-negotiables in the first two sentences: subject, the objects that must exist, their spatial relations.
* One element per sentence, not a comma chain burying three requirements in one clause.
* Cap at about eight discrete elements; generate the base scene and add the rest in a second call.
* Put text you want rendered inside double quotes. That is ByteDance's stated technique, not folklore\[3].
* Read the prompt back as a checklist and tick each item off in the output.
ByteDance's prompt guide, which covers the lite, 4.5 and 4.0 models and predates 5.0 Pro, is explicit about the shape that works: subject plus action plus environment in connected natural language, with style, color, light and composition as supplementary phrases. Its own counter-example marks "a girl, holding an umbrella, tree-lined street, oil-painting-like delicate brushwork" as the version to avoid, and it states that concise precise prompts generally beat stacking ornate vocabulary\[3].
## Symptom: extra limbs and warped proportions
This is the one with no spec-sheet fix, so treat it as a probability to manage rather than a bug to close.
The strongest field evidence available is a comment from a tester who had run more than 1,000 generations on an unlimited Seedream 5.0 Pro plan: anatomy "can sometimes be an issue", with heads that come out too big and characters that "end up with three hands", noted specifically in character work\[5]. At a thousand renders you have seen the distribution, not one unlucky sample.
Failures are not uniformly distributed. They cluster where the limb count is ambiguous in your description, because when a joint is hidden nothing anchors which shin belongs to which hip and a plausible spare fills the gap.
| High-risk setup | Why the count breaks | Prompt language that lowers the risk |
| ---------------------------------------- | ------------------------------------------- | ------------------------------------------------------------------- |
| Hips or knees under loose fabric | Hidden joints stop anchoring the legs | "two lower legs emerge from under the drape, crossed at the ankles" |
| Crossed legs, tucked feet | Overlapping calves blur which shin is which | "legs crossed at the ankles, both feet visible" |
| Interlocked hands, hands behind the back | Occluded fingers invite extras | "both hands rest flat on the table, fingers visible" |
| Two figures standing close | Limbs get assigned to the wrong body | "her arm on his shoulder, his hands in his pockets" |
"Seated elegantly" leaves every one of those counts open. The dull explicit version reads like stage direction because that is what it is.
**The process rule that matters more than the prompt.** Three failures in five runs means the description is the problem and you rewrite it. One failure in five is variance and the cheapest response is another run. Without that ratio you end up rewriting prompts that were fine.
**The repair that beats rerolling.** Interactive editing lets you mark a region on a finished image, either by drawing on the input or by writing `` / `` coordinate tags into the prompt, and edits inside that region only\[2]. A keeper with one bad hand does not need a full regeneration. Check the pixels just outside the marked area afterward, because a local edit can nudge its neighbors. This is exclusive to 5.0 Pro among the Seedream models, and it is not part of reAPI's current parameter set\[2]\[6].
## Symptom: the request fails before anything generates
Good news: this pile is deterministic. Enumerate it once, encode it in your own validation, stop hitting it.
| What you send | Result | Why |
| ---------------------------------------------------------- | ----------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- |
| `size: "512x512"` | Rejected | Below the 921,600-pixel floor; the docs use this exact value as the invalid example\[1] |
| Dimensions above 4,624,220 total pixels | Rejected | That is the 5.0 Pro ceiling; tiers are 1K and 2K and nothing higher\[2] |
| An 11th reference image | Rejected | 5.0 Pro caps at 10; lite, 4.5 and 4.0 take 14\[1] |
| A 40 MB reference photo | Rejected | References must be under 30 MB and 36 million pixels\[1] |
| A reference under 15 px on a side, or outside 1:16 to 16:1 | Rejected | Documented input bounds\[1] |
| A base64 `data:` image | Rejected at the gateway | Public HTTP(S) URLs only, platform-wide on reAPI\[6] |
| An `n`, `seed` or `negative_prompt` field | Rejected | Strict schema; unknown keys fail validation\[6] |
| A request expecting four images | One image returned | Group output unsupported on 5.0 Pro; reAPI returns exactly one per call\[2]\[6] |
| A streaming request, or one asking for web search | Not available | Both supported on 5.0 lite, unsupported on 5.0 Pro\[2] |
| A reference URL ByteDance cannot fetch | Rejected | Reference URLs must be publicly reachable\[1] |
Two behaviors in the same family that are not rejections but bite the same way.
**URLs expire.** On ByteDance's own API the retention is 24 hours\[2]; on reAPI it is 72\[6]. A pipeline treating the returned URL as permanent storage starts serving dead links, and the failure surfaces days later in whatever database you saved it to.
**A blocked output stays charged.** On reAPI a prompt or reference rejected before generation refunds the task. An image that generates and is then blocked by the output check returns a content-policy error, hides the image, and remains charged, because the generation already happened upstream\[6].
## Symptom: a watermark you didn't ask for
One boolean. On ByteDance's Ark API `watermark` defaults to `true`, which stamps an AI-generated mark in the bottom-right corner of every image\[1]. People integrate, ship, and notice the badge when a client points at it.
reAPI ships the watermark off, and the `watermark` field is accepted but ignored, kept for callers who integrated against the earlier request shape\[6].
Two relatives worth checking while you are in there. `output_format` on Ark defaults to `jpeg`, so PNG has to be asked for, and on reAPI's current surface that field is also accepted and ignored\[1]\[6]. And `size` defaults to 2K, the more expensive tier, so anyone who omitted it while iterating has been paying the higher rate for drafts\[2].
One knob does exist and is worth knowing about, because I have seen it described under invented names: `optimize_prompt_options.mode` on Ark, taking `standard` by default and `fast` for lower latency at some quality cost\[2]. That is the only documented mode switch on this model.
## Two ways to run a retest suite
Anatomy failures are random, so one generation tells you nothing and you cannot separate a bad prompt from bad luck without repeats. At $0.032 for a 1K image, ten runs of a suspect prompt is 32 cents\[7], which makes "run it ten times and count" the first thing you do rather than the last.
Keep a failure suite either way: five to ten prompts that genuinely broke, saved verbatim with the aspect ratio and tier they ran at.
### Method 1 — reproduce by hand at the draft tier
1. Re-run each saved prompt five times at 1K, unchanged, and record the pass rate. This is your baseline.
2. Change exactly one thing: shorten the prompt, front-load the key elements, or add explicit count language.
3. Re-run the same five and compare pass rates.
4. Keep the winner, then render the approved prompt once at 2K for the deliverable.
Hold the tier constant while you measure. Changing resolution changes the failure surface, and a suite where half the runs are 1K and half are 2K tells you nothing actionable.
### Method 2 — batch the suite through the API
Once the fixes stabilize, a ten-prompt suite is a loop instead of an afternoon of clicking. Submit each prompt N times and collect task ids:
```bash
for i in $(seq 1 5); do
curl -s https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer $REAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "doubao-seedream-5-0-pro",
"prompt": "",
"aspect_ratio": "3:4",
"quality": "basic"
}' | jq -r '.id'
done
```
Then poll each id until it settles. Polling is free, so a suite of fifty runs costs exactly fifty draft images and nothing else\[6]:
```bash
curl -s https://reapi.ai/api/v1/tasks/$TASK_ID \
-H "Authorization: Bearer $REAPI_API_KEY" | jq '.status, .output.image_urls[0]'
```
Two operational notes. The documented rate limit is 500 images per minute per account per model version, which is generous but a real ceiling if you fan out a large suite at once\[2]. And re-run the whole suite when the model ID changes: ByteDance ships version-stamped snapshots, the current one being `doubao-seedream-5-0-pro-260628`\[4], and behavior at this layer shifts quietly between them.
## FAQ
### Does Seedream 5.0 Pro have a hand problem specifically?
No public benchmark isolates anatomy errors per model, so any answer is field evidence. The strongest report available is from a tester past 1,000 generations who saw oversized heads and three-handed characters in character work\[5]. The inviting conditions are occluded joints, crossed limbs, and fabric over hips or knees.
### How long can a prompt be?
reAPI accepts up to 4,000 characters\[6], but ByteDance recommends under 300 Chinese characters or 600 English words and explains that exceeding it causes dropped elements\[1]. Count discrete elements rather than words; about eight is where I stop.
### Can I fix one bad hand without regenerating the image?
On ByteDance's API, yes, through interactive editing with a marked region or `` / `` tags\[2]. It is exclusive to 5.0 Pro among the Seedream models and is not in reAPI's current parameter set\[6].
### Why did my 512×512 request fail?
It is below the minimum. Seedream 5.0 Pro requires between 921,600 and 4,624,220 total pixels, and the documentation names 512×512 as an invalid example for exactly this reason\[1].
### Can I get four variations in one call?
Not on this model. Group-image output is unsupported for 5.0 Pro though 5.0 lite, 4.5 and 4.0 have it\[2]. Send four calls; the API is async so they run in parallel.
### Does the model inherit my reference photo's camera angle?
That has been reported, and it is consistent with how reference conditioning works. ByteDance's guide addresses the general case: state explicitly what should be taken from the reference and, separately, what the generated scene should be\[3].
### Do these problems make it the wrong model?
No. The documented failures have documented fixes, the rejections are enumerable, and the random one is cheap to reroll around. Users in the same threads that report the anatomy issue also rate the output as the most naturally photographic of the models they compared\[5].
## Fixing what you can, rerolling what you can't
Read the parameter documentation once, because most of what gets filed as a Seedream 5.0 Pro bug is a limit or a default sitting in plain text. Then build the failure suite for the part that isn't. The long-prompt problem has a number attached and a fix that costs nothing. The rejection list is finite and belongs in your own validation. The watermark is one boolean. What remains is anatomy variance, and the defense there is explicit counting language plus reruns at three cents each, which is a workable amount of friction for a model this good at typography and layout.
## References
1. Volcano Engine. *Image generation API — prompt guidance, reference-image limits, parameter defaults.* Retrieved July 2026 from [volcengine.com/docs/82379/1541523](https://www.volcengine.com/docs/82379/1541523)
2. Volcano Engine. *Doubao Seedream 5.0 Pro guide — capability matrix, interactive editing, resolution tiers, rate limit, retention.* Retrieved July 2026 from [volcengine.com/docs/82379/2582774](https://www.volcengine.com/docs/82379/2582774)
3. Volcano Engine. *Seedream 4.0-5.0 prompt guide — prompt structure, quoted text rendering, reference-image instructions.* Retrieved July 2026 from [volcengine.com/docs/82379/1829186](https://www.volcengine.com/docs/82379/1829186)
4. Volcano Engine. *Model release announcements — doubao-seedream-5-0-pro-260628.* Retrieved July 2026 from [volcengine.com/docs/82379/1159178](https://www.volcengine.com/docs/82379/1159178)
5. Reddit r/GenAIGallery. *New Seedream 5.0 Pro vs GPT Image 2. What do you prefer?* Retrieved July 2026 from [reddit.com/r/GenAIGallery/comments/1uxdtea](https://www.reddit.com/r/GenAIGallery/comments/1uxdtea/new_seedream_50_pro_vs_gpt_image_2_what_do_you/)
6. reAPI. *Seedream 5.0 Pro API reference — request body, limits, errors, retention.* Retrieved July 2026 from [reapi.ai/docs/seedream-5-0-pro](/docs/seedream-5-0-pro)
7. reAPI. *Seedream 5.0 Pro — live pricing table.* Retrieved July 2026 from [reapi.ai/models/seedream-5-0-pro](/models/seedream-5-0-pro)
### Further reading
* reAPI. *Seedream 5.0 Pro Seedance 2.5 Workflow: What Actually Ships.* [reapi.ai/blog/seedream-5-0-pro-seedance-2-5-workflow](/blog/seedream-5-0-pro-seedance-2-5-workflow)
* reAPI. *Seedream 5.0 Pro model page.* [reapi.ai/models/seedream-5-0-pro](/models/seedream-5-0-pro)
---
# Seedream 5.0 Pro Price: The Pixel Line That Decides It (https://reapi.ai/blog/seedream-5-0-pro-price)
The Seedream 5.0 Pro price has two tiers, and the thing that decides which one you land in is not the tier name. It is a pixel count: **2.36 million**\[1].
That detail is worth more than the headline rate, because it means a large 16:9 frame at 2048×1152 stays in the cheaper tier while a smaller-sounding 2048×2048 square does not. Sizing exports against that boundary on purpose is the difference between paying the low rate and the high one for output that looks the same on a page.
On reAPI the two tiers are **$0.032 and $0.063 per image**, against ByteDance's published $0.045 and $0.09.
## TL;DR
* **Billing is per image, split at 2.36M pixels**, not by a resolution label\[1].
* **reAPI rates: $0.032 at 1K, $0.063 at 2K.** Roughly 29% below the published rate.
* **A 16:9 frame at 2048×1152 bills at the cheap tier**, because it is under the pixel line.
* **Up to 10 reference images**, with in-image text across 15 languages\[1].
* **No 4K at launch.** If 4K is a hard requirement, this is the wrong tier\[1].
* **reAPI carries an output-check-disabled variant** at $0.034 and $0.067, API only.
## The pixel boundary, not the label

The tiers are labeled 1K and 2K, and those labels are shorthand for a pixel budget rather than a literal dimension. The largest square in each is 1536×1536 and 2048×2048, and the bill follows the 2.36M-pixel line\[1].
| Tier | 1:1 | 4:3 | 16:9 |
| -------------------------- | --------- | --------- | --------------- |
| **Cheap tier**, ≤ 2.36M px | 1536×1536 | 1776×1328 | **2048×1152** |
| **High tier**, > 2.36M px | 2048×2048 | 2360×1770 | up to 2752×1536 |
Read the 16:9 row twice. At 2048×1152 you get a wide, genuinely large frame for the cheaper rate, because 2,359,296 pixels sits just under the line. Push the same aspect ratio to its maximum and the long edge reaches about 2752, which is roughly 2.7K, at the higher rate.
The practical consequence: for web heroes, thumbnails, social crops, and slide covers, specifying 2048×1152 rather than "2K" keeps a large asset in the cheap tier. That is a sizing decision, not a quality compromise.
## What it costs on reAPI
| Variant | Per image |
| ----------------------------------- | ---------- |
| 1K image (≤ 2.36M px) | **$0.032** |
| 2K image (> 2.36M px) | **$0.063** |
| 1K, output check disabled, API only | $0.034 |
| 2K, output check disabled, API only | $0.067 |
The published rate for the same model is $0.045 and $0.09\[1], so the gateway rate lands roughly 29% lower on both tiers.
The second pair is a variant with the output check disabled, available through the API only, at about a 6% premium. It exists for pipelines where an automated content check on the returned image is not wanted in the loop. If that is not your situation, the standard rows are the ones to use.
## Where it sits against the other image models on reAPI
Comparing against the models on the same gateway is more useful than a cross-vendor table, because these are the choices you can actually route between with one key.
| Model | Per image at 1K |
| -------------------- | --------------- |
| Nano Banana 2 Lite | $0.020 |
| **Seedream 5.0 Pro** | **$0.032** |
| Nano Banana Pro | $0.042 |
Seedream 5.0 Pro sits in the middle: 60% more than Lite, and about 24% less than Nano Banana Pro. That position is the interesting part. It is priced near the draft tier while being a full-quality model rather than a speed-optimized one.
The routing implication is that Seedream 5.0 Pro can be a default rather than an escalation. Where the Lite-to-Pro pattern is draft-cheap-then-render-expensive, a Seedream-centered pipeline can generate at near-draft cost and skip the second render for anything that does not need 4K or Nano Banana Pro's reasoning rating.
## What the price does not buy
Two limits worth knowing before you standardize on it\[1].
**No 4K at launch.** The model is positioned as 2K-class by long edge. If print-adjacent or large-crop output is a requirement today, this tier does not cover it.
**Reference images are capped at 10.** Generous for most work, but below what a heavy multi-character consistency pipeline might want.
What it does cover: in-image text rendering across 15 native languages, which matters for localized creative, and up to 10 references for character and product consistency.
## Calling it
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedream-5-0-pro",
"prompt": "editorial hero for a coffee brand, warm morning light, hand-lettered headline",
"size": "2048x1152"
}'
```
Specifying `2048x1152` rather than a 2K preset is the sizing decision described above: a wide hero frame that bills at the cheaper tier.
Two platform rules. Media inputs are **public http(s) URLs only**, no base64 on any model. And rates move, so the live table on [reapi.ai/models/seedream-5-0-pro](/models/seedream-5-0-pro) is canonical rather than any number quoted here. Request shapes are in the [reapi.ai/docs/seedream-5-0-pro](/docs/seedream-5-0-pro) reference.
## FAQ
### How much does Seedream 5.0 Pro cost per image?
On reAPI, $0.032 for output at or below 2.36M pixels and $0.063 above it. The published rate for the same model is $0.045 and $0.09\[1].
### What decides which tier I am billed at?
Output pixel count, not the tier name. The line is 2.36 million pixels\[1].
### Can I get a large 16:9 image at the cheap rate?
Yes. 2048×1152 is 2,359,296 pixels, just under the boundary, so it bills at the lower tier.
### Does Seedream 5.0 Pro support 4K?
Not at launch. It is positioned as 2K-class by long edge\[1].
### How many reference images can it take?
Up to 10, with in-image text supported across 15 native languages\[1].
### What is the output-check-disabled variant?
A variant available through the API only, at $0.034 and $0.067, for pipelines that do not want an automated check on the returned image in the loop. It carries about a 6% premium.
### How does the price compare to the other image models on reAPI?
It sits between Nano Banana 2 Lite at $0.020 and Nano Banana Pro at $0.042, at $0.032 for 1K.
### Where is the authoritative price?
The [Seedream 5.0 Pro model page](/models/seedream-5-0-pro). Rates change, so treat article figures as planning numbers.
## Sizing against the boundary
The useful thing about the Seedream 5.0 Pro price is not that it is low, though at $0.032 on reAPI it is. It is that the billing rule is a pixel count you can design around rather than a tier you are assigned to.
Decide the output dimensions deliberately. If the asset is a web hero, a social crop, or a slide cover, 2048×1152 gives you a large wide frame in the cheap tier. Reserve the higher tier for the frames that genuinely need the pixels, and check the live rate rather than the one you remember, because the Seedream 5.0 Pro price is a moving number in a table, not a constant.
## References
1. BytePlus. *ModelArk — Seedream 5.0 Pro pricing, resolution tiers, and reference-image limits.* Retrieved July 2026 from [byteplus.com/en/modelark](https://www.byteplus.com/en/modelark)
2. reAPI. *Seedream 5.0 Pro model page — live per-image rates.* Retrieved July 2026 from [reapi.ai/models/seedream-5-0-pro](/models/seedream-5-0-pro)
### Further reading
* reAPI. *Nano Banana Pro vs Nano Banana 2.* [reapi.ai/blog/nano-banana-pro-vs-nano-banana-2](/blog/nano-banana-pro-vs-nano-banana-2)
* reAPI. *How to use Nano Banana 2 Lite.* [reapi.ai/blog/how-to-use-nano-banana-2-lite](/blog/how-to-use-nano-banana-2-lite)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# Seedream 5.0 Pro Seedance 2.5 Workflow: What Actually Ships (https://reapi.ai/blog/seedream-5-0-pro-seedance-2-5-workflow)
The Seedream 5.0 Pro Seedance 2.5 workflow has a scheduling problem: half of it is not callable. Seedream 5.0 Pro shipped as `doubao-seedream-5-0-pro-260628` and sits in ByteDance's public model list today\[3]. Seedance 2.5 does not. As of July 25, 2026, Volcano Engine's documentation portal still answers a search for it with a card saying the Doubao video generation model 2.5 is "releasing in July, stay tuned"\[4], and the model list carries no `doubao-seedance-2-5` entry at all\[3].
The architecture is still right, so this is a build guide against the models that answer HTTP today. Two phases, one worked shot with the costs accumulating as we go, and a note at each step about what changes when 2.5 opens.
## TL;DR
* **Phase 1 is Seedream 5.0 Pro, live now.** 1K and 2K tiers, up to 10 reference images, 14 languages of native in-image text, and exactly one image per call\[1].
* **Phase 2 is the shipping Seedance generation, not 2.5.** Duration 4 to 15 seconds, 480p through 4K, mode decided by which media fields you set\[9].
* **The handoff needs no storage of your own** if your chain runs inside 72 hours: Seedream's output URL is public and Seedance accepts public URLs\[9].
* **Image references do not cut the per-second video rate.** The cheaper row on reAPI's pricing table applies to uploaded reference *video*\[9]. The saving is that you iterate on a $0.032 still instead of a $1.025 clip.
* **One finished shot costs about $1.28** including eight rejected compositions. Eight text-to-video attempts at the same resolution would be $8.20\[7]\[8].
* **Three of the specs circulating about this pipeline are wrong**, including the resolution tiers and the transparent-PNG layer export.
## Phase 1: the keyframe that gets approved
The job of this phase is to produce one still that a human signs off on, as cheaply as possible. Everything else follows from that.
### Which tier to draft at
Seedream 5.0 Pro has two resolution tiers and no more: 1K and 2K, selected on reAPI with `quality: "basic"` or `quality: "high"`\[1]\[6]. At 16:9 that is 1424×800 and 2816×1584, inside a total-pixel envelope of 921,600 to 4,624,220 and an aspect range from 1:16 to 16:1\[1]\[2].
Draft at 1K. It costs $0.032 against $0.063, and 1424×800 is already above 720p, so it is a perfectly adequate video reference on its own\[7]. Re-render the approved frame at 2K only if the still itself ships as a poster or thumbnail.
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer $REAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "doubao-seedream-5-0-pro",
"prompt": "Wide shot of a violinist on a rain-slicked Berlin street at dusk, sodium streetlights, shallow depth of field, poster headline NACHTMUSIK in the upper left",
"aspect_ratio": "16:9",
"quality": "basic"
}'
```
Submission is async. Poll `GET /api/v1/tasks/{id}` until `status` is `completed`, then read `output.image_urls[0]`. Polling is free\[6].
### Prompt a frame, not a scene
The image prompt describes a static composition: subject, framing, lighting, materials, and any copy you want rendered in the image. Seedream renders in-image text well enough to skip a typesetting pass, and ByteDance's own technique for that is to put the exact words in double quotes\[1].
What does not belong here is motion. Save the camera move for phase two, because describing it now produces a still that looks like a frame grab from a pan.
### What to lock before you animate
Three things are cheap to fix in image space and expensive in video: framing, wardrobe, and light direction. Get all three signed off at 1K before a single video credit moves. If your workflow has a human reviewer, this is the gate they should be standing at.
One capability to know about and not expect on reAPI: Seedream 5.0 Pro is the only model in its family with interactive editing, where you mark a region on an input image or write `` / `` tags into the prompt and it edits inside that region only\[1]. reAPI's surface accepts prompt, reference images, aspect ratio and quality tier, and nothing else, so region-level repair is not available there\[6].
## Phase 2: the motion pass
### The handoff is one URL
This is the part I did not expect to like. Seedream's finished image URL stays valid for 72 hours, and Seedance accepts public HTTP(S) URLs in `image_urls`\[9]. If your chain runs inside that window, the entire handoff is copying one string from the image task's output into the video task's input. No download, no bucket, no re-upload. Mirror the asset yourself only if it needs to outlive three days.
```bash
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer $REAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "doubao-seedance-2.0-face",
"prompt": "Slow push-in. The violinist lifts the bow and begins to play; rain streaks catch the streetlight; steam rises behind her",
"image_urls": ["https://cdn.reapi.ai/...jpg"],
"resolution": "720p",
"duration": 5,
"size": "adaptive"
}'
```
### Mode routing, duration, resolution
Mode is implicit. A prompt alone is text-to-video; adding `image_urls` makes the same call image-to-video; a first and last frame interpolates between them; reference videos or audio drive style and sound\[9]. There is no mode parameter to get wrong.
Four fields earn their place in that body. `size: "adaptive"` inherits the still's aspect ratio instead of re-cropping the composition you just approved. `duration` accepts 4 to 15 seconds and defaults to 5. `resolution` defaults to 720p and ranges to 4K. And the `-face` variant is the one that accepts identifiable real people, which is a one-word change to `model` rather than a redesign\[9].
Write this prompt about change, not about the scene. Camera move, what enters, what leaves, what the light does. Re-describing the composition here fights the reference image you just paid to approve, and the model splits the difference badly.
### Getting past fifteen seconds
Fifteen seconds is the ceiling for one generation, so chain instead of stretching. Set `return_last_frame: true`, take the returned frame URL, and pass it as `image_urls` on the next segment\[9]. Each segment starts on the exact pixels the previous one ended on, so there is no drift at the seam.
Chaining also buys something a single long generation cannot: segment three is independently re-rollable. A flaw at second twenty-two costs you one segment, not the twenty-one seconds before it. This is the honest substitute for the 30-second single pass that 2.5 advertises, and I expect plenty of production work will keep choosing it after 2.5 ships.
## One shot, start to finish
Here is the whole thing with the meter running. Rates are reAPI's as of July 25, 2026\[7]\[8].
| Step | Call | Cost | Running total |
| ---------------------------- | ----------------------- | ----------- | ------------- |
| Draft the frame, attempt 1–8 | 1K image × 8 | $0.032 each | $0.256 |
| Test the motion | 480p × 5s, fast variant | $0.078/s | $0.646 |
| Final render | 720p × 5s | $0.205/s | $1.671 |
Drop the optional motion test and the same shot lands at $1.281. Compare that with discovering your framing problem in video: eight text-to-video attempts at 720p is $8.20, and you would still be holding a clip nobody approved as a frame.
The ratio underneath is the thing to remember. A 720p second costs $0.205; a 1K still costs $0.032. So one second of finished video is worth six draft stills, and a five-second clip is worth thirty-two of them. Every composition question you answer in video instead of image space is answered at roughly thirty times the necessary price.
One correction to make loudly, because it is the most common false claim about this pipeline. reAPI's Seedance pricing table has two rows per resolution, and the cheaper one reads "720P & uploaded videos" at $0.125 per second against $0.205 for the plain row. That cheaper row triggers on an uploaded reference **video**. A first-frame image, a first/last-frame pair, and reference images all bill on the plain row\[9]. Going image-first buys control and cheap iteration. It does not buy a lower per-second rate.
## Three specs in the layered-workflow story that don't hold up
**"Layer separation returns transparent PNGs for subject, background, props, and text."** ByteDance's capability table has exactly one line touching this, in the output-count row for 5.0 Pro: "supports generating a single image or multiple layers"\[1]. No documented parameter requests layers, no response field returns them, and no example appears anywhere in the Seedream tutorials. The same table marks group-image output unsupported on this model, and reAPI returns exactly one image per call\[6]. Plan for stills.
**"Resolution tiers are 1.5K and 2K."** They are 1K and 2K\[1]. 1.5K appears in neither the model comparison table, the resolution mapping table, nor the pixel-range spec.
**"Seedance 2.5 brings native 4K at 10-bit."** The shipping 2.0 model already outputs 480p through 4K with 4K in 10-bit H.265, and the 2.5 promotion page mentions no resolution at all\[5]. Carrying an existing capability forward as a headline is how spec sheets get invented.
## FAQ
### Can I use Seedance 2.5 through any API today?
No. It has no model ID in ByteDance's model list and the documentation portal describes it as releasing in July with no further detail\[3]\[4]. Tutorials showing 2.5 API calls are showing a shape, not a working endpoint.
### Does feeding a Seedream image into Seedance lower the video cost?
Not the per-second rate. reAPI's cheaper reference row applies to uploaded reference video only, so a first-frame image bills at the same rate as text-to-video\[9]. The saving is upstream: $0.032 per rejected composition instead of $1.025.
### How many reference images can each phase take?
Seedream 5.0 Pro accepts up to 10, fused into one output\[1]. Seedance 2.0 on reAPI accepts up to 9 images, plus up to 3 reference videos and 3 audio clips, with videos and audio each capped at 15 seconds in total\[9].
### Do I need to re-host the image between phases?
Not within 72 hours\[9]. Base64 and `data:` URIs are rejected, so anything from local disk needs a public URL first.
### Which resolution should the keyframe be?
Draft at 1K, deliver at 2K. A 1K 16:9 still is 1424×800, already above 720p and adequate as a video reference; 2K is 2816×1584\[2].
### What happens if a generation fails?
Charges are placed on submit and refunded automatically when a task ends `failed`, and polling never costs credits\[9].
## Moving the workflow onto 2.5 when it opens
Build both phases now and the migration is small, because the expensive parts of this integration are not model-specific. Async submit and poll, task-lifecycle handling, storage for approved frames, the retry and refund path, and the review gate deciding whether a still deserves to be animated all carry forward untouched. What changes is a model string, probably a wider reference array, and a duration ceiling moving from 15 seconds to 30.
Ignore the version number and build the split. The reason the Seedream 5.0 Pro Seedance 2.5 workflow is worth adopting is not that either model is new. It is that deciding your composition costs three cents and deciding your motion costs a dollar, and any pipeline that puts those two decisions in the same API call is paying video prices to answer image questions.
## References
1. Volcano Engine. *Doubao Seedream 5.0 Pro guide — capability matrix, interactive editing, resolution tiers, reference limits.* Retrieved July 2026 from [volcengine.com/docs/82379/2582774](https://www.volcengine.com/docs/82379/2582774)
2. Volcano Engine. *Image generation guide — Seedream resolution-to-pixel mapping and model comparison.* Retrieved July 2026 from [volcengine.com/docs/82379/1824121](https://www.volcengine.com/docs/82379/1824121)
3. Volcano Engine. *Model list — Ark model IDs for the Seedance and Seedream families.* Retrieved July 2026 from [volcengine.com/docs/82379/1330310](https://www.volcengine.com/docs/82379/1330310)
4. Volcano Engine. *Documentation search — Doubao video generation model 2.5, "releasing in July, stay tuned".* Retrieved July 2026 from [volcengine.com/docs/search?q=Seedance 2.5](https://www.volcengine.com/docs/search?q=Seedance%202.5)
5. ModelArk. *Doubao Seedance 2.5 — 30-Second long narrative, full-mode reference expansion.* Retrieved July 2026 from [ark.volcengine.com/promotion?modelName=seedance-2-5](https://ark.volcengine.com/promotion?modelName=seedance-2-5)
6. reAPI. *Seedream 5.0 Pro API reference — request body, tiers, one-image-per-call, retention.* Retrieved July 2026 from [reapi.ai/docs/seedream-5-0-pro](/docs/seedream-5-0-pro)
7. reAPI. *Seedream 5.0 Pro — live pricing table.* Retrieved July 2026 from [reapi.ai/models/seedream-5-0-pro](/models/seedream-5-0-pro)
8. reAPI. *Seedance 2.0 — live pricing table.* Retrieved July 2026 from [reapi.ai/models/seedance-2-0](/models/seedance-2-0)
9. reAPI. *Seedance 2.0 API reference — mode routing, duration and resolution limits, billing, frame chaining.* Retrieved July 2026 from [reapi.ai/docs/seedance-2-0](/docs/seedance-2-0)
### Further reading
* reAPI. *Seedance 2.5 Release Status: What's Real, What's Keynote.* [reapi.ai/blog/seedance-2-5-release-status](/blog/seedance-2-5-release-status)
* reAPI. *Seedance 2.0 character consistency guide.* [reapi.ai/blog/seedance-2-0-character-consistency-guide](/blog/seedance-2-0-character-consistency-guide)
* reAPI. *Seedance 2.5 model page.* [reapi.ai/models/seedance-2-5](/models/seedance-2-5)
---
# Is Seedream 5 Pro Uncensored? NSFW Checker and API Control (2026) (https://reapi.ai/blog/seedream-5-pro-nsfw-checker)
**Seedream 5 Pro is not literally uncensored. On reAPI, direct API callers
can set `nsfw_checker: false` to skip an additional final-image output check,
but the switch does not remove the model's upstream policy, acceptable-use
rules, copyright, consent requirements, or the moderation your own product
needs.** The accurate description is a less-restricted delivery path, not a
policy-free image model.\[1]
That distinction answers the confusion behind searches for “Seedream 5 Pro
uncensored,” “Seedream 5 Pro NSFW,” and “Seedream 5 Pro censorship.” A blocked
request can come from four different layers, and `nsfw_checker` controls only
one of them. It also changes the billing outcome in a way production teams need
to understand before adding an automatic retry.
## TL;DR
* **The safe default is on.** `nsfw_checker` defaults to `true`.
* **Direct API callers have a control.** Send `"nsfw_checker": false` to skip
the additional final-image output check. The hosted playground keeps it on.
* **It does not override the model.** An upstream policy refusal can still
stop the request before an image is delivered.
* **Timing determines billing.** A request rejected before generation is
refunded. An image generated and then hidden by the enabled output check
remains charged.\[1]
* **The disabled-check path has its own SKU.** Check the live model page rather
than hard-coding a price from an article.\[2]
* **“Uncensored” does not mean unrestricted.** Provider terms, applicable law,
rights, consent, and downstream platform policies still apply.
## What “Seedream 5 Pro uncensored” actually means
There is no standard definition of an uncensored image API. Different hosts
use the word for different product decisions, so the useful question is not
“Is it uncensored?” but “Which check can be disabled?”
| Label | What it may mean | What it does not prove |
| -------------------- | ----------------------------------------------------- | ----------------------------------------------------- |
| No prompt filter | The host adds no extra keyword gate before submission | The upstream model accepts every prompt |
| NSFW checker off | One post-generation output check is skipped | Every safety or policy layer is gone |
| Relaxed moderation | Fewer lawful borderline images become false positives | Prohibited material is permitted |
| Unrestricted editing | A host accepts a wider range of reference-image edits | Consent and image rights are optional |
| Fully uncensored | A marketing claim of no checks at all | A commercial API actually has no terms or enforcement |
For reAPI, the product behavior is specific: `nsfw_checker` controls the
additional final-image check. It is enabled by default, locked on in the
browser playground, and configurable by direct API requests.\[1]
This control can reduce false positives for lawful work such as editorial
fashion, stylized portraits, clinical diagrams, theatrical makeup, or product
photography that a broad classifier misreads. It is not a prompt-obfuscation
tool and does not grant permission to create or distribute prohibited content.
## Seedream 5 Pro censorship has four separate layers

The same prompt can behave differently across products because each product
can add checks around the model. A useful diagnosis starts by identifying the
layer that returned the refusal.
### 1. Upstream model policy
The model provider decides what the generation system will attempt. A request
can be rejected before useful output is produced, regardless of the host's
additional settings. `nsfw_checker: false` does not rewrite that upstream
policy.
BytePlus maintains acceptable-use terms for ModelArk generative AI services.
The fact that an API exposes a Boolean output control does not replace those
terms.\[4]
### 2. The additional output check
This is the layer reAPI exposes as `nsfw_checker`. With the default `true`, the
check runs after generation. If it flags the image, the task returns content
policy error `80006`, the image URL is hidden, and the completed generation is
still charged.\[1]\[5]
With `false`, that additional final-image check is skipped. The upstream model
can still refuse the request, and your application may still need to review
the resulting image before displaying or publishing it.
### 3. Host or platform rules
A marketplace, image editor, or API reseller may run another prompt filter,
output classifier, account gate, or category block. Two sites can expose the
same nominal Seedream model while applying different policies around it.
That is why community reports often conflict. One user is describing the
model; another is describing a host; a third is describing an app-store-safe
mobile client. “Seedream is censored” compresses three different systems into
one sentence.
### 4. Your application's client settings
Your own SaaS may restrict a route by workspace, user role, age gate,
distribution channel, or feature flag. Those controls happen after the model
choice and remain your responsibility.
An internal creative tool and a public image feed should not inherit the same
moderation configuration automatically. Decide the policy per surface, not
once for the entire account.
## What `nsfw_checker: false` changes—and what it does not
The field name invites a larger promise than the API makes. The practical
contract is narrower.
| Behavior | `true` or omitted | `false` |
| ------------------------------- | ----------------- | ---------------------------------------- |
| Additional final-image check | Enabled | Skipped |
| Browser playground | Always enabled | Not available to normal playground users |
| Direct API | Supported | Supported when explicitly sent |
| Upstream model policy | Still applies | Still applies |
| Number of returned images | Exactly one | Exactly one |
| Rights and acceptable-use rules | Still apply | Still apply |
| Pricing | Standard SKU | Separate disabled-check SKU |
It does not increase resolution, improve prompt adherence, make editing more
accurate, or change the one-image-per-call limit. It also does not promise that
an image will be generated. The flag affects moderation behavior, not the
model's core image capabilities.
The browser playground intentionally keeps the check enabled. Teams that need
the API-only behavior should use a server-side integration, keep the API key
private, and decide which trusted workflows are allowed to send the flag.
## The billing difference: when the image was blocked matters
Seedream 5 Pro moderation failures do not all settle the same way. The key
question is whether the image had already been generated when the block
occurred.
| Outcome | Was generation completed? | Image delivered? | Billing result |
| ---------------------------------------------- | :-----------------------: | :--------------: | ------------------------------ |
| Prompt or reference rejected before generation | No | No | Refunded |
| Provider or infrastructure failure | No usable output | No | Refunded |
| Enabled output check flags the completed image | Yes | No | Charged |
| Successful request with the check enabled | Yes | Yes | Charged |
| Successful request with the check disabled | Yes | Yes | Charged under the matching SKU |
A charged, hidden result is not the same as a failed provider call. Compute was
used and an image was produced before the final check withheld it. This is why
a retry loop that treats every missing URL as a transient error can spend money
on the same outcome repeatedly.
Log the task's structured status and error code. Retry infrastructure errors
according to your normal backoff policy, but treat `80006` as a content-policy
decision rather than a signal to resubmit indefinitely.\[5]
## How to call the Seedream 5 Pro API with the check disabled
The current public model ID is `doubao-seedream-5-0-pro`. Requests use
`aspect_ratio` and `quality`, not the older display-name model ID and arbitrary
size field found in some launch-week examples.\[1]
This safe example demonstrates the parameter without using an explicit or
policy-sensitive prompt:
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer $REAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "doubao-seedream-5-0-pro",
"prompt": "Editorial poster for an imaginary perfume called LUMEN, cobalt glass bottle on cream paper, dramatic red side light, precise readable title, no real brands",
"aspect_ratio": "16:9",
"quality": "high",
"nsfw_checker": false
}'
```
Submission is asynchronous. Poll the shared task endpoint until it completes:
```bash
curl https://reapi.ai/api/v1/tasks/TASK_ID \
-H "Authorization: Bearer $REAPI_API_KEY"
```
A completed task returns one URL in `output.image_urls`. Result URLs remain
valid for 72 hours, so production systems should mirror accepted assets to
their own storage.\[1]
For image-to-image work, add `image_urls` with up to ten public HTTP(S)
references. Base64 and `data:` URLs are not accepted. The official BytePlus
image generation guide likewise documents text, single-image, and multi-image
Seedream workflows.\[3]
## Seedream 5.0 Pro limits that stay the same
Moderation controls do not change the model contract:
| Item | Current reAPI contract |
| ------------------- | ----------------------------------- |
| Output | Exactly one image per request |
| Quality | `basic` for 1K; `high` for 2K |
| 4K tier | Not exposed |
| Aspect ratios | 1:1, 4:3, 3:4, 16:9, 9:16, 2:3, 3:2 |
| Reference images | Up to 10 public HTTP(S) URLs |
| Response mode | Async task polling |
| Result URL lifetime | 72 hours |
The absence of `n` is easy to miss. If a workflow needs four candidates, send
four requests and track four task IDs. Do not add an unsupported batch count
and expect the API to fan out automatically.
Pricing is per output image and varies by quality, reference count, and the
matching moderation SKU. Use the [live Seedream 5.0 Pro model page](/models/seedream-5-0-pro)
for current numbers rather than copying a percentage into application code.\[2]
## Why a safe Seedream prompt can still be blocked
False positives happen because classifiers make decisions from incomplete
context. A fully clothed editorial portrait, skin-toned packaging, clinical
terminology, theatrical makeup, or an unusually tight crop can resemble the
features a broad classifier is designed to catch.
For lawful, clearly safe material, troubleshoot the layer instead of trying to
hide the prompt's meaning:
1. Confirm that the prompt and every reference asset are original, licensed,
or used with permission.
2. Keep the prompt, quality, aspect ratio, and references identical.
3. Run one request with `nsfw_checker: true` and one with `false` through a
trusted direct API workflow.
4. Compare structured errors and task timing, not only whether a URL appears.
5. If both requests fail upstream, stop retrying; the optional output check was
not the cause.
When wording is genuinely ambiguous, add context that makes the lawful intent
clear. Do not use misspellings, encoded phrases, or other filter-evasion tricks.
Those techniques make logs harder to review and do not solve an upstream
policy decision.
## Building a less-restricted image route responsibly
The API flag should be one part of a product policy, not the policy itself.
For a SaaS integration, a sound baseline is:
* allow the disabled-check path only for authenticated, trusted workflows;
* keep task-to-user, prompt, reference-asset, and policy-outcome logs;
* require rights and consent attestations where people or protected assets are
involved;
* rate-limit repeated policy failures and suspicious request bursts;
* scan or review output before it reaches a public gallery or end user;
* provide reporting, takedown, and account-enforcement controls;
* apply the destination platform's rules before publication.
This design lets an internal team reduce costly false positives without making
the same configuration available to every anonymous user. It also preserves a
clear audit trail when an image is withheld or later reported.
## FAQ
### Is Seedream 5 Pro uncensored?
No. reAPI lets direct API callers skip an additional final-image NSFW check by
sending `nsfw_checker: false`, but upstream model policy, acceptable-use terms,
rights, law, and downstream moderation still apply.
### Does Seedream 5 Pro allow NSFW content?
The API field should not be read as blanket approval for a content category.
It controls one output check. The model provider can still reject a request,
and callers remain responsible for what they generate and distribute.
### How do I turn off the Seedream safety filter?
Direct reAPI requests can send `"nsfw_checker": false`. The hosted playground
keeps the field enabled. This skips the additional final-image check; it does
not turn off every policy layer.\[1]
### Am I charged when a Seedream image is blocked?
It depends on when the block occurs. A request rejected before generation is
refunded. A completed image hidden by the enabled output check is still
charged because the generation already ran.
### Does the disabled-check route cost more?
It uses a separate SKU. Prices can change independently, so consult the live
model page for the current 1K and 2K rows instead of relying on an old fixed
percentage.\[2]
### Can I disable `nsfw_checker` in the playground?
No. It is locked on for normal browser playground users. Use a direct API call
from a server-side integration if your trusted workflow needs the optional
setting.
### Why does the same prompt work on one Seedream platform and fail on another?
Hosts and apps add their own prompt filters, output checks, account rules, and
client restrictions. A different result does not prove the underlying model
changed; first identify which layer returned the block.
## “Less restricted” is the defensible Seedream label
The opportunity behind the Seedream 5 Pro uncensored keyword is real, but the
answer needs precision. `nsfw_checker: false` gives direct API users documented
control over one additional output layer. It does not erase upstream policy or
shift responsibility away from the caller.
Use the default for ordinary requests. Enable the API-only path selectively
for trusted, lawful workflows where false positives justify it. Then log the
task outcome, distinguish pre-generation refunds from post-generation billed
blocks, and moderate the asset for its actual destination. The
[Seedream 5.0 Pro API reference](/docs/seedream-5-0-pro) lists the current
request contract, and the [model page](/models/seedream-5-0-pro) carries live
pricing.
**Disclosure:** reAPI publishes this article and operates the Seedream 5.0 Pro
endpoint described above. reAPI parameter and billing behavior comes from our
public API documentation. Official model capabilities and acceptable-use
requirements come from BytePlus documentation. “Uncensored” is discussed as a
search and marketing term, not as a promise that prohibited content is
accepted.
## References
1. reAPI. *Seedream 5.0 Pro API reference — `nsfw_checker`, output behavior, billing, parameters, and limits.* Retrieved August 1, 2026. [reapi.ai/docs/seedream-5-0-pro](/docs/seedream-5-0-pro)
2. reAPI. *Seedream 5.0 Pro model page and live pricing.* Retrieved August 1, 2026. [reapi.ai/models/seedream-5-0-pro](/models/seedream-5-0-pro)
3. BytePlus ModelArk. *Image generation tutorial — Seedream text, single-image, and multi-image workflows.* Retrieved August 1, 2026. [docs.byteplus.com/api/docs/ModelArk/1824121](https://docs.byteplus.com/api/docs/ModelArk/1824121)
4. BytePlus. *GenAI Acceptable Use Policy.* Updated July 11, 2026. [docs.byteplus.com/en/docs/legal/acceptable\_use\_policy\_byteplus\_genai](https://docs.byteplus.com/en/docs/legal/acceptable_use_policy_byteplus_genai)
5. reAPI. *API error codes and content-policy failure responses.* Retrieved August 1, 2026. [reapi.ai/docs/api/errors](/docs/api/errors)
### Further reading
* reAPI. *Seedream 5.0 Pro Price: The Pixel Line That Decides It.* [reapi.ai/blog/seedream-5-0-pro-price](/blog/seedream-5-0-pro-price)
* reAPI. *Seedream 5.0 Pro Bugs: What's Documented, What's Random.* [reapi.ai/blog/seedream-5-0-pro-bugs-and-fixes](/blog/seedream-5-0-pro-bugs-and-fixes)
* reAPI. *Seedream to Seedance: The Keyframe Handoff Workflow.* [reapi.ai/blog/seedream-seedance-handoff](/blog/seedream-seedance-handoff)
---
# Seedream to Seedance: The Keyframe Handoff Workflow (https://reapi.ai/blog/seedream-seedance-handoff)
Text-to-video has a structural problem that no prompt fixes: changing one detail rewrites the whole scene. Adjust the lighting and the character's face shifts. Fix the face and the background moves. Every iteration is a fresh roll.
The workaround is to stop asking one model to do two jobs. Lock the frame first with an image model, then hand that approved frame to a video model as reference. That is the Seedream Seedance workflow, and its value is not higher quality per generation. It is that approval becomes durable.
## TL;DR
* **Two phases, two models.** Seedream 5.0 Pro settles what the scene looks like. Seedance handles what moves.
* **The handoff is a reference image**, not a re-description. The approved frame goes in as `image_urls`.
* **Seedream takes up to 10 references** with anchored edits and multilingual typography at 2K.
* **Iterate cheaply on the still**, where a regeneration costs cents and seconds rather than dollars and minutes.
* **Approve motion before details.** Once the camera move is right, remaining fixes are local.
* **Some layered-workflow claims do not survive checking.** We went through them separately in [what actually ships](/blog/seedream-5-0-pro-seedance-2-5-workflow).
## Why decoupling works

A video generation is expensive in both senses: it costs more per call than an image, and it takes long enough that a bad result wastes real time. Pushing every creative decision through that loop is why teams report burning most of their schedule on retakes.
Splitting the pipeline changes what each iteration costs.
**Phase one is cheap and fast.** Composition, lighting, product geometry, typography, brand color. These are still-frame decisions, and settling them on an image model means each attempt costs a few cents and returns in seconds.
**Phase two inherits a settled frame.** Motion generation starts from something already approved rather than re-deriving the scene from a paragraph. Fewer variables move, so fewer things break.
The economic argument is straightforward: the expensive call should run once, after the cheap calls have removed the uncertainty.
## Phase one: settle the frame
Seedream 5.0 Pro is the asset step. What matters for the handoff:
**Up to 10 reference images** with anchored edits, so a product or character can be pinned rather than described.
**Multilingual typography at 2K**, preserving exact color and layout, which matters when the frame carries brand text that must survive into the video.
**Pixel-budget billing**, split at 2.36M pixels. A 16:9 frame at 2048×1152 lands in the cheaper tier, which is the right size for most keyframes anyway. Details in [the pricing breakdown](/blog/seedream-5-0-pro-price).
The output of this phase is not a folder of options. It is **one approved frame per shot**, plus the prompt that produced it.
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedream-5-0-pro",
"prompt": "product hero on a stone counter, soft window light from camera left, brand wordmark legible on the label",
"image_urls": ["https://example.com/packshot.jpg"],
"size": "2048x1152"
}'
```
## Phase two: hand the frame forward
The handoff is the part people get wrong. The instinct is to describe the approved frame again in the video prompt. That re-introduces exactly the ambiguity phase one removed.
Pass the frame itself:
```bash
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "seedance-2-5",
"prompt": "0-5s: [locked medium] no camera movement, steam rises from the cup. 5-12s: [slow dolly in] 0.3 m/s, product stays centered. Do not add text or logos.",
"image_urls": ["https://cdn.example.com/approved-keyframe.png"],
"resolution": "1080p",
"duration": 12
}'
```
Note what the video prompt does and does not contain. It describes **motion and constraints**. It does not re-describe the lighting, the composition, or the product, because the reference frame already carries those.
That division is the discipline. Anything visible in the keyframe belongs in the keyframe, not in the motion prompt.
## What each phase owns
| Decision | Phase | Why |
| ------------------------- | ----- | --------------------------------------------- |
| Composition and framing | Image | Cheap to iterate, expensive to fix in motion |
| Lighting direction | Image | Drifts if only described in a video prompt |
| Product geometry and logo | Image | Anchored by reference rather than generated |
| Typography and brand text | Image | Where text rendering is strongest |
| Camera path and velocity | Video | Motion is the video model's job |
| Timing and beats | Video | Timestamped structure belongs with the motion |
| Negative constraints | Both | Text suppression matters in each phase |
The rows that generate the most confusion are lighting and framing. Both feel like they belong in the video prompt because that is where the final output comes from. Putting them there means re-deciding them on every generation.
## Where the savings actually come from
The 60%-faster-iteration claims floating around this workflow are marketing figures, not measurements. The mechanism underneath them is real though, and it is worth stating plainly rather than as a percentage.
**A rejected still costs one image generation.** A rejected video costs a video generation plus the wait.
**An approved still is reusable.** The same keyframe can drive several motion variants, different durations, different camera moves, without regenerating the scene.
**Local edits stay local.** Fixing a highlight on the keyframe does not disturb a camera move you already approved, because the move has not been generated yet.
That third point is the one teams discover late. Sequencing approval so the cheap decisions settle first is not a productivity hack, it is what makes the expensive call worth making.
## Where the pipeline breaks
**Re-describing the frame in the video prompt.** The most common error. It gives the video model permission to reinterpret what was already settled.
**Approving a keyframe at the wrong aspect ratio.** Generate the still at the aspect ratio the video will use. Cropping a 1:1 keyframe into a 16:9 shot reintroduces framing decisions.
**Skipping the still for "simple" shots.** Simple shots are exactly where the still is cheapest, so the saving ratio is highest.
**Treating the reference as a suggestion.** If the output drifts from the keyframe, add explicit constraints tying the video to it rather than rewriting the description.
## FAQ
### Why use two models instead of one text-to-video call?
Because a text-to-video call re-derives the entire scene each time, so changing one detail changes everything. Settling the frame first makes approval durable across iterations.
### What exactly gets passed between the phases?
The approved keyframe as a reference image, plus a motion prompt. The visual description does not get repeated.
### Should the video prompt describe the lighting?
No. Anything visible in the keyframe belongs in the keyframe. The motion prompt covers camera path, timing, and constraints.
### What aspect ratio should the keyframe be?
The same one the video will use. Generating a still at a different ratio and cropping reintroduces the framing decision you already made.
### Can one keyframe drive several videos?
Yes, and that is much of the value. Different durations, camera moves, and cutdowns can share an approved frame without regenerating the scene.
### Does this work for multi-shot sequences?
Yes. Approve one keyframe per shot, then generate each shot's motion from its own reference. Consistency across shots comes from the stills sharing references, not from one long prompt.
### Are the layered-output claims about this workflow accurate?
Several widely repeated specs do not hold up against the official capability tables. We checked them in [what actually ships](/blog/seedream-5-0-pro-seedance-2-5-workflow).
## Sequencing approval, not just generation
The Seedream Seedance workflow gets described as a quality technique, and it is not really. Both models produce what they produce whether or not you chain them.
What chaining changes is the order in which decisions become final. Composition, lighting, and brand detail get settled on a surface where being wrong costs cents. Motion gets generated once, against a frame nobody is still arguing about. The pipeline is worth building for that reason alone, and the teams who report the largest gains are the ones who moved approval earlier, not the ones who found a better prompt.
## References
1. Volcano Engine. *Seedream and Seedance model families.* Retrieved July 2026 from [volcengine.com](https://www.volcengine.com/)
### Further reading
* reAPI. *Seedream 5.0 Pro Seedance 2.5 workflow: what actually ships.* [reapi.ai/blog/seedream-5-0-pro-seedance-2-5-workflow](/blog/seedream-5-0-pro-seedance-2-5-workflow)
* reAPI. *Seedance 2.5 camera control.* [reapi.ai/blog/seedance-2-5-camera-control](/blog/seedance-2-5-camera-control)
* reAPI. *Seedream 5.0 Pro price.* [reapi.ai/blog/seedream-5-0-pro-price](/blog/seedream-5-0-pro-price)
---
# How to Stop Claude Code Asking for Permission Every Time (https://reapi.ai/blog/stop-claude-code-permission-prompts)
Claude Code asks before it acts. That is the design, and on a first pass through an unfamiliar repository it is the right default. Twenty prompts into a refactor it stops feeling like safety and starts feeling like friction.
There is an official answer, and it is not a flag that turns checking off. Claude Code ships **six permission modes**, and the one built for this problem routes actions through a separate classifier model instead of through you.
## TL;DR
* **Six modes**: `default` (shown as **Manual**), `acceptEdits`, `plan`, `auto`, `dontAsk`, `bypassPermissions`\[1].
* **`auto` is the one you want** for long tasks. A classifier reviews each action; you stop seeing routine prompts\[1].
* **`Shift+Tab` cycles modes mid-session.** The status bar shows which one is live\[1].
* **Permission rules are the surgical fix.** Pre-approve the specific commands you keep approving, with `allow`, and keep everything else prompting\[2].
* **A real trap**: `defaultMode: "auto"` is ignored in `.claude/settings.json`. It has to live in `~/.claude/settings.json`\[1].
* **`bypassPermissions` is not the answer** outside isolated containers and VMs\[1].
## The six modes

| Mode | What runs without asking | Built for |
| ---------------------- | -------------------------------------------------------------------- | --------------------------------------- |
| `default` (**Manual**) | Reads only | Getting started, sensitive work |
| `acceptEdits` | Reads plus file edits | Editing sessions you are watching |
| `plan` | Reads, plus classifier-approved commands when auto mode is available | Exploring before changing |
| **`auto`** | **Everything, with background safety checks** | **Long tasks, reducing prompt fatigue** |
| `dontAsk` | Only pre-approved tools | Locked-down CI and scripts |
| `bypassPermissions` | Everything | Isolated containers and VMs only |
Source is Anthropic's permission-modes documentation\[1]. Note the naming: the mode that reviews every action is labeled **Manual** in the CLI and the extensions, while its config value stays `default`. `manual` works as an alias wherever you type the value, on v2.1.200 and later.
## Switching mid-session
Press **`Shift+Tab`** to cycle `default` → `acceptEdits` → `plan`. The status bar shows the active mode: `⏸ manual mode on`, `⏵⏵ accept edits on`, `⏵⏵ auto mode on`, `⏵⏵ don't ask on`, or `⏵⏵ bypass permissions on`\[1].
Not every mode is in that cycle by default:
* **`auto`** appears once your account meets its requirements, and cycling into it does not ask for confirmation.
* **`bypassPermissions`** only appears if you started with `--permission-mode bypassPermissions`, `--dangerously-skip-permissions`, `--allow-dangerously-skip-permissions`, or set it as `defaultMode`.
* **`dontAsk`** never appears. Set it with `--permission-mode dontAsk`.
At startup, pass it as a flag:
```bash
claude --permission-mode auto
```
## Auto mode, and what it actually checks
Auto mode does not remove review. It moves review off your keyboard and onto a separate classifier model that inspects actions before they run, blocking anything that escalates beyond what you asked for, targets unrecognized infrastructure, or looks driven by hostile content Claude read\[1].
Explicit `ask` rules still force a prompt, so anything you have deliberately marked as needing confirmation keeps confirming.
**What it blocks by default**\[1]:
* Downloading and executing code, such as `curl | bash`
* Sending sensitive data to external endpoints
* Production deploys and migrations
* Mass deletion on cloud storage
* Granting IAM or repo permissions
* Force push
* `git reset --hard`, `git checkout -- .`, `git restore .`, `git clean -fd`, `git stash drop`, `git stash clear`
* `git commit --amend` when the HEAD commit was not created in this session, or has already been pushed
* `terraform destroy`, `pulumi destroy`, `cdk destroy`, `terragrunt destroy`
* Irreversibly destroying files that existed before the session
* Committing or pushing a change that would send secrets outside the repository, or widen what a deploy exposes
The classifier trusts your working directory and the git remotes configured when the session started. A remote added or repointed mid-session with `git remote add` or `git remote set-url` is **not** trusted\[1].
Anthropic states the limit plainly: auto mode reduces prompts but does not guarantee safety. It is for tasks where you trust the general direction, not a substitute for reviewing sensitive operations\[1].
### Requirements
Auto mode is available only when all of these hold\[1]:
* **Plan**: all plans.
* **Owner**: on Team and Enterprise, an Owner must enable it in Claude Code admin settings first. Admins can also force it off with `permissions.disableAutoMode: "disable"` in managed settings.
* **Model**: on the Anthropic API, Claude Opus 4.6 or later, Sonnet 4.6 or later, or Fable 5. On Bedrock, Google Cloud's Agent Platform, and Microsoft Foundry, only Sonnet 5, Opus 4.7 or later, and Fable 5. Older models including Sonnet 4.5, Opus 4.5, Haiku, and claude-3 are unsupported everywhere.
* **Provider**: available by default on the Anthropic API, Claude Platform on AWS, Bedrock, Agent Platform, Foundry, and signed-in Claude apps gateway sessions.
If Claude Code reports auto mode unavailable, one of those is unmet. It is not a transient outage.
## The settings trap
This one wastes real time. If you set `defaultMode: "auto"` and the session still starts in Manual with no error, the setting is probably in the wrong file.
From v2.1.142, Claude Code **ignores `auto` in `.claude/settings.json` and `.claude/settings.local.json`**, so a repository cannot grant itself auto mode\[1]. It has to be in your user settings:
```json
// ~/.claude/settings.json
{
"permissions": {
"defaultMode": "auto"
}
}
```
Settings files hot-reload, so `permissions` changes apply to a running session without a restart\[2].
## The surgical fix: permission rules
Modes set a baseline. Rules are what you reach for when the same three commands keep prompting and everything else is fine as it is.
```json
// ~/.claude/settings.json
{
"permissions": {
"allow": ["Bash(git diff *)", "Bash(npm test *)"],
"ask": ["Bash(git push *)"],
"deny": ["Read(./.env)", "Read(./secrets/**)", "Bash(curl *)"]
}
}
```
Three things worth knowing about how these behave\[2]:
**Rules merge across scopes rather than override.** Unlike most settings, where a project value replaces a user value, permission rules from user, project, local, and managed settings all stay in effect.
**`deny` and explicit `ask` apply in every mode**, including `bypassPermissions`. That makes `deny` the right place for secrets: `Read(./.env)` holds regardless of what mode someone switches into.
**Local allow rules skip the workspace-trust step.** Rules in your own `.claude/settings.local.json` take effect without the trust prompt that `.claude/settings.json` allow rules require, because that file is yours rather than the repository's. If the repo commits the file, trust applies again.
## Why not just bypass everything
`bypassPermissions` exists and it does what the name says. Anthropic scopes it to isolated containers and VMs, and the reason is structural rather than cautious.
Writes to protected paths are never auto-approved in any mode except `bypassPermissions`\[1]. Those protections guard repository state and Claude's own configuration against accidental corruption. Turning them off on a machine that holds anything you care about removes the last thing standing between a bad command and your working tree.
If prompt fatigue is the problem, `auto` solves it while keeping a classifier in the loop. Reach for `bypassPermissions` only when the whole environment is disposable.
## Picking a setup
**Working in an unfamiliar repo**: stay in Manual. The prompts are doing their job.
**Editing session you are actively watching**: `acceptEdits` via `Shift+Tab`. File edits stop prompting, commands still do.
**Long autonomous task**: `auto`, set as `defaultMode` in `~/.claude/settings.json` if you want it every session.
**The same command prompting twenty times**: an `allow` rule for that command, not a mode change.
**CI or a script**: `dontAsk` with an explicit allowlist, so anything unlisted fails rather than waits.
**Disposable container**: `bypassPermissions`, and only there.
## FAQ
### How do I stop Claude Code asking for permission?
Switch to auto mode, which routes actions through a classifier instead of prompting you. Press `Shift+Tab` to cycle to it in-session, start with `claude --permission-mode auto`, or set `defaultMode: "auto"` in `~/.claude/settings.json`\[1].
### Why is my `defaultMode: "auto"` being ignored?
Because it is in `.claude/settings.json` or `.claude/settings.local.json`. Claude Code ignores `auto` from those files so a repository cannot grant itself the mode. Move it to `~/.claude/settings.json`\[1].
### Does auto mode mean no safety checks?
No. A separate classifier reviews each action and blocks escalation, unrecognized infrastructure, destructive git operations, production deploys, and more. Explicit `ask` rules still prompt\[1].
### Why is auto mode unavailable for me?
One of the requirements is unmet: plan, an Owner enabling it on Team or Enterprise, a supported model, or a supported provider. It is not a transient failure\[1].
### How do I stop one specific command from prompting?
Add an `allow` rule for it rather than changing mode: `"allow": ["Bash(npm test *)"]`\[2].
### What is the difference between Manual and default?
Nothing. `default` is the config value; **Manual** is the label shown in the CLI and extensions. `manual` works as an alias on v2.1.200 and later\[1].
### Can I block Claude from reading `.env`?
Yes, with a `deny` rule. Deny rules apply in every mode including `bypassPermissions`\[2].
### Is `bypassPermissions` safe on my laptop?
No. It is scoped to isolated containers and VMs, and it is the only mode where writes to protected paths are auto-approved\[1].
## Matching the mode to the risk
The instinct when prompts pile up is to look for the switch that turns checking off. That switch exists, and it is the wrong one for a machine you care about.
The better framing is that Claude Code gives you three separate dials: a **mode** that sets the baseline, **rules** that carve out specific tools in either direction, and a **classifier** that reviews what the mode would otherwise wave through. Prompt fatigue on a long task is a mode problem, solved by `auto`. The same command asking twenty times is a rules problem, solved by one `allow` line. Knowing how to stop Claude Code asking for permission is mostly knowing which of those two you actually have.
## References
1. Anthropic. *Claude Code permission modes — the six modes, auto mode requirements, and what the classifier blocks.* Retrieved July 2026 from [docs.claude.com/en/docs/claude-code/permission-modes](https://docs.claude.com/en/docs/claude-code/permission-modes)
2. Anthropic. *Claude Code settings — permission rules, scopes, and precedence.* Retrieved July 2026 from [docs.claude.com/en/docs/claude-code/settings](https://docs.claude.com/en/docs/claude-code/settings)
### Further reading
* reAPI. *How to use Claude Code.* [reapi.ai/blog/how-to-use-claude-code](/blog/how-to-use-claude-code)
* reAPI. *How to use Claude Opus 5.* [reapi.ai/blog/how-to-use-claude-opus-5](/blog/how-to-use-claude-opus-5)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# Unbelievable: Run Kimi K3 — 2.8 Trillion Parameters on a Single 4GB GPU (https://reapi.ai/blog/unbelievable-run-kimi-k3-2-8-trillion-parameters-on-a-single-4gb-gpu)
**Can you run Kimi K3 on a 4GB GPU? Not
in the ordinary meaning of “run locally.”** Kimi K3's native MXFP4 weights
need roughly 1.4 TB before metadata and runtime overhead. A 4GB card can only
act as a small staging device if software streams tiny pieces of the model
from disk or host memory. That is technically interesting, but it is not the
same as loading Kimi K3 into 4GB of VRAM, and no current mainstream Kimi K3
recipe makes it a practical setup.\[1]\[2]
The distinction matters because three true facts are often combined into one
misleading conclusion: Kimi K3 is sparse, only 104B parameters are active per
token, and layer-offloading tools have run smaller 70B models on 4GB GPUs. None
of those facts makes a 2.8T checkpoint fit in 4GB. This guide does the memory
math, explains what offloading really changes, checks the current software
support, and shows the low-hardware route that works today.
## TL;DR
* **Kimi K3 really has 2.8T total parameters.** It is a sparse Mixture-of-Experts
model with 104B activated parameters per token, 93 layers, 896 experts, and a
1M-token context window.\[1]
* **MXFP4 does not make it small.** Four bits per parameter puts the raw weight
floor at about 1.4 TB decimal, or 1.27 TiB, before scales, metadata, non-4-bit
tensors, and runtime state.
* **104B active is a compute figure, not a storage figure.** The router can
choose different experts on the next token, so all experts must remain
accessible somewhere.
* **A 4GB GPU can only be a staging area.** Disk or CPU offloading moves model
pieces through VRAM; it does not remove the need to store the full model.
* **AirLLM does not currently document Kimi K3 support.** Its published 4GB
example targets Llama 3 70B, while Kimi K3 uses a new custom multimodal MoE
architecture.\[4]
* **The practical route is an API.** A low-end laptop can call Kimi K3 through
an OpenAI-compatible endpoint while the model runs on remote infrastructure.
## Kimi K3 hardware reality at a glance
| Item | Kimi K3 |
| ---------------------- | ------------------------------------------------ |
| Total parameters | 2.8 trillion |
| Activated per token | 104 billion |
| Architecture | Sparse MoE with KDA and Gated MLA attention |
| Routed experts | 896, with 16 selected per token |
| Layers | 93 |
| Native weight format | MXFP4 weights |
| Activation format | MXFP8 |
| Context window | 1,048,576 tokens |
| Raw 4-bit weight floor | About 1.4 TB decimal / 1.27 TiB |
| 4GB GPU verdict | Cannot hold the model; experimental staging only |
Moonshot describes Kimi K3 as an open-weight, native multimodal agentic model
built on Kimi Delta Attention, Attention Residuals, and Stable LatentMoE. It
activates 16 of 896 routed experts for each token and includes a MoonViT-V2
vision encoder.\[1]\[2]
Those architectural choices reduce compute and long-context cost. They do not
turn a trillion-scale checkpoint into a consumer-GPU model.
## Why 2.8 trillion MXFP4 parameters still need about 1.4 TB
The first calculation is simple:
```text
2.8 trillion parameters × 4 bits
= 11.2 trillion bits
= 1.4 trillion bytes
= about 1.27 TiB
```
This is a lower bound, not a complete deployment estimate. Real checkpoints
also carry quantization scales, indexes, configuration, embeddings, tensors
stored at other precisions, and the 401M-parameter vision encoder. Runtime adds
activations, selected-expert working memory, attention state, CUDA or NPU
kernels, and a KV or recurrent-state budget that grows with concurrency and
context settings.\[1]
The official Hugging Face checkpoint is split into many large safetensor
shards; individual listed shards are measured in tens of gigabytes. One shard
can exceed the entire capacity of a 4GB card before the inference engine has
allocated a single activation buffer.\[3]
For comparison, a 4GB card can theoretically hold around 8 billion plain
4-bit parameters if every byte were available for weights. In practice it can
hold less because the runtime needs memory too. Kimi K3 is about 350 times
larger by total parameter count.
## Why “104B active parameters” does not mean a 52GB model
MoE sparsity reduces arithmetic, not the checkpoint you must keep available.
Kimi K3 routes each token through a small subset of its 896 experts, so only
104B of the 2.8T parameters participate in that token's forward pass. At four
bits each, 104B parameters alone represent a rough 52GB of weight data before
runtime overhead.
Even that 52GB estimate should not be mistaken for a static mini-checkpoint.
The next token may choose a different set of experts. Unless the workload pins
expert choices—which would change model behavior—the runtime needs access to
the complete expert pool across the sequence.
There are three separate numbers:
* **2.8T total parameters** determine full checkpoint storage.
* **104B active parameters** approximate per-token compute and data access.
* **4GB VRAM** is only the amount that can reside on the GPU at one moment.
Sparse activation makes Kimi K3 more efficient than a dense 2.8T model. It
does not make the full model a 104B download, and 104B is still far beyond a
4GB GPU.
## How layer-by-layer offloading can use a 4GB GPU

Layer offloading changes *where* weights wait, not how many weights exist. A
basic offload loop looks like this:
1. Keep most model weights on SSD or in system RAM.
2. Load the next required layer or expert chunk into GPU memory.
3. Run that part of the forward pass.
4. Evict the chunk and load the next one.
5. Repeat the sequence for every layer and every generated token.
This is how a model larger than VRAM can execute at all. AirLLM popularized
the pattern with a Llama 3 70B demonstration on a 4GB GPU, decomposing a model
into layer-wise shards and overlapping loading with compute.\[4]
Kimi K3 makes the pattern much harder. It has 93 layers, hundreds of possible
experts, a custom KDA/Gated-MLA attention stack, native multimodality, and more
than a terabyte of quantized weights. A full MoE layer can itself be larger
than 4GB, so a compatible engine would need sub-layer or expert-level
streaming—not merely ordinary layer offload.
The performance bottleneck then becomes data movement. Generating one token
may trigger many random or semi-random expert reads across dozens of layers.
Even a fast NVMe SSD is orders of magnitude slower than accelerator memory,
and the same process repeats for the next token. Offloading can make an
experiment start; it does not make it interactive.
## Can you run Kimi K3 on a 4GB GPU with AirLLM today?
**Not according to its published support and examples as of August 1, 2026.**
AirLLM documents Llama, Mixtral, Qwen, ChatGLM, Baichuan, Mistral, InternLM,
and related model families. Its headline 4GB result is Llama 3 70B, not Kimi
K3.\[4]
This absence matters. Kimi K3 is not a larger Llama checkpoint that an existing
loader can identify automatically. A working implementation must understand:
* Kimi K3's custom model configuration and tensor names;
* Stable LatentMoE routing across 896 experts;
* native MXFP4 weights and MXFP8 activations;
* KDA and periodic Gated MLA attention layers;
* preserved reasoning output and the multimodal vision tower;
* expert-aware partitioning small enough for the available VRAM.
Moonshot currently recommends vLLM, SGLang, and TokenSpeed for Kimi K3
deployment. It does not list AirLLM as a supported engine.\[1]
That does not prove a community 4GB port is impossible. It means a copied
AirLLM Llama example is not a reproducible Kimi K3 tutorial today. A credible
claim should provide a public branch, exact commit, storage and RAM details,
prompt, output, tokens per second, and proof that the full official weights
were used.
## What a validated Kimi K3 deployment looks like instead
Production Kimi K3 recipes operate at cluster scale. One current vLLM-Ascend
guide validates a 131K-context deployment across four Atlas 800 A3 nodes, each
with sixteen 64GB NPUs. Its 1M-context configuration requires at least eight
such nodes.\[5]
That is not a universal minimum—different accelerators, engines, concurrency,
and context limits change the requirement—but it is a useful reality check.
The validated configuration measures aggregate accelerator memory in
terabytes, not gigabytes.
| Goal | Sensible route |
| ------------------------------------------------ | ------------------------------------------------------------- |
| Test model behavior from a low-end PC | Use the hosted API |
| Run production inference | Follow official vLLM, SGLang, or TokenSpeed cluster recipes |
| Research extreme offloading | Expect custom engine work, >1.4TB storage, and very low speed |
| Run fully offline on consumer hardware | Choose a much smaller model |
| Use a 4GB GPU for an interactive local assistant | Use a 3B–7B quantized model, not Kimi K3 |
The right hardware answer depends on whether the goal is proof of execution,
interactive use, multi-user serving, or production throughput. The phrase
“runs on 4GB” is meaningless without that target.
## A checklist for evaluating any 4GB Kimi K3 claim
Before following a tutorial, look for evidence that answers these questions:
1. **Is it the official 2.8T checkpoint?** A distillation, proxy, or smaller
model carrying the Kimi name is not Kimi K3.
2. **Where are the full weights stored?** The answer should account for well
over a terabyte of local or network storage.
3. **How much system RAM is required?** “4GB GPU” says nothing about 512GB or
1TB of host memory sitting beside it.
4. **Which inference-engine commit supports K3?** A generic `pip install`
command is not enough for a new architecture.
5. **Is vision supported, or language only?** Skipping the 401M vision encoder
changes the tested model surface.
6. **What context length was used?** A 64-token demo and a 1M-token session
have radically different runtime needs.
7. **What is measured throughput?** Require tokens per second—or per minute—
plus time to first token.
8. **Was the answer verified?** A process starting successfully is not proof
that it loaded the correct weights or generated coherent K3 output.
If a post reports only VRAM, it has left out the resources that make the trick
possible.
## The practical way to use Kimi K3 from a 4GB GPU computer
The route that works today is to keep inference remote and use the low-end
computer as the client. The local GPU is irrelevant; all it needs is a network
connection and an API key.
reAPI exposes Kimi K3 through an OpenAI-compatible Chat Completions endpoint:
```bash
curl https://api.reapi.ai/v1/chat/completions \
-H "Authorization: Bearer $REAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k3",
"messages": [
{
"role": "user",
"content": "Explain how expert routing changes memory access in a sparse MoE model."
}
],
"reasoning_effort": "high",
"stream": true
}'
```
This is not local inference, and it should not be marketed as such. It is the
practical answer for developers who want Kimi K3 capability without acquiring
and operating a multi-node accelerator cluster. The full request contract is
in the [Kimi K3 API documentation](/docs/kimi-k3), with live rates on the
[model page](/models/kimi-k3).\[6]
## If you still want to experiment with local offloading
Treat it as systems research, not a one-command installation. A realistic
preflight list is:
* at least 1.5–2TB of free fast storage for the checkpoint, caches, and
conversion artifacts;
* enough bandwidth and patience to download many large shards;
* an inference-engine branch that explicitly supports Kimi K3's architecture
and MXFP4 format;
* a plan for host RAM, memory mapping, page cache, and SSD endurance;
* language-only mode if the experimental runtime has not implemented vision;
* short context and tiny outputs for the first validation;
* instrumentation for disk reads, GPU utilization, time to first token, and
output correctness.
Do not invent a working command by replacing a Llama model ID with
`moonshotai/Kimi-K3`. Until the offloading engine explicitly supports K3, the
most likely result is an unsupported configuration, tensor mismatch, or an
out-of-memory failure during conversion.
## FAQ
### Can Kimi K3 really run on a 4GB GPU?
Not as a self-contained, practical local model. A future specialized runtime
could use 4GB VRAM as a staging buffer while streaming weights from much larger
storage or RAM, but the full model does not fit and current mainstream 4GB
recipes do not document Kimi K3 support.
### How large are Kimi K3's weights?
The theoretical floor for 2.8T parameters at four bits each is about 1.4 TB
decimal, or 1.27 TiB. The real checkpoint and runtime require more because of
quantization metadata, non-4-bit tensors, the vision encoder, and execution
state.
### Why does Kimi K3 say only 104B parameters are active?
Kimi K3 is a sparse MoE model. Each token uses a subset of experts, reducing
compute, but future tokens can choose different experts. The complete 2.8T
expert pool still needs to remain accessible.
### Does MXFP4 mean any GPU with 4GB can run it?
No. MXFP4 means the main weights use roughly four bits per value. Four bits
times 2.8 trillion is still about 1.4 TB before overhead.
### Can AirLLM run Kimi K3?
AirLLM does not currently list Kimi K3 among its documented model families or
provide a Kimi K3 recipe. Its 4GB example is for Llama 3 70B. Support could be
added later, but it should not be assumed from the generic model loader.
### What is the cheapest practical way to use Kimi K3?
For occasional or development usage, use a hosted per-token API. Self-hosting
only becomes rational when control, sustained volume, or data-location needs
justify multi-node hardware and operations work.
### Can I run a smaller version of Kimi K3 locally?
Community distillations may appear, but they are separate models with different
weights and capability. If the requirement is a 4GB local assistant, choose a
model designed for that memory budget and label it accurately.
## The honest meaning of running Kimi K3 on a 4GB GPU
Running Kimi K3 on a single 4GB GPU is believable only under a narrow
definition: the GPU holds a small piece while the rest of a roughly 1.4TB
checkpoint lives elsewhere and streams through it. That may become a valuable
research demonstration, but it is not a practical local deployment today.
The title is unbelievable because it omits the machine around the GPU: SSD,
system RAM, custom runtime, transfer time, and often remote infrastructure.
Count all of those resources before judging the claim. If the goal is to use
Kimi K3 rather than study extreme offloading, the API is the route that works
on a 4GB laptop now.
**Disclosure:** reAPI publishes this article and offers hosted Kimi K3 API
access. Architecture and quantization facts come from Moonshot's official
repository and technical report. The 4GB assessment is derived from those
specifications, current engine documentation, and basic storage arithmetic; it
is not a claim that reAPI reproduced a full Kimi K3 generation on a 4GB GPU.
## References
1. Moonshot AI. *Kimi K3 official repository — architecture, model summary, native MXFP4, and recommended inference engines.* Retrieved August 1, 2026. [github.com/MoonshotAI/Kimi-K3](https://github.com/MoonshotAI/Kimi-K3)
2. Kimi Team. *Kimi K3: Open Frontier Intelligence.* Published July 2026. [arxiv.org/abs/2607.24653](https://arxiv.org/abs/2607.24653)
3. Moonshot AI. *Kimi K3 official weights and model card.* Retrieved August 1, 2026. [huggingface.co/moonshotai/Kimi-K3](https://huggingface.co/moonshotai/Kimi-K3)
4. AirLLM. *Supported model families and 4GB Llama 3 70B layer-offloading example.* Retrieved August 1, 2026. [github.com/lyogavin/airllm](https://github.com/lyogavin/airllm)
5. vLLM Ascend. *Validated Kimi K3 multi-node deployment guide.* Retrieved August 1, 2026. [docs.vllm.ai/projects/ascend/tutorials/models/Kimi-K3](https://docs.vllm.ai/projects/ascend/zh-cn/v0.23.0/tutorials/models/Kimi-K3.html)
6. reAPI. *Kimi K3 API reference and current model page.* Retrieved August 1, 2026. [reapi.ai/docs/kimi-k3](/docs/kimi-k3) and [reapi.ai/models/kimi-k3](/models/kimi-k3)
### Further reading
* reAPI. *Kimi K3: The Complete Guide to Moonshot's 2.8T Flagship.* [reapi.ai/blog/kimi-k3-complete-guide](/blog/kimi-k3-complete-guide)
* reAPI. *Best Open-Source AI Video Models for Local GPUs.* [reapi.ai/blog/best-open-source-ai-video-models-local-gpu-2026](/blog/best-open-source-ai-video-models-local-gpu-2026)
* Moonshot AI. *Kimi K3 official repository.* [github.com/MoonshotAI/Kimi-K3](https://github.com/MoonshotAI/Kimi-K3)
---
# Uncensored AI Image API? How Content Filters Work (2026) (https://reapi.ai/blog/uncensored-ai-image-api-content-filters)
An **uncensored AI image API** is usually a misleading promise. Some gateways
let developers disable an additional `nsfw_checker`, but that switch does not
erase provider policies, model alignment, upstream request inspection, legal
restrictions, or application-level obligations. It changes one layer in a
larger moderation system.
The more useful question is: **which safety layer blocked the request, and which
controls are actually configurable?** That framing helps teams handle benign
false positives—medical illustration, fine art, swimwear, health education—
without pretending prohibited content has become acceptable.
## TL;DR
* No production image API should be treated as unconditionally uncensored.
* `nsfw_checker: false` can disable an additional reAPI output check on selected
routes. It does not disable upstream moderation.
* Qwen Image 2 and Seedream 5.0 Pro currently expose this direct-API control;
their public playgrounds keep it enabled.
* GPT Image 2 and Seedream 5.0 Lite do not expose the same reAPI field.
* Diagnose rejections by stage: before task creation, during provider
generation, or after an output is returned.
* “Less restrictive” still requires consent, age controls, abuse prevention,
and compliance with the provider's acceptable-use rules.
## Why “uncensored AI image API” is the wrong technical model
Image safety is a pipeline, not one Boolean value. A request can pass one
classifier and fail at the next stage. Even when a model has comparatively
light alignment, the host, gateway, storage provider, or application may apply
separate controls.
A typical production request moves through four layers:
1. **Application policy.** Your own product validates users, prompts, source
images, permissions, and intended use.
2. **Gateway moderation.** An API gateway may run an independent prompt or
output classifier.
3. **Provider and model controls.** The upstream service can reject inputs,
refuse generation, or filter results under its own policies.
4. **Output review and distribution.** The application decides whether a
generated asset can be stored, shown, shared, or published.
OpenAI, for example, documents separate moderation models for classifying text
and image inputs, while its platform data controls still require customers to
follow usage policies even when approved for modified abuse monitoring or zero
data retention.\[1]\[2]
Retention settings, moderation behavior, and generation policy are related but
not interchangeable controls.

## What `nsfw_checker` actually controls on reAPI
On supported reAPI routes, `nsfw_checker` controls an additional independent
output-classification layer. With the value set to `true`, the finished image is
checked before it is returned. With `false`, that additional reAPI check is
skipped.
It does **not**:
* remove the upstream provider's prompt rules;
* retrain or “unalign” the model;
* guarantee that a job will be accepted;
* allow illegal, exploitative, non-consensual, or rights-violating content;
* make the API anonymous or exempt it from logging and abuse controls.
This field can matter for permitted content that generic classifiers sometimes
misread. A medical anatomy diagram, a breastfeeding education graphic, or a
museum-style figure study may be legal and non-exploitative yet still trigger a
broad nudity classifier. In that situation, a configurable final checker can
reduce false positives, but the application should replace it with a more
specific review policy rather than removing review entirely.
## Current image API filter controls compared
| reAPI route | `nsfw_checker` field | Playground behavior | What remains upstream |
| ----------------- | ------------------------------ | ------------------- | -------------------------------------------- |
| Qwen Image 2 | Direct callers may set `false` | Locked on | Alibaba/Qwen policy and model controls |
| Seedream 5.0 Pro | Direct callers may set `false` | Locked on | ByteDance/provider policy and model controls |
| Seedream 5.0 Lite | Not exposed | No separate toggle | Provider and model controls |
| GPT Image 2 | Not exposed | No separate toggle | OpenAI policy and safety systems |
This table describes reAPI's current public request fields as of August 1,
2026\. It is not a ranking of which model accepts the most sensitive material.
Acceptance changes by category, provider revision, source image, jurisdiction,
and account status.
For Qwen Image 2, the direct API control is documented in the
[Qwen Image 2 API guide](/blog/qwen-image-2-api-guide). Seedream's route and
billing behavior are covered in
[Is Seedream 5 Pro Uncensored?](/blog/seedream-5-pro-nsfw-checker).
## A safe request with the additional checker disabled
Direct callers can set the field on a supported route. The prompt below uses a
permitted educational context and explicitly excludes sexualization.
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer $REAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen-image-2",
"prompt": "Non-sexual adult anatomy illustration for a physical therapy textbook. Neutral standing pose, simplified muscle groups, clinical labels, no erotic framing, no minors, no graphic injury.",
"image_size": "3:4",
"output_format": "png",
"nsfw_checker": false
}'
```
Disabling the extra output check should be an application policy decision, not
a user-controlled passthrough. Restrict the option to reviewed use cases, record
why the workflow needs it, and apply a domain-appropriate classifier or human
review before distribution.
## How to diagnose image API safety rejections
### Rejected before a task ID exists
The request likely failed in application validation, gateway prompt moderation,
authentication, or upstream input inspection. Capture the HTTP status, provider
error code, model ID, and which configurable checks were enabled.
Do not automatically mutate and resubmit the prompt. First determine whether the
request is malformed, benign but ambiguous, or outside policy.
### Task created but generation failed
The provider or model may have rejected the prompt or source image after deeper
inspection. Some systems also fail when a reference image cannot be downloaded,
is the wrong format, or violates a size limit. Keep safety errors separate from
technical media errors in your telemetry.
### Generation completed but the output is hidden
This pattern points to an output-classification layer. On routes with a separate
`nsfw_checker`, the output can be generated and billed even when the final image
is withheld. A retry can therefore repeat the charge without changing the
underlying problem.
### Output returned but your application blocks it
That is your product policy working as designed. Model output is untrusted user
content until it passes your publication and rights checks. A provider accepting
an image does not mean your marketplace, ad network, school, or app store must
accept it.
## Billing when an image is filtered
Filtering does not always cancel inference cost. If a provider generates an
image and a later classifier hides it, the expensive model work has already
occurred. The route may therefore charge the request even though no usable image
is returned.
Production cost reports should separate:
* rejected before inference;
* failed during inference;
* generated but filtered;
* returned but rejected by application policy;
* accepted and published.
That funnel shows whether the problem is prompt design, a generic checker,
upstream policy, or your own acceptance criteria. A single “failed” counter
cannot support a useful moderation or budget decision.
## How to choose a less restrictive image API responsibly
Look for **control and documentation**, not the word uncensored.
### Prefer explicit request fields
A documented field such as `nsfw_checker` is easier to govern than a provider
that vaguely promises “no filters.” You can test it, restrict it by account,
record configuration, and detect schema changes.
### Require a clear acceptable-use policy
The provider should define prohibited categories, appeals, data handling, and
enforcement. Missing rules are operational risk, not creative freedom.
### Test a benign boundary set
Build a small evaluation set covering permitted medical, art, fashion, and
health scenarios that are relevant to your product. Record acceptance,
false-positive stage, latency, and billing. Do not include illegal or
exploitative material in this test set.
### Keep identity and consent checks separate
A general nudity classifier does not solve non-consensual intimate imagery,
face misuse, impersonation, copyright, or model-release consent. These need
their own rules and sometimes human review.
### Preserve provider traceability
Store the model ID, provider route, policy version, prompt hash, source-asset
provenance, safety configuration, and decision result. Minimize retained
personal data, but keep enough structured evidence to investigate abuse and
appeals.
## Claims that should make buyers cautious
Treat the following marketing phrases as warning signs:
* “100% uncensored” without an acceptable-use policy;
* “no logs” without retention documentation or contractual terms;
* “all content allowed” without jurisdiction or age restrictions;
* “filter off” without explaining whether it affects the gateway, provider, or
model;
* “private by default” without describing storage and abuse monitoring.
The same caution applies to unauthorized account automation. A third-party
wrapper can appear permissive because it hides whose account or endpoint it is
using. That is not a stable API contract. The article on
[Midjourney API availability](/blog/does-midjourney-have-an-api) explains why
authorization matters even when a JSON endpoint appears to work.
## Building an application-level moderation policy
A practical image application needs more than one universal threshold:
1. Define prohibited content that is never accepted.
2. Define restricted content allowed only in specific contexts or age groups.
3. Detect source-image identity, consent, and rights risks separately.
4. Review prompts and outputs, because either side can carry policy risk.
5. Give users a reason code and an appeal path for permitted edge cases.
6. Rate-limit repeated rejected requests and investigate adversarial patterns.
7. Re-test whenever the model, provider, or classifier version changes.
OpenAI's moderation endpoint can classify both text and image inputs, but its
categories and thresholds are only one possible policy component.\[1]
Your final rules must match the product's audience, geography, distribution
channel, and risk tolerance.
For video, the same layered principle applies. Read
[Seedance 2.0 safety filters and API control](/blog/seedance-2-0-uncensored-safety-filter-api)
instead of assuming an image-model toggle transfers to video routes.
## FAQ
### Is there a truly uncensored AI image API?
Not in the unconditional sense implied by the phrase. Hosted APIs operate under
provider policies, law, infrastructure controls, and application rules even
when one optional checker can be disabled.
### What happens when `nsfw_checker` is false?
On supported reAPI routes, it skips an additional output checker. It does not
remove upstream provider moderation or model alignment.
### Which reAPI image models expose `nsfw_checker`?
Qwen Image 2 and Seedream 5.0 Pro currently expose it to direct API callers.
Their public playgrounds keep the checker enabled.
### Can a filtered image still be billed?
Yes. If generation finishes before a final output checker blocks the result,
the inference cost has already been incurred.
### How should medical or fine-art applications handle false positives?
Use reviewed prompts, adult-only and non-sexual context where relevant, a
domain-specific moderation policy, reason codes, and human appeals. Do not rely
on disabling every safety layer.
## Conclusion
The best answer to “uncensored AI image API” is not a provider leaderboard. It
is a map of the moderation pipeline. Choose routes with explicit controls,
document which layer each control affects, keep upstream policy assumptions
visible, and replace broad false-positive filters with a more precise safety
system—not with no safety system at all.
## References
1. OpenAI. *Moderations API reference for text and image inputs.* [platform.openai.com/docs/api-reference/moderations](https://platform.openai.com/docs/api-reference/moderations)
2. OpenAI. *Data controls, abuse monitoring, and customer responsibilities.* [platform.openai.com/docs/models/default-usage-policies-by-endpoint](https://platform.openai.com/docs/models/default-usage-policies-by-endpoint)
3. Alibaba Cloud Model Studio. *Image generation and editing model documentation.* [alibabacloud.com/help/en/model-studio/image-model](https://www.alibabacloud.com/help/en/model-studio/image-model)
---
# Unlimited AI Video Generator Plans: Run the Math First (https://reapi.ai/blog/unlimited-ai-video-generator-cost-test)
Runway's pricing page carried two sentences on 9 August 2026. "Get 7 days of unlimited Seedance 2.5 when it launches if you sign up for a Max plan now," and "unlimited Seedance 2.5 on new Max plans until August 14th."\[1] Read them together and you are being asked to pay now for unlimited access to a model that has not arrived, inside a window that closes in five days.
That is not a Runway problem. Every unlimited AI video generator offer resolves into the same arithmetic, and the arithmetic never closes, because video inference costs real money per second and a subscription does not. What follows is a four-step test you can run on any of these plans in about ten minutes, using only numbers the vendor publishes about itself, plus the eight places the cap hides once you know it has to be somewhere.
## TL;DR
* **Divide the plan price by the vendor's own output conversion.** Runway publishes one: Max at $76 a month buys 9,500 credits, which they translate as 791 seconds of Gen-4.5. That is $0.096 a second.\[1]
* **The same product charges 2.4x more per second at the entry tier.** Standard is $12 for 52 seconds, or $0.231.\[1] Credits hide that spread.
* **Compare against a metered rate.** Published per-second video rates on reAPI run $0.055 to $0.3685 depending on model and resolution.\[2]\[3]\[4]
* **Count the clips you throw away.** One operator who says he spends over $300,000 a month on models puts real iteration at 3 to 10 attempts per usable video.\[5]
* **The caps are real and sometimes documented.** Higgsfield states that "Unlimited generations run in the standard queue, while credit-based generations always run in the priority queue," and that video bundles run "1 generation at a time, up to 15 seconds."\[6]
* **If your honest monthly volume costs more than about 3x the plan price, the plan is capping something.** Every unlimited AI video generator has eight places to put that cap, listed below.
## Step 1: find the plan's own implied price per second
Every subscription is a quota wearing a costume. The vendor knows what a second of video costs and has priced the plan around an assumed volume. Your job is to recover that assumption.
The good case is a vendor who publishes the conversion. Runway does, on the pricing page, for each tier:\[1]
| Runway plan | Annual price | Credits/month | Their stated Gen-4.5 output | Implied rate |
| ----------- | ------------ | ------------- | --------------------------- | ------------ |
| Standard | $12/mo | 625 | 52 s | $0.231 /s |
| Pro | $28/mo | 2,250 | 187 s | $0.150 /s |
| Max | $76/mo | 9,500 | 791 s | $0.096 /s |
Two things fall out immediately. The entry tier costs 2.4 times more per second than the top tier for identical output, which the credit abstraction is doing a good job of obscuring. And the top tier's $0.096 a second is a real number you can now compare to anything.
The bad case is a vendor who quotes videos rather than seconds. Higgsfield's plans state "\~15 Seedance 2.0 Fast videos", "\~53 Seedance 2.0 videos" and "\~133 Seedance 2.0 videos" at $19, $47 and $99 a month annually.\[7] That divides to $1.27, $0.89 and $0.74 per video, but the page never says how long those videos are or at what resolution, so the number cannot be checked against anything.
**That is your first finding, not a dead end.** A vendor who publishes seconds is telling you their unit economics. A vendor who publishes "videos" has chosen a unit nobody can falsify.
## Step 2: get a metered rate for the same output
Now price the same second on a pay-as-you-go API. Match the spec, because resolution moves the rate more than model choice does.
Published per-second rates on reAPI, August 2026:
| Model | Rate band per second |
| ------------ | ----------------------------------------------------- |
| MiniMax H3 | $0.055 – $0.1825\[4] |
| Seedance 2.5 | $0.0712 – $0.2668\[2] |
| Kling 3.0 | $0.077 – $0.3685\[3] |
One trap to avoid. Not every model bills per second. Veo 3.1 on the same platform is priced per generation, $0.0805 to $1.725 depending on tier.\[8] Comparing a per-generation price to a per-second price produces nonsense in whichever direction you happen to be hoping for, so convert one to the other before you compare anything.
Notice that Runway's own Max-tier implied rate of $0.096 a second sits inside the metered band rather than below it. A subscription is not buying you cheaper inference. It is buying you a different risk profile.
## Step 3: multiply by what you actually generate, including the throwaways
This is the step people skip, and it is the one that decides the answer.
The unit that matters is not cost per render. It is cost per usable clip. Tibo, who runs the video tool Revid and states he spends more than $300,000 a month on AI models, puts normal iteration at "3-10 times to get one."\[5] He sells a competing product, so treat that as an interested estimate, but it matches what anyone who has tried to hit a specific shot already knows. Discarded generations cost exactly what kept ones cost.
The formula:
```
monthly cost = keepers × attempts_per_keeper × seconds_per_clip × rate
```
Run it at five attempts per keeper, ten-second clips, and a mid-band $0.15 a second:
| Usable clips per month | Attempts | Billable seconds | Cost |
| ---------------------- | -------- | ---------------- | ---- |
| 10 | 50 | 500 | $75 |
| 40 | 200 | 2,000 | $300 |
| 100 | 500 | 5,000 | $750 |
Ten keepers a month is a light hobbyist. That is already at the price of Runway's Max plan. Forty is one short-form video a day with weekends off, and it is four times Max. A hundred is an agency, and it is ten times.
Now put your own number in. If you cannot estimate keepers per month, you are not ready to buy either option.
## Step 4: if the gap is more than about 3x, find the cap
Compare step 3 against the plan price. A subscription that costs less than roughly a third of your metered equivalent is not generous, it is constrained, and the constraint exists whether or not you can see it yet.
No unlimited AI video generator is absorbing a 4x to 10x loss per active account. Multiply a $76 plan by a few hundred subscribers who each consume $300 to $750 of inference, and the shortfall stops being a marketing budget and becomes the company. So when the arithmetic does not close, the honest conclusion is not that the vendor is generous. It is that you have not found the cap yet.
## The eight places an unlimited AI video generator hides its cap
Ranked roughly by how often they show up, with what each looks like in the wild.
**1. Model substitution.** The unlimited badge lands on a cheaper variant. Higgsfield's entry plan restricts access to Seedance 2.0 Fast and 2.0 Mini rather than the flagship.\[7] A creator who bought a 7-day unlimited Seedance offer found the unlimited access was to the Fast tier only.\[9]
**2. Resolution.** The same account documented an 8-second, 720p ceiling on the unlimited path, with those limits lifting only if you switched unlimited off and spent credits instead.\[9] A commenter in the same thread reported another platform running 1080p on its unlimited tier for a couple of weeks before dropping it to 720p.\[9]
**3. Duration.** Eight seconds when the model does thirty. Cheapest possible cap to apply and the least likely to be in the headline.
**4. Queue time.** The main lever, and usually the one buyers discover last. It is also, to Higgsfield's credit, documented: "Unlimited generations run in the standard queue, while credit-based generations always run in the priority queue. During peak traffic hours, generation speed and concurrency on Unlimited may temporarily vary."\[6] The same page describes the alternative in one sentence: switching Unlimited off "deducts credits and runs in the priority queue at maximum speed."\[6]
Read those together and the product is legible. Unlimited is the slow lane, credits are the fast lane, and the fee you avoided is the fee that buys speed back. Users of past unlimited offers report around ten minutes per generation, and one described waits growing from five minutes to forty.\[9]
**5. Concurrency, and a duration cap hiding inside it.** Higgsfield allows 2 parallel videos on Starter and 8 on Ultra.\[7] Unlimited bundles bought through their marketplace are tighter still, and the published table is worth reading twice: video bundles run "1 generation at a time, up to 15 seconds."\[6]
Fifteen seconds. On a model family whose headline capability is a 30-second single pass. An unlimited video bundle can therefore be unlimited and incapable of producing the thing the model is advertised for, at the same time, without either statement being false.
**6. Surface.** The allocation often runs in the browser only. Higgsfield states unlimited models "are not accessible on MCP/CLI, Canvas or Supercomputer,"\[7] and puts it more bluntly in the help centre: "Unlimited access applies only on higgsfield.ai: outside it, generations always deduct credits."\[6] If you are automating anything, unlimited does not apply to you at all.
**7. The window, and the subscription holding it.** "Unlimited" routinely means seven days rather than a month, and sometimes seven days that have not started. Runway's August offer is a 7-day window, on a model described as launching later, inside a promotion expiring 14 August.\[1] Higgsfield's help centre says the pattern is deliberate: most included models carry 365-day access, but "the newest flagship models may have shorter windows (for example, 7 days)."\[6]
The long windows come with their own string attached. A 365-day unlimited purchase "only works while your subscription is active," and if the subscription ends the access "stops working, even if its period hasn't finished yet."\[6] You have not bought a year of anything; you have bought a year of eligibility conditional on continuing to pay.
**8. Your account.** The cap of last resort. A user in r/seedance reported being banned from a video platform, along with others, after heavy use of an annual unlimited subscription, with the appeal denied and the payment kept: "Paid for a year and got a month."\[10] Individual reports are not audits, but the incentive is structural. If a subscriber's marginal usage costs more than the subscription earns, every hour they work is a loss.
## Running the test on an offer that is live right now
Take Runway Max at $76 a month annually, with 7 days of unlimited Seedance 2.5 attached.\[1]
Step 1 gives the plan's own implied rate: $0.096 a second. Step 2 puts metered Seedance 2.5 at $0.2668 a second for 720p.\[2] Step 3: to consume $76 of metered generation you need 285 seconds of output, which is 57 five-second clips.
So the unlimited week pays for the whole month the moment you produce 57 clips. At a ten-minute queue with two concurrent slots, that is under five hours of wall-clock. Entirely achievable, which is exactly why the queue is ten minutes rather than one.
That is the test working. The offer is not fake, and for a week of concentrated work it can genuinely beat metered pricing. It is bounded by caps 1, 2, 4, 5 and 7 simultaneously, and the moment you want a 30-second take, a specific resolution, an API call, or a second week, every one of those bounds is load-bearing.
## FAQ
### Is any unlimited AI video generator actually unlimited?
No. Video inference has a real per-second cost and a fixed monthly fee does not, so every such plan caps something. The useful question is which of the eight caps applies to your workflow.
### How do I calculate whether a plan beats pay-as-you-go?
Divide the plan price by the vendor's published output conversion to get their implied per-second rate, then multiply your real monthly volume, including failed attempts, by a metered rate for the same spec. If your metered figure is several times the plan price, the plan is capping something.
### Why do unlimited plans get so slow?
Queue time is the rationing mechanism, and it is the only cap that never needs to be published as a number. At least one vendor discloses it in general terms as "dynamic speed adjustments during high-traffic periods."\[7]
### Can I use an unlimited plan through an API?
Usually not. Higgsfield explicitly excludes MCP and CLI from its unlimited models.\[7] Automated or scheduled work generally needs metered access.
### What is a realistic cost per finished video?
At five attempts per keeper, ten-second clips and a mid-band $0.15 per second, about $7.50 per usable clip. Raise or lower the attempt count first, since it moves the total more than the rate does.
### Are credits a fair way to price this?
They are fair when the vendor publishes the conversion to seconds, as Runway does.\[1] They are not when a plan is quoted in "videos" of unstated length and resolution, because that number cannot be checked.
### Should I never buy an unlimited plan?
Buy one when your work fits inside the caps: browser-based, burst rather than steady, tolerant of a queue, and happy with the variant on offer. Avoid it when you need a specific model, a specific resolution, an API, or predictable throughput.
## Deciding without taking anyone's word
The four steps take ten minutes and they work on any offer, from any vendor, in any month, because they only use numbers the vendor has already published about itself. Implied rate, metered rate, honest volume with waste included, then hunt the cap.
What you should not do is treat the word unlimited as information. It describes a billing structure, not a capacity, and the capacity is set by whichever of the eight caps binds first for your particular workflow. An unlimited AI video generator plan can still be the right purchase after you have found that cap and confirmed you fit inside it. It is never the right purchase before.
If your work is programmatic, scheduled, or latency-sensitive, metered per-second billing is the only structure that survives step 4, because the rate is published before you pay and stays the same number next month. reAPI's video model rates are on the [model pages](/models), and the [Seedance 2.5 page](/models/seedance-2-5) carries the per-second band used in the worked example above.
## References
1. Runway. *Pricing — plan tiers, credit-to-seconds conversions and the Seedance 2.5 unlimited offer.* Retrieved 9 August 2026 from [runwayml.com/pricing](https://runwayml.com/pricing)
2. reAPI. *Seedance 2.5 — published per-second rate band.* Retrieved 9 August 2026 from [reapi.ai/models/seedance-2-5](/models/seedance-2-5)
3. reAPI. *Kling 3.0 — published per-second rate band.* Retrieved 9 August 2026 from [reapi.ai/models/kling-3-0](/models/kling-3-0)
4. reAPI. *MiniMax H3 — published per-second rate band.* Retrieved 9 August 2026 from [reapi.ai/models/minimax-h3](/models/minimax-h3)
5. Tibo (@tibo\_maker), founder of Revid. *Post on per-second model cost and iteration counts behind one usable video.* Posted 8 August 2026, retrieved August 2026 from [x.com/tibo\_maker/status/2085999680186450326](https://x.com/tibo_maker/status/2085999680186450326)
6. Higgsfield. *What are Unlimited models, and which plans include them? — help centre, queue behaviour, bundle concurrency and access conditions.* Updated 3 August 2026, retrieved 9 August 2026 from [higgsfield.ai/creator-hub/help-center/credits-and-usage/what-are-unlimited-models-and-which-plans-include-them](https://higgsfield.ai/creator-hub/help-center/credits-and-usage/what-are-unlimited-models-and-which-plans-include-them)
7. Higgsfield. *Pricing — individual plans, unlimited model list and terms.* Retrieved 9 August 2026 from [higgsfield.ai/pricing](https://higgsfield.ai/pricing)
8. reAPI. *Veo 3.1 — published per-generation rate band.* Retrieved 9 August 2026 from [reapi.ai/models/veo3-1](/models/veo3-1)
9. r/Seedance\_AI. *Details on "unlimited" Seedance2 deals (Higgsfield and Loova).* Posted 25 April 2026, retrieved August 2026 from [reddit.com/r/Seedance\_AI/comments/1svgicx](https://www.reddit.com/r/Seedance_AI/comments/1svgicx)
10. r/seedance. *There are multiple versions of Seedance 2. Here's what each website provides and whether they employ additional censorship.* Posted 29 May 2026, retrieved August 2026 from [reddit.com/r/seedance/comments/1trfu94](https://www.reddit.com/r/seedance/comments/1trfu94)
### Further reading
* reAPI. *Seedance 2.5 Unlimited Plans: The Math That Can't Work.* [reapi.ai/blog/seedance-2-5-unlimited-plans-math](/blog/seedance-2-5-unlimited-plans-math)
* reAPI. *AI Video Generation API: Why the Prices Don't Compare.* [reapi.ai/blog/ai-video-generation-api-pricing](/blog/ai-video-generation-api-pricing)
---
# Use ChatGPT Without Sounding Like AI: A Human-First Workflow (https://reapi.ai/blog/use-chatgpt-without-sounding-like-ai)
The best way to use ChatGPT without sounding like AI is not to request a complete article and then swap a few synonyms. Give the model a real argument, real source material, and a narrow drafting job. Then edit the result as an author, not a proofreader.
That is the gap between AI-assisted writing and AI-finished writing. In the first, software helps with structure, alternatives, and compression while a person remains responsible for the thesis, evidence, voice, and final claims. In the second, a generic prompt produces a generic page and the writer becomes its delivery system.
This workflow is designed for publishable blog content. It does not promise to “beat” AI detectors. Its goal is more useful: produce work a reader would choose to finish.
## TL;DR
* Start with a **point of view and a reader decision**, not “write 2,000 words about X.”
* Build a **claim ledger** from primary sources before drafting. ChatGPT should not invent the reporting layer.
* Give the model a voice sample and explicit anti-goals, but do not ask it to impersonate a living writer.
* Draft one section at a time with a defined job, evidence packet, and output limit.
* Perform separate **substance, voice, and line edits**. One vague “make it human” prompt cannot replace them.
* Verify every number, quote, date, product specification, and external link against the source.
* Treat detector results as one quality signal. Optimize for specificity and trust, not a guaranteed score.
## Why “write me an article” produces AI-sounding prose
A broad prompt gives the model almost no information about what makes this article yours. It still has to fill the requested space, so it reaches for statistically safe defaults: a scene-setting introduction, balanced sections, conventional transitions, generalized benefits, and an optimistic conclusion.
The output is not generic because ChatGPT has a secret vocabulary. It is generic because the assignment is generic.
Usman's essay, “This is Why Nobody Can Tell I Used ChatGPT,” makes the useful distinction between creators who paste the first output and creators who barely use the tool at all.\[1] There is a practical middle: use ChatGPT heavily where it is good, while keeping authorship decisions outside the model.
| Let ChatGPT help with | Keep under human control |
| -------------------------- | --------------------------------------- |
| Outline variants | Thesis and point of view |
| Counterarguments | Which evidence is trustworthy |
| Compression and reordering | First-hand experience |
| Headline alternatives | Ethical and disclosure decisions |
| Awkward sentence repair | Final factual responsibility |
| Consistency checks | What the conclusion actually recommends |
## The seven-stage human-first writing workflow

### 1. Write the one-sentence thesis yourself
Before opening ChatGPT, complete this sentence:
```text
After reading this article, [specific reader] should understand that
[contestable claim], and therefore [decision or action].
```
“This article explains email marketing” is a topic, not a thesis. A useful version is: “After reading, a solo SaaS founder should understand that a smaller behavior-based onboarding sequence outperforms a large calendar-based sequence when product usage is sparse, and should launch three triggered emails before building ten scheduled ones.”
The claim narrows the research. The decision gives the conclusion somewhere to go.
### 2. Build a claim ledger, not a pile of tabs
Create a small table before the outline:
| Claim | Best source | Evidence | Limitation | Status |
| ------------------------------- | ---------------------- | ----------------------------------- | ----------------- | -------------------- |
| Product supports feature X | Official documentation | Exact specification and date | Beta availability | Verified |
| Workflow reduces editing time | Your timed test | 42 vs 67 minutes across five drafts | Small sample | Verified with caveat |
| Industry is adopting approach Y | None yet | — | — | Remove or research |
The ledger prevents the most common failure in AI-assisted content: smooth sentences appearing before anyone has earned the claim inside them. It also makes fact-checking finite. At the end, every consequential sentence should trace back to one row.
For SEO content, prefer primary sources: official documentation for product behavior, public rate cards for pricing, original papers for research findings, and your own disclosed test for first-hand claims. Use commentary articles to find questions, not as the final authority on technical facts.
### 3. Make a voice brief from your real writing
Do not ask for “a human tone.” That phrase is too broad to constrain anything. Build a short voice brief from two or three pieces you actually wrote and like.
Record observable choices:
* average paragraph length;
* whether openings start with a claim, scene, or question;
* preferred level of formality;
* how often you use first person;
* words or constructions you avoid;
* how you qualify uncertain claims;
* whether humor is dry, warm, rare, or absent;
* what a typical conclusion does.
Then include a 300–600 word sample you own. Ask ChatGPT to identify patterns first, and approve the analysis before it drafts. OpenAI's prompt guidance recommends clear instructions, explicit context, and examples of the desired output format; those principles apply directly to voice work.\[2]
Do not ask it to write “exactly like” a living author. Borrowing structural techniques is one thing. Passing off a recognizable person's voice is another.
### 4. Design an outline around reader questions
Each section needs a job. A practical outline labels four fields:
```text
Section: Why generic prompts create generic prose
Reader question: What exactly causes the problem?
Claim: Missing editorial constraints push the model toward safe defaults.
Evidence: Before/after prompt and output comparison.
Exit: The reader can diagnose the prompt before blaming vocabulary.
```
This exposes filler before it becomes paragraphs. If two headings answer the same reader question, combine them. If a section has no evidence or decision, cut it.
### 5. Use ChatGPT as a constrained section drafter
Now give the model one section, not the whole article. A strong working prompt looks like this:
```text
You are helping me draft one section of an article. I remain the author and
will verify and edit the result.
ARTICLE THESIS
[one-sentence thesis]
SECTION JOB
[reader question, claim, and intended exit]
SOURCE PACKET
[only the verified notes and links relevant to this section]
VOICE
[brief plus a short sample I own]
CONSTRAINTS
- 300–450 words
- start with the claim, not scene-setting
- distinguish source facts from my inference
- no invented examples, quotations, statistics, or personal experience
- use headings only if they help scanning
- avoid inflated adjectives and generic future-facing conclusions
- flag missing evidence in [brackets] instead of filling the gap
Return the draft, then list every factual claim that needs verification.
```
The last line turns the model into a preliminary auditor. It does not make the claims true, but it creates a checklist for the next pass.
### 6. Edit in three separate passes
Trying to fix facts, structure, voice, and punctuation simultaneously makes each pass shallow.
**Substance edit.** Check whether the argument follows, the strongest evidence receives the most space, and every section changes what the reader knows or decides. Remove fabricated certainty and unsupported generalization.
**Voice edit.** Replace phrases you would never say. Add the distinction, objection, or hard-earned detail that came from your judgment. Vary paragraph shape where the logic demands it. Do not manufacture typos or awkwardness; humans are not defined by lower quality.
**Line edit.** Cut repeated setup, throat-clearing, redundant conclusions, ornamental metaphors, and transitions that explain obvious movement. Read the page aloud. The sentence that makes you speed up to get through it usually needs to be split or deleted.
The [reAPI Humanize API](/docs/humanize) can support this pass at scale by rewriting AI-sounding text, but it should receive a voice brief and a verified draft. Rewriting cannot repair a missing thesis or turn a false claim true.
### 7. Run a publication integrity check
Before publishing:
1. Open every citation and confirm it supports the adjacent claim.
2. Search the draft for every number, date, superlative, and quotation.
3. Remove personal stories the model supplied; it has no lived experience.
4. Confirm that examples are real, clearly hypothetical, or reproducible.
5. Check names, model IDs, pricing units, and version dates.
6. Verify internal and external links.
7. Follow the destination's AI-use and disclosure rules.
Google says generative AI can be useful for research and structure, while warning that many automated pages without added value can violate scaled-content policies. Its current advice emphasizes accuracy, quality, relevance, and useful context about automation.\[3] That is a sensible publication standard even when search traffic is not the goal.
## Before-and-after: improving the assignment, not disguising the output
### Weak request
```text
Write a 2,000-word SEO article about AI writing. Make it engaging and human.
```
This request supplies a length, topic, and vague adjective. The model must invent the reader, thesis, evidence, level of expertise, and editorial position.
### Strong request
```text
Draft the “detector limitations” section for editors at SaaS companies.
The claim is that detector scores should trigger review, not automatic rejection.
Use only the attached OpenAI classifier page and Liang et al. study.
Explain what each source found and state that the evidence does not evaluate
every current detector. End with a five-line review policy. 450 words maximum.
```
The stronger prompt is not a trick for hiding AI. It is a better brief. A human editor could hand the same assignment to another human and expect a more useful draft.
## Five edits that recover the author's presence
### Replace category language with observed detail
“Businesses can improve efficiency” says nothing. “The editor reduced source-checking from 14 tabs to a five-row ledger” gives the reader an object and a result.
### State where your inference begins
Write “The documentation confirms X; my inference is Y” when moving beyond a source. This small boundary creates more trust than confident fluency.
### Keep the exception next to the rule
Do not hide limitations three sections later. If a benchmark used one language, a price excludes caching, or a workflow was tested on five drafts, say so where the result appears.
### Let important points occupy unequal space
A model likes balance. An author chooses. Give the decisive issue six paragraphs and the minor caveat two sentences if that is what the evidence warrants.
### End with a decision
Do not conclude that “AI will continue to transform writing.” Tell the specific reader what to do on Monday and what not to delegate.
## What about AI detectors?
You can run a draft through the [reAPI AI text detector](/docs/ai-text-detector) to identify passages worth another look, especially in a high-volume workflow. But do not edit toward a promised zero score. Classifiers vary, and a 2023 peer-reviewed study found false-positive risks for non-native English writing.\[4]
A better target is an editorial scorecard:
| Dimension | Passing question |
| ----------- | -------------------------------------------------------------- |
| Specificity | Could these examples appear only in this article? |
| Evidence | Can every important claim be traced to a source? |
| Voice | Does the reasoning sound like the named author? |
| Utility | Can the reader make a decision or perform a task? |
| Integrity | Are assistance, uncertainty, and limitations handled honestly? |
If the page passes those tests, it has become better writing rather than merely less detectable writing.
## FAQ
### How do I make ChatGPT writing sound less AI-generated?
Start with your own thesis and evidence, provide a concrete voice brief, draft in constrained sections, and edit separately for substance, voice, and line quality. Do not rely on synonym swaps or instructions to “sound human.”
### Can ChatGPT copy my writing style?
It can identify and follow observable patterns from samples you provide, such as tone, paragraph length, pacing, and degree of formality. Review the result carefully, and use samples you own rather than requesting imitation of a living writer.
### Should I edit the entire ChatGPT draft at once?
For short copy, perhaps. For a researched article, section-by-section drafting is easier to constrain and verify. It also makes it less likely that an early factual error will be repeated throughout the piece.
### Can a humanizer replace editing?
No. A humanizer can improve phrasing, rhythm, and stylistic variation. It cannot supply genuine experience, choose a defensible thesis, or verify a source. Use it inside an editorial workflow, not in place of one.
### Is AI-assisted content bad for SEO?
Google's published guidance focuses on helpfulness, originality, accuracy, and value rather than a blanket ban on AI assistance. Large-scale low-value generation is the risk.\[3]
## Keep the author in the loop where judgment compounds
ChatGPT is most useful between blank page and final page. It can propose structures, expose counterarguments, compress research notes, and produce a section worth editing. Those are substantial contributions.
The author still has to decide what is true, what matters, what came from experience, which caveat changes the result, and what the reader should do next. Keep those decisions human and the writing stops feeling like a model trying to fill space. It begins to feel like someone had a reason to publish.
## References
1. Usman. *This is Why Nobody Can Tell I Used ChatGPT.* Write A Catalyst, March 2026. [medium.com](https://medium.com/write-a-catalyst/this-is-why-nobody-can-tell-i-used-chatgpt-db3e1a2053ac)
2. OpenAI. *Best practices for prompt engineering with the OpenAI API.* Updated 2026. [help.openai.com](https://help.openai.com/en/articles/6654000-comprehensive-step-by-step-guide-to-prompt-engineering-with-chatgpt)
3. Google Search Central. *Guidance on using generative AI content on your website.* Updated December 2025. [developers.google.com](https://developers.google.com/search/docs/fundamentals/using-gen-ai-content)
4. Weixin Liang et al. *GPT detectors are biased against non-native English writers.* Patterns, 2023. [pubmed.ncbi.nlm.nih.gov](https://pubmed.ncbi.nlm.nih.gov/37521038/)
### Further reading
* reAPI. *Humanize API.* [reapi.ai/docs/humanize](/docs/humanize)
* reAPI. *AI text detector API.* [reapi.ai/docs/ai-text-detector](/docs/ai-text-detector)
* reAPI. *AI essay writer API.* [reapi.ai/docs/ai-essay-writer](/docs/ai-essay-writer)
* reAPI. *Can you tell if a writer used ChatGPT?* [reapi.ai/blog/can-you-tell-if-writer-used-chatgpt](/blog/can-you-tell-if-writer-used-chatgpt)
---
# Veo 3.1 vs Seedance 2.0: Picking a Video Model in 2026 (https://reapi.ai/blog/veo-3-1-vs-seedance-2-0-2026)
Two of the strongest text-to-video models in 2026 ship from very different directions. Google's Veo 3.1 plays the precision-and-resolution game: 4K, first/last-frame interpolation, tight commercial control. ByteDance's Seedance 2.0 plays the multimodal game: 15 seconds in one shot with native audio, lip-sync in 8 languages, and a reference pipeline that accepts up to 9 images plus 3 videos plus 3 audio clips simultaneously.
If you're picking Veo 3.1 vs Seedance 2.0 in 2026, the answer depends on what you're shipping. This piece walks through every capability that meaningfully differs, with prices anchored to each provider's own listing and capability claims sourced from ByteDance and Google's release pages.
## TL;DR
* **Origin and release.** Veo 3.1 from Google DeepMind, Gemini API since October 2025; Seedance 2.0 from ByteDance, released February 12, 2026\[1].
* **Resolution ceiling.** Veo 3.1 supports 4K (3840×2160) on widescreen aspects; Seedance 2.0 caps at 1080p\[2]\[3].
* **Duration.** Veo 3.1 official tier exposes 4 / 6 / 8 second outputs; Seedance 2.0 supports 4–15 seconds in a single generation, with multi-shot cuts inside one clip\[1].
* **Multimodal references.** Seedance 2.0 takes up to 9 images + 3 video clips + 3 audio clips in one request; Veo 3.1 accepts up to 3 reference images on its alt tier or first/last-frame anchoring on the official tier\[1].
* **Audio.** Seedance 2.0 generates joint audio-video natively (lip-sync, BGM, ambient) and accepts reference audio. Veo 3.1 bundles audio at synthesis time on Google direct, optional on per-second tiers across gateways.
* **Cheapest 720p 5-second clip with audio.** Veo 3.1 Lite at $0.25 on Google direct ($0.05/s × 5s)\[3]; Seedance 2.0 Fast in reference mode at \~$0.43 on reAPI ($0.0865/s × 5s)\[4].
* **The split.** Veo for hero / 4K / character consistency. Seedance for one-shot multi-cut storyboards, real lip-synced audio, multimodal reference compositions.
## Where each model comes from
Veo 3.1 launched on the Gemini API in October 2025 and added the Lite variant on March 31, 2026\[5]. It runs on Google's Vertex AI and Gemini API, and is also routed by every major AI inference gateway including fal.ai, Replicate, OpenRouter, and reAPI.
Seedance 2.0 launched on February 12, 2026 from ByteDance's Seed research group\[1]. ByteDance's own description: "next-generation video creation model" with a "unified multimodal audio-video joint generation architecture" supporting text, image, audio, and video inputs\[1]. The model went viral in China for photorealistic clips of named celebrities, and Disney sent ByteDance a cease-and-desist letter on February 13, 2026 over training-data concerns\[6]. Seedance 2.0 ships with C2PA watermarking by default, so provenance signaling is baked into every output.
Both models are accessible through reAPI on the same OpenAI-compatible endpoint (`POST /api/v1/videos/generations`). The request shape barely changes between them; the difference is what you set as `model`.
## What each model can actually do
| Capability | Veo 3.1 | Seedance 2.0 |
| ------------------------------ | --------------------------------------- | ------------------------------------------------------------------------------------------------------- |
| Text-to-video | yes | yes |
| Image-to-video (single ref) | yes (alt tier, ≤3 images) | yes |
| Image-to-video (multi-ref) | up to 3 images | up to 9 images |
| First/last-frame interpolation | yes (official tier only) | yes (`image_with_roles` field) |
| Reference video | no | up to 3 clips, ≤15s combined |
| Reference audio | no | up to 3 clips, ≤15s combined |
| Audio synthesis | yes (bundled) | yes (native joint generation, 8+ languages, phoneme-level lip-sync)\[1] |
| Multi-shot in single output | no | yes, multiple cuts in one generation\[1] |
| 4K output | yes (widescreen aspects only) | no (1080p ceiling) |
| Duration options | 4 / 6 / 8s (official) or fixed 8s (alt) | any 4–15s |
| Negative prompts | official tier | not exposed |
| Seed reproducibility | official tier | yes |
| Aspect ratios | 16:9, 9:16 | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, adaptive |
The biggest single differentiator: **Seedance 2.0 generates multi-shot sequences inside one 15-second clip**. Send one prompt, get back what feels like an edited storyboard with natural cuts and transitions\[1]. That's not a feature Veo 3.1 has. Veo's outputs are single continuous shots.
The other big asymmetry: reference inputs. Seedance 2.0's reference pipeline (9 images, 3 videos, 3 audio) is built for tightly directed brand spots. Feed it product shots, a style reference clip, and a music bed, and it composes against all three. Veo 3.1's image reference cap is 3 frames on the alt tier, with first/last-frame anchoring on the official tier. Meaningful control, narrower aperture.
## Quality positioning
Independent benchmarks place Seedance 2.0 ahead of Veo 3.1 on aggregate composite scores covering visual fidelity, motion smoothness, prompt alignment, and temporal consistency. The leaderboard gap is small enough that prompt and use case matter more than the headline number\[2].
Where each tends to win:
**Seedance 2.0 is stronger on:**
* Photorealism in skin texture and surface detail
* Multi-subject scenes (groups, crowds, complex compositions)
* Camera motion that respects physics (crane shots, tracking shots)\[2]
**Veo 3.1 is stronger on:**
* Human face rendering with less uncanny-valley artifacting
* Text legibility within the video frame (signage, captions)
* Character/product consistency across multiple clips in a series\[2]
If you're producing a 30-second product spot in 4–5 cuts and need the protagonist to look like the same person in every cut, Veo 3.1 is the safer bet. If you're producing a single 15-second hero clip with multiple beats and don't need to chain it with other clips, Seedance 2.0's multi-shot output gives you what would otherwise take a 4-call Veo workflow.
## Audio
Both models output video with audio. The mechanics differ.
**Veo 3.1.** On Google direct, audio is bundled in every per-second cell. You can't strip it to save money. On per-second tiers via gateways like reAPI's Fast Official channel, audio becomes a `generate_audio` toggle. Veo's audio is synthesis-time: the model generates ambient sound, music, and voice based on the prompt.
**Seedance 2.0.** Audio is decoupled into two orthogonal controls. `generate_audio: true` triggers native joint audio-video synthesis — the model generates the audio track as part of the same forward pass, not as a post-process. ByteDance's claim is phoneme-level lip-sync across 8+ languages with dual-channel audio output\[1]. Separately, `audio_urls` accepts up to 3 reference audio clips that the model aligns to (feed it a music bed and the generated video matches the rhythm).
For lip-synced dialogue, Seedance 2.0 has a real architectural advantage. For ambient soundtracks on cinematic scenes, both produce comparable output. Veo 3.1's audio quality is solid, and the bundled-by-default convenience matters when you're not optimizing for cost.
## Price math
Per-second rates differ by tier, resolution, and audio toggle. Below: cheapest 720p 5-second clip with audio across providers, May 2026.
| Provider | Model | Tier | 5s 720p with audio |
| ----------------- | ------------------------ | ------- | ----------------------------------------- |
| Google Gemini API | Veo 3.1 Lite | n/a | $0.25\[3] |
| Google Gemini API | Veo 3.1 Fast | n/a | $0.50\[3] |
| Google Gemini API | Veo 3.1 Standard | n/a | $2.00\[3] |
| reAPI | Veo 3.1 Fast Official | per-sec | $0.69\[7] |
| reAPI | Seedance 2.0 (text) | per-sec | $0.90\[4] |
| reAPI | Seedance 2.0 Fast (text) | per-sec | $0.72\[4] |
| reAPI | Seedance 2.0 Fast (ref) | per-sec | $0.43\[4] |
| fal.ai | Seedance 2.0 Standard | per-sec | $1.52\[2] |
| fal.ai | Seedance 2.0 Fast | per-sec | $1.21\[2] |
Seedance 2.0's pricing has a quirk worth knowing: **reference mode (any of `image_urls`, `video_urls`, `audio_urls` set) bills at a lower per-second rate than text mode**\[4]. If your workflow always feeds at least one reference image, the effective cost drops by roughly 35%.
For the cheapest possible Veo 3.1 path at 720p with audio, Google direct's Lite tier at $0.05/s wins outright. For Seedance 2.0 at the cheapest verifiable rate, reAPI's Fast variant in reference mode at $0.04/s undercuts every fal.ai cell and lands in the same neighborhood as Veo Lite. The catch with Seedance: 720p ceiling, no 4K available.
## Veo 3.1 vs Seedance 2.0 in practice
Two clean decision rules.
**Pick Veo 3.1 when:**
* You need 4K output (Seedance caps at 1080p)
* Character or product consistency across multiple clips matters more than per-clip wow
* Your workflow uses first-then-last-frame anchoring or chained sequels
* The budget tier matters (Veo Lite at $0.05/s is the cheapest verifiable rate with audio)
* Hero shots, commercial spots, anything where Google's QA-baked face rendering matters
**Pick Seedance 2.0 when:**
* A single 15-second multi-shot output replaces what would otherwise be a stitched 4-clip Veo workflow
* The scene needs lip-synced dialogue in non-English languages
* You're feeding the model multiple reference modalities (product images + style video + audio bed)
* 21:9 cinematic ultrawide or non-standard aspect ratios are required
* Skin texture and physically grounded motion are dealbreakers
Neither is universally better. Veo 3.1 vs Seedance 2.0 only resolves once you know your output spec.
## FAQ
### Is Seedance 2.0 free to use?
Not at the API level. Seedance 2.0 is paid-tier on every provider that exposes it (fal.ai, Replicate, reAPI, Volcengine direct). ByteDance's consumer products (Dreamina, CapCut) include some free Seedance 2.0 quota for end users, but those aren't API-accessible.
### Does Veo 3.1 have multi-shot output?
No. Veo 3.1 generates single continuous shots. To stitch multiple shots together, generate clips separately and edit in post, or use Veo 3.1's first/last-frame interpolation to chain shorter pieces. Seedance 2.0 generates multi-shot sequences inside one 15-second clip natively\[1].
### Which model handles real people better?
Both have policies. Google direct exposes a `person_generation` enum on Veo 3.1's official tier with values `allow_adult` and `disallow`. Seedance 2.0 has dedicated face-aware variants (`-face`, `-fast-face`) on platforms that surface them. Uploading identifiable real-person reference assets to the non-face variants is rejected upstream. Both models add C2PA watermarking by default in 2026.
### Can I use Seedance 2.0 with reference video and audio together?
Yes, and that's its headline use case. Send a request with `prompt` + `image_urls` + `video_urls` + `audio_urls` populated, and the model composes against all three modalities. Combined reference video duration is capped at 15 seconds; combined reference audio at 15 seconds\[8].
### Does Veo 3.1 support 21:9 aspect ratio?
No. Veo 3.1 only exposes 16:9 and 9:16. Seedance 2.0 supports 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, and adaptive (matches input ratio)\[8].
### What about Seedance 2.0 and Hollywood IP?
Real risk worth flagging. ByteDance received a Disney cease-and-desist on February 13, 2026 over claims that Seedance 2.0 was trained on Disney works without permission\[6]. Avoid prompts that target named films, characters, or studio styles you don't have rights to. C2PA watermarking is on by default on Seedance 2.0 outputs, so derivative work carries provenance signals downstream.
### Which model is faster end-to-end?
Comparable. Both take 60–180 seconds for an 8-second 1080p clip on standard tiers. Fast variants on either model cut that by roughly 30–40% with a quality trade. The dominant factor in wall-clock time is queue depth at the underlying provider, not the model's intrinsic speed.
### Can I switch between them with one code change?
On reAPI, yes. Both run on `POST /api/v1/videos/generations` with the same envelope. Switching from Veo 3.1 to Seedance 2.0 means changing `"model": "veo3.1-fast"` to `"model": "doubao-seedance-2.0"` and adjusting fields that don't apply (Veo's `aspect_ratio` becomes Seedance's `size`; Veo's `first_frame_image` becomes Seedance's `image_with_roles[].role: "first_frame"`).
## So which video model wins
For most product workloads in 2026, the answer to Veo 3.1 vs Seedance 2.0 is "both, in different positions." Veo 3.1 carries the budget tier and 4K resolution; Seedance 2.0 carries the multi-shot, multi-modal, lip-sync work. A typical shipping pipeline runs both behind one OpenAI-compatible endpoint and routes per-request based on what the output needs.
If forced into one for everything, I'd pick Veo 3.1 for any project where the output gets shown to paying customers. Google's QA on faces, character consistency, and resolution ceiling matters more in commercial contexts than Seedance 2.0's multi-shot trick. For high-volume drafting, social-first creative, and anything that benefits from native audio-video joint generation, Seedance 2.0 wins.
Veo 3.1 vs Seedance 2.0 really comes down to whether your workflow is producing single hero clips (Veo) or multi-beat one-shots (Seedance). Pick the model that matches your output spec, not the model with the better leaderboard score.
## References
1. ByteDance Seed. *Official Launch of Seedance 2.0.* February 12, 2026. [seed.bytedance.com/en/blog/official-launch-of-seedance-2-0](https://seed.bytedance.com/en/blog/official-launch-of-seedance-2-0)
2. fal.ai. *Seedance 2.0 vs. Veo 3.1: What's The Difference?* Retrieved May 2026. fal.ai/learn/tools/seedance-2-0-vs-veo-3-1
3. Google. *Gemini API pricing — Veo 3.1 per-second rates by tier and resolution.* Retrieved May 2026 from [ai.google.dev/gemini-api/docs/pricing](https://ai.google.dev/gemini-api/docs/pricing)
4. reAPI. *Seedance 2.0 — Model page (live pricing).* Retrieved May 2026 from [reapi.ai/models/seedance-2-0](/models/seedance-2-0)
5. Google. *Build with Veo 3.1 Lite, our most cost-effective video generation model.* The Keyword (Google blog), March 31, 2026. [blog.google/innovation-and-ai/technology/ai/veo-3-1-lite](https://blog.google/innovation-and-ai/technology/ai/veo-3-1-lite/)
6. Wikipedia contributors. *Seedance 2.0.* Retrieved May 2026 from [en.wikipedia.org/wiki/Seedance\_2.0](https://en.wikipedia.org/wiki/Seedance_2.0)
7. reAPI. *Veo 3.1 — Model page (live pricing).* Retrieved May 2026 from [reapi.ai/models/veo3-1](/models/veo3-1)
8. reAPI. *Seedance 2.0 — API reference.* Retrieved May 2026 from [reapi.ai/docs/seedance-2-0](/docs/seedance-2-0)
### Further reading
* ByteDance Seed. *Seedance 2.0 product page.* [seed.bytedance.com/en/seedance2\_0](https://seed.bytedance.com/en/seedance2_0)
* Hugging Face. *Seedance 2.0: Advancing Video Generation for World Complexity (paper).* [huggingface.co/papers/2604.14148](https://huggingface.co/papers/2604.14148)
* TechCrunch. *ByteDance's new AI video generation model, Dreamina Seedance 2.0, comes to CapCut.* March 26, 2026. [techcrunch.com/2026/03/26/bytedances-new-ai-video-generation-model-dreamina-seedance-2-0-comes-to-capcut](https://techcrunch.com/2026/03/26/bytedances-new-ai-video-generation-model-dreamina-seedance-2-0-comes-to-capcut/)
* reAPI. *Cheapest Veo 3.1 API in 2026.* [reapi.ai/blog/cheapest-veo-3-1-api-2026](/blog/cheapest-veo-3-1-api-2026)
---
# How to Verify an LLM API Serves the Model It Claims (https://reapi.ai/blog/verify-llm-api-model-fingerprint)
Ask a language model to pick a random number between 1 and 100 enough times and something strange happens: the answers are not random. One model keeps landing on 42 and 73. Another favors 47 and 57. The pattern is stable, reproducible, and different for every model.
A July 2026 paper turned that quirk into a verification method. *One Token Is Enough: Fingerprinting and Verifying Large Language Models from Single-Token Output Distributions* shows that these answer distributions form a reliable behavioral fingerprint, enough to check from the outside whether an API endpoint is serving the model it claims\[1]. No weights, no logits, no insider access. A few hundred one-word answers.
This guide explains why a `model` field is a claim rather than a guarantee, how single-token fingerprinting works, what the numbers mean, and where the method's limits are.
## TL;DR
* **The `model` field in an API response is unverifiable by the protocol.** Nothing proves your request for a flagship model was served by it\[1].
* **Language models cannot produce uniform randomness.** Across the paper's probe set, the median answer distribution carries about **1.0 bit of entropy** against the 6.64 bits a uniform 1-to-100 pick would have\[1].
* **That failure is consistent**, so it works as a signature. Same model, same skew, every time.
* **The scale has usable daylight**: same model against itself lands near JSD 0.14, same model across two providers near 0.23, two genuinely different models near 0.46\[1].
* **Accuracy is strong, not perfect.** Equal error rate is 10.6% at 8 probe units and 7.3% at the full 40, with an AUC of 0.971\[1].
* **A mismatch is evidence, not proof.** Quantization, a silent version update, or a hidden system prompt can all shift a distribution without anyone lying.
## You cannot see what is behind an API
When you call a chat-completions endpoint you send text and get text back. The `model` field in the response says whatever the server chooses to put there. Nothing in the protocol proves the request was served by the model named in it, rather than by something a tenth the price\[1].
That gap matters more as more of the market sits behind intermediaries: aggregators reselling hundreds of models through one endpoint, regional resellers offering flagship access below official pricing, and third-party hosts serving open-weight models with quantization and serving-stack choices you never see\[1].
Most of these businesses are legitimate. The incentive to cheat is still obvious, and a separate audit of 17 shadow-API operators found several endpoints that did not pass verification against the models being advertised\[2].
It is also not only about fraud. A provider can quantize a model to cut serving costs, roll out a silent version update, or route traffic across mixed backends. If your product's quality depends on a specific model, "what am I actually getting?" should be answerable with evidence rather than trust.
## Why random numbers give the game away
The method rests on a well-documented weakness. Language models do not compute, they predict, so when asked for a random number they reproduce the biases of their training data and preference tuning\[1].

Three clusters dominate:
* **42** is massively over-represented, the Hitchhiker's Guide answer echoing through decades of internet text.
* **7** carries centuries of cultural weight as the lucky pick humans reach for.
* **37, 47, 73** and other two-digit primes *feel* random to humans, so they dominate the human-generated "random" examples the model learned from.
The paper quantifies the collapse: the median answer distribution carries roughly **1.0 bit of entropy**, where a fair pick from 1 to 100 would carry **6.64 bits**\[1]. Where a fair die has a hundred faces, most models behave like a slightly weighted coin.
The key move is that the failure is *consistent*. The same model produces the same skewed distribution every time, and different models, including sibling versions in one family, produce measurably different ones. A bug becomes a signature.
## The protocol, in four steps
**1. Probe.** Ask the endpoint a battery of one-word questions: pick a random number between 1 and 100, name a random color, flip a coin. The paper uses 10 tasks across 4 languages for 40 probe units, sampling each 30 times at temperature 1.0 with `max_tokens=16` and reasoning disabled\[1].
**2. Fingerprint.** For each probe unit, tally the answers into an empirical distribution. The collection of those distributions is the endpoint's behavioral fingerprint.
**3. Compare.** Measure the distance between that fingerprint and a trusted reference for the claimed model, using Jensen-Shannon divergence in base 2, so the scale runs from 0 (identical) to 1 (disjoint).
**4. Decide.** Small distance means consistent with the claim. Large distance means the endpoint is behaviorally a different animal.
Two properties make this hard to game. It needs no special access, since anything answering chat completions can be fingerprinted. And there is no magic string to filter, because every probe is an ordinary harmless question drawn from paraphrase pools, so a dishonest middlebox cannot special-case the test without breaking normal traffic\[1].
## Reading the distance number
A divergence value means nothing without the reference points, and this is where the method becomes practical.

| Comparison | Typical JSD |
| ----------------------------------- | ----------- |
| Same model against itself | \~0.14 |
| Same model, two different providers | \~0.23 |
| Two genuinely different models | \~0.46 |
There is real daylight between "same" and "different"\[1]. Note what the middle row implies: even an honest deployment of the same model by a different host drifts measurably, because quantization, serving stack, and hidden system prompts all leave marks.
Accuracy scales with probe count. At 8 probe units the equal error rate is 10.6%; at the full 40 it falls to 7.3%, with an AUC of 0.971\[1].
## What the paper found in the wild
**The palmyra-x5 case.** The paper's most striking result concerns a model offered as a proprietary flagship whose fingerprint sat at a JSD of 0.141 from an open-weight 235B model, statistically indistinguishable from the \~0.140 you get comparing a model against itself. Behaviorally, the paper concludes, the endpoint was serving something functionally identical to the open-source model\[1].
**Same model, different providers, sometimes suspiciously different.** Of 34 same-model pairs served across different providers, 10 diverged beyond the 5th percentile of the impostor distribution\[1]. Some official-model third-party deployments drift far enough to look like different models. Verification is not paranoia even when nobody is lying about the name.
The research artifacts are open: the fingerprint dataset is published under CC-BY-4.0 and the reproduction code under MIT, both on Zenodo\[3]\[4].
## Running the check yourself
The protocol is simple enough to implement directly. The shape of it:
```python
import collections, math
from openai import OpenAI
client = OpenAI(api_key="...", base_url="https://your-endpoint/v1")
def probe(model, prompt, n=25):
counts = collections.Counter()
for _ in range(n):
r = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": prompt}],
temperature=1.0,
max_tokens=16,
)
counts[r.choices[0].message.content.strip()] += 1
total = sum(counts.values())
return {k: v / total for k, v in counts.items()}
def jsd(p, q):
keys = set(p) | set(q)
m = {k: 0.5 * (p.get(k, 0) + q.get(k, 0)) for k in keys}
def kl(a):
return sum(a[k] * math.log2(a[k] / m[k]) for k in a if a[k] > 0)
return 0.5 * kl(p) + 0.5 * kl(q)
```
Four practical notes on doing this properly.
**Hold the sampling conditions fixed.** Temperature 1.0, a small `max_tokens`, reasoning disabled. A fingerprint collected under different settings is not comparable to one collected under the paper's.
**Use more than one probe.** A single question is noisy. The error rate roughly halves going from 8 probe units to 40.
**Compare against a reference collected the same way.** The cleanest reference is the official endpoint for the same model, fingerprinted in the same session under identical settings, rather than a published table gathered months earlier.
**Watch the auxiliary signals too.** A fingerprint that disagrees with its own split halves suggests multi-backend routing. One-word answers that bill hundreds of completion tokens suggest padding. Inflated prompt-token counts suggest a large hidden system prompt.
A standard check of 8 probe units at 25 samples is about 200 tiny requests, which costs a fraction of a cent on a small model.
## What a result does and does not mean
Be precise about the claims this method supports.
**A mismatch is evidence, not proof.** The method has an inherent error rate, around 10.6% EER at 8 probe units. Aggressive quantization, a silent model update, a stale reference, or a serving-side system prompt can all shift distributions with no intent to deceive. Treat a red result as a reason to re-run with more probes, test a second reference, and ask questions\[1].
**A match is strong but not absolute.** A sophisticated impostor could in principle mimic another model's distributions, though doing so across dozens of paraphrased multilingual probes while serving normal traffic correctly is harder than it sounds.
**Reasoning models need care.** Fingerprints are collected with reasoning disabled. Where an endpoint cannot disable it, confidence drops.
**A distance is a property of an endpoint at a moment in time**, not a verdict on a business.
## FAQ
### What is LLM fingerprinting?
Measuring the distribution of a model's answers to a battery of one-word questions, then comparing that distribution against a reference for the model an endpoint claims to serve\[1].
### Why can't language models produce random numbers?
They predict rather than compute, so a request for randomness returns the biases of training data and preference tuning. Culturally loaded values like 42 and 7 dominate, collapsing the distribution to about 1.0 bit of entropy against a uniform ideal of 6.64\[1].
### How much of a difference counts as a mismatch?
Use the paper's baselines rather than a fixed threshold: about 0.14 for a model against itself, 0.23 for the same model across providers, and 0.46 for genuinely different models\[1].
### How many requests does a check need?
The paper's full protocol is 40 probe units sampled 30 times each. A lighter 8-unit check at 25 samples is roughly 200 requests and raises the equal error rate from 7.3% to 10.6%\[1].
### Can a provider detect and defeat the test?
Not easily. Probes are ordinary harmless questions drawn from paraphrase pools, so special-casing them without breaking normal traffic is difficult\[1].
### Does a failed check mean a provider is cheating?
No. Quantization, silent version updates, hidden system prompts, and stale references all produce drift without deception. A distance is a statistical observation warranting closer inspection.
### Is the dataset available?
Yes. The fingerprint dataset is on Zenodo under CC-BY-4.0 and the reproduction code under MIT\[3]\[4].
## Treating the model field as a claim
The useful shift here is small and specific. The `model` field in an API response is a claim, and there is now a cheap, open, statistically grounded way to check that claim from the outside, built on nothing more exotic than the fact that language models cannot say a random number to save their lives.
If you buy capacity through any intermediary, the right place for this is next to uptime monitoring: a periodic check against a reference you collected yourself, with the auxiliary signals watched alongside the headline distance. And the right way to read a red result is as the beginning of a conversation, not the end of one. To verify an LLM API is to gather evidence about an endpoint at a moment in time, which is worth doing precisely because the alternative is assuming.
## References
1. Bruckner, Tomáš. *One Token Is Enough: Fingerprinting and Verifying Large Language Models from Single-Token Output Distributions.* arXiv, July 2026. [arxiv.org/abs/2607.10252](https://arxiv.org/abs/2607.10252)
2. CISPA researchers. *Real Money, Fake Models — an audit of shadow LLM API operators.* arXiv. [arxiv.org/abs/2603.01919](https://arxiv.org/abs/2603.01919)
3. *LLM fingerprint dataset (models × tasks × languages).* Zenodo, CC-BY-4.0. [zenodo.org](https://zenodo.org/)
4. *Reproduction code for One Token Is Enough.* Zenodo, MIT License. [zenodo.org](https://zenodo.org/)
### Further reading
* reAPI. *How to use Claude Opus 5.* [reapi.ai/blog/how-to-use-claude-opus-5](/blog/how-to-use-claude-opus-5)
* reAPI. *How to use GPT-5.6.* [reapi.ai/blog/how-to-use-gpt-5-6](/blog/how-to-use-gpt-5-6)
---
# Wan 2.7 Video API: 1080p, Audio, Pricing, and Limits (2026) (https://reapi.ai/blog/wan-2-7-video-api-guide)
The **Wan 2.7 Video API** covers more than text-to-video. The same model family
can generate from text, animate first and last frames, continue an existing
clip, use reference images or videos, accept driving audio, and edit video. It
supports 720p and 1080p output with integer durations from 2 to 15 seconds.
That breadth creates one practical problem: request shape determines the mode,
and mode affects what is billable. A five-second text-to-video job is not always
priced like a five-second reference-to-video or edit job.
## TL;DR
* Wan 2.7 supports **720p and 1080p**, five aspect ratios, and **2–15 second**
outputs.\[1]
* Text-to-video accepts an optional audio file; without one, the model can
generate matching background music or sound effects.
* reAPI currently displays **$0.074/s at 720p** and **$0.121/s at 1080p**.
* The reAPI umbrella ID is `wan2.7-video`; inputs determine whether the job is
text, image, reference, continuation, or edit.
* Generated URLs are temporary. Download results to durable storage rather
than serving the provider URL in a product.
## Wan 2.7 Video API specifications
| Specification | Wan 2.7 Video |
| --------------------- | -------------------------------------- |
| reAPI model ID | `wan2.7-video` |
| Output resolution | 720p, 1080p |
| Duration | Any integer from 2 to 15 seconds |
| Aspect ratios | `16:9`, `9:16`, `1:1`, `4:3`, `3:4` |
| Frame rate | 30 fps |
| Text prompt limit | 5,000 characters |
| Negative prompt limit | 500 characters |
| Custom audio | WAV or MP3, 2–30 seconds, up to 15MB |
| Prompt extension | Optional, enabled by default |
| Watermark | Optional, disabled by default |
| API pattern | Asynchronous task creation and polling |
Alibaba's official HTTP API requires asynchronous invocation for Wan 2.7.
After submission, the task ID and result URL remain available for 24 hours.\[1]
reAPI normalizes task submission behind its video endpoint, but your application
still needs to persist the returned job ID, poll with backoff, and copy the
finished asset into permanent storage.
## Wan 2.7 Video API pricing by resolution
Wan 2.7 uses per-second billing. The current reAPI rate card applies a 10%
platform margin to a discounted upstream rate and rounds the completed request
to whole credits.
| Resolution | Current reAPI display rate | Upstream list rate | 5s reAPI job | 15s reAPI job |
| ---------- | -------------------------: | -----------------: | -----------: | ------------: |
| 720p | $0.074/s | $0.083/s | $0.366 | $1.096 |
| 1080p | $0.121/s | $0.137/s | $0.603 | $1.809 |
The 5- and 15-second examples use aggregate request rounding, which is why they
can be slightly lower than multiplying the displayed rounded rate by duration.
Rates were checked on August 1, 2026.
For text-to-video and ordinary image-to-video, the billable count is the output
duration. Reference-to-video and video editing can also include probed source
video duration. Budget those modes from the estimator or returned usage rather
than assuming the output slider is the entire bill.

## How to call Wan 2.7 through reAPI
A text-to-video request is the smallest useful example:
```bash
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer $REAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "wan2.7-video",
"prompt": "Generate a single shot. A paper boat drifts down a rain-soaked city street at dusk. Slow low-angle tracking camera, warm shop reflections, natural rain and distant traffic audio.",
"negative_prompt": "text, logo, watermark, camera shake",
"size": "16:9",
"resolution": "1080P",
"duration": 5,
"prompt_extend": true,
"watermark": false,
"seed": 31415
}'
```
The [Wan 2.7 Video model page](/models/wan-2-7-video) exposes the available
modes, while the [API documentation](/docs/wan-2-7-video) should be used for
the current payload schema.
The mode is implicit. Do not add a fictional `mode: "text-to-video"` field.
Instead, supply the media fields for the workflow you want.
## Choosing the correct Wan 2.7 input mode
### Text-to-video
Provide `prompt` and common output controls. You can also pass `audio_url` for a
custom soundtrack. Alibaba says that when audio is omitted, Wan 2.7 can produce
matching background music or effects.\[1]
### First-frame or first-and-last-frame video
Use `image_urls` for a simple first-frame workflow or `image_with_roles` when
the images need explicit `first_frame` and `last_frame` roles. The first image
usually determines output composition and approximate aspect ratio.
### Video continuation
Provide one source through `video_urls`. The source clip is continued until the
requested total duration. For example, a 3-second source with `duration: 10`
produces roughly seven new seconds and a ten-second combined result.
### Reference-to-video
Use `reference_image_urls`, `reference_video_urls`, or `reference_voice_urls`
to preserve appearance, motion, or voice characteristics. The current reAPI
surface allows up to five URLs in each reference category.
### Video editing
Provide `video_url` and an edit prompt. Up to four supporting images can guide
replacement or styling. `audio_setting: "origin"` favors source audio;
`"auto"` lets the workflow decide how audio is handled.
## How to write Wan 2.7 prompts
Wan 2.7 interprets shot structure from the prompt. The older `shot_type` field
does not control Wan 2.7 text-to-video. To request one continuous shot, say
“Generate a single shot.” For multi-shot work, name each shot and give it a time
range.\[1]
```text
Generate a multi-shot video.
Shot 1 [0–4s]: wide shot of a quiet train platform before sunrise.
Shot 2 [4–8s]: medium tracking shot as a traveler walks beside the train.
Shot 3 [8–12s]: close-up of a paper ticket bending in the wind.
Natural station ambience, no dialogue, restrained camera motion.
```
Keep the timeline inside the requested duration. If three shots each imply ten
seconds but the API asks for a six-second output, the model has to compress or
drop instructions. The official prompt guide also recommends naming references
as `Image 1` or `Video 1` when several assets are present.\[3]
## Audio behavior and edge cases
Custom audio can drive timing, dialogue, or background sound, but the durations
need to match. If audio is longer than the requested video, it is truncated. If
it is shorter, the remainder can be silent. Pre-trim the audio when the exact
ending matters.
For lip-sync or voice reference work, use clean speech with limited background
noise. A soundtrack that mixes music, overlapping dialogue, and environmental
noise gives the model less reliable timing evidence than a dedicated voice
track.
Also remember that audio changes rights and consent obligations. A technically
valid reference voice is not evidence that you have permission to clone or
publish it.
## Common integration mistakes
### Treating generation as synchronous
Video jobs can take minutes. Store the task ID, poll with exponential backoff,
handle failure states, and make the UI resumable after a page refresh.
### Serving the temporary result URL
Alibaba's result URLs expire after 24 hours. Download successful results and
upload them to your own object storage before returning a permanent asset URL.
### Assuming 1080p is always the best test setting
Prototype prompts at 720p, then rerun selected shots at 1080p. That keeps prompt
iteration cheaper while preserving final delivery quality.
### Ignoring billable input duration
Reference-video and edit workflows may include source duration in the bill.
Log request mode, source length, output length, resolution, and final charge.
For a broader comparison of incompatible billing units, read
[AI video generation API pricing](/blog/ai-video-generation-api-pricing). If
your workflow starts from a still image and needs a different motion model, the
[Grok Imagine Video 1.5 API guide](/blog/grok-imagine-video-1-5-api) explains
its image-to-video-only surface.
## FAQ
### How long can Wan 2.7 videos be?
Wan 2.7 accepts integer durations from 2 to 15 seconds for its primary text and
image workflows.
### Does Wan 2.7 support 1080p?
Yes. It supports 720p and 1080p. Resolution affects per-second price.
### Can Wan 2.7 generate audio?
Yes. Without a custom audio file, the text-to-video route can generate matching
background music or sound effects. It also accepts WAV or MP3 input.
### Does Wan 2.7 support first and last frames?
Yes. Its image-to-video API supports first-frame and first-plus-last-frame
workflows, along with continuation from an existing clip.
### How much does a five-second Wan 2.7 video cost on reAPI?
At the rates checked on August 1, 2026, an ordinary five-second job is about
$0.366 at 720p or $0.603 at 1080p. Input-billed reference and edit modes can
cost more.
## Conclusion
The Wan 2.7 Video API is valuable because one family covers generation,
continuation, reference control, audio, and editing. The integration works best
when the application determines the request mode explicitly, budgets resolution
and billable duration correctly, and treats job polling and asset persistence as
first-class parts of the product.
## References
1. Alibaba Cloud Model Studio. *Wan2.7 text-to-video API reference.* [alibabacloud.com/help/en/model-studio/text-to-video-api-reference](https://www.alibabacloud.com/help/en/model-studio/text-to-video-api-reference)
2. Alibaba Cloud Model Studio. *Wan2.7 image-to-video API reference.* [alibabacloud.com/help/en/model-studio/image-to-video-general-api-reference](https://www.alibabacloud.com/help/en/model-studio/image-to-video-general-api-reference)
3. Alibaba Cloud Model Studio. *Wan video prompt guide.* [alibabacloud.com/help/en/model-studio/text-to-video-prompt](https://www.alibabacloud.com/help/en/model-studio/text-to-video-prompt)
---
# WaveSpeed AI Pricing: Free Trial, Tiers, and PAYG (https://reapi.ai/blog/wavespeed-ai-pricing-free-trial-and-payg-2026)
**WaveSpeed AI is pay as you go. There is no standard monthly membership: you add prepaid credits, then WaveSpeed deducts the cost of each generation from that balance.** New accounts receive a $1 trial without a credit card, but some premium models are excluded, and WaveSpeed's API documentation says an API key needs a successful top-up before it will work\[1]\[3].
That last detail changes how developers should evaluate the “free trial.” It is useful for trying supported models in WaveSpeed's web interface. It should not be treated as a permanent free API tier.
This article was checked against WaveSpeed's official pricing and account documentation on July 30, 2026. Disclosure: reAPI publishes this site and also sells pay-as-you-go AI API access. All WaveSpeed claims below link to WaveSpeed's own documentation rather than to our marketing pages.
## WaveSpeed pricing at a glance
| Question | Short answer |
| --------------------------------------------- | ---------------------------------------------------------------------- |
| Is there a monthly membership? | No standard consumer subscription; most accounts use prepaid credits |
| Is billing pay as you go? | Yes; the charge varies by model and generation settings |
| Is there a free trial? | Yes, $1 for new accounts, with no card required |
| Does trial credit cover every model? | No; some premium models require a paid balance |
| Can trial credit activate an API key? | WaveSpeed's authentication docs say API keys require a top-up |
| Do purchased credits expire? | No |
| Are purchased credits refundable? | No, although eligible system-side inference failures are credited back |
| What do Bronze, Silver, Gold, and Ultra mean? | They are rate-limit levels, not recurring membership plans |
WaveSpeed's public pricing page describes the service as pay per use: there are no monthly fees or commitments, and users pay for generated output\[1]. Its billing documentation adds two conditions that matter when budgeting: purchased credits do not expire, and they cannot be converted back to cash\[4].
## How the free trial actually works
A new WaveSpeed account receives $1 in trial credit without entering a credit card. The trial reaches most of the catalog, but not every endpoint. WaveSpeed explicitly warns that some premium models are unavailable until the account has a paid balance\[1].
There are therefore two separate tests:
1. **Web test:** create an account and use the trial balance on an eligible model in the WaveSpeed interface.
2. **API test:** create an API key for an application or script.
The second step has an extra condition. WaveSpeed's authentication page says that keys generated without a top-up will not work and instructs users to make at least one top-up to activate API access\[3]. Taken together, the official pages indicate that the no-card trial is a product trial, not a promise of free API calls.
This distinction is easy to miss if you land on the pricing page first. If your evaluation requires queues, webhooks, retry handling, or production latency measurements, budget for a small paid top-up rather than assuming the trial can run that test.
WaveSpeed's refund policy says account-to-account credit transfers can be requested through support. It does not describe a self-service transfer feature or separately define trial-credit transferability, so confirm any transfer request with support before relying on it\[5].
## The account tiers are throughput levels
WaveSpeed calls Bronze, Silver, Gold, and Ultra “account levels.” They do not provide a monthly bundle of generations. Their main purpose is to set how many jobs an account can start per minute and how many can run at once\[2].
| Account level | How it is activated | Predictions per minute | Concurrent predictions |
| ------------- | ------------------------------------------------ | ---------------------: | ---------------------: |
| Bronze | Default for a new account | 2 | 2 |
| Silver | A successful single top-up below $1,000 | 500 | 300 |
| Gold | A successful single top-up from $1,000 to $4,999 | 3,000 | 3,000 |
| Ultra | A successful single top-up of at least $5,000 | 5,000 | 10,000 |
The upgrade is based on the size of one successful top-up, not cumulative lifetime spend. It only moves upward: WaveSpeed says an existing higher level is not automatically downgraded\[2].
That makes Silver the practical production threshold for many small applications. Bronze permits only two running tasks, while a successful paid top-up below the Gold threshold moves an unrestricted lower-level account to Silver. The top-up still becomes spendable credit; it is not a separate membership fee.
Gold and Ultra make sense when concurrency is the constraint. They do not make an expensive model cheaper by themselves. Volume pricing is a separate sales conversation, and WaveSpeed directs high-volume customers to contact its team for discounts or limits beyond Ultra\[1]\[2].
## What determines the charge for a generation
There is no single WaveSpeed rate that can be multiplied by every request. The final amount depends on the selected model and the parameters sent with the job. WaveSpeed lists four common pricing factors\[4]:
* model complexity;
* output resolution;
* video duration;
* batch size or number of outputs.
This is why a model card may say “from” a particular price. A low-resolution image, a long high-resolution video, and a multi-image batch use different billing units even when they sit behind the same account balance.
WaveSpeed shows an estimated price on the Run button before a web generation, while warning that the final charge can differ slightly from the estimate\[4]. For an API workflow, its pricing endpoint accepts a model ID and intended inputs so a project can estimate the task before submitting it\[6].
That estimate should be part of the application, not a number copied into a spreadsheet once. Model rates and supported parameters can change. The live model page or pricing endpoint is safer than an article—including this one—for the final cost of a specific request.
## Failed jobs, timeouts, and refunds
WaveSpeed says eligible inference failures caused by system errors are automatically returned to the credit balance\[5]. A browser or client timeout is different.
If a synchronous HTTP request disconnects, the prediction may continue running on WaveSpeed's servers. A completed prediction can still be billable even though the client never received the initial response. WaveSpeed advises retaining the prediction ID and retrieving the result rather than assuming a timeout means the job failed\[5].
For production budgeting, count accepted outputs and server-side prediction status—not only successful HTTP responses. A client that blindly resubmits after every timeout can pay for duplicate work.
## How to keep PAYG spending predictable
Pay-as-you-go billing is easiest to control when every test changes one variable at a time.
1. **Begin at the lowest useful resolution.** WaveSpeed recommends testing video prompts at lower resolution before generating the final version\[6].
2. **Use one output while refining a prompt.** Increase batch size only after the prompt is stable.
3. **Ask the pricing endpoint before expensive jobs.** Include the same duration, resolution, and batch settings you intend to submit.
4. **Store prediction IDs.** Poll or use a webhook after a client timeout instead of immediately creating a duplicate.
5. **Separate experimentation from production.** WaveSpeed supports multiple API keys, so different projects can have separate credentials even though billing remains at account level\[3].
6. **Review successful-output cost.** Total spend divided by accepted outputs is more useful than the lowest advertised model rate.
Credits do not expire, so an irregular workload does not need to race a monthly reset. The trade-off is that purchased balance is non-refundable. Top up for the workload you can reasonably forecast, not for a tier label alone.
## When WaveSpeed's pay-as-you-go model fits
PAYG is a good fit when usage changes from week to week, the application routes across several image or video models, or the team wants to test without taking on a monthly commitment. It also works well when a project can tolerate queueing at first and upgrade throughput only after real traffic arrives.
It is less comfortable when finance needs an identical invoice every month or when a production launch requires high concurrency from day one. In those cases, the relevant question is not whether WaveSpeed has a “Pro membership.” It is whether the necessary single top-up, enterprise agreement, or reviewed monthly credit line matches the project's cash-flow and capacity requirements.
WaveSpeed documents monthly credit lines only for reviewed AWS-login accounts. Approval is case by case, generally tied to a verifiable product or business, payment history, and repayment capacity. Standard Google and GitHub accounts continue to use prepaid top-ups\[2]\[7].
## WaveSpeed vs reAPI for an API workload
reAPI, the company publishing this article, is another pay-as-you-go AI API provider. That makes us commercially interested in the comparison. The WaveSpeed limits above come from WaveSpeed's first-party pages; the reAPI side points to live product pages so both columns can be checked before funding an account.
| Decision point | WaveSpeed | reAPI |
| -------------------------- | -------------------------------------------------- | ------------------------------------------------------------------------------------- |
| Product shape | Broad, speed-first inference catalog | Curated multi-model gateway for product teams |
| Getting an API key working | Official docs say a successful top-up is required | Create one key and use the available account balance across chat and media |
| Billing | Prepaid PAYG; price varies by model and parameters | PAYG credits; each [model page](/models) shows its current billing unit |
| Failed media jobs | Eligible system-side failures are credited back | A task that finishes in `failed` is automatically refunded |
| Integration | OpenAI-compatible surface plus model APIs | OpenAI-compatible chat plus one asynchronous task pattern for media |
| Best fit | Teams prioritizing a very broad catalog and speed | SaaS teams prioritizing a focused catalog, one balance, and predictable task handling |
The table is not a claim that one provider is always cheaper. Run the same model, prompt, inputs, resolution, duration, and retry policy on both, then compare total spend per accepted output. For reAPI, start with the [live model catalog](/models) and use the [API quickstart](/docs/api/quickstart) for the first request.
If billing is only one part of the decision, the separate [WaveSpeed alternatives guide](/blog/best-wavespeed-alternatives) covers differences in model access, infrastructure control, and free testing. Those comparison questions are intentionally outside the scope of this pricing guide.
## FAQ
### How much does WaveSpeed AI cost?
WaveSpeed does not have one platform-wide generation price. It charges per use, with the amount determined by the selected model and settings such as resolution, duration, and batch size. Check the live model page or pricing endpoint before submitting a job.
### Is WaveSpeed AI free?
New accounts receive $1 in free trial credit without a card. Some premium models cannot use trial credit, and the official authentication documentation says an API key requires a top-up before it works.
### Does WaveSpeed have a monthly membership?
No standard monthly consumer membership is listed. Most users add prepaid credits that never expire. Bronze, Silver, Gold, and Ultra are rate-limit levels rather than recurring subscription plans.
### What is the WaveSpeed Silver tier?
Silver allows 500 predictions per minute and 300 concurrent predictions. WaveSpeed says a successful single top-up below $1,000 upgrades an eligible lower-level account to Silver.
### Do WaveSpeed credits expire?
Purchased credits do not expire. They are generally non-refundable, so unused balance stays in the account rather than converting back to cash.
### Does WaveSpeed refund failed generations?
WaveSpeed says eligible system-side inference failures are returned to the account balance. A client timeout does not necessarily qualify because the prediction may continue and finish on the server.
### Can I use the WaveSpeed API with free trial credit?
Do not assume so. Although new accounts receive trial credit, WaveSpeed's API authentication page separately says keys need a top-up to activate. Confirm the current status in your account before building an evaluation around free API access.
### Is WaveSpeed pay as you go?
Yes. Standard accounts prepay credits, and WaveSpeed deducts usage as generations complete. There are no standard monthly fees or commitments on the public pricing page.
### When should I choose reAPI instead of WaveSpeed?
Choose reAPI when you are building a product around a curated set of models and value one credit balance, live per-model billing pages, and a consistent asynchronous task flow more than maximum catalog size. Choose WaveSpeed when its broader catalog or speed-focused infrastructure is the requirement. Test the exact shared models before deciding.
## References
1. WaveSpeed AI. *Pricing.* Pay-per-use billing, no monthly commitment, free trial restrictions, account-level summary, and enterprise pricing. Retrieved July 30, 2026 from wavespeed.ai/pricing
2. WaveSpeed AI. *Account Levels & Rate Limits.* Tier activation rules, predictions per minute, concurrency, upgrades, and monthly credit-line overview. Retrieved July 30, 2026 from wavespeed.ai/docs/account-levels
3. WaveSpeed AI. *Authentication.* API-key activation, multiple-key support, and credential guidance. Retrieved July 30, 2026 from wavespeed.ai/docs/api-authentication
4. WaveSpeed AI. *How Pricing Works.* Pricing factors, prepaid-credit policy, and volume discounts. Retrieved July 30, 2026 from wavespeed.ai/docs/how-pricing-works
5. WaveSpeed AI. *Refund Policy.* Credit transfers, system-side failure eligibility, and client-timeout behavior. Retrieved July 30, 2026 from wavespeed.ai/docs/refund-policy
6. WaveSpeed AI. *How to Reduce Costs.* Lower-resolution testing, batch sizing, and the model-pricing endpoint. Retrieved July 30, 2026 from wavespeed.ai/docs/reduce-costs
7. WaveSpeed AI. *Payment Methods.* Prepaid balance, non-expiring credits, and reviewed monthly credit lines. Retrieved July 30, 2026 from wavespeed.ai/docs/payment-methods
---
# What Can reAPI Do for You? Image, Video & LLM Use Cases (https://reapi.ai/blog/what-can-reapi-do)
Most AI projects do not need one model. They need a few: a chat model for reasoning, an image model for assets, a video model for clips, maybe a voice model on top. **reAPI** puts all of them behind one key, one balance, and one base URL, priced 20-50% below the providers' official rates. This is a practical look at what reAPI does today, who it fits, and how to put it to work.
## What is reAPI?
reAPI is a unified API for generative AI. One endpoint, `https://reapi.ai/api/v1`, serves 200+ models across chat, image, video, and audio. The chat surface is OpenAI-compatible, so existing OpenAI code runs against it by changing the base URL and key, while image and video run as asynchronous jobs under the same credential.
### One API for text, image, video, and audio
The point of reAPI is breadth without sprawl. A single integration covers:
* **Text and reasoning:** GPT-5, Claude Opus 4.8, Gemini.
* **Image:** GPT-Image-2, Gemini 3 Pro Image, Imagen 4, Seedream 5.0.
* **Video:** Veo 3.1, Seedance 2.0, Wan 2.7, Kling.
* **Audio and music:** Mureka V9 and a set of voice tools.
You move between them by changing a model string, not by onboarding a new vendor.
### OpenAI-compatible by design
reAPI follows OpenAI's conventions where it makes sense, so most OpenAI clients work by changing only the base URL and key. That lowers the switching cost to almost nothing: the code you already wrote against OpenAI keeps working, and you reach Claude, Gemini, and the media models through the same client.
### Fits into automation pipelines
Because the chat surface speaks the OpenAI format, reAPI drops into any tool that accepts an OpenAI-compatible endpoint, from agent frameworks to no-code automation builders. Point the tool at reAPI's base URL, paste a key, and it can call any model in the catalog. The media endpoints follow a submit-then-poll pattern that fits a queue or a scheduled job.
## Who reAPI is for
* **SaaS startups** adding AI features without a separate vendor integration per capability.
* **Enterprise teams** that want one balance, one invoice, and central usage visibility.
* **AI automation builders** wiring models into workflows and no-code platforms.
* **AI product developers** who need to swap models as new ones ship.
* **Agencies and marketing teams** producing images and video at a predictable per-output cost.
## Five things you can build with reAPI today
### 1. Customer-facing AI features in a SaaS product
Add a chat assistant, an image generator, and a short-video feature to your app through one key. Because chat is OpenAI-compatible, the assistant is a base-URL change away, and the image and video features call the task endpoints. One balance covers all three, so finance tracks a single line item instead of three vendor invoices. When a better model ships, you point the same call at it without a new contract or integration, which keeps a shipping product current with little engineering cost.
### 2. Content and marketing production
Generate product images, social assets, and short video clips at a flat per-output rate. GPT-Image-2 starts at $0.0066 per image and Seedance 2.0 at $0.0506 per video, so a campaign's cost is quotable in advance rather than a surprise at month end. A team can storyboard with a chat model, render the stills with an image model, and animate the hero shots with a video model, all from the same key and the same budget line.
### 3. Agency work across many clients
Run every client's generation through one reAPI balance and tag usage per key. Pricing sits 20-50% below official rates, which becomes margin on fixed-bid creative work, and the single balance removes the overhead of reconciling several provider accounts.
### 4. Research and prototyping
Compare models without throwaway accounts. A/B test GPT-5 against Claude Opus 4.8 by changing one string, then try a different image model the same way. New accounts start with free credits, so evaluation costs nothing up front.
### 5. AI automation and no-code workflows
Connect reAPI to an automation builder through its OpenAI-compatible endpoint and let a workflow call models on a trigger: summarize an inbound document, generate a thumbnail, render a clip. The submit-then-poll media pattern fits scheduled and event-driven jobs.
## A typical rollout
Teams usually start small and widen. The first week is a single feature behind one key: a chat assistant pointed at reAPI's base URL, or an image endpoint wired into an existing form. Because the chat surface is OpenAI-compatible, that first integration is a base-URL change, not a rewrite.
From there, the same key picks up a second modality. A support tool that summarizes tickets adds a thumbnail generator; a content pipeline that writes copy adds a video step. Nothing new gets provisioned, because the models already share one balance. By the time three or four models are in play, the consolidation that looked optional at the start is the reason the integration stayed simple.
The flat per-output pricing helps here too. As volume grows, cost scales linearly and stays quotable, so a feature that worked in a pilot does not turn into a billing surprise at scale.
## reAPI vs OpenRouter vs CometAPI
All three put many models behind one OpenAI-style key. They differ on modality depth, pricing, and how you start.
| | reAPI | OpenRouter | CometAPI |
| ---------------- | ----------------------------------- | ------------------------------ | ------------------------- |
| Modalities | Image, video, audio, chat | Mostly text, some multimodal | Text, image, video, audio |
| Pricing | 20-50% below official | Pass-through + 5.5% credit fee | \~20% below official |
| Video generation | Curated (Veo, Seedance, Wan, Kling) | Some, routed | Yes (Sora, Veo, Kling) |
| Free to start | Free credits | Small allowance + free models | Test credits |
| Best for | Media plus LLMs, flat pricing | Text-model breadth | Unified discounted access |
OpenRouter passes the provider's rate through and adds a 5.5% ($0.80 minimum) fee on credit purchases\[2]; CometAPI prices around 20% below official\[3]. reAPI's edge is curated media depth, especially video, alongside frontier LLMs at a flat per-output rate.
## When reAPI is the right call
reAPI fits best when a project spans more than one modality or more than one model, and you want predictable cost and a single integration. If you only ever call a single text model and already run it direct, the consolidation matters less. The moment a second model or a second modality enters the picture, one key and one balance start paying off.
## Setup checklist for your first integration
1. Create an account at [reapi.ai](https://reapi.ai) and generate a key under API Keys.
2. For chat, point an OpenAI client at `https://reapi.ai/api/v1` and set the key.
3. For image and video, POST to the generation endpoints and poll `GET /api/v1/tasks/{id}`.
4. Watch usage and top up the credit balance from the dashboard.
5. Add retry handling on your side; reAPI does not deduplicate, and failed tasks refund automatically.
## FAQ
### What can reAPI do that a single-vendor API cannot?
It covers multiple modalities and multiple providers' models through one key and one balance. You reach text, image, video, and audio, and swap between GPT-5, Claude, and Gemini, without a separate account or integration for each.
### Is reAPI good for video generation?
Yes. reAPI carries a curated video lineup, including Veo 3.1, Seedance 2.0, Wan 2.7, and Kling, billed at a flat rate per video so the cost is known before you render.
### Can reAPI replace my OpenAI integration?
For chat, yes. reAPI is OpenAI-compatible, so most code works by changing the base URL to `https://reapi.ai/api/v1` and the key. You then also get Claude, Gemini, and the media models through the same client.
### Does reAPI work with no-code and automation tools?
Any tool that accepts an OpenAI-compatible endpoint can call reAPI's chat models by pointing at the base URL. Media generation uses submit-then-poll endpoints that fit scheduled or event-driven workflows.
### How do I keep costs predictable?
Media is billed flat per output and chat per token, both 20-50% below official rates, from one pay-as-you-go balance. Failed tasks refund automatically, so you only pay for output you receive.
## Putting reAPI to work
What reAPI does, in one line, is collapse a multi-vendor AI stack into one key, one balance, and one OpenAI-compatible base URL, across image, video, audio, and chat, at 20-50% below official rates. Pick the use case closest to yours, create a key, and the first calls run on free starting credits. That is the fastest way to see what reAPI can do for your own workload.
## Further reading
* [reapi.ai/models](/models) — browse every model across image, video, audio, and chat.
* [What is reAPI?](/blog/what-is-reapi) — the platform, pricing, and how it works.
* [Best fal.ai alternatives](/blog/best-fal-ai-alternatives) — how reAPI compares in the media space.
## References
1. reAPI. *API overview — base URL, authentication, and asynchronous tasks.* Retrieved May 2026 from [reapi.ai/docs/api](/docs/api)
2. OpenRouter. *Docs FAQ — pass-through pricing and credit fees.* Retrieved May 2026 from openrouter.ai/docs/faq
3. CometAPI. *Pricing — below-official model rates.* Retrieved May 2026 from cometapi.com/pricing
---
# What Is Claude Opus 4.8? Anthropic's New Model Explained (https://reapi.ai/blog/what-is-claude-opus-4-8)
Anthropic shipped **Claude Opus 4.8** on May 28, 2026, an upgrade to Opus 4.7 that it describes as a modest but tangible improvement\[1]. It is Anthropic's most capable model for complex reasoning and long-horizon agentic coding, and it launched at the same price as the model it replaces. The headline change is not a benchmark leap but reliability: Claude Opus 4.8 is roughly four times less likely than Opus 4.7 to let a flaw in code it wrote pass unremarked\[1].
This guide explains what Claude Opus 4.8 is, what changed from 4.7, its published benchmarks, what it costs, and how to call it through an OpenAI-compatible API. Every number below comes from Anthropic's own announcement, model overview, and pricing pages, retrieved on May 30, 2026.
## What is Claude Opus 4.8?
Claude Opus 4.8 is the latest model in Anthropic's Opus line, the tier aimed at the hardest reasoning and agentic-coding work. The API identifier is `claude-opus-4-8`, with no date suffix\[2].
The core specs:
* **Context window:** 1M tokens, billed at the standard per-token rate across the whole window\[2].
* **Max output:** 128K tokens on the synchronous API, up to 300K on the Batch API with a beta header\[2].
* **Input modalities:** text and image. Output is text. Vision is supported; Anthropic does not list audio or video input\[2].
* **Knowledge cutoff:** January 2026\[2].
## What's new compared to Opus 4.7
Anthropic frames Opus 4.8 as building on 4.7 rather than fully replacing it, and ships it at the same input and output price\[1]. Three changes stand out.
* **Honesty.** Opus 4.8 is about four times less likely than its predecessor to let a flaw in its own code pass without flagging it, and it is more willing to mark uncertainty instead of asserting unsupported claims\[1].
* **A cheaper, faster fast mode.** Fast mode now runs at $10 per million input tokens and $50 per million output, which Anthropic says is three times cheaper and 2.5 times faster than the previous generation's fast tier\[1].
* **Effort control.** The model defaults to `high` effort and can be pushed to `extra` or `max`, spending more tokens for better results on hard problems\[2].
## Core capabilities of Claude Opus 4.8
### A 1M-token context window
Opus 4.8 reads up to a million tokens in one request, enough for a large codebase or a long document set, and a 900K-token request is billed at the same per-token rate as a 9K-token one\[2]. On Microsoft Foundry the window is capped at 200K\[2].
### Agentic coding and reasoning
The model leads Anthropic's published comparison on agentic coding and multidisciplinary reasoning. It scores 69.2% on SWE-Bench Pro and 57.9% on Humanity's Last Exam with tools, both ahead of Opus 4.7\[1]. Full numbers are in the benchmark table below.
### Effort control and adaptive thinking
Rather than a single extended-thinking toggle, Opus 4.8 uses adaptive thinking plus an effort setting. Leaving effort at the default `high` suits most work; raising it to `extra` or `max` trades more tokens for deeper reasoning on the hardest tasks\[2].
### Vision, tool use, and prompt caching
Opus 4.8 accepts image input alongside text, supports tool use and function calling, and works with prompt caching to cut the cost of repeated context\[2]. These are the building blocks for agents that read screenshots, call tools, and carry long context across a session.
## Claude Opus 4.8 benchmarks
These are Anthropic's own published results, comparing Opus 4.8 with Opus 4.7, GPT-5.5, and Gemini 3.1 Pro\[1]. As with any vendor-run benchmark, treat them as a directional signal rather than a neutral audit.
| Benchmark | Opus 4.8 | Opus 4.7 | GPT-5.5 | Gemini 3.1 Pro |
| -------------------------------------------- | -------- | -------- | ------- | -------------- |
| Agentic coding (SWE-Bench Pro) | 69.2% | 64.3% | 58.6% | 54.2% |
| Terminal coding (Terminal-Bench 2.1) | 74.6% | 66.1% | 78.2% | 70.3% |
| Reasoning (Humanity's Last Exam, no tools) | 49.8% | 46.9% | 41.4% | 44.4% |
| Reasoning (Humanity's Last Exam, with tools) | 57.9% | 54.7% | 52.2% | 51.4% |
| Computer use (OSWorld-Verified) | 83.4% | 82.8% | 78.7% | 76.2% |
| Knowledge work (GDPval-AA) | 1890 | 1753 | 1769 | 1314 |
| Financial analysis (Finance Agent v2) | 53.9% | 51.5% | 51.8% | 43.0% |
Opus 4.8 tops the table on every row except Terminal-Bench 2.1, where GPT-5.5 leads at 78.2%\[1].
## Claude Opus 4.8 pricing
Pricing is per million tokens (MTok), unchanged from Opus 4.7\[1]\[3]:
| Item | Price / MTok |
| ---------------------- | ------------ |
| Input | $5 |
| Output | $25 |
| Cache read | $0.50 |
| Cache write (5-minute) | $6.25 |
| Cache write (1-hour) | $10 |
| Batch input | $2.50 |
| Batch output | $12.50 |
| Fast mode input | $10 |
| Fast mode output | $50 |
The Batch API runs at half the standard input and output price, and prompt caching makes repeated context cheap to reread at $0.50 per million tokens\[3].
## How to access Claude Opus 4.8
You can call Claude Opus 4.8 directly through Anthropic, or reach it through reAPI's OpenAI-compatible gateway at Anthropic's standard rates. See the [reapi.ai/models/claude-opus-4-8](/models/claude-opus-4-8) page for live rates and a playground. The advantage of the gateway is one key and one client across providers: the same integration that calls GPT-5 and Gemini, plus reAPI's image and video models, also calls Claude Opus 4.8.
Point the OpenAI SDK at reAPI and set the model:
```python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_REAPI_KEY",
base_url="https://api.reapi.ai/v1",
)
resp = client.chat.completions.create(
model="claude-opus-4-8",
messages=[
{"role": "user", "content": "Refactor this function and explain the change."},
],
extra_body={"group": "default"},
)
print(resp.choices[0].message.content)
```
The native Anthropic `/v1/messages` surface is available too, so SDKs written for either format work without a rewrite. Input runs $5 per million tokens and output $25, the same rates Anthropic charges directly.
## Who should use Claude Opus 4.8
Opus 4.8 is built for work where a mistake is expensive and context is large:
* **Agentic coding.** Long-horizon tasks where the model plans, edits, and verifies its own work, and where the honesty gains reduce silent bugs.
* **Large-context analysis.** Reading an entire codebase or document set in one request, up to a million tokens.
* **Tool-using agents.** Workflows that call tools, read screenshots, and carry state across many steps.
* **High-stakes reasoning.** Research, financial analysis, and other work where flagged uncertainty beats a confident wrong answer.
For high-volume, low-complexity calls, a smaller and cheaper model is usually the better fit; Opus is the tier you reach for when capability matters more than cost per token.
## Why Claude Opus 4.8 matters in 2026
The interesting part of Opus 4.8 is not a single benchmark. It is that Anthropic chose to spend a release on reliability, cutting how often the model lets its own mistakes slide, rather than chasing a leaderboard. For agentic coding, where a model edits real code across many steps, that kind of honesty compounds: fewer silent bugs to catch downstream. Paired with a million-token window and a cheaper fast mode, Opus 4.8 is aimed squarely at teams building agents they have to trust.
## FAQ
### When was Claude Opus 4.8 released?
May 28, 2026, as an upgrade to Claude Opus 4.7, at the same price\[1].
### What is the Claude Opus 4.8 context window?
1M tokens, billed at the standard per-token rate across the full window. Maximum output is 128K tokens, or up to 300K on the Batch API\[2].
### How much does Claude Opus 4.8 cost?
$5 per million input tokens and $25 per million output, with cache reads at $0.50 and the Batch API at half price\[3]. Fast mode is $10 input and $50 output\[1].
### Is Claude Opus 4.8 multimodal?
It accepts text and image input and returns text, with vision supported. Anthropic does not list audio or video input for the model\[2].
### How is Claude Opus 4.8 different from Opus 4.7?
Same price, with higher honesty (about four times less likely to let its own code flaws pass), a cheaper and faster fast mode, and small gains across Anthropic's benchmark suite\[1].
### How do I call Claude Opus 4.8 with the OpenAI SDK?
Point the OpenAI client at `https://api.reapi.ai/v1`, set `model` to `claude-opus-4-8`, and send a normal chat request. reAPI exposes Claude through an OpenAI-compatible endpoint at Anthropic's standard rates.
## The short version
Claude Opus 4.8 is an incremental but real upgrade: the same price as Opus 4.7, a million-token context, and a deliberate focus on honesty that matters most for agentic coding. If you are building agents that edit code or reason over large context, Claude Opus 4.8 is the current top of Anthropic's line, and you can reach it through an OpenAI-compatible API alongside every other model you call.
## Further reading
* [reapi.ai/docs/claude-opus-4-8](/docs/claude-opus-4-8) — endpoint reference and SDK examples.
* [What is reAPI?](/blog/what-is-reapi) — the unified API behind one key for chat, image, and video.
* [reapi.ai/models](/models) — browse the full model catalog.
## References
1. Anthropic. *Introducing Claude Opus 4.8.* Retrieved May 2026 from [anthropic.com/news/claude-opus-4-8](https://www.anthropic.com/news/claude-opus-4-8)
2. Anthropic. *Models overview — context, output, modality, and capabilities.* Retrieved May 2026 from [platform.claude.com/docs/en/about-claude/models/overview](https://platform.claude.com/docs/en/about-claude/models/overview)
3. Anthropic. *Pricing — token, cache, and batch rates.* Retrieved May 2026 from [platform.claude.com/docs/en/about-claude/pricing](https://platform.claude.com/docs/en/about-claude/pricing)
---
# What Is reAPI? Models, Pricing, and How to Use It in 2026 (https://reapi.ai/blog/what-is-reapi)
**reAPI** is one API for generative AI across every modality. A single key and base URL reach 200+ image, video, audio, and chat models, the chat side speaks the OpenAI format, and every model is priced 20-50% below the providers' official rates. Instead of opening an account with each vendor, wiring up four billing systems, and maintaining four SDKs, you integrate once.
This guide covers what reAPI is, the models it carries, what it costs, and how to make your first call. Every endpoint and code sample below is the real API; you can paste them and run them after creating a key.
## What is reAPI?
reAPI is a unified API gateway for AI models. One endpoint, `https://reapi.ai/api/v1`, serves text and reasoning models, image generation, video generation, and audio, billed from a single pay-as-you-go credit balance.
Two request styles sit behind that one base URL:
* **Chat models are synchronous and OpenAI-compatible.** Point an OpenAI client at reAPI's base URL and your existing code works.
* **Image, video, and audio are asynchronous.** You submit a job, get a task id back immediately, and poll until it finishes.
The result is one integration that covers GPT-5, Claude Opus 4.8, and Gemini for text, plus Veo 3.1, Seedance 2.0, and GPT-Image-2 for media, without four separate vendor relationships.
## Core features and benefits
* **One key for 200+ models.** Switch models by changing a string, not by onboarding a new vendor.
* **OpenAI-compatible chat.** A drop-in for existing OpenAI code; change the base URL and key.
* **20-50% below official rates.** The same frontier models, priced under what the providers charge directly.
* **Every modality.** Text, image, video, and audio under one balance.
* **Pay-as-you-go credits.** No subscription, no minimum spend; 1 credit equals $0.001, and new accounts start with free credits.
* **Automatic refunds on failure.** A failed generation refunds its credits the moment the job ends in failure, with no support ticket.
## How reAPI works
Authentication is a Bearer token on every request:
```http
Authorization: Bearer rk_live_xxxxxxxxxxxx
```
Create a key in the reAPI dashboard under API Keys. The key works across chat and media, so one credential covers the whole platform.
Chat requests return in one round trip, the same as calling OpenAI. Media requests are asynchronous because image and video take seconds to minutes to render:
```
POST /api/v1/images/generations → { task_id, status: "processing" }
│
▼
GET /api/v1/tasks/{task_id} → processing → completed / failed
│
▼
output.image_urls
```
reAPI calls the upstream model exactly once per request you send. There is no idempotency deduplication: every successful POST creates a task and charges credits, so a request is never silently skipped, and a failed task is refunded automatically.
## Call a chat model in the OpenAI format
Because the chat surface is OpenAI-compatible, the official OpenAI SDK works with two changes, the base URL and the key:
```python
from openai import OpenAI
client = OpenAI(
base_url="https://reapi.ai/api/v1",
api_key="rk_live_xxx",
)
resp = client.chat.completions.create(
model="gpt-5.5",
messages=[
{"role": "user", "content": "Tell me what reAPI does in one sentence."},
],
)
print(resp.choices[0].message.content)
```
Swap `model` for `claude-opus-4-8` or a Gemini model and the rest of the call is identical. The same pattern works in the Node SDK and any OpenAI-compatible tooling.
## Generate an image
Submit the job, then poll the task:
```bash
curl https://reapi.ai/api/v1/images/generations \
-H "Authorization: Bearer rk_live_xxx" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-image-2",
"prompt": "a cute red panda eating bamboo, photorealistic",
"size": "1:1"
}'
```
The response comes back immediately with a task id and `status: "processing"`. Poll until it is done:
```bash
curl https://reapi.ai/api/v1/tasks/task_018f5a3a1b6e7d9f8c2b4d6e8f0a2c4e \
-H "Authorization: Bearer rk_live_xxx"
```
When `status` is `completed`, the image is at `output.image_urls[0]`, rehosted on reAPI's CDN. Poll about once every one to two seconds; image tasks usually finish within seconds.
## Generate a video
Video uses the same submit-then-poll pattern, on the videos endpoint:
```bash
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer rk_live_xxx" \
-H "Content-Type: application/json" \
-d '{
"model": "doubao-seedance-2.0",
"prompt": "a drone shot flying over a coastal city at sunset",
"resolution": "720p"
}'
```
Video runs longer than image, so pace polling accordingly; the task reaches `completed` with the result at `output.video_urls[0]`. Per-model request fields are documented on each model's page.
## What reAPI costs
reAPI is pay-as-you-go with no subscription and no minimum spend. Credits are the unit: 1 credit equals $0.001, and you spend them per call. New accounts start with free credits, so the first request costs nothing.
Media is billed at a flat rate per output, which means a render costs the same whether the GPU was warm or cold:
| Model | Type | From |
| ------------ | ----- | ------------------- |
| GPT-Image-2 | Image | $0.0066 / image |
| Seedance 2.0 | Video | $0.0506 / video |
| Veo 3.1 Fast | Video | $0.207 / generation |
Chat models bill per token at the same 20-50% discount below official rates. A failed task is refunded automatically and in full, so you only ever pay for output you actually receive.
## Reliability and billing, in plain terms
A few platform rules are worth knowing before you build on reAPI.
Billing is one credit equals $0.001, deducted per call from a single balance, with no monthly fee and no minimum spend. A failed task is refunded automatically and atomically the moment the job ends in failure, before any poll sees the failed status, and the refund happens once and only once.
reAPI calls the upstream model exactly once for every request you send, with no idempotency deduplication. That keeps billing honest: you are charged for the calls you make, never silently double-charged and never silently skipped. If you need retry safety, send each payload once on your side, or lean on the automatic refund and resubmit.
Requests are rate-limited per user at five per second, polling included, so pace polling to about once every one to two seconds. Generated files are rehosted on reAPI's CDN; copy anything you need long-term into your own storage once a task completes.
## reAPI vs calling each provider direct
Going direct means one account, key, SDK, and invoice per vendor. reAPI collapses that into one of each, at a lower rate.
| | Calling providers direct | reAPI |
| ----------------- | -------------------------- | ---------------------- |
| Accounts and keys | One per vendor | One |
| Billing | One invoice per vendor | One balance |
| Pricing | Official rates | 20-50% below official |
| Switching models | New integration per vendor | Change a string |
| Failed calls | Handled per vendor | Refunded automatically |
## Models on reAPI
The catalog spans every modality. A sample of what is live:
* **Chat and reasoning:** GPT-5.5, GPT-5.4, Claude Opus 4.8, Claude Sonnet 4.6, Gemini.
* **Image:** GPT-Image-2, Gemini 3 Pro Image, Imagen 4, Seedream 5.0.
* **Video:** Veo 3.1, Seedance 2.0, Wan 2.7, Kling, HappyHorse 1.0, PixVerse V6, Vidu Q3.
* **Audio and music:** Mureka V9, plus voice tools for cleanup, separation, and conversion.
Browse the full set on the [reapi.ai/models](/models) directory, each with its request schema and pricing.
## Who uses reAPI
The platform fits any team that touches more than one model:
* **SaaS products** adding AI features without standing up a vendor integration per capability.
* **Content and marketing teams** generating images and video at a flat per-output cost.
* **Agencies** billing client work against one predictable balance.
* **Researchers** A/B testing models by changing a string instead of an account.
## FAQ
### Is reAPI OpenAI-compatible?
Yes, for chat. Point an OpenAI client at `https://reapi.ai/api/v1` with your reAPI key and existing code works. Image and video use reAPI's asynchronous task endpoints under the same base URL and key.
### How much does reAPI cost?
It is pay-as-you-go with no subscription. Credits are 1 credit = $0.001, models run 20-50% below the providers' official rates, and new accounts start with free credits. Media is flat per output, for example GPT-Image-2 from $0.0066 per image.
### What models does reAPI support?
200+ models across chat, image, video, and audio, including GPT-5, Claude Opus 4.8, Gemini, Veo 3.1, Seedance 2.0, Wan 2.7, and GPT-Image-2. The full list is on [reapi.ai/models](/models).
### Do I need separate keys for chat and media?
No. One reAPI key works across chat, image, video, and audio, billed from a single credit balance.
### What happens if a generation fails?
The credits are refunded automatically, in full, the moment the task ends in failure. The refund is one-shot and happens before any poll observes the failure, so you never pay for a failed render.
### Does reAPI deduplicate repeated requests?
No. Every successful POST to a generation endpoint creates a new task and charges credits, by design, so no upstream call is ever silently skipped. Handle retry safety on your side.
### How do I get started?
Create an account at [reapi.ai](https://reapi.ai), make a key under API Keys, and send the chat or image request above. The first call runs on your free starting credits.
## Further reading
* [reapi.ai/docs/api](/docs/api) — full API conventions and the task reference.
* [What can reAPI do for you?](/blog/what-can-reapi-do) — use cases across image, video, and LLMs.
* [Best fal.ai alternatives](/blog/best-fal-ai-alternatives) — how reAPI compares to a media-only API.
## Getting started with reAPI
reAPI exists to remove the busywork between you and a model: one key, one balance, and one base URL for 200+ image, video, audio, and chat models, priced below what the providers charge directly. Create a key, point your OpenAI client at `https://reapi.ai/api/v1`, and the first request runs on free credits. That is the whole of what it takes to start using reAPI.
---
# What Is Seedance 2.0 and How to Use It (2026 Guide) (https://reapi.ai/blog/what-is-seedance-2-0-and-how-to-use-it)
Seedance 2.0 is ByteDance's flagship AI video model: text, images, audio, or clips in, up to 15 seconds of 480p-to-4K video with synchronized sound out\[1]. Since its February 12, 2026 launch\[2] it has topped the major blind-test leaderboards\[3]\[4], and "how do I actually use this" became one of the most-searched questions in AI video.
Fair question, because the access story is genuinely confusing. The model goes by three names across ByteDance's own properties, there is no single "seedance.com," and the top search results for the name are squatter sites that admit in their own terms of service they have nothing to do with ByteDance. This guide sorts it out: what the model is, which sites are real, every place you can use Seedance 2.0 today, and what calling the API looks like.
## TL;DR
* **Seedance 2.0 is ByteDance's video generation model**: 480p to 4K output, 4 to 15 seconds per clip, references of up to 9 images plus 3 videos plus 3 audio clips, with audio generated alongside the video\[1].
* **It leads the leaderboards**: #1 on Artificial Analysis text-to-video (Elo 1,222, the largest lead on the board) and #1 on LMArena image-to-video (1,474 on \~82,000 votes)\[3]\[4].
* **There is no single official website.** ByteDance splits it three ways: seed.bytedance.com for model info, Dreamina (international) and Jimeng (China) for consumer creation, BytePlus and Volcano Engine for the enterprise API\[2].
* **The top search results are not ByteDance.** Sites like seedance2.ai and seedance.ai sell subscriptions while disclaiming any affiliation in their own fine print. Check before you pay.
* **You cannot run it locally, and ChatGPT does not have it.** The weights are closed, and OpenAI has never offered Seedance\[5].
* **API access runs from $0.0400/s** on reAPI's Seedance 2.0 lineup, with per-second billing and no subscription\[6].
## What Seedance 2.0 actually is
Seedance 2.0 is the second-generation video model from ByteDance's Seed research group, released on February 12, 2026\[2]. The pitch that separates it from most rivals is multimodal referencing: alongside a text prompt you can attach up to 9 images, 3 video clips, and 3 audio tracks in one generation, and the model composes them into a single scene\[1]. That is how creators keep the same character across a whole series of clips, clone a camera move from a reference video, or drive a performance from a song.
The published spec sheet, per ByteDance's own model documentation: output from 480p up to 4K (the 4K tier renders in 10-bit H.265), durations of 4 to 15 seconds at 24fps, and audio generated jointly with the picture rather than bolted on afterward\[1]. Lighter Fast and Mini variants trade the top resolutions for lower cost.
The blind-test numbers back the hype up. On Artificial Analysis, Seedance 2.0 ranks #1 in text-to-video at Elo 1,222, and the 71-point gap to second place is the largest lead anywhere on that board; it holds #1 in image-to-video too\[3]. On LMArena it is #1 in image-to-video at 1,474 across roughly 82,000 votes, ahead of Veo 3.1 at 1,391\[4]. Five months after launch, it is still the model everyone else is measured against, which is also why ByteDance has already announced a successor, Seedance 2.5, for July.
One asterisk on availability: after copyright disputes with Hollywood studios in March 2026, ByteDance paused the global rollout and resumed it everywhere except the United States\[7]. US access remains restricted on ByteDance's own consumer surfaces; third-party platforms have varying policies.
## The real official websites (and the fake ones)
The single most-searched Seedance question after "what is it" is some version of "what is the real official website." Here is the verified answer.
ByteDance splits the model across three official doors, and the model info page's own buttons prove the mapping: the "Try Now" button on seed.bytedance.com routes to Dreamina, and "Get API" routes to BytePlus\[2].
| You want to | Official destination |
| ------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| Read about the model | seed.bytedance.com (ByteDance Seed research site)\[2] |
| Make videos as a consumer | Dreamina (dreamina.capcut.com, international) or Jimeng (China)\[8] |
| Call it programmatically | BytePlus ModelArk (global) or Volcano Engine (China)\[1]\[9] |
Adding to the confusion, ByteDance itself uses three names for the same model: Volcano Engine calls it Doubao Seedance 2.0, while BytePlus and Dreamina call it Dreamina Seedance 2.0. Same model, different storefronts.
Now the fake ones. When I checked in early July 2026, the #1 organic result for "seedance," "seedance 2.0," and "seedance ai" was in every case a lookalike domain, with ByteDance's own pages ranking second to fifth. A few receipts from the squatters' own pages: seedance2.ai and seedance.ai both state in their terms that they are "not affiliated with ByteDance" while charging $15 to $100 a month; one operator runs multiple domains with word-for-word identical pricing copy; another is registered to an anonymous Wyoming shell company; one left "\[Your Country or State]" template placeholders in its legal page. None of these sites is ByteDance, and paying them buys you, at best, resold generations at a markup.
The safe rule: if the domain is not bytedance.com, capcut.com, byteplus.com, volcengine.com, or a platform you already know by reputation, assume it is not official.
## Everywhere you can use Seedance 2.0 today
I verified each of these on July 3, 2026 by visiting the platform directly. Access modes vary a lot, and two popular assumptions are wrong: OpenArt does not carry the 2.0 model (its newest Seedance is 1.5 Pro), and ChatGPT has never offered Seedance at all\[5].
| Platform | Seedance 2.0? | Access model |
| ----------------------------------------------------- | ------------------------------------------------ | ----------------------------------------------------------------------------------------------------- |
| Dreamina\[8] | Yes (2.0 + Fast) | Free login + credit system |
| Jimeng (China) | Yes | Membership + credits |
| BytePlus ModelArk\[9] | Yes (2.0 / Fast / Mini) | Enterprise API, token billing |
| Higgsfield\[10] | Yes (2.0 + Fast) | Subscription; the $19 Starter tier does NOT include 2.0, full access starts at Plus (\~$47/mo annual) |
| Krea\[11] | Yes (2.0 / Fast / Mini) | Subscription compute units; Basic $9/mo covers roughly 20 videos |
| Pollo.ai\[12] | Yes | Subscription from $15/mo (Lite, watermarked) |
| fal.ai\[13] | Yes | API, per-second: $0.30/s at 720p, $0.68/s at 1080p |
| Replicate\[14] | Yes (official listing) | API, per-second: $0.18/s at 720p, $0.45/s at 1080p, $1/s at 4K |
| reAPI\[6] | Yes (2.0 / Fast / Mini) | API, per-second from $0.0400/s, plus Mini from $0.03/s; no subscription |
| OpenArt | No (tops out at Seedance 1.5 Pro) | — |
| ChatGPT | Never had it\[5] | — |
Two other questions from the search data worth answering flat: you cannot install Seedance 2.0 locally, because ByteDance has not released the weights; anything labeled "Seedance local install" is a different model wearing the name. And the GPT Store listings named "Seedance" are user-made wrappers, not the model. OpenAI is in fact winding its own video product down; Sora's web app closed in April 2026 and its API retires in September\[5].
## How to use Seedance 2.0, both routes
**The consumer route.** Sign into Dreamina with a free account, pick Seedance 2.0 or 2.0 Fast in the video tool, write a prompt, optionally attach reference images, and spend credits per generation\[8]. This is the right door for trying the model, with two caveats: daily free credits are modest, and Dreamina blocks real-person faces on 2.0 uploads.
**The API route.** For anything repeatable, whether a content pipeline, an app feature, or bulk e-commerce clips, you want per-second API billing rather than subscription credits. On reAPI the call is one POST:
```bash
curl -X POST https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "doubao-seedance-2.0",
"prompt": "A kitten yawning at the camera",
"size": "16:9",
"resolution": "720p",
"duration": 5
}'
```
The request returns a task id; poll `GET /api/v1/tasks/:id` until the clip is ready. Reference images, video clips, and audio tracks attach as public URLs, up to the model's 9 image / 3 video / 3 audio caps\[1], and reference-mode generations actually bill lower than pure text-to-video. Full parameter docs live at [reapi.ai/docs/seedance-2-0](/docs/seedance-2-0).
Which route is yours: if you are exploring what the model can do, Dreamina costs nothing to try. The moment you generate on a schedule or inside a product, per-second API pricing beats subscription math, and switching tiers (Standard, Fast, Mini) is a one-word change to the model string.
## What Seedance 2.0 costs, in one paragraph
Official API pricing is token-based and lands around $3.89 for a 5-second 4K clip on BytePlus\[9]. Per-second platforms are easier to reason about: Replicate charges $0.18/s at 720p\[14], fal.ai $0.30/s\[13], and reAPI runs Seedance 2.0 from $0.0400/s to $0.4048/s depending on tier, resolution, and reference mode, with the Mini tier from $0.03/s\[6]. Subscriptions (Higgsfield, Krea, Pollo) make sense mainly if you also use those platforms' editors. The [live pricing table](/models/seedance-2-0#pricing) has current per-second rates; a full platform-by-platform cost breakdown is its own article.
## FAQ
### What is Seedance 2.0 in one sentence?
ByteDance's flagship video generation model: prompts plus optional image, video, and audio references in; up to 15 seconds of 480p-to-4K video with synchronized audio out\[1].
### What is the real Seedance 2.0 official website?
There isn't one single site. Model information lives at seed.bytedance.com, consumer creation at Dreamina (dreamina.capcut.com) internationally and Jimeng in China, and the official API at BytePlus and Volcano Engine\[2]. Domains built around the "seedance" name, like seedance2.ai and seedance.ai, are unaffiliated resellers by their own terms of service.
### Is Seedance 2.0 free to use?
Trying it is free: Dreamina gives logged-in users daily credits\[8], and several API platforms include signup credit. Sustained free usage is not a thing; the "unlimited Seedance" annual deals some platforms advertised have been quietly withdrawn or capped.
### Can I install Seedance 2.0 locally?
No. The weights are closed and ByteDance offers no download. Local "Seedance" installs are other open models relabeled. If you need self-hosted video generation, look at open-weight models; if you need Seedance itself, the API is the only programmatic route.
### Does ChatGPT have Seedance 2.0?
No, and it never has\[5]. GPT Store items named "Seedance" are user-created wrappers. OpenAI's own video model, Sora, is being retired through 2026.
### Why can't I use Seedance 2.0 in the US?
After copyright disputes with studios, ByteDance paused and then resumed the global rollout everywhere except the United States\[7]. Some third-party platforms serve US users anyway under their own arrangements; ByteDance's consumer apps do not.
### How long can Seedance 2.0 videos be?
4 to 15 seconds per generation\[1]. The announced successor, Seedance 2.5, claims 30-second single-pass generations; see our [Seedance 2.5 pre-launch guide](/blog/seedance-2-5-what-we-know-2026).
### When did Seedance 2.0 come out?
February 12, 2026, announced by ByteDance Seed\[2]. Global consumer rollout followed in late March, and API access opened in mid-April 2026.
## Picking your door into Seedance 2.0
Sorted by intent: read about the model at ByteDance Seed, play with it free on Dreamina, and build on it through a per-second API. Skip anything with "seedance" in the domain that is not ByteDance's, skip the GPT Store clones, and do not wait for a local version that is not coming. If you are heading down the API route, [Seedance 2.0 on reAPI](/models/seedance-2-0) runs the full Standard, Fast, and Mini lineup from $0.03/s with signup credits to test on, so the distance from reading this to your first generated clip with Seedance 2.0 is about one curl command.
## References
1. Volcano Engine (ByteDance). *Doubao Seedance 2.0 — model specifications (resolutions, durations, reference limits).* Retrieved July 2026 from [volcengine.com/docs/82379/1330310](https://www.volcengine.com/docs/82379/1330310)
2. ByteDance Seed. *Seedance 2.0 — official model page and launch announcement.* Retrieved July 2026 from [seed.bytedance.com/en/seedance2\_0](https://seed.bytedance.com/en/seedance2_0)
3. Artificial Analysis. *Text to Video and Image to Video Leaderboards.* Retrieved July 2026 from [artificialanalysis.ai/video/leaderboard/text-to-video](https://artificialanalysis.ai/video/leaderboard/text-to-video)
4. LMArena. *Image to Video Leaderboard.* Retrieved July 2026 from [arena.ai/leaderboard/image-to-video](https://arena.ai/leaderboard/image-to-video)
5. OpenAI. *Sora sunset notice (help center).* Retrieved July 2026 from [help.openai.com/en/articles/20001152](https://help.openai.com/en/articles/20001152)
6. reAPI. *Seedance 2.0 — model page and live pricing.* Retrieved July 2026 from [reapi.ai/models/seedance-2-0](/models/seedance-2-0)
7. Reuters. *ByteDance suspends launch of video AI model after copyright disputes.* Retrieved July 2026 from [reuters.com/technology/bytedance-suspends-launch-video-ai-model](https://www.reuters.com/technology/bytedance-suspends-launch-video-ai-model-after-copyright-disputes-information-2026-03-14/)
8. Dreamina (CapCut). *AI video generation — Seedance model picker.* Retrieved July 2026 from [dreamina.capcut.com/ai-tool/home](https://dreamina.capcut.com/ai-tool/home?type=video)
9. BytePlus. *ModelArk — Seedance 2.0 documentation and pricing.* Retrieved July 2026 from [docs.byteplus.com/en/docs/ModelArk/1330310](https://docs.byteplus.com/en/docs/ModelArk/1330310)
10. Higgsfield. *Pricing.* Retrieved July 2026 from higgsfield.ai/pricing
11. Krea. *Video generation.* Retrieved July 2026 from krea.ai/video
12. Pollo.ai. *Seedance 2.0.* Retrieved July 2026 from pollo.ai/m/seedance-2-0
13. fal.ai. *Seedance 2.0 — Text to Video.* Retrieved July 2026 from fal.ai/models/bytedance/seedance-2.0/text-to-video
14. Replicate. *bytedance/seedance-2.0.* Retrieved July 2026 from replicate.com/bytedance/seedance-2.0
### Further reading
* reAPI. *Seedance 2.5: What We Know Before the Public Launch.* [reapi.ai/blog/seedance-2-5-what-we-know-2026](/blog/seedance-2-5-what-we-know-2026)
* reAPI. *Veo 3.1 vs Seedance 2.0.* [reapi.ai/blog/veo-3-1-vs-seedance-2-0-2026](/blog/veo-3-1-vs-seedance-2-0-2026)
* reAPI. *Seedance 2.0 API documentation.* [reapi.ai/docs/seedance-2-0](/docs/seedance-2-0)
---
# Where to Use Seedance 2.0: Consumer Apps vs the API (https://reapi.ai/blog/where-to-use-seedance-2-0)
Where can I use Seedance 2.0 is a different question from what Seedance 2.0 is, and it has a more annoying answer: it depends on whether you want to click a button or write a request.
Both routes exist. They cost differently, they cap differently, and the one that fits depends on whether you are making one video or a thousand.
## TL;DR
* **Consumer apps** put Seedance behind a subscription with credits. Fastest to try, hardest to forecast.
* **The API route** bills per second of output with no subscription. On reAPI that is $0.095/s at 480p and $0.205/s at 720p from a text prompt.
* **Uploading a source video drops the rate** to $0.058/s and $0.125/s respectively.
* **Duration is clamped to 4–15 seconds** per call.
* **Nano Banana runs the same split**: consumer surfaces inside Google products, or the API for volume.
* **Pick by volume**, not by preference. Under a dozen clips a month, an app is fine. Above that, credits become the constraint.
## The two routes

**Consumer apps** wrap the model in an editor. You sign in, type a prompt, get a video, and the cost is a monthly plan that converts into credits or points. Good for trying it, and for one-off work where the interface saves you more time than the pricing costs you.
**The API** gives you the model and nothing else. No editor, no timeline, no export presets. You send a request and poll for a URL. Billing is per second of generated output, so a 5-second clip costs five times the per-second rate and nothing else happens to your account.
The pattern that catches people out is starting on the app route because it is easier, then hitting a credit wall mid-project and discovering that the per-clip cost was never visible.
## What the API route costs
Rates below are per **second of output video**, so multiply by clip length\[1].
| Resolution | From a text prompt | With an uploaded video |
| ---------- | ------------------ | ---------------------- |
| 480p | $0.095 | **$0.058** |
| 720p | $0.205 | **$0.125** |
| 1080p | $0.510 | **$0.310** |
| 4K | $1.040 | **$0.640** |
A 5-second 720p clip is **$1.03**. A 10-second one is **$2.05**. There is no subscription underneath those numbers and no credit conversion in between.
One detail worth knowing before you model costs: the cheaper column applies **only when the request carries an uploaded source video**. Images and first/last-frame references stay on the base rate, which is a common budgeting error. We covered the full billing model in [Seedance 2.0 cost per second](/blog/seedance-2-0-cost-per-second).
Duration is clamped between **4 and 15 seconds** per call, so a 3-second request bills as 4, and longer pieces are assembled from multiple generations.
## Using it
Submit returns a `task_id`; poll until the video is ready.
```bash
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "doubao-seedance-2.0",
"prompt": "a slow dolly through a neon-lit alley after rain, cinematic",
"resolution": "720p",
"duration": 5
}'
```
Two platform rules. Media inputs are **public http(s) URLs only**, so reference material has to be hosted before the call, with no base64 accepted. And live rates are on [reapi.ai/models/seedance-2-0](/models/seedance-2-0), which is the canonical source rather than any figure in an article.
Full request and response shapes are in the [reapi.ai/docs/seedance-2-0](/docs/seedance-2-0) reference.
## Where Nano Banana fits
Nano Banana follows the same split. Google surfaces it inside consumer products, and the model is also available as an API for volume work\[2].
On the API side the tiers matter more than with video, because they are priced differently and capped differently:
| Tier | Resolution ceiling | Per image on reAPI |
| ------------------ | ------------------ | ------------------ |
| Nano Banana 2 Lite | 1K | **$0.020** |
| Nano Banana Pro | up to 4K | **$0.042** |
Lite is the draft tier: roughly four seconds per image, ten aspect ratios, 1K ceiling. Pro is the only tier Google rates high on reasoning, which matters when a composition has to be factually correct rather than merely attractive\[2].
The comparison in full is in [Nano Banana Pro vs Nano Banana 2](/blog/nano-banana-pro-vs-nano-banana-2).
## Choosing by volume
**Under a dozen clips a month**: use a consumer app. The editor is worth more than the pricing difference, and you will not hit a credit wall.
**Dozens to hundreds**: the API. At $1.03 for a 5-second 720p clip, a hundred clips is about $103 with no plan underneath it, and the number is knowable before you start.
**Anything embedded in a product**: the API, necessarily. A consumer app cannot generate video inside software you ship to other people.
**Iterating on look and motion**: draft at 480p on the fast tier, then render the approved shot once at the resolution you actually deliver. The 480p-to-4K spread is roughly 11x.
## FAQ
### Where can I use Seedance 2.0?
Through consumer apps that wrap it in an editor, or through an API that bills per second of output with no subscription. The API route is at [reapi.ai/models/seedance-2-0](/models/seedance-2-0).
### Do I need a subscription to use Seedance 2.0?
Not on the API route. Billing is per second of generated video, so there is no plan and no credit conversion.
### How much does a Seedance 2.0 video cost?
A 5-second 720p clip from a text prompt is about $1.03; the same clip driven by an uploaded video is about $0.63\[1].
### How long can a single Seedance 2.0 generation be?
Between 4 and 15 seconds. Shorter requests bill as 4 seconds, and longer pieces are assembled from several generations.
### Where can I use Nano Banana?
Inside Google's consumer surfaces, or through an API for volume work. On reAPI the Lite tier is $0.020 per image and Pro is $0.042\[2].
### Can I use these models inside my own product?
Only through an API. Consumer apps cannot generate on behalf of your users.
### Why is my image-to-video generation billed at the higher rate?
Because the cheaper rate applies only when an uploaded source **video** is present. Image and first/last-frame references stay on the base rate\[1].
### Do these accept base64 image input?
No. Every model on the platform takes public http(s) URLs only for media inputs.
## Picking the entry point that matches the job
The honest answer to where can I use Seedance 2.0 is that there are two doors and they are built for different people. An app sells you an interface and charges for a plan. An API sells you the model and charges for seconds.
The decision is volume, not preference. One video a week and the interface is worth paying for. A hundred videos a month, or any video generated inside a product you ship, and the per-second arithmetic is the only version of the question that has an answer you can plan around.
## References
1. reAPI. *Seedance 2.0 model page — live per-second rates by resolution, tier, and input mode.* [reapi.ai/models/seedance-2-0](/models/seedance-2-0)
2. reAPI. *Nano Banana Pro vs Nano Banana 2 — tiers, resolution ceilings, and per-image rates.* [reapi.ai/blog/nano-banana-pro-vs-nano-banana-2](/blog/nano-banana-pro-vs-nano-banana-2)
### Further reading
* reAPI. *Seedance 2.0 cost per second.* [reapi.ai/blog/seedance-2-0-cost-per-second](/blog/seedance-2-0-cost-per-second)
* reAPI. *How to use Nano Banana 2 Lite.* [reapi.ai/blog/how-to-use-nano-banana-2-lite](/blog/how-to-use-nano-banana-2-lite)
* reAPI. *Model catalog.* [reapi.ai/models](/models)
---
# 10 AI Cost Optimization Strategies for 2026 (https://reapi.ai/blog/cost-optimization-strategies)
Your AI bill usually doesn't explode all at once. It creeps. A chat feature launches with a premium model because shipping mattered more than tuning. A media workflow retries jobs after timeouts. Product asks for region-specific handling. One team batches nothing because "**real time**" sounds safer, while another burns budget on duplicate generations. A month later, finance has questions and engineering has very few clean answers.
That's the moment teams typically start searching for cost optimization strategies. The problem is that generic advice rarely helps when you're running multimodal workloads across chat, image, video, and music APIs. The useful work happens lower in the stack, inside routing rules, retry behavior, prompt design, queue topology, spend guardrails, and model policy.
The good news is that the biggest wins usually don't require a rewrite. A combination of five core tactics, model selection, prompt optimization, caching, API aggregation, and batching, can cut total AI API spend by [40% to 70% according to Crazy Router's 2026 optimization guide](https://crazyrouter.com/en/blog/ai-api-cost-optimization-complete-guide-2026). In practice, teams get there by tightening architecture and governance together, not by treating cost as a finance-only problem.
This guide is written from the engineering side. It focuses on controls you can implement directly, with reAPI as a concrete example of how a unified AI API layer makes these strategies easier to enforce across providers and teams.
## Table of Contents
- [1. Intelligent Request Routing and Load Balancing](#1-intelligent-request-routing-and-load-balancing)
- [Route on policy, not instinct](#route-on-policy-not-instinct)
- [2. Request Deduplication and Idempotent Retries](#2-request-deduplication-and-idempotent-retries)
- [3. Per-Key and Team-Level Spend Caps and Alerts](#3-per-key-and-team-level-spend-caps-and-alerts)
- [Build hard stops before you need them](#build-hard-stops-before-you-need-them)
- [4. Model-Specific Cost Optimization and Selection Strategy](#4-model-specific-cost-optimization-and-selection-strategy)
- [Match quality tier to business value](#match-quality-tier-to-business-value)
- [5. Batch Processing and Request Aggregation](#5-batch-processing-and-request-aggregation)
- [Move non-urgent work off the hot path](#move-non-urgent-work-off-the-hot-path)
- [6. Geographic Region Pinning and Data Locality Optimization](#6-geographic-region-pinning-and-data-locality-optimization)
- [Latency problems often become cost problems](#latency-problems-often-become-cost-problems)
- [7. Model Caching and Response Reuse Strategy](#7-model-caching-and-response-reuse-strategy)
- [8. Async Processing and Queue-Based Architecture](#8-async-processing-and-queue-based-architecture)
- [Design the queue around cost decisions](#design-the-queue-around-cost-decisions)
- [9. Provider Negotiation and Volume Discounts](#9-provider-negotiation-and-volume-discounts)
- [Consolidation changes the negotiation](#consolidation-changes-the-negotiation)
- [10. Input Optimization and Efficient Prompt Engineering](#10-input-optimization-and-efficient-prompt-engineering)
- [10-Strategy Cost Optimization Comparison](#10-strategy-cost-optimization-comparison)
- [Build a Culture of Cost-Aware AI Development](#build-a-culture-of-cost-aware-ai-development)
## 1. Intelligent Request Routing and Load Balancing
Routing is where cost control becomes operational. If every request goes to the same premium model or the same vendor by default, you're paying for convenience, not outcomes. In multimodal systems, that gets expensive fast because availability, latency, and price shift independently across providers.
reAPI makes this practical because one integration can route traffic across providers and model families instead of forcing your app to hardcode a single choice. A video workflow can send jobs across Veo, Seedance, Kling, Runway, and Happyhorse. An image workflow can switch between Flux Pro, GPT Image 2, Seedream, or Qwen Image based on the kind of output you need. Chat traffic can move between Claude, GPT-class models, and Gemini when one route degrades.

### Route on policy, not instinct
The best routing rules start with request classes. Draft image generation, user-generated video, support chat, and executive report writing shouldn't share the same default route. Industry benchmarks show that moving routine tasks from flagship models to lighter tiers within the same family can cut bills by [70% to 85% for those tasks, with the benchmark discussed in oFox's cost reduction guide](https://ofox.ai/blog/how-to-reduce-ai-api-costs-2026/).
A few routing rules matter more than is commonly assumed:
- **Separate draft from final output:** Send exploratory or preview jobs to cheaper models. Reserve premium routes for publish-ready assets.
- **Add fallback chains:** If a preferred provider slows down, route to the next acceptable vendor before clients start retrying.
- **Pin by region:** Keep EU traffic in EU routes and US traffic in US routes when compliance or latency requires it.
- **Reweight often:** Vendor economics change. Your routing policy should too.
> **Practical rule:** Don't optimize for the "best model." Optimize for the cheapest model that reliably clears the quality bar for that request type.
What doesn't work is static routing based on team preference. Engineers remember the model that saved a launch. Finance remembers the invoice that followed.
## 2. Request Deduplication and Idempotent Retries
A timeout during an image or video generation request often means only one thing. Your client gave up waiting. It does not mean the provider stopped processing, and it definitely does not mean the bill disappeared.
That gap creates a common cost leak. The app retries, the backend retries again, and the same user action turns into two or three paid generations. I have seen this happen in mobile clients on weak connections, in webhook consumers after a 5xx, and in queue workers that treat "no response yet" as failure. For high-cost media jobs, a small retry bug can become a real line item on the invoice.
The fix is architectural, not cosmetic. Assign one stable identity to each logical operation and carry it through every retry path. With reAPI, that means pairing your retry logic with idempotency controls and spend visibility from the start. The [reAPI balance and spend controls documentation](https://reapi.ai/docs/api/balance) is useful here because billing protection is easier to enforce when request identity and spend tracking live in the same platform.
A practical pattern looks like this:
- **Generate the request ID at the edge:** Create it in the client, gateway, or first API hop before any provider call is attempted.
- **Reuse the same ID on every retry:** Retries should represent the same job, not a fresh generation attempt.
- **Persist final state:** Store success, failure, and provider job references so later retries can return the known result.
- **Check for duplicates before provider submission:** Suppress the second call in your own system instead of hoping the upstream vendor catches it.
This approach is particularly effective for mobile apps, browser sessions, and webhook-driven backends where dropped connections are routine. It also matters in internal batch systems, where one stuck worker can replay the same payload many times before anyone notices.
There is a trade-off. Deduplication adds state management, key design, and expiration rules. If your request fingerprint is too broad, you can suppress legitimate reruns. If it is too narrow, duplicates slip through and you pay anyway. For reAPI workloads, I usually treat the idempotency key as a business event key, not a raw payload hash. "Generate preview image for asset 482, version 7" is safer than hashing the whole request body and hoping formatting changes do not create accidental misses.
> If your team cannot answer "did we already submit this exact job?" with one query, retries are still a cost risk.
Client-side retry libraries help user experience. They do not protect margins. Billing-safe retries require server-side deduplication, stored job state, and idempotent request handling all the way to the provider boundary.
## 3. Per-Key and Team-Level Spend Caps and Alerts
Every fast-growing AI product eventually discovers the same truth. Visibility without enforcement is just a better way to watch overruns happen. Teams need budgets attached to keys, environments, and business units before a bug, prompt loop, or misconfigured worker turns into a bad week.
This is one of the reasons unified platforms are useful. With reAPI, you can mint separate keys for development, staging, and production, assign limits at the key level, and layer team ceilings on top. That gives engineering managers and platform teams a real control plane instead of a spreadsheet ritual.
### Build hard stops before you need them
Start with loose limits if you have to, but start. A development key should never have the same spending freedom as a production pipeline. A creative generation feature should never consume the same shared pool as customer support chat.
A practical setup usually looks like this:
- **Development keys:** Lower daily limits, because test loops and prompt experiments are unpredictable.
- **Staging keys:** Enough headroom for load tests, but still isolated from production.
- **Production keys:** Higher ceilings, paired with alerting to multiple owners.
- **Team budgets:** Separate caps for product areas so one feature can't starve another.
You can manage and inspect these controls through reAPI's [balance and spend controls documentation](https://reapi.ai/docs/api/balance).
The gap many teams miss is unit-level accountability. Gartner highlights an underserved need for cost visibility tied to features such as cost per generated video, per chat turn, or per song, especially when a single capability can consume [60% to 80% of AI spend](https://www.gartner.com/en/insights/cost-optimization). Aggregate monthly spend won't tell you what to fix first. Per-feature spend usually will.
What doesn't work is a single org-wide budget with one alert email. That's not governance. That's delayed surprise.
## 4. Model-Specific Cost Optimization and Selection Strategy
A team ships a feature with one high-end model behind every request. The demo looks great. Thirty days later, finance asks why a low-risk workflow like FAQ drafting costs almost as much per user as the flagship experience.
That usually happens because model selection was never turned into an engineering policy. It stayed a product preference. Cost control starts when each workload gets a quality tier, an allowed model set, and a fallback path inside the API layer.
### Match quality tier to business value
The practical question is simple. Does better output quality change the business result enough to justify the price? If the answer is no, the expensive model should not be in the default path.
In reAPI, that decision can be enforced across text, image, audio, and video instead of handled ad hoc by each feature team. An e-commerce workflow might use a lower-cost image model for internal merchandising mockups, then switch to a premium route only for launch assets that reach customers. A video product might send first-pass user experiments to a cheaper generation path and reserve the top model for paid exports, where quality affects conversion and refund rates. The win is not standardizing on one model. The win is standardizing the rule for when each model is allowed.
Teams usually save real money without hurting the product by differentiating task requirements. Simple classification, translation, moderation, metadata extraction, and support-answer drafting are strong candidates for lower-cost models. Executive reports, legal review support, homepage creative, and high-visibility customer outputs usually justify a higher tier.
A workable policy often looks like this:
- **Draft tier:** Lowest acceptable cost. Used for internal previews, first passes, and disposable outputs.
- **Standard tier:** Default for production features where users need reliable quality but not the best available model.
- **Premium tier:** Restricted to flows where better quality improves revenue, retention, approval rates, or risk reduction.
The implementation pattern matters as much as the policy. Put the tier in code and config, not in a wiki. Route by use case, attach the expected cost profile, and log which tier each request used. If a team wants premium by default, require them to show why the output changes an important metric.
For video teams, pricing moves fast enough that this review should happen on a schedule, not only during incidents. reAPI's analysis of [the cheapest Seedance 2.0 options in 2026](https://reapi.ai/blog/cheapest-seedance-2-0-2026) shows the kind of model-by-model comparison that platform teams should maintain for their own approved catalog.
One more pattern is worth adopting early. Keep manual model pickers out of the main product unless your users are experts. End users rarely have the context to balance quality, latency, and cost well. The platform should make that choice for them, with exceptions handled through policy rather than guesswork.
## 5. Batch Processing and Request Aggregation
At 2 p.m., every request feels urgent. At 2 a.m., the same workload is usually a queue.
That distinction matters more than many teams admit. I have seen AI platforms pay real-time prices for nightly summaries, catalog enrichment, back-office classification, and media jobs that no user was waiting on. The result is predictable. Higher unit cost, noisier traffic patterns, and more operational work for jobs that could have been scheduled and grouped.

### Move non-urgent work off the hot path
Batching cuts cost in two ways. Some providers price asynchronous batch work below interactive inference. You also reduce the hidden overhead around burst handling, retry storms, and overprovisioned worker capacity.
In a reAPI-style setup, this is straightforward to implement because the control plane already sees traffic across providers and model families. That gives engineering teams one place to tag requests as interactive or deferred, set batch windows, and enforce different retry and timeout policies. The governance benefit is easy to miss. Once background jobs are classified explicitly, teams stop routing them through expensive low-latency paths by default.
Good candidates for batching include:
- **Overnight catalog generation:** Product descriptions, alt text, tagging, and image analysis across large inventories.
- **Scheduled reporting:** Daily summaries, customer health notes, and internal analytics writeups.
- **Background media processing:** User uploads that can complete minutes later without hurting the product experience.
- **Bulk classification:** Moderation backfills, taxonomy assignment, entity extraction, and language detection.
A simple financial model helps here. If a workflow handles 500,000 non-urgent jobs per day, even a modest per-request discount or a small drop in duplicate processing adds up fast over a month. The exact number depends on the provider and payload size, but this is one of the few optimizations that usually lowers both spend and operational noise at the same time.
> Batch the work users will not wait for. Reserve real-time capacity for the requests that change the live product experience.
The implementation detail that separates a real batch system from a fake one is aggregation. Do not just push single requests onto a queue and call it optimized. Group work by model, tenant, or task type so workers can submit larger jobs efficiently, apply consistent rate limits, and retry partial failures without replaying the entire backlog.
One caution: partial batching often disappoints. If the frontend polls every few seconds, workers keep tiny batch sizes, and product teams promise near-real-time delivery anyway, you inherit queue complexity without getting much of the savings. Set clear SLAs for deferred work, return job IDs immediately, and let completion happen through webhooks or event-driven updates instead of constant polling.
## 6. Geographic Region Pinning and Data Locality Optimization
Region pinning usually enters the conversation as a compliance feature. That's correct, but incomplete. It's also a cost control feature because latency problems create retries, long-held connections, and failed jobs that users resubmit.
In reAPI, region pinning across EU, US, and APAC gives you a straightforward way to keep traffic where it belongs. For regulated products, that's table stakes. For everyone else, it's often the simplest way to reduce timeout-driven waste.
### Latency problems often become cost problems
This is especially visible in multimodal apps. A chat call might degrade gracefully under extra latency. A large image upload, audio transcription, or video generation request often won't. If clients give up, workers retry, and users click again, your locality mistake turns into duplicate spend.
A few region rules are worth codifying:
- **Route by user or data origin:** Don't pin everything to one region because it's operationally convenient.
- **Keep storage and inference close:** Uploading in one geography and processing in another creates unnecessary friction.
- **Use region-specific health checks:** Degradation is often local, not global.
- **Fail over intentionally:** Cross-region failover should preserve policy, not bypass it accidentally.
The sequencing matters too. Gamayaa's analysis of AI cost optimization warns against the "optimization sequencing trap" and argues that teams should eliminate idle and retry spend first. Doing the order wrong can lock teams into [20% to 40% higher unit costs than optimized peers](https://gamayaa.com/it-cost-optimization-strategies-2/). Region and retry behavior sit right in that first bucket.
What doesn't work is assuming regional optimization is only legal or procurement work. Engineers usually create the retry patterns that make geography expensive.
## 7. Model Caching and Response Reuse Strategy
A support bot goes live, traffic climbs, and the bill follows the same curve. Then the team inspects request logs and finds the same intents showing up thousands of times: refund policy, password reset steps, plan limits, invoice questions, and the same long system prompt attached to every call. That pattern is common, and it is usually one of the fastest places to cut spend.
Caching works because a meaningful share of AI traffic is repetitive even when the product feels dynamic. In reAPI-style deployments, repetition shows up in three places: long static context, identical requests, and near-duplicate requests that ask for the same answer with slightly different wording. Each one needs a different control.
Start with provider-side prompt caching for large repeated context. If every request carries the same policy block, product catalog, or tool instructions, paying to reprocess that text on every call is avoidable. This is often the lowest-effort win because the application flow stays mostly intact. The trade-off is scope. Provider caching only helps with the repeated portion of the request, and the savings depend on how much of each call is stable.
Then add an application cache for exact matches. This catches requests before they hit the model at all. It is effective for FAQ bots, translation of repeated strings, moderation checks on duplicated content, and template-based generations where deterministic output is acceptable. Teams using reAPI often put this behind a normalized request key that includes model, prompt template version, locale, and any policy flags. Without versioning in the key, the cache saves money and invisibly serves stale answers.
Semantic caching is the next layer. Instead of asking whether two prompts are identical, it asks whether they are similar enough to reuse the same answer. That matters for high-volume support and knowledge workflows where users phrase the same request in different words. A semantic cache adds retrieval cost and a quality risk, so the threshold needs tuning. Set it too loose and users get mismatched answers. Set it too tight and the hit rate stays low enough that the extra infrastructure is hard to justify.

A practical cache stack usually looks like this:
- **Provider cache:** Use for repeated system prompts, reference docs, and other long static context.
- **Application cache:** Use for exact duplicate requests where the same input should return the same output.
- **Semantic cache:** Use for high-volume, meaning-similar queries in support, search, and internal knowledge tools.
- **Object storage cache:** Use for generated images, audio, and other media assets that are requested again later.
The governance part matters as much as the cache itself. Cache keys should include anything that changes the answer: model version, prompt version, user tier, language, policy state, and time sensitivity where relevant. Set TTLs by use case, not by convenience. Product documentation can sit longer than pricing answers or compliance guidance. Track hit rate, stale-response rate, and cache savings per route so the team knows which layer is paying for itself.
The financial impact is usually straightforward. Every cache hit removes a model call or shrinks the token volume of one. In products with repeated prompts, that can move monthly spend fast. In products with highly personalized requests, cache complexity can exceed the savings. The right question is not whether caching is good. It is whether a given route has enough repetition, enough determinism, and enough margin pressure to justify another layer in the stack.
## 8. Async Processing and Queue-Based Architecture
A customer submits 20,000 document translation jobs at 4:45 PM and expects status by morning. If every request goes through the same synchronous path, the platform pays peak pricing for work that had hours of scheduling slack. A queue changes the economics. It lets the system separate urgent work from deferrable work, smooth provider spikes, and route jobs only when the latency requirement justifies the cost.
That control matters more than teams expect. In AI systems, a queue is part of the pricing layer. Once a job is queued, the platform can attach policy before any model call happens: priority, max cost, allowed providers, retry budget, region, and deadline. With reAPI, that means a real-time support answer can go to a low-latency model immediately, while nightly summarization, bulk transcription, or translation batches wait for a cheaper route that still meets the SLA.
### Design the queue around cost decisions
The useful fields are operational and financial at the same time:
- **Request ID and idempotency key:** Prevent duplicate charges when workers retry or clients resubmit jobs.
- **Priority and deadline:** Reserve premium paths for user-facing flows and keep background jobs off them.
- **Provider and model constraints:** Limit expensive models to routes that require them.
- **Retry metadata:** Cap retries by job value, not by habit.
- **Estimated size:** Use token or file-size estimates to decide whether to batch, defer, or split work.
A practical reAPI setup usually has at least two lanes. One serves interactive traffic with tight timeout budgets. The other handles deferred jobs, where the system can wait for lower-cost capacity or use a slower model family. That split is simple to implement and often one of the fastest ways to cut spend without changing product quality.
Streaming still has a place here, especially for long-running text generation where users benefit from partial output. But the bigger cost win usually comes earlier. Teams stop forcing every job through an expensive synchronous path.
The supporting pieces are familiar, but the tuning matters:
- **Webhook result delivery:** Better than polling for long-running image, audio, or document jobs.
- **Dead-letter queues:** Keep malformed payloads and repeated provider failures from draining budget.
- **Visibility timeouts:** Match them to realistic model runtimes so workers do not duplicate expensive calls.
- **Priority routing rules:** Tie premium latency to revenue, user tier, or contractual SLA.
One trade-off is product complexity. Async flows need progress states, cancellation handling, and clearer user messaging. They also require finance-minded defaults. A low-value background job should not retry five times against the most expensive provider. Set retry ceilings by route, define expiration rules, and drop stale work that no longer has business value.
This pattern is especially useful in translation and localization pipelines, where deadlines vary widely across content types. Teams building multilingual apps can combine queued AI jobs with [cost-effective Django translation strategies](https://translatebot.dev/en/blog/cost-of-translation-services/) so only user-visible content stays on the fastest path.
The teams that benefit most are the ones with mixed workloads: some requests need seconds, others can wait minutes or hours. If every request is interactive, a queue adds operational overhead. If even 20 to 30 percent of traffic is deferrable, queue-based scheduling usually pays for itself quickly because it gives the platform room to make cheaper routing decisions before tokens start burning.
## 9. Provider Negotiation and Volume Discounts
By the time a team thinks about negotiation, it usually has already made the hardest part difficult. Spend is fragmented across providers, business units, and model-specific accounts. Nobody has a clean view of total volume, so procurement walks into vendor conversations without bargaining power.
Consolidation fixes that. One of the practical advantages of an aggregator like reAPI is that it centralizes traffic and spend reporting across a broad model catalog. That changes both economics and operations. Smaller teams can benefit from aggregated purchasing power, and larger teams get a cleaner basis for direct negotiation.
### Consolidation changes the negotiation
This strategy matters most after you've removed obvious waste. If you negotiate too early, you can lock in inefficient baselines. Gartner's broader cost-optimization view is useful here because reporting quality determines negotiation quality. If stakeholders can't trust the numbers, contract decisions drift toward guesswork. Better internal reporting habits matter long before procurement signs anything, which is why strong teams invest in [improving cloud spend reporting](https://serverscheduler.com/blog/stakeholder-reporting) alongside technical optimization.
A few negotiation rules hold up well in practice:
- **Bring consolidated usage data:** Vendors respond better to actual routed volume than rough forecasts.
- **Separate stable from volatile demand:** Commit only on the baseline you can defend.
- **Ask about aggregator economics:** Sometimes the better rate is through the platform, not around it.
- **Protect flexibility:** Pricing terms that block model switching can erase later savings.
For teams dealing with multilingual or localized products, related operational choices also influence unit economics. This is visible in adjacent areas such as [cost-effective Django translation strategies](https://translatebot.dev/en/blog/cost-of-translation-services/), where workflow design matters as much as the nominal model rate.
What doesn't work is chasing discounts while leaving duplicate retries, poor routing, and bloated prompts untouched. That only makes waste cheaper.
## 10. Input Optimization and Efficient Prompt Engineering
A common failure pattern looks like this. The team picks a reasonably priced model, sets spend caps, and still gets a monthly bill that feels wrong. Then you inspect the traffic and find 600-token system prompts wrapped around tasks that only need a short policy block, plus responses that ramble because nobody set output limits. Cost control often breaks at the prompt layer.
Prompt engineering is a production discipline, not a creative exercise. In cost terms, it is input compression plus output control. Every repeated instruction, oversized example, and loose formatting request increases spend and often hurts reliability at the same time.
The trade-off is real. Shorter prompts are cheaper, but over-compression can remove the context that keeps accuracy stable. The goal is not "smallest possible prompt." The goal is the shortest prompt that still produces the right answer shape on the first try.
Three prompt changes usually pay back fast:
- **Trim system prompts to hard requirements only:** Keep role, policy, and strict constraints. Move commentary and duplicate guidance out.
- **Use structured outputs instead of prose instructions:** A schema usually costs fewer tokens than explaining formatting in paragraphs, and it cuts retry rates when downstream code expects fixed fields.
- **Set explicit length targets:** Ask for five bullets, three labels, or a JSON object with named keys. Do not ask for a "detailed response" unless the product requires one.
At reAPI, this is easy to test in the [prompt playground for multi-model prompt iteration](https://reapi.ai/prompt). Run the same task across routes, compare token counts, and keep the cheapest version that still hits your quality bar. That workflow matters because prompt savings only count if they survive contact with production traffic.
A practical pattern is to separate prompts into layers. Keep a minimal global system prompt. Add task-specific instructions per endpoint. Inject user context only when it changes the answer. Teams that collapse all guidance into one giant prompt usually pay more and debug more.
What fails in practice is adding examples every time output quality drops. One example that matches the target format often helps. Five examples can cost more than switching to a better schema or tightening the task definition. If the model keeps missing, fix the contract first: clarify the task, define the output, and cap the response length.
## 10-Strategy Cost Optimization Comparison
| Strategy | 🔄 Implementation complexity | ⚡ Resource requirements | 📊 Expected outcomes | 💡 Ideal use cases | ⭐ Key advantages |
|---|---:|---|---|---|---|
| Intelligent Request Routing and Load Balancing | Medium, multi-vendor rules & monitoring | High, vendor integrations, real-time metrics, ops | 20–40% cost reduction; ~99.96% availability; lower latency | Multi-vendor deployments; latency-sensitive apps; failover scenarios | Maximizes uptime, reduces per-request cost, avoids vendor lock-in |
| Request Deduplication and Idempotent Retries | Medium‑High, distributed cache & tracking | Medium, distributed cache, request fingerprinting | 10–25% savings from avoided duplicate expensive ops; fewer wasted jobs | Video/image generation; flaky networks; mobile apps | Prevents duplicate charges and reduces waste; improves reliability |
| Per-Key and Team-Level Spend Caps and Alerts | Low, budget rules & alerts | Low, billing controls, alerting hooks | 30–60% reduction by preventing uncontrolled spend; real‑time visibility | Startups, multi-team orgs, cost governance | Granular budget control; prevents cost spikes and enables chargeback |
| Model-Specific Cost Optimization and Selection | Medium, benchmarking & routing logic | Medium, multiple model access, A/B testing | 40–70% savings by right-sizing models; maintains quality for critical tasks | Mixed-quality pipelines; preview vs. final content; cost-sensitive tasks | Significant cost reduction while preserving quality where needed |
| Batch Processing and Request Aggregation | Medium, queueing & batch job infra | Medium, job queues, schedulers, result aggregation | 50–90% cost reduction for batchable workloads; higher throughput | Non-real-time processing, overnight jobs, analytics | Large per-call discounts; efficient high-volume processing |
| Geographic Region Pinning and Data Locality Optimization | Low‑Medium, region routing config | Low, region endpoints, compliance checks | 10–30% gains via reduced retries/latency; compliance assurance | GDPR/CCPA compliance, regional performance optimization | Ensures data residency, lowers latency and retry-related costs |
| Model Caching and Response Reuse Strategy | Medium, cache layers & invalidation | High, CDN/edge, storage, cache coordination | 30–70% savings for cache-friendly workloads; instant responses on hits | Popular prompts, templates, frequently requested content | Eliminates cost on cache hits; large latency and load reductions |
| Async Processing and Queue-Based Architecture | Medium‑High, queues and worker logic | Medium, message queues, workers, DLQs | 25–50% savings via batching/deduplication; improved resiliency | Background jobs, retry-prone tasks, bulk processing | Decouples clients, enables batching/dedup/retries, improves throughput |
| Provider Negotiation and Volume Discounts | Low, contract negotiation & consolidation | Low‑Medium, procurement effort, consolidated billing | 15–40% savings with volume/commitment discounts; predictable pricing | High-volume customers, enterprise procurement, aggregators | Direct negotiated savings; predictable costs and better terms |
| Input Optimization and Efficient Prompt Engineering | Low‑Medium, prompt design & testing | Low, engineering time, A/B testing tools | 10–40% token/cost reduction; fewer retries; improved outputs | Token-billed models, chatbots, content generation | Reduces token costs and retries while improving quality without infra changes |
## Build a Culture of Cost-Aware AI Development
The strongest cost optimization strategies don't live in a spreadsheet and they don't survive on good intentions alone. They live in the product architecture, in the API gateway, in worker policies, in budget controls, and in the model-selection rules your team enforces. That's why engineering has to own a large part of AI cost management. Finance can flag the problem. Procurement can help with contracts. But engineers decide whether a retry duplicates work, whether a task gets batched, whether prompts are bloated, and whether a cheap model is even allowed to compete.
What works is a layered system. Start by eliminating pure waste. Deduplicate retries, fix timeout behavior, and add caching where repetition is obvious. Then route requests by quality tier instead of habit. Then move non-urgent jobs into queues and batches. After that, tighten budgets, alerts, and feature-level visibility so overruns are caught by the people who can fix them. Finally, negotiate from a position of clean data and stable baselines.
reAPI fits well into that stack because it reduces the operational friction of doing the right thing. One key, one dashboard, and one routing layer make it easier to enforce spend caps, compare providers, pin by region, fail over cleanly, and test multiple models without multiplying integration work. That matters because fragmented tooling is one of the main reasons optimization gets postponed. Teams don't avoid cost discipline because they dislike savings. They avoid it because the implementation surface is messy.
There's also a cultural shift that has to happen. Teams need to stop evaluating AI features only on quality and latency. They need to ask what each chat turn, each image generation, each video render, and each music request contributes to the business. If a feature drives value, support it with better routing, better batching, and better contracts. If it doesn't, prune it or lower its service tier. That's not anti-innovation. It's what lets innovation continue.
The most effective teams I've seen treat AI cost the same way they treat reliability. They instrument it early, review it often, and design for failure modes before those failure modes hit production. They know that an expensive model can be the right choice in the right place. They also know that "works" is not the same as "works economically."
If you adopt that mindset, cost optimization stops being a cleanup project. It becomes a product capability. And once that happens, scaling multimodal AI gets a lot less chaotic.
---
reAPI gives engineering teams a practical way to apply these cost optimization strategies without stitching together separate providers, dashboards, and failover logic by hand. If you're shipping chat, image, video, music, or code features and want one API layer for routing, governance, region pinning, and spend control, [reAPI](https://reapi.ai) is a strong place to start.
*Powered by the Outrank tool*