GPT Image 2.5 is live — OpenAI's newest image model, targeted edits that leave the rest of the frame alone
Opus 5.5 Video Generation: Directing Seedance 2.5 Shots
2026/09/29

Opus 5.5 Video Generation: Directing Seedance 2.5 Shots

Opus 5.5 video generation, done right: Claude writes the shot plan, GPT Image 2.5 locks the character, Seedance 2.5 renders it. Prompts, code and real costs.

Opus 5.5 video generation now shows up in YouTube and Brave autocomplete, and the phrase hides a misunderstanding. Claude Opus 5.5 does not render video. It is a language model with a 1M-token context window, served through a chat endpoint[1]. What it does unusually well is the part most AI clips get wrong: deciding who is in the frame, where the camera sits, what happens second by second, and what the audio should be.

The clearest public example is a September 24 breakdown on r/Seedance_AI. The author said they "didn't write a single prompt manually": Opus 5.5 turned one concept and a few style screenshots into character sheets, start frames, camera moves, dialogue timing and sound, and the pixels came from GPT Image 2.5 and Seedance 2.5[2]. This guide takes that pipeline apart, checks each step against the vendors' own documentation, and shows how to run all three models from one API key.

TL;DR

  • Opus 5.5 plans, it does not render. In this workflow it writes every prompt; GPT Image 2.5 draws the stills and Seedance 2.5 makes the moving, talking shot[2].
  • Describe the camera operator, not "realism." The breakdown's key trick is telling the model who holds the camera and where they sit, which produces handheld sway, autofocus hunting and exposure shifts[2].
  • GPT Image 2.5 is the identity anchor. OpenAI says it is "better at preserving the subjects in your reference photos"[3], which is exactly what a character sheet needs.
  • Seedance 2.5 wants structure. ByteDance's guide asks for numbered asset bindings (Image 1, Image 2), gap-free timestamps, and explicit negative controls such as "No subtitles"[4].
  • Cost is dominated by video seconds. Two 10-second 1080p Seedance 2.5 shots cost $9.24 on reAPI; the two stills and the Opus 5.5 planning run add well under a dollar[5][6][7].

Why Opus 5.5 video generation starts with a script

Video models reward specificity and punish ambiguity. ByteDance says so directly in its Seedance 2.5 guide: prompt writing "primarily affects instruction following, material consistency, and generation controllability," and it does not change the model's underlying limits[4]. The same guide warns that too little plot in a time range lets the model improvise, while too much produces excessive cuts or dropped beats[4].

That is a writing problem, and it is where a strong language model earns its cost. A single Seedance 2.5 shot in the Reddit breakdown runs to more than 1,200 words: a reference block, a camera lock, four timed beats with dialogue and emotion cues, a film-look section, an audio section, and a list of negatives[2]. Keeping all of that consistent across several shots by hand is tedious. Asking Opus 5.5 to hold the character bible in context and emit each prompt is not.

Anthropic's own pitch for Opus 5.5 is about exactly this kind of discipline. It says the model "puts the most important information up front" and "follows the writing rules you give it"[8]. For prompt generation, following your template every time matters more than flair.

One commenter on the breakdown called it using a language model to do the video model's job. Another answered the practical point: "these prompts are massive, and it can be tough to hold everything and format it correctly in each prompt"[2]. The same commenter noted the cost of this approach: Claude tends to add "a bunch of gunk" when you ask it for corrections, so edit the template, not the output.

The four-stage pipeline, step by step

StageModelInputOutput
1. DirectorClaude Opus 5.5One concept, style screenshots, your prompt templateCharacter bible, image prompts, timed video prompts
2. Character sheetGPT Image 2.5Photo of the subject + style references + sheet promptOne 16:9 sheet: turnaround, expressions, detail panels
3. Start frameGPT Image 2.5Character sheet + scene promptThe exact first frame of the shot
4. ShotSeedance 2.5Character sheet + start frame + timed prompt10-second clip with dialogue and ambient sound

The breakdown's character sheet prompt is a good template to copy. It asks for a six-angle turnaround row and a row of five expressions plus a detail panel, and states "Identity is locked, no drift between panels," while marking the style images as "STYLE ONLY"[2]. That split matters: a style reference that is not labeled can leak its person into your character.

The start frame is where the look gets baked in. The author put the early-2000s point-and-shoot grade into the image prompt itself, not into a post filter: "35mm film, 28mm wide lens with slight barrel distortion," "lifted blacks," "fine visible film grain"[2]. Seedance 2.5 then treats that frame as the master reference for location, light and grade.

For stages 2 and 3, GPT Image 2.5 comes in two versions. OpenAI positions Flare as "the default choice for most applications" and Sunburst for "premium visual workflows that benefit from tighter control across edits"[3]. A character sheet you will reuse for every shot is a reasonable place to spend the extra $0.002 per image on Sunburst.

The camera-operator trick, and what the Seedance guide adds

The detail commenters singled out was the author's second point: do not ask for a "realistic video." Say who is holding the camera, where they are sitting, and how they are moving[2]. The author reports that one line, "my friend is filming from the back seat," was enough for Opus 5.5 to write micro-shake, delayed framing, autofocus hunting and exposure shifts into the prompt.

This works because Seedance 2.5 understands plain camera language. ByteDance lists shot sizes, moves such as "push in," "orbit," and "handheld shake," angles such as "first-person perspective," and techniques such as one-shot long takes as terms you can write directly[4]. A camera operator with a location and a body gives the model a reason to use them consistently.

Five more rules from the official guide belong in the template you hand to Opus 5.5:

  1. Number and bind every asset. Refer to uploads as Image 1, Image 2 in upload order, and state each binding in the text. ByteDance advises against relying on labels drawn inside the image[4].
  2. Keep timestamps continuous. Write "0-3 seconds... 3-7 seconds..." and avoid gaps; do not use timestamps for rapid repeated actions like "shake your head three times per second"[4].
  3. Say what you do not want in the audio. "No BGM; generate only environmental sounds and action sounds" and "No subtitles" are both supported negatives[4].
  4. Put spoken lines in double quotes. reAPI's Seedance 2.5 reference notes that quoted lines steer the generated speech[9].
  5. Budget the plot per time range. Each stage should hold one clear event and an end state[4].

ByteDance also publishes an official Seedance 2.5 prompt-optimization skill. Its guide recommends installing it with npx --yes skills@latest add "https://arkdocs-en.tos-ap-southeast-1.volces.com/skills/" --skill sd25-pe --yes, then typing /sd25-pe followed by your prompt in an AI chat box[4]. Opus 5.5 can draft the shot, and the skill can then check it against ByteDance's own conventions.

A director prompt for Opus 5.5

Every Opus 5.5 video generation run starts from the same instruction block. This is a starting template, not a transcript of the Reddit author's session. Replace the bracketed parts.

You are the director and prompt writer for a short live-action AI video.
Concept: [one paragraph]
Style references: Image 2 and Image 3 are STYLE ONLY (grade, lens, grain).
Subject: Image 1 is the person. Identity must never drift.
Camera operator: [who films, where they sit, what device, how they move].

Produce, in order:
1. A character bible: face, hair, wardrobe, jewelry, props. Fixed wording.
2. A GPT Image 2.5 prompt for a 16:9 character sheet (turnaround,
   five expressions, detail panels, "Identity is locked").
3. A GPT Image 2.5 prompt for the start frame of shot 1.
4. A Seedance 2.5 prompt per shot, 10 seconds each, with:
   - Asset bindings: "Image 1 = identity, Image 2 = start frame"
   - A camera section written from the operator's point of view
   - Continuous timestamps (0-3s, 3-5s, 5-8.5s, 8.5-10s), one event each
   - Dialogue in double quotes, with emotion cues per phrase
   - Audio: on-camera mic only, list ambient sounds, "No BGM" if needed
   - Negatives: "No subtitles", no cuts unless requested, no morphing
Reuse the character bible wording verbatim in every prompt.

The last line is the one that stops drift. If each prompt describes the necklace slightly differently, the video model will too.

Running Opus 5.5, GPT Image 2.5 and Seedance 2.5 from one key

All three models sit behind the same reAPI key. Opus 5.5 uses the chat completions endpoint with model ID claude-opus-5-5[1]. GPT Image 2.5 and Seedance 2.5 are asynchronous: you submit, receive a task ID, and poll GET /api/v1/tasks/{id} until the status is completed[9][10].

import time, requests

BASE = "https://reapi.ai/api/v1"
H = {"Authorization": "Bearer YOUR_API_KEY"}

def wait(task_id):
    while True:
        t = requests.get(f"{BASE}/tasks/{task_id}", headers=H).json()
        if t["status"] in ("completed", "failed"):
            return t
        time.sleep(10)

# 1. Opus 5.5 writes the prompts (DIRECTOR_PROMPT = the template above)
plan = requests.post(f"{BASE}/chat/completions", headers=H, json={
    "model": "claude-opus-5-5",
    "messages": [{"role": "user", "content": DIRECTOR_PROMPT}],
    "reasoning_effort": "high",
    "max_tokens": 16000,
}).json()["choices"][0]["message"]["content"]
# parse SHEET_PROMPT, FRAME_PROMPT, SHOT_PROMPT out of `plan`

# 2. Character sheet from the subject photo and style references
sheet = wait(requests.post(f"{BASE}/images/generations", headers=H, json={
    "model": "gpt-image-2.5-sunburst",
    "prompt": SHEET_PROMPT,
    "image_urls": [SUBJECT_URL, STYLE_URL_1, STYLE_URL_2],
}).json()["id"])["output"]["image_urls"][0]

# 3. Start frame, anchored to the sheet
frame = wait(requests.post(f"{BASE}/images/generations", headers=H, json={
    "model": "gpt-image-2.5-sunburst",
    "prompt": FRAME_PROMPT,
    "image_urls": [sheet],
}).json()["id"])["output"]["image_urls"][0]

# 4. The 10-second shot. Image 1 = sheet, Image 2 = start frame
shot = wait(requests.post(f"{BASE}/videos/generations", headers=H, json={
    "model": "doubao-seedance-2.5-face",
    "prompt": SHOT_PROMPT,
    "image_urls": [sheet, frame],
    "duration": 10,
    "resolution": "1080p",
    "size": "16:9",
}).json()["id"])["output"]["video_urls"][0]
print(shot)

Three details from the reference docs are worth knowing before you run it. Every media field must be a public HTTP(S) URL; base64 and data: URIs are rejected[9]. Seedance 2.5 accepts up to 30 reference images, so a sheet, a start frame and a few prop shots fit easily[9]. And Opus 5.5 always thinks adaptively; reasoning_effort changes how deeply, and the native default is medium[1].

What one Opus 5.5 video generation run costs

Prices below are reAPI's published rates on September 29, 2026[5][6][7]. The Opus 5.5 line assumes a planning session of 30,000 input tokens and 10,000 output tokens. That figure is an illustration, not a measurement; your template length and effort level move it.

ItemRateQuantityCost
Opus 5.5 input$3.20 per 1M tokens30,000$0.096
Opus 5.5 output$16.00 per 1M tokens10,000$0.160
GPT Image 2.5 Sunburst$0.025 per image2$0.05
Seedance 2.5, 1080p$0.462 per second20 s$9.24
Totalabout $9.55

Video seconds are over 96% of the bill. The levers that matter are all on the Seedance side: 720p runs $0.267 per second ($5.34 for the same 20 seconds), and the Seedance 2.5 Eco tier bills 1080p at $0.347 per second ($6.94) with the same request body[7][9]. A sensible routine is to iterate at 720p until the performance and camera are right, then render the keeper at 1080p.

For comparison, Anthropic lists Opus 5.5 at $4 input and $20 output per million tokens, and says it costs about 40% less than Opus 5 on typical workloads[8]. Either way, the director is the cheap part.

FAQ

Can Claude Opus 5.5 generate videos by itself?

No. Opus 5.5 is a language model called through chat completions; it returns text[1]. In an Opus 5.5 video generation workflow it writes the prompts, and a video model such as Seedance 2.5 renders the clip.

What is the Seedance 2.5 skill?

It is ByteDance's official prompt-optimization skill for Seedance 2.5, named sd25-pe. The BytePlus guide installs it with npx skills and invokes it by typing /sd25-pe plus your prompt in an AI chat box[4].

How long can a Seedance 2.5 video be?

Up to 30 seconds per generation. On reAPI, duration accepts 4 to 30 seconds, or -1 to let the model choose[9]. The Reddit workflow used 10-second shots and cut them together[2].

How much does an Opus 5.5 and Seedance 2.5 clip cost?

About $9.55 for two 10-second 1080p shots, two GPT Image 2.5 Sunburst stills and one planning session at the token counts assumed above. Rendering at 720p brings the same run to roughly $5.65[5][6][7].

Should I use GPT Image 2.5 Flare or Sunburst for the character sheet?

Sunburst, if the sheet will anchor several shots. OpenAI describes it as offering "tighter control across edits"; Flare is the faster default[3]. On reAPI the difference is $0.023 versus $0.025 per image[6].

Do I need Claude Code for this workflow?

No. The script above calls Opus 5.5 over plain HTTP. Anthropic does offer Opus 5.5 in Claude Code, including a fast mode[8], and that is a comfortable place to iterate on a template or run the sd25-pe skill. The production pipeline only needs an API key.

Can Opus 5.5 plan a music video?

It can write the shot list, timing and sound direction. Seedance 2.5 generates speech, sound effects and background music with the clip, and it also accepts reference audio (up to 10 tracks, 30 seconds combined), so a song can drive the edit[9].

Treat Opus 5.5 as the director and Seedance 2.5 as the crew

The workflow that went around r/Seedance_AI succeeds for a boring reason: every prompt is long, specific and consistent, and a language model is better than a tired human at producing that. Give Opus 5.5 a fixed template, a character bible it must reuse word for word, and a camera operator with a seat and a device. Let GPT Image 2.5 lock the face, and spend your iteration budget on Seedance 2.5 seconds at 720p before the final 1080p render. That is Opus 5.5 video generation that actually ships: one model to think, two to render, and one API key for all three on reAPI.

References

  1. reAPI. Claude Opus 5.5 API: model ID, parameters and pricing. Retrieved September 2026 from reapi.ai/docs/claude-opus-5-5
  2. r/Seedance_AI. Full Breakdown: How I used Claude Opus & Seedance 2.5 to achieve photorealistic AI videos (All Prompts Included). Posted September 24, 2026. Retrieved September 2026 from reddit.com/r/Seedance_AI/comments/1wop0ju
  3. OpenAI. Introducing ChatGPT Images 2.5. September 8, 2026. Retrieved September 2026 from openai.com/index/introducing-chatgpt-images-2-5
  4. BytePlus ModelArk. Dreamina Seedance 2.5 prompt guide. Retrieved September 2026 from docs.byteplus.com/en/docs/modelark/seedance-2-5-prompt-guide
  5. reAPI. Claude Opus 5.5 model page and pricing. Retrieved September 29, 2026 from reapi.ai/models/claude-opus-5-5
  6. reAPI. GPT Image 2.5 model page and pricing. Retrieved September 29, 2026 from reapi.ai/models/gpt-image-2-5
  7. reAPI. Seedance 2.5 model page and pricing. Retrieved September 29, 2026 from reapi.ai/models/seedance-2-5
  8. Anthropic. Introducing Claude Opus 5.5. Retrieved September 2026 from anthropic.com/claude-opus-5-5
  9. reAPI. Seedance 2.5 API: parameters, modes and billing. Retrieved September 2026 from reapi.ai/docs/seedance-2-5
  10. reAPI. GPT Image 2.5 API reference. Retrieved September 2026 from reapi.ai/docs/gpt-image-2-5