GPT Image 2.5 is live — OpenAI's newest image model, targeted edits that leave the rest of the frame alone
GPT-6 Sol Pricing: API Price, Token Cost and Cache Rates
2026/10/07

GPT-6 Sol Pricing: API Price, Token Cost and Cache Rates

GPT-6 Sol pricing per million tokens: input, output, cache reads and writes, the 272K long-context rule, Batch, Flex and Fast mode, with worked cost examples.

GPT-6 Sol pricing starts at $2 per million input tokens and $10 per million output tokens on OpenAI's API, for prompts up to 272K input tokens[1]. Those two numbers are only part of the bill. OpenAI also charges separately for cached input and cache writes, doubles the input rate when a prompt goes over 272K tokens, and offers Batch, Flex and Fast mode at different multiples[2].

This article lists every rate OpenAI publishes for GPT-6 Sol as of October 7, 2026, works through four costed examples with their assumptions stated, and answers the question we see searched alongside it: whether the current price is a promotion.

TL;DR

  • Standard rates: $2 input, $0.20 cached input, $2.50 cache write and $10 output per million tokens, for prompts up to 272K input tokens[1].
  • Long prompts: above 272K input tokens the whole request moves to $4 input, $0.40 cached, $5 cache write and $15 output[1][2].
  • Reasoning tokens bill as output. The effort level you choose changes the output count, and with it the cost[3].
  • Processing tiers: Batch and Flex cost 50% of Standard, Fast mode costs 2x, and regional processing adds 10%[2].
  • No promotion found: the promotional price on OpenAI's pages belongs to GPT-5.6 Sol, not GPT-6 Sol[1][4].
  • On reAPI: billed per token on the same four dimensions, with rates listed on the GPT-6 Sol model page.

GPT-6 Sol pricing per million tokens

OpenAI's pricing page splits each rate into a short-context column (up to 272K input tokens) and a long-context column (over 272K)[1]:

GPT-6 Sol, Standard, per 1M tokensUp to 272K inputOver 272K input
Input$2.00$4.00
Cached input$0.20$0.40
Cache writes$2.50$5.00
Output$10.00$15.00

The model page explains the ratios. Cached input is 10% of the uncached input rate, cache writes are 1.25 times it, and prompts above 272K are "priced at 2x input and cache rates and 1.5x output for the full request"[2]. The words "for the full request" are the important ones. Crossing the threshold re-prices every token in the request, not just the tokens above 272K.

Two limits frame what a single request can cost. GPT-6 Sol accepts up to 922,000 input tokens and returns up to 128,000 output tokens, inside a 1,050,000-token context window[2].

What changes a GPT-6 Sol bill

Reasoning tokens. GPT-6 Sol is a reasoning model. OpenAI's reasoning guide says reasoning tokens are not visible through the API but "are billed as output tokens"[3]. The effort setting runs from none to max, with medium as the default[2], so a request that leaves effort unset still pays for a reasoning pass at the $10 output rate.

Caching. For GPT-5.6 and later models, OpenAI bills a cache write at 1.25x the input rate and a cache read at 0.1x. A cached prefix stays available for at least 30 minutes after its latest write or reuse, and the minimum cacheable prefix is 1,024 input tokens[5]. Cache writes replace the normal input rate for those tokens rather than adding to it[5]. In practice, writing a prefix once and reusing it once costs 1.35x its uncached price, against 2x without caching[5].

Processing tier. The same tokens cost different amounts depending on how you send them[1]:

GPT-6 Sol, per 1M tokens, up to 272KInputCached inputCache writesOutput
Standard$2.00$0.20$2.50$10.00
Batch$1.00$0.10$1.25$5.00
Flex$1.00$0.10$1.25$5.00
Fast mode$4.00$0.40$5.00$20.00

Flex uses Batch rates in exchange for slower responses and occasional resource unavailability, and OpenAI describes it as a beta[6]. Regional processing (data residency) adds 10% where it is available, and Fast mode is not available with EU data residency for GPT-6 Sol[2][7].

GPT-6 Sol cost examples

All four examples use OpenAI's list prices above. They are illustrations with stated assumptions, not measurements of real traffic. Token counts for your prompts will differ, and output counts include reasoning tokens.

Example A: one chat request. 2,000 input tokens and 800 output tokens, Standard, no cache. 2,000 × $2 / 1M = $0.004 for input and 800 × $10 / 1M = $0.008 for output, so $0.012 per request, or $12 per 1,000 requests.

Example B: a shared prefix with caching. 40 requests reuse the same 50,000-token block of instructions and reference material, each adds 2,000 new input tokens and gets 1,000 output tokens back. Assume one explicit cache breakpoint after the shared block and all 40 requests inside the cache lifetime. The new input after the breakpoint is billed at the plain input rate with no write charge[5].

LineCalculationCost
First cache write50,000 × $2.50 / 1M$0.125
39 cache reads39 × 50,000 × $0.20 / 1M$0.390
New input40 × 2,000 × $2 / 1M$0.160
Output40 × 1,000 × $10 / 1M$0.400
Total with caching$1.075
Total without caching40 × 52,000 × $2 / 1M + $0.40$4.56

In this scenario caching cuts the bill by about three quarters, because most of each prompt is the repeated block.

Example C: crossing 272K. One request with 300,000 input tokens and 5,000 output tokens moves to the long-context rates: 300,000 × $4 / 1M = $1.20, plus 5,000 × $15 / 1M = $0.075, for $1.275. Trim the same request to 270,000 input tokens and it costs 270,000 × $2 / 1M = $0.54 plus 5,000 × $10 / 1M = $0.05, or $0.59. Cutting 10% of the input more than halves the cost, so check prompt length before you send whole repositories or document sets.

Example D: the same job on each tier. Example A run 10,000 times costs $120 on Standard, $60 on Batch or Flex, and $240 in Fast mode (2,000 × $4 / 1M + 800 × $20 / 1M = $0.024 per request).

To estimate your own workload, multiply each token type by its rate: uncached input × input rate, cache reads × cached rate, cache writes × write rate, and output (including reasoning) × output rate. Use the long-context column when any single request goes over 272K input tokens.

Is GPT-6 Sol on promotional pricing?

We found no indication that it is. OpenAI's pricing page carries one promotional note: "GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026"[1]. The GPT-6 Sol model page, the September 22 changelog entry that announced GPT-6 Sol at $2 input and $10 output, and the launch post contain no promotional label or end date for GPT-6 Sol[2][4][8]. The launch post uses "promotional" only to describe what GPT-6 Sol was compared against: Sol and Luna prices fell "by 50% compared with their GPT‑5.6 promotional pricing"[8].

If OpenAI later attaches a time limit to GPT-6 Sol's price, it would most likely appear on the same pricing page and changelog cited here.

How GPT-6 Sol pricing compares with other OpenAI models

Standard, short-context rates per million tokens on OpenAI's pricing page[1]:

ModelInputCached inputCache writesOutput
GPT-6 Astra$10.00$1.00$12.50$50.00
GPT-6.1 Sol$2.00$0.10$2.50$10.00
GPT-6 Sol$2.00$0.20$2.50$10.00
GPT-6 Luna$0.10$0.01$0.125$0.50
GPT-5.6 Sol$4.00$0.40$5.00$20.00

GPT-6.1 Sol, released on September 29, 2026, has the same input, output and cache-write rates as GPT-6 Sol but half the cached-input rate, 5% of input instead of 10%[4][5]. For cache-heavy agent loops that difference adds up, although GPT-6.1 Sol also drops the none effort and tool calling on Chat Completions[9]. Our GPT-6 Sol vs 5.6 Sol comparison covers the migration side.

GPT-6 Sol pricing on reAPI

reAPI serves GPT-6 Sol through its Chat Completions endpoint with the model ID gpt-6-sol, billed per token from your reAPI credit balance. The billing dimensions follow OpenAI's: input, output, cache read and cache write each have their own rate, reasoning tokens bill as output, and prompts over 272K input tokens re-price the whole request at the long-context rates[10]. reAPI's docs state that every rate sits below OpenAI's published per-token rate[10]. Current numbers, and a cost estimator, are on the GPT-6 Sol model page.

FAQ

GPT-6 Sol: how long at promotional pricing?

We found no promotional period for GPT-6 Sol on OpenAI's pricing page, model page, changelog or launch post[1][2][4][8]. The promotion OpenAI documents is for GPT-5.6 Sol, at $4 input and $20 output, available at least through November 21, 2026[1].

How much does chat gpt 6 cost?

In the API, GPT-6 Sol costs $2 per million input tokens and $10 per million output tokens on Standard processing. GPT-6 Astra costs $10 and $50, and GPT-6 Luna $0.10 and $0.50[1]. In ChatGPT, GPT-6 Sol is available in ChatGPT Work and Codex on Plus, Pro, Business, Enterprise and Edu plans[8]; plan prices are outside the scope of this article.

GPT-6 Sol fast pricing

Fast mode for GPT-6 Sol costs 2x Standard: $4 input, $0.40 cached input, $5 cache write and $20 output per million tokens up to 272K, and $8, $0.80, $10 and $30 above it[1][2].

GPT-6 Sol token cost

Per token, GPT-6 Sol costs $0.000002 per input token and $0.00001 per output token at Standard short-context rates, which is $2 and $10 per million[1]. Cached input tokens cost a tenth of the input rate[2].

GPT-6 Sol pricing calculator

Multiply each token type by its rate and add them up: input × $2, cached input × $0.20, cache writes × $2.50 and output × $10, all per million tokens. Switch to $4, $0.40, $5 and $15 when one request exceeds 272K input tokens[1]. The GPT-6 Sol model page has an estimator for reAPI rates.

GPT 5.6 Sol vs 6 Sol price

GPT-6 Sol costs half of GPT-5.6 Sol on every line: $2 against $4 for input, $10 against $20 for output, $0.20 against $0.40 for cached input, and $2.50 against $5 for cache writes[1]. GPT-5.6 Sol's $4 and $20 are a promotional price that runs at least through November 21, 2026[1].

How much does a GPT API cost?

It depends on the model. On OpenAI's current pricing page, Standard input rates per million tokens range from $0.10 for GPT-6 Luna to $10 for GPT-6 Astra among the GPT-6 models, with output priced higher than input for each of them[1].

Budgeting for GPT-6 Sol

For most workloads, the first lever is the output count, because output costs five times as much as input and includes reasoning. Pick the lowest effort that passes your quality checks. The second lever is caching any block of instructions or documents you send repeatedly; in Example B it cut cost by about three quarters. The third is staying under 272K input tokens per request. If results can wait, Batch halves the bill. GPT-6 Sol pricing on OpenAI's list is $2 input and $10 output per million tokens, with no promotional end date published, and reAPI's current per-token rates are on the GPT-6 Sol model page.

References

  1. OpenAI. API pricing. Retrieved October 2026 from developers.openai.com/api/docs/pricing
  2. OpenAI. GPT-6 Sol model page. Retrieved October 2026 from developers.openai.com/api/docs/models/gpt-6-sol
  3. OpenAI. Reasoning models. Retrieved October 2026 from developers.openai.com/api/docs/guides/reasoning
  4. OpenAI. API changelog. Retrieved October 2026 from developers.openai.com/api/docs/changelog
  5. OpenAI. Prompt caching. Retrieved October 2026 from developers.openai.com/api/docs/guides/prompt-caching
  6. OpenAI. Flex processing. Retrieved October 2026 from developers.openai.com/api/docs/guides/flex-processing
  7. OpenAI. Fast mode. Retrieved October 2026 from developers.openai.com/api/docs/guides/fast-mode
  8. OpenAI. Introducing GPT-6 Sol and Luna. Retrieved October 2026 from openai.com/index/introducing-gpt-6-sol-and-luna
  9. OpenAI. GPT-6.1 Sol model page. Retrieved October 2026 from developers.openai.com/api/docs/models/gpt-6.1-sol
  10. reAPI. GPT-6 Sol API docs. Retrieved October 2026 from reapi.ai/docs/gpt-6-sol