GPT Image 2.5 is live — OpenAI's newest image model, targeted edits that leave the rest of the frame alone
GPT-6 Luna API Pricing: Token Rates and Cost Calculator
2026/10/07

GPT-6 Luna API Pricing: Token Rates and Cost Calculator

GPT-6 Luna API pricing from OpenAI's price page: input, output, cache reads and writes, the 272K long-context tier, Batch and Fast, plus worked cost examples.

GPT-6 Luna API pricing starts at $0.10 per million input tokens and $0.50 per million output tokens on OpenAI's Standard tier, the lowest rates in the GPT-6 lineup[1]. Those two numbers are only part of the bill, though. Cached input, cache writes, a separate rate for prompts over 272K tokens, and reasoning tokens that bill as output can all move the real cost.

This guide lists every GPT-6 Luna rate from OpenAI's pricing page, retrieved October 7, 2026, explains how a request is priced, and works through five cost scenarios with their assumptions written out, so you can redo the arithmetic with your own token counts.

TL;DR

  • Standard rates per 1M tokens: $0.10 input, $0.01 cached input, $0.125 cache writes, $0.50 output[1].
  • Over 272K input tokens, the whole request moves to $0.20 input, $0.02 cached input, $0.25 cache writes and $0.75 output[1][2].
  • Batch and Flex cost half of Standard; Fast mode costs twice; regional processing adds 10%[2].
  • Reasoning tokens bill as output, so reasoning effort is the biggest lever after volume[3].
  • A short request is cheap: 500 input and 10 output tokens cost $0.000055 at Standard rates, or $55 per million such requests (our arithmetic, shown below).
  • Against the rest of GPT-6: Astra's Standard rates are exactly 100 times Luna's, and GPT-6 Sol's input and output are 20 times[1].

GPT-6 Luna API pricing for every token type

OpenAI prices GPT-6 Luna by token type and by prompt length. "Short context" is up to 272K input tokens; above that, the long-context rates apply to the full request[1].

Per 1M tokens (Standard)Up to 272K inputOver 272K input
Input$0.10$0.20
Cached input$0.01$0.02
Cache writes$0.125$0.25
Output$0.50$0.75

Source: OpenAI API pricing, Standard tier[1].

The model page states the multipliers behind these numbers. Cached input is 10% of the uncached input rate, cache writes are 1.25 times it, and prompts with more than 272K input tokens are "priced at 2x input and cache rates and 1.5x output for the full request"[2].

Batch, Flex, Fast and regional processing

The processing tier changes every rate by a fixed factor. OpenAI's model page says "Batch and Flex are priced at 50% of Standard rates. Fast mode is priced at 2x the applicable rates," and regional processing "adds a 10% premium where available"[2].

Per 1M tokens, up to 272K inputInputCached inputCache writesOutput
Standard$0.10$0.01$0.125$0.50
Batch$0.05$0.005$0.0625$0.25
Flex$0.05$0.005$0.0625$0.25
Fast$0.20$0.02$0.25$1.00

Source: OpenAI API pricing, Standard, Batch, Flex and Fast tables[1].

Batch fits offline jobs such as nightly classification or backfills. Fast mode is OpenAI's option for latency-sensitive traffic, and OpenAI says it is not available with EU data residency for GPT-6 Luna[4].

How a GPT-6 Luna request is priced

For a prompt of 272K input tokens or fewer, at Standard rates:

cost = ( uncached_input × 0.10
       + cached_input   × 0.01
       + cache_write    × 0.125
       + output         × 0.50 ) / 1,000,000

Above 272K input tokens, swap in 0.20, 0.02, 0.25 and 0.75 for the whole request. Four rules decide which bucket each token lands in.

Reasoning tokens are output. OpenAI says reasoning tokens "occupy space in the model's context window and are billed as output tokens"[3]. GPT-6 Luna supports six effort levels, none to max, with medium as the default[2]. A two-word classification label can carry hundreds of billed reasoning tokens behind it.

Cache charges replace the input rate rather than adding to it. OpenAI's caching guide says cache-write pricing "is not an additive fee: input tokens use the uncached-input, cached-input, or cache-write rate"[5].

Short prefixes never cache. The minimum cacheable prompt length is 1,024 tokens for GPT-5.6 and later models, and a cached prefix stays eligible for at least 30 minutes after its last write or reuse[5].

The 272K line is per request. One prompt of 300K input tokens is billed entirely at long-context rates, output included[2].

GPT-6 Luna API cost calculator: five worked examples

Every row below uses OpenAI's list rates above. Token counts are assumptions chosen for illustration, not measurements. Replace them with numbers from the usage object of your own responses.

ScenarioAssumptionsArithmeticTotal
A. Ticket tagging, effort none1,000,000 requests × 500 input, 10 output500M × $0.10 + 10M × $0.50 = $50 + $5$55.00
B. Same job, effort mediumAs A, plus an assumed 300 reasoning tokens per request500M × $0.10 + 310M × $0.50 = $50 + $155$205.00
C. Scenario A through BatchAs A, at Batch rates500M × $0.05 + 10M × $0.25 = $25 + $2.50$27.50
D. Support assistant with a cached prefix100,000 requests; 4,000-token stable prefix written 1,000 times and read 99,000 times; 1,000 new input and 400 output tokens each396M × $0.01 + 4M × $0.125 + 100M × $0.10 + 40M × $0.50 = $3.96 + $0.50 + $10 + $20$34.46
E. Long documents1,000 requests × 400,000 input, 2,000 output (long-context tier)400M × $0.20 + 2M × $0.75 = $80 + $1.50$81.50

Three readings stand out.

Effort costs more than volume here. In B, the assumed reasoning pass makes output tokens the larger bill, and the job costs 3.7 times scenario A. Test whether none or low holds your accuracy target before paying for medium.

Caching halves scenario D. Without a cache, the same traffic costs 500M × $0.10 + 40M × $0.50 = $70.00. Cache reads cut the prefix from $40 to $4.46 including writes.

The 272K line doubles input. If each document in E could be split into two 200,000-token requests with 2,000 output tokens each, the bill would be 400M × $0.10 + 4M × $0.50 = $42.00. Splitting changes what the model sees at once, so check quality before chasing the saving.

How GPT-6 Luna pricing compares with other GPT models

All Standard rates per 1M tokens for prompts up to 272K input tokens[1]:

ModelInputCached inputCache writesOutput
GPT-6 Luna$0.10$0.01$0.125$0.50
GPT-5.6 Luna$0.20$0.02$0.25$1.20
GPT-6 Sol$2.00$0.20$2.50$10.00
GPT-6.1 Sol$2.00$0.10$2.50$10.00
GPT-6 Astra$10.00$1.00$12.50$50.00

OpenAI attributes Luna's low price to "improvements in caching and inference," passed on as a cut from GPT-5.6 pricing[6]. The model-by-model trade-offs are covered in GPT-6 Luna vs 5.6 Luna and GPT-6 Sol vs Luna; Sol's own rates are in GPT-6 Sol pricing.

Paying for GPT-6 Luna through reAPI

reAPI serves GPT-6 Luna on its Chat Completions endpoint with the model id gpt-6-luna, billed per token from your reAPI credit balance. Input, output, cache read and cache write each have their own rate, reasoning tokens bill as output, and prompts over 272K input tokens re-price the whole request[7]. The usage object in each response reports prompt_tokens, completion_tokens (reasoning included) and cached_tokens, so you can reconcile the bill with the formula above[7].

reAPI's current per-token rates and a cost estimator for input and output tokens are on the GPT-6 Luna model page. One request-shape note: sending temperature to GPT-6 Luna returns 400, covered in GPT-6 Luna temperature.

FAQ

How much does chat gpt 6 cost?

In the API, GPT-6 is priced per token by model: GPT-6 Luna at $0.10 input and $0.50 output per 1M tokens, GPT-6 Sol and GPT-6.1 Sol at $2 and $10, and GPT-6 Astra at $10 and $50[1]. In ChatGPT, OpenAI made GPT-6 Sol and Luna available to Plus, Pro, Business, Enterprise and Edu users, and Free and Go users can use GPT-6 Luna in the desktop app[6].

gpt 6 luna pricing for 10000 tokens

At Standard rates, 10,000 input tokens cost $0.001 and 10,000 output tokens cost $0.005[1]. If those 10,000 input tokens were cache reads, they would cost $0.0001.

gpt 6 luna output token cost

Output costs $0.50 per 1M tokens for prompts up to 272K input tokens and $0.75 above that, at Standard rates[1]. Reasoning tokens are included in billed output[3].

gpt 6 luna pricing discount

OpenAI lists two: Batch and Flex at 50% of Standard rates, and cached input at 10% of the uncached input rate[2]. Fast mode and regional processing go the other way, at 2x and +10%.

gpt 6 luna free

In the API, GPT-6 Luna is billed per token at the rates above[1]. In ChatGPT, OpenAI says Free and Go users can access GPT-6 Luna in the desktop app[6].

gpt 6 luna api cost vs gpt 6 astra

GPT-6 Astra's Standard rates are $10 input, $1 cached input, $12.50 cache writes and $50 output per 1M tokens, exactly 100 times each GPT-6 Luna rate[1]. Astra also offers an Ultrafast tier at higher prices; Luna does not appear in that table[1].

How much does a GPT API cost?

It depends on the model. In OpenAI's current Standard table, rates run from GPT-6 Luna's $0.10 input and $0.50 output per 1M tokens up to $30 input and $180 output for GPT-5.5 Pro and GPT-5.4 Pro[1].

Keeping GPT-6 Luna costs predictable

Most of a GPT-6 Luna bill is set by three choices: the reasoning effort you send, whether your stable prompt prefix is long enough to cache, and whether any single prompt crosses 272K input tokens. Measure each from real usage data on a sample of traffic, then plug those counts into the formula. With that in place, GPT-6 Luna API pricing is easy to forecast: $0.10 in, $0.50 out, and every exception is written down.

References

  1. OpenAI. API pricing. Retrieved October 2026 from developers.openai.com/api/docs/pricing
  2. OpenAI. GPT-6 Luna model page. Retrieved October 2026 from developers.openai.com/api/docs/models/gpt-6-luna
  3. OpenAI. Reasoning models. Retrieved October 2026 from developers.openai.com/api/docs/guides/reasoning
  4. OpenAI. Fast mode. Retrieved October 2026 from developers.openai.com/api/docs/guides/fast-mode
  5. OpenAI. Prompt caching. Retrieved October 2026 from developers.openai.com/api/docs/guides/prompt-caching
  6. OpenAI. Introducing GPT-6 Sol and Luna. Retrieved October 2026 from openai.com/index/introducing-gpt-6-sol-and-luna
  7. reAPI. GPT-6 Luna API docs. Retrieved October 2026 from reapi.ai/docs/gpt-6-luna