Seedance 2.5 is live — 30-second cinematic video with native audio & real-person referencesfrom $0.071/s
rreAPI Docs

kimi-k3

Kimi K3 — Moonshot's open-weight flagship for long-horizon coding and knowledge work. OpenAI-compatible /v1/chat/completions on api.reapi.ai with a 1,048,576-token context window, output up to the whole window, always-on reasoning, and native image and video input.

Moonshot's Kimi K3 — the flagship for long-horizon coding and end-to-end knowledge work — exposed through api.reapi.ai as a drop-in OpenAI-compatible Chat Completions endpoint. A 1,048,576-token context window, output defaulting to 131,072 tokens and raisable to the full window, reasoning that is always on with a three-rung reasoning_effort dial, native image and video input, tool calling with a required tool choice, strict JSON Schema output, and automatic context caching. The wire model id is kimi-k3. Current rates live on the model page and on api.reapi.ai/pricing.

Quick example

curl https://api.reapi.ai/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "kimi-k3",
    "messages": [
      { "role": "user", "content": "Refactor this module and add tests." }
    ],
    "reasoning_effort": "high",
    "stream": true
  }'
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.reapi.ai/v1",
)

stream = client.chat.completions.create(
    model="kimi-k3",
    messages=[{"role": "user", "content": "Refactor this module and add tests."}],
    reasoning_effort="high",
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta
    reasoning = getattr(delta, "reasoning_content", None)
    if reasoning:
        print(reasoning, end="")
    if delta.content:
        print(delta.content, end="")
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: "YOUR_API_KEY",
  baseURL: "https://api.reapi.ai/v1",
});

const stream = await client.chat.completions.create({
  model: "kimi-k3",
  messages: [{ role: "user", content: "Refactor this module and add tests." }],
  reasoning_effort: "high",
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
package main

import (
    "bytes"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
)

func main() {
    body, _ := json.Marshal(map[string]any{
        "model": "kimi-k3",
        "messages": []map[string]string{
            {"role": "user", "content": "Refactor this module and add tests."},
        },
        "reasoning_effort": "high",
    })
    req, _ := http.NewRequest("POST",
        "https://api.reapi.ai/v1/chat/completions", bytes.NewReader(body))
    req.Header.Set("Authorization", "Bearer YOUR_API_KEY")
    req.Header.Set("Content-Type", "application/json")

    resp, _ := http.DefaultClient.Do(req)
    defer resp.Body.Close()
    out, _ := io.ReadAll(resp.Body)
    fmt.Println(string(out))
}

Authentication

Chat models are served from the api.reapi.ai gateway, which has its own console and its own key. Create the key there, then send it as a bearer token:

Authorization: Bearer YOUR_API_KEY

The gateway key is not the same credential as a media-generation key used against reapi.ai/api/v1. Sign in at api.reapi.ai to create one.


Endpoint

POST /v1/chat/completions

Base URL https://api.reapi.ai. The wire format is OpenAI-compatible, so the same SDKs (openai-python, openai-node, openai-go, …) work once you swap the base URL, the key and the model string. Moonshot's own quickstart drives this model through the OpenAI SDK, so the compatibility is the vendor's design rather than a translation layer.


Request body

model — string, required

Must be kimi-k3 exactly — no date suffix, no version decimal.

messages — array, required

Conversation history, each entry an object with role and content. Text, image and video parts are supported — see Visual input.

Echo the assistant message back unchanged. In multi-turn conversations and tool-call loops, the complete assistant message returned by the API has to go back into messages as-is, including reasoning_content and tool_calls. Keeping only content — the habit most OpenAI-shaped code has — breaks this model's Preserved Thinking.

max_completion_tokens — integer

Upper bound on output tokens for this response. Moonshot documents a default of 131,072, raisable to 1,048,576 — the size of the whole context window. Reasoning runs before the answer, so budget for both.

reasoning_effort — string, default "max"

Top-level field controlling how hard the model thinks. See Reasoning.

stream — boolean, default false

When true, tokens arrive as server-sent events. Reasoning and answer come as separate deltasreasoning_content and content — so a client can render the reasoning live and still parse the answer on its own.

tools / tool_choice — optional

Tool definitions and the strategy for picking among them. tool_choice accepts auto, none and required. See Tool calling.

response_format — object, optional

A JSON Schema with strict: true constrains the answer field. Parse that field, not reasoning_content, which is prose.

group — string, default "default"

Token group on the gateway. Leave it at default unless your account has been given another group to route through.


Reasoning

This model always reasons. Moonshot's own FAQ answers "how do I turn off the chain of thought" with a flat you can't. There is no thinking toggle; the lever is effort.

reasoning_effort takes three rungs and defaults to the most expensive one:

LevelUse it for
maxThe default. The hardest reasoning, where correctness outweighs cost.
highA step down that still reasons hard.
lowShort, scoped, latency-sensitive work — and the answer to "the reasoning is taking too long".

There is no medium.

Pick the rung before the conversation starts. Moonshot documents that switching effort levels invalidates prefix-cache hits, so flipping it mid-session silently throws away your cache.

The K2-era thinking object is not a parameter on this model. If you are migrating from kimi-k2.6 or kimi-k2.7-code, delete it and use top-level reasoning_effort instead. Code that already sends OpenAI's reasoning_effort needs no change beyond the accepted values.


Parameters that are fixed

Moonshot pins all five of these on this model. Passing any other value returns an error, so leave them out of the request entirely:

ParameterFixed at
temperature1.0
top_p0.95
n1
presence_penalty0
frequency_penalty0

Steer with prompting and the reasoning_effort rung instead.


Tool calling

Two controls are specific to this generation.

tool_choice: "required" forces at least one tool call on the turn — the way to stop an agent answering from memory when it was supposed to look something up. The K2 models reject this value; this one accepts it.

Dynamic tool loading introduces a tool partway through a conversation: put its complete definition in a system message that carries tools and no content. It becomes available from that position onward.

{
  "role": "system",
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "calculate",
        "description": "Evaluate an arithmetic expression",
        "parameters": {
          "type": "object",
          "properties": { "expression": { "type": "string" } },
          "required": ["expression"]
        }
      }
    }
  ]
}

The server does not retain a dynamically loaded tool. Keep that system message in your request history or the tool disappears on the next turn.

Include the complete name, description and parameters in the definition, and after executing calls, append one tool message per call with the matching tool_call_id.


Visual input

ModalitySupported
Text
Image
Video

Output is text only. Vision is native to this model rather than a separate variant, which is what makes screenshot-driven work — frontend layout, game development, CAD — practical: it can look at what it just produced.

Public image URLs are not accepted. Moonshot documents base64 data URIs, or a file reference from its own upload API, as the two ways in. content must also be an array of parts, not a serialized string.

{
  "role": "user",
  "content": [
    { "type": "image_url", "image_url": { "url": "data:image/png;base64,iVBORw0..." } },
    { "type": "text", "text": "Why does this layout break at 768px?" }
  ]
}

Note that file references come from Moonshot's Files API, which the Chat Completions endpoint does not itself provide.


Context caching

Caching is automatic — no cache id, no TTL, no extra parameter. Keep a long prefix (system prompt, knowledge document, tool definitions) stable and repeat requests attempt a hit on their own. Two conditions decide whether that happens:

  • The previous request's prompt tokens must exceed 256. Below that the request is not cached at all.
  • The reasoning_effort rung must not change — switching it invalidates the prefix cache.

Put the fixed bulk at the front of messages and append the varying question after it.


Pricing dimensions

Billing is per token, in USD, against your api.reapi.ai balance, with separate input and output rates. Three things to keep in mind:

  • Reasoning tokens bill as output. Since this model always reasons, that is the usual reason a bill exceeds an estimate built from visible answer length.
  • The rate is flat across context length. Moonshot documents no long-context tier, so a million-token prompt costs the same per token as a short one.
  • Effort moves the bill. It is the only reasoning lever, and it defaults to the most expensive rung — so it is the first thing to tune.

Both token rates sit below Moonshot's published per-token rate, output by a full fifth. Current numbers are on the model page and api.reapi.ai/pricing — those tables are the canonical source, not this page.

Chat models bill in USD on the gateway balance. They do not draw down the integer credits used by the media-generation endpoints on reapi.ai/api/v1.


Response shape

Non-streaming (stream: false)

{
  "id": "chatcmpl-...",
  "object": "chat.completion",
  "created": 1785000000,
  "model": "kimi-k3",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "reasoning_content": "...",
        "content": "..."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 1204,
    "completion_tokens": 860,
    "total_tokens": 2064
  }
}

reasoning_content is a sibling of content, not a replacement for it — and it is the field you must carry into the next request.

Streaming (stream: true)

Server-sent events, each data: line carrying a chat.completion.chunk. Reasoning arrives in choices[0].delta.reasoning_content and the answer in choices[0].delta.content, terminated by data: [DONE].


Errors

Failures use the gateway's standard envelope. See the errors catalog for the full code list. Common cases:

TriggerWhat to do
Missing or invalid bearer tokenCreate a key in the api.reapi.ai console
Unknown model valueSend kimi-k3 exactly
temperature / top_p / n / penalty params presentRemove them; they are fixed on this model
K2-era thinking object presentRemove it and use reasoning_effort
reasoning_effort set to mediumNot a valid rung — use low, high or max
A public image URL in an image_url partSend a base64 data URI or an uploaded file reference
content sent as a string for a multimodal turnMake it an array of parts
max_completion_tokens above the ceilingLower it; the documented ceiling is 1,048,576
Insufficient balanceTop up on the gateway
Upstream rate limitRetry with backoff, or route through another model

Moonshot flags its own web_search tool as being updated and does not recommend using it in the near term. Treat web search as unavailable on this model for now and supply fresh context yourself.


Tips

  • Give the whole task up front. This model is built for long-horizon work under minimal supervision; a complete specification beats a drip-feed of short turns.
  • Drop your sampling parameters first. The single most common migration error here is a carried-over temperature.
  • Tune effort before anything else. The default is max; low exists precisely for the case where reasoning is costing more than the answer is worth.
  • Budget for reasoning. If responses truncate on hard prompts, the reasoning pass consumed the allowance — raise max_completion_tokens.
  • Stream anything interactive, and render reasoning_content separately from the answer rather than concatenating them.
  • Keep the assistant message intact in history. Stripping reasoning_content is silent — it does not error, it just degrades multi-turn behaviour.
  • Front-load the stable bulk of your prompt so automatic caching can hit, and do not change the effort rung mid-conversation.

Table of Contents