GPT Image 2.5 is live — OpenAI's newest image model, targeted edits that leave the rest of the frame alone
rreAPI Docs

claude-haiku-5-5

Use Claude Haiku 5.5 with native Messages, streaming, structured JSON, tool calls, document citations, and prompt caching on reAPI.

Call Claude Haiku 5.5 with the native Messages format; see the model page for current pricing.

Quick example

curl --no-buffer https://reapi.ai/api/v1/messages \
  -H "x-api-key: YOUR_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{"model":"claude-haiku-5-5","max_tokens":4096,"messages":[{"role":"user","content":"Explain the next step clearly."}],"output_config":{"effort":"medium"},"stream":true}'
import json
import urllib.request

payload = {
    "model": "claude-haiku-5-5",
    "max_tokens": 4096,
    "messages": [{"role": "user", "content": "Explain the next step clearly."}],
    "output_config": {"effort": "medium"},
}
request = urllib.request.Request(
    "https://reapi.ai/api/v1/messages",
    data=json.dumps(payload).encode(),
    headers={"x-api-key": "YOUR_API_KEY", "anthropic-version": "2023-06-01",
             "Content-Type": "application/json"},
    method="POST",
)
with urllib.request.urlopen(request, timeout=120) as response:
    message = json.load(response)
    print(json.dumps(message, indent=2))
const response = await fetch("https://reapi.ai/api/v1/messages", {
  method: "POST",
  headers: {
    "x-api-key": "YOUR_API_KEY",
    "anthropic-version": "2023-06-01",
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "claude-haiku-5-5",
    max_tokens: 4096,
    messages: [{ role: "user", content: "Explain the next step clearly." }],
    output_config: { effort: "medium" },
  }),
});
const body = await response.json();
if (!response.ok) throw new Error(JSON.stringify(body));
console.log(body);
package main

import (
    "bytes"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "time"
)

func main() {
    body, err := json.Marshal(map[string]any{
        "model": "claude-haiku-5-5",
        "max_tokens": 4096,
        "messages": []map[string]string{{"role": "user", "content": "Explain the next step clearly."}},
        "output_config": map[string]string{"effort": "medium"},
    })
    if err != nil { panic(err) }
    req, err := http.NewRequest("POST", "https://reapi.ai/api/v1/messages", bytes.NewReader(body))
    if err != nil { panic(err) }
    req.Header.Set("x-api-key", "YOUR_API_KEY")
    req.Header.Set("anthropic-version", "2023-06-01")
    req.Header.Set("Content-Type", "application/json")
    client := &http.Client{Timeout: 120 * time.Second}
    resp, err := client.Do(req)
    if err != nil { panic(err) }
    defer resp.Body.Close()
    out, err := io.ReadAll(resp.Body)
    if err != nil { panic(err) }
    if resp.StatusCode >= 400 { panic(string(out)) }
    fmt.Println(string(out))
}

Endpoint and headers

POST https://reapi.ai/api/v1/messages

HeaderRequirement
x-api-keyYour reAPI key; required unless using the existing Bearer authentication option
Content-Typeapplication/json
anthropic-versionOptional; defaults to 2023-06-01
anthropic-betaOptional; forward the exact feature header required by the official documentation

For an Anthropic SDK, use https://reapi.ai/api as the base URL; the SDK appends /v1/messages. Disable client-side automatic retries when one attempt per action is required. reAPI makes one upstream generation attempt and does not retry automatically.

Parameters

Native Messages fields and unknown extensions pass through unchanged. There is no field allowlist or parameter-count restriction. The upstream determines acceptance and behavior; HTTP 200 alone does not establish that a parameter took effect. Existing authentication, request-size, and billing admission checks still apply.

FieldType / requiredDefault and behavior
modelstring; requiredclaude-haiku-5-5
messagesnon-empty array; requiredNative conversation content, including tool results and signed thinking blocks
max_tokensinteger; required0–128,000 including thinking; zero is forwarded for cache warmup
systemstring or text-block array; optionalTop-level instructions; blocks may include cache controls
streamboolean; optionalfalse; true returns native SSE
output_configobject; optionalEffort defaults to medium; low, medium, high, xhigh, max; format selects structured JSON; beta extensions remain intact
thinkingobject; optionalAdaptive by default; display: "summarized" opts into a summary; type: "disabled" is documented at low, medium, or high effort
toolsarray; optionalNative definitions with name, input_schema, optional strict, and other official fields
tool_choiceobject; optionalSee the model-specific thinking and tool guidance below
stop_sequencesstring array; optionalStop generation on a matched sequence
cache_controlobject or null; optionaltype: "ephemeral"; TTL 5m or 1h; minimum cacheable prefix is 512 tokens
temperature, top_p, top_knumber; optionalForwarded; omit sampling settings in new integrations with this model
metadataobject; optionalRequest metadata
containerstring, object, or null; optionalApplicable container reuse and skill settings
context_managementobject or null; optionalContext edits; nested options are preserved
compactionobject or null; optionalOfficial on-demand compaction uses type: "summarize" and the compact-2026-09-04 beta header; the current probe returned ordinary text, without a signed compaction block
output_formatobject or null; optionalLegacy beta field; prefer current output_config.format for new integrations
mcp_serversarray; optionalOfficial MCP definitions and required beta headers; service availability depends on the upstream
diagnosticsobject or null; optionalIncludes previous_message_id; passed through
service_tierstring; optionalOfficial selection such as auto or standard_only; availability depends on the model
inference_geostring or null; optionalRequested inference location; passed through
speedstring or null; optionalPassed through; fast mode is not documented for Claude Haiku 5.5
fallbacksstring, array, or null; optionalOfficial beta options are retained; not behavior-tested here
fallback_credit_tokenstring, object, or null; optionalRetained unchanged; not behavior-tested here

Per-message controls such as output_config and clear_at, and nested beta thinking options, also pass through. Send the appropriate beta header when required by the official documentation. Inputs with image or PDF content must use public HTTP(S) URLs under reAPI's media policy; base64 and data URIs are not accepted. Plain-text document blocks can contain inline text.

Thinking and tools

Claude Haiku 5.5 defaults to medium effort and adaptive thinking. max_tokens includes thinking, so leave enough allowance for the final answer. Preserve the original assistant content array when sending a tool result back, including signatures and block ordering.

Haiku supports thinking: {"type":"disabled"} at low, medium, or high effort. It also supports tool_choice: {"type":"any"} and a named {"type":"tool","name":"..."}. Forced selection does not produce an initial thinking block.

Structured JSON

{
  "output_config": {
    "format": {
      "type": "json_schema",
      "schema": {
        "type": "object",
        "properties": {"answer": {"type": "string"}},
        "required": ["answer"],
        "additionalProperties": false
      }
    }
  }
}

Send the complete schema; reAPI does not remove constraints or create missing required fields. Document citations and structured JSON are separate features and cannot be combined in the same request under the official contract.

Observed capabilities

Small synthetic probes on October 9, 2026 established the following behavior. They are not a throughput benchmark or a maximum-context test.

FeatureObservation
Text and system instructionsCorrect response to the synthetic marker and system override
Native streamingNative message/content events and cumulative output usage observed
Structured JSONOutput matched the requested schema even when the prompt requested plain text
Strict client toolsTool name and schema-shaped input observed; the next response used the supplied tool result
Stop sequencesMatched stop reason; output after the marker was absent
Prompt cachingNonzero cache-creation usage, then nonzero cache-read usage on repeated input
Document citationsText-document answer contained a citation block
Effort levelsAll five official values accepted; relative reasoning depth was not measured
Thinking modeModel-specific mode accepted with no thinking block in that response; internal computation was not measured
URL imagesSingle-image and two-image requests returned 400 on the tested provider path; image interpretation was not reached
PDFNot yet verified with a permitted PDF fixture; the text-document citation test does not establish PDF support
Web searchBasic web_search_20250305 returned an actual search call, matching results and a citation
Web fetchBasic web_fetch_20250910 returned 400 on the tested provider path
Code executioncode_execution_20260521 returned 400 before execution
MCPBoth tested official MCP beta versions returned 400; the reachable synthetic server received no provider tool-discovery or tool-call callback
On-demand compactionRequest accepted with 200, but returned end_turn and no signed compaction block; effect unproven
Signed thinkingA signed block and unchanged-history replay were observed with empty transformation diagnostics. Edited-prefix controls returned explicit drop records, including when error was requested; strict error behavior is not established

These are dated observations of the tested path, not universal model limitations. An in-response tool error is distinct from HTTP transport success. Requests and extension fields still pass through; these observations do not introduce a local capability gate.

Other settings, including legacy thinking, sampling, and assistant prefill, were recorded as upstream response observations. Acceptance does not certify their effect and is not a reason for local parameter rejection.

Pricing dimensions and billing

Input, output, cache use, and applicable server tools can affect cost. Thinking counts toward output usage. See the model page for current rates; the response's native usage reports token categories separately.

For this endpoint, settlement uses the trusted per-request USD cost with the model's configured multiplier: billed_usd = reported_request_cost_usd × multiplier. It is token-based LLM billing, not an asynchronous image/video credit estimate. Failed delivery or missing trusted usage/cost follows the existing no-charge settlement path.

Output and streaming

A non-streaming response has the native shape below; values are illustrative.

{
  "id": "msg_example",
  "type": "message",
  "role": "assistant",
  "model": "claude-haiku-5-5",
  "content": [{"type": "text", "text": "The next step is to verify the input."}],
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {
    "input_tokens": 20,
    "output_tokens": 12,
    "cache_creation_input_tokens": 0,
    "cache_read_input_tokens": 0
  }
}

Native SSE uses message_start, content_block_start, content_block_delta, content_block_stop, message_delta, and message_stop. Handle ping and error events too. Thinking signatures and partial tool JSON remain in their native events; Messages does not end with [DONE].

A tool loop sends the previous full content as an assistant message and the matching tool_result as a user content block. Your application executes client tools; the playground does not execute them. Inspect stop_reason for incomplete answers, tool calls, and refusals even when HTTP status is 200.

Existing Chat Completions clients

POST https://reapi.ai/api/v1/chat/completions remains available with Bearer authentication. It maps reasoning_effort to output_config.effort, function tools to native definitions, stop to stop_sequences, and JSON schemas to output_config.format. Explicit native options take precedence. Use Messages for the official native request and response format.

The compatibility response retains choices, tool calls, and [DONE] streaming. Full native assistant content is available as choices[].message.anthropic_content; send it back unchanged as the assistant's anthropic_content. Streams carry anthropic_event so clients can reconstruct the original content blocks. Existing compatibility handling of OpenAI-only fields remains unchanged.

Errors

StatusMeaning
400Invalid request or an upstream-rejected parameter combination
401 / 403Authentication or permission failure
402Balance below admission requirements
413Request body larger than 32 MiB
429Rate or concurrency limit
502 / 503Service, pricing, or response delivery unavailable

Messages errors use {"type":"error","error":{"type":"invalid_request_error","message":"..."},"request_id":"..."}. See the error catalog. Unknown parameters do not trigger a local whitelist error.

Tips

  • Start with one task whose result you can check, then add context or tool steps.
  • Use output_config.format for schema-constrained JSON and explicit document citations for source-based answers.
  • Keep conversation content intact across tool turns; a displayed summary is not a replacement for signed blocks.
  • Use documented model controls as a starting point and evaluate effects on your own workload.

Official references, checked October 9, 2026: Messages, model overview, migration guidance, structured output, prompt caching.

Table of Contents