claude-haiku-5-5
Use Claude Haiku 5.5 with native Messages, streaming, structured JSON, tool calls, document citations, and prompt caching on reAPI.
Call Claude Haiku 5.5 with the native Messages format; see the model page for current pricing.
Quick example
curl --no-buffer https://reapi.ai/api/v1/messages \
-H "x-api-key: YOUR_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{"model":"claude-haiku-5-5","max_tokens":4096,"messages":[{"role":"user","content":"Explain the next step clearly."}],"output_config":{"effort":"medium"},"stream":true}'import json
import urllib.request
payload = {
"model": "claude-haiku-5-5",
"max_tokens": 4096,
"messages": [{"role": "user", "content": "Explain the next step clearly."}],
"output_config": {"effort": "medium"},
}
request = urllib.request.Request(
"https://reapi.ai/api/v1/messages",
data=json.dumps(payload).encode(),
headers={"x-api-key": "YOUR_API_KEY", "anthropic-version": "2023-06-01",
"Content-Type": "application/json"},
method="POST",
)
with urllib.request.urlopen(request, timeout=120) as response:
message = json.load(response)
print(json.dumps(message, indent=2))const response = await fetch("https://reapi.ai/api/v1/messages", {
method: "POST",
headers: {
"x-api-key": "YOUR_API_KEY",
"anthropic-version": "2023-06-01",
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "claude-haiku-5-5",
max_tokens: 4096,
messages: [{ role: "user", content: "Explain the next step clearly." }],
output_config: { effort: "medium" },
}),
});
const body = await response.json();
if (!response.ok) throw new Error(JSON.stringify(body));
console.log(body);package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"net/http"
"time"
)
func main() {
body, err := json.Marshal(map[string]any{
"model": "claude-haiku-5-5",
"max_tokens": 4096,
"messages": []map[string]string{{"role": "user", "content": "Explain the next step clearly."}},
"output_config": map[string]string{"effort": "medium"},
})
if err != nil { panic(err) }
req, err := http.NewRequest("POST", "https://reapi.ai/api/v1/messages", bytes.NewReader(body))
if err != nil { panic(err) }
req.Header.Set("x-api-key", "YOUR_API_KEY")
req.Header.Set("anthropic-version", "2023-06-01")
req.Header.Set("Content-Type", "application/json")
client := &http.Client{Timeout: 120 * time.Second}
resp, err := client.Do(req)
if err != nil { panic(err) }
defer resp.Body.Close()
out, err := io.ReadAll(resp.Body)
if err != nil { panic(err) }
if resp.StatusCode >= 400 { panic(string(out)) }
fmt.Println(string(out))
}Endpoint and headers
POST https://reapi.ai/api/v1/messages
| Header | Requirement |
|---|---|
x-api-key | Your reAPI key; required unless using the existing Bearer authentication option |
Content-Type | application/json |
anthropic-version | Optional; defaults to 2023-06-01 |
anthropic-beta | Optional; forward the exact feature header required by the official documentation |
For an Anthropic SDK, use https://reapi.ai/api as the base URL; the SDK appends /v1/messages. Disable client-side automatic retries when one attempt per action is required. reAPI makes one upstream generation attempt and does not retry automatically.
Parameters
Native Messages fields and unknown extensions pass through unchanged. There is no field allowlist or parameter-count restriction. The upstream determines acceptance and behavior; HTTP 200 alone does not establish that a parameter took effect. Existing authentication, request-size, and billing admission checks still apply.
| Field | Type / required | Default and behavior |
|---|---|---|
model | string; required | claude-haiku-5-5 |
messages | non-empty array; required | Native conversation content, including tool results and signed thinking blocks |
max_tokens | integer; required | 0–128,000 including thinking; zero is forwarded for cache warmup |
system | string or text-block array; optional | Top-level instructions; blocks may include cache controls |
stream | boolean; optional | false; true returns native SSE |
output_config | object; optional | Effort defaults to medium; low, medium, high, xhigh, max; format selects structured JSON; beta extensions remain intact |
thinking | object; optional | Adaptive by default; display: "summarized" opts into a summary; type: "disabled" is documented at low, medium, or high effort |
tools | array; optional | Native definitions with name, input_schema, optional strict, and other official fields |
tool_choice | object; optional | See the model-specific thinking and tool guidance below |
stop_sequences | string array; optional | Stop generation on a matched sequence |
cache_control | object or null; optional | type: "ephemeral"; TTL 5m or 1h; minimum cacheable prefix is 512 tokens |
temperature, top_p, top_k | number; optional | Forwarded; omit sampling settings in new integrations with this model |
metadata | object; optional | Request metadata |
container | string, object, or null; optional | Applicable container reuse and skill settings |
context_management | object or null; optional | Context edits; nested options are preserved |
compaction | object or null; optional | Official on-demand compaction uses type: "summarize" and the compact-2026-09-04 beta header; the current probe returned ordinary text, without a signed compaction block |
output_format | object or null; optional | Legacy beta field; prefer current output_config.format for new integrations |
mcp_servers | array; optional | Official MCP definitions and required beta headers; service availability depends on the upstream |
diagnostics | object or null; optional | Includes previous_message_id; passed through |
service_tier | string; optional | Official selection such as auto or standard_only; availability depends on the model |
inference_geo | string or null; optional | Requested inference location; passed through |
speed | string or null; optional | Passed through; fast mode is not documented for Claude Haiku 5.5 |
fallbacks | string, array, or null; optional | Official beta options are retained; not behavior-tested here |
fallback_credit_token | string, object, or null; optional | Retained unchanged; not behavior-tested here |
Per-message controls such as output_config and clear_at, and nested beta thinking options, also pass through. Send the appropriate beta header when required by the official documentation. Inputs with image or PDF content must use public HTTP(S) URLs under reAPI's media policy; base64 and data URIs are not accepted. Plain-text document blocks can contain inline text.
Thinking and tools
Claude Haiku 5.5 defaults to medium effort and adaptive thinking. max_tokens includes thinking, so leave enough allowance for the final answer. Preserve the original assistant content array when sending a tool result back, including signatures and block ordering.
Haiku supports thinking: {"type":"disabled"} at low, medium, or high effort. It also supports tool_choice: {"type":"any"} and a named {"type":"tool","name":"..."}. Forced selection does not produce an initial thinking block.
Structured JSON
{
"output_config": {
"format": {
"type": "json_schema",
"schema": {
"type": "object",
"properties": {"answer": {"type": "string"}},
"required": ["answer"],
"additionalProperties": false
}
}
}
}Send the complete schema; reAPI does not remove constraints or create missing required fields. Document citations and structured JSON are separate features and cannot be combined in the same request under the official contract.
Observed capabilities
Small synthetic probes on October 9, 2026 established the following behavior. They are not a throughput benchmark or a maximum-context test.
| Feature | Observation |
|---|---|
| Text and system instructions | Correct response to the synthetic marker and system override |
| Native streaming | Native message/content events and cumulative output usage observed |
| Structured JSON | Output matched the requested schema even when the prompt requested plain text |
| Strict client tools | Tool name and schema-shaped input observed; the next response used the supplied tool result |
| Stop sequences | Matched stop reason; output after the marker was absent |
| Prompt caching | Nonzero cache-creation usage, then nonzero cache-read usage on repeated input |
| Document citations | Text-document answer contained a citation block |
| Effort levels | All five official values accepted; relative reasoning depth was not measured |
| Thinking mode | Model-specific mode accepted with no thinking block in that response; internal computation was not measured |
| URL images | Single-image and two-image requests returned 400 on the tested provider path; image interpretation was not reached |
| Not yet verified with a permitted PDF fixture; the text-document citation test does not establish PDF support | |
| Web search | Basic web_search_20250305 returned an actual search call, matching results and a citation |
| Web fetch | Basic web_fetch_20250910 returned 400 on the tested provider path |
| Code execution | code_execution_20260521 returned 400 before execution |
| MCP | Both tested official MCP beta versions returned 400; the reachable synthetic server received no provider tool-discovery or tool-call callback |
| On-demand compaction | Request accepted with 200, but returned end_turn and no signed compaction block; effect unproven |
| Signed thinking | A signed block and unchanged-history replay were observed with empty transformation diagnostics. Edited-prefix controls returned explicit drop records, including when error was requested; strict error behavior is not established |
These are dated observations of the tested path, not universal model limitations. An in-response tool error is distinct from HTTP transport success. Requests and extension fields still pass through; these observations do not introduce a local capability gate.
Other settings, including legacy thinking, sampling, and assistant prefill, were recorded as upstream response observations. Acceptance does not certify their effect and is not a reason for local parameter rejection.
Pricing dimensions and billing
Input, output, cache use, and applicable server tools can affect cost. Thinking counts toward output usage. See the model page for current rates; the response's native usage reports token categories separately.
For this endpoint, settlement uses the trusted per-request USD cost with the model's configured multiplier: billed_usd = reported_request_cost_usd × multiplier. It is token-based LLM billing, not an asynchronous image/video credit estimate. Failed delivery or missing trusted usage/cost follows the existing no-charge settlement path.
Output and streaming
A non-streaming response has the native shape below; values are illustrative.
{
"id": "msg_example",
"type": "message",
"role": "assistant",
"model": "claude-haiku-5-5",
"content": [{"type": "text", "text": "The next step is to verify the input."}],
"stop_reason": "end_turn",
"stop_sequence": null,
"usage": {
"input_tokens": 20,
"output_tokens": 12,
"cache_creation_input_tokens": 0,
"cache_read_input_tokens": 0
}
}Native SSE uses message_start, content_block_start, content_block_delta, content_block_stop, message_delta, and message_stop. Handle ping and error events too. Thinking signatures and partial tool JSON remain in their native events; Messages does not end with [DONE].
A tool loop sends the previous full content as an assistant message and the matching tool_result as a user content block. Your application executes client tools; the playground does not execute them. Inspect stop_reason for incomplete answers, tool calls, and refusals even when HTTP status is 200.
Existing Chat Completions clients
POST https://reapi.ai/api/v1/chat/completions remains available with Bearer authentication. It maps reasoning_effort to output_config.effort, function tools to native definitions, stop to stop_sequences, and JSON schemas to output_config.format. Explicit native options take precedence. Use Messages for the official native request and response format.
The compatibility response retains choices, tool calls, and [DONE] streaming. Full native assistant content is available as choices[].message.anthropic_content; send it back unchanged as the assistant's anthropic_content. Streams carry anthropic_event so clients can reconstruct the original content blocks. Existing compatibility handling of OpenAI-only fields remains unchanged.
Errors
| Status | Meaning |
|---|---|
| 400 | Invalid request or an upstream-rejected parameter combination |
| 401 / 403 | Authentication or permission failure |
| 402 | Balance below admission requirements |
| 413 | Request body larger than 32 MiB |
| 429 | Rate or concurrency limit |
| 502 / 503 | Service, pricing, or response delivery unavailable |
Messages errors use {"type":"error","error":{"type":"invalid_request_error","message":"..."},"request_id":"..."}. See the error catalog. Unknown parameters do not trigger a local whitelist error.
Tips
- Start with one task whose result you can check, then add context or tool steps.
- Use
output_config.formatfor schema-constrained JSON and explicit document citations for source-based answers. - Keep conversation content intact across tool turns; a displayed summary is not a replacement for signed blocks.
- Use documented model controls as a starting point and evaluate effects on your own workload.
Related
Official references, checked October 9, 2026: Messages, model overview, migration guidance, structured output, prompt caching.