API Reference
Every endpoint, parameter and error the API returns.
Base URL
https://reroute.wtf/api/v1Every response is JSON. Errors share one shape: an error object with a numeric code equal to the HTTP status and a human-readable message.
Chat completions
POST /api/v1/chat/completions — OpenAI Chat Completions compatible. Requires a bearer key.
Request body
| Field | Type | Description |
|---|---|---|
messages | array, required | Conversation so far. Must be non-empty, otherwise 400. |
model | string | Catalog model ID. Omitted or reroute/auto uses openai/gpt-oss-20b. |
stream | boolean | Server-sent events when true. Default false. |
any other field | — | Forwarded to the carrier unchanged: temperature, max_tokens, top_p, tools, response_format, stop, … Check a model's supported_parameters. |
POST https://reroute.wtf/api/v1/chat/completions
Authorization: Bearer <REROUTE_API_KEY>
Content-Type: application/json
{
"model": "anthropic/claude-sonnet-4.5",
"messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "user", "content": "Explain request routing in one line." }
],
"temperature": 0.7,
"max_tokens": 256,
"stream": false
}Response
The upstream completion, with id replaced by an Reroute generation ID, model set to the model you asked for, provider naming the carrier that served it, and usage.cost set to the USD deducted from your balance.
{
"id": "gen-3y8Rq0vZ1kXbT2mN4pLc",
"object": "chat.completion",
"created": 1791411439,
"model": "anthropic/claude-sonnet-4.5",
"provider": "Anthropic",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Moving goods from where they are to where they're needed." },
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 24,
"completion_tokens": 14,
"total_tokens": 38,
"cost": 0.000282
}
}Streaming
With stream: true the response is text/event-stream. Each data: line is a chunk with the same id, model and provider. The last chunk before data: [DONE] carries usage when the carrier reports it. The generation ID is also sent as the X-Generation-Id header so you can look the request up before the stream ends.
HTTP/1.1 200 OK
Content-Type: text/event-stream; charset=utf-8
X-Generation-Id: gen-3y8Rq0vZ1kXbT2mN4pLc
data: {"id":"gen-3y8Rq0vZ1kXbT2mN4pLc","model":"anthropic/claude-sonnet-4.5","provider":"Anthropic","choices":[{"index":0,"delta":{"content":"Moving"}}]}
data: {"id":"gen-3y8Rq0vZ1kXbT2mN4pLc","model":"anthropic/claude-sonnet-4.5","provider":"Anthropic","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":24,"completion_tokens":14}}
data: [DONE]Errors
{
"error": {
"code": 402,
"message": "Insufficient credits. Add more using https://reroute.wtf/settings/credits"
}
}| Status | When |
|---|---|
| 400 | Body is not JSON; messages missing or empty; model ID not in the catalog. |
| 401 | No key, unknown key, or the key is disabled. |
| 402 | Paid model with a balance at or below zero, or the key has reached its credit limit. |
| 404 | Generation, model or route not found. |
| 429 | The carrier rate-limited the request. Retry with backoff. |
| 502 | The carrier returned an error or did not respond. The message names the carrier's reason when it gives one. |
| 503 | No carrier can serve this model right now. Nothing is charged. |
Requests that fail are recorded in your activity log with a cost of $0.
Get a generation
GET /api/v1/generation?id=<gen-id> — metadata for one of your own requests. Latency is time to first token in milliseconds; generation_time is the full request duration.
curl "https://reroute.wtf/api/v1/generation?id=gen-3y8Rq0vZ1kXbT2mN4pLc" \
-H "Authorization: Bearer $REROUTE_API_KEY"
{
"data": {
"id": "gen-3y8Rq0vZ1kXbT2mN4pLc",
"model": "anthropic/claude-sonnet-4.5",
"provider_name": "Anthropic",
"app": "Your App",
"streamed": true,
"created_at": "2026-10-08T00:17:19.148Z",
"latency": 612,
"generation_time": 1840,
"tokens_prompt": 24,
"tokens_completion": 14,
"native_tokens_reasoning": 0,
"total_cost": 0.000282,
"finish_reason": "stop",
"status": 200
}
}Get current key
GET /api/v1/key — label, usage, limit, remaining limit, free-tier flag and disabled flag for the key you authenticate with. See Authentication.
Get credits
GET /api/v1/credits — returns { "data": { "total_credits": number, "total_usage": number } } in USD for the key's account.
List models
GET /api/v1/models — public. The full catalog. GET /api/v1/models/<author>/<slug>/endpoints — public. The carriers serving one model, with pricing, context, quantization and uptime.
List providers
GET /api/v1/providers — public. Every provider with headquarters, datacenters and policy links.