Infer API · live
One API. Four routing options.
Infer provides Responses-compatible inference plus native image and video generation with private supply, locked prices, spending limits, a prepaid balance, and clear receipts.
Infer Setup connects Codex, Claude Code, OpenCode, and more on macOS, Windows, and Linux. The open-source CLI keeps your existing settings and checks the connection without paid model calls.
The compatibility record covers exact Codex CLI and OpenCode versions plus receipt-backed live requests for Pydantic AI, LangChain OpenAI, and the Vercel AI SDK, while keeping setup, protocol observations, and current availability separate.
The reviewed OpenAPI 3.1 description covers the public catalog/status resources, authenticated text-inference endpoints, and native image/video generation endpoints. Its publication is not a route-availability signal.
https://infer.flow7.org/v1
Quickstart
Go from a new account to the first API request:
- Create an account, then confirm the one-time verification link sent by Infer.
- Open API keys, create an API key, and copy it when shown. Set the copied value as
INFER_API_KEYin your local shell or secret manager; never put it in a URL, repository, or chat. - Open Wallet and add at least $20 of wallet credit through checkout before sending metered inference.
- Send the Responses request below. The selected model must show as available on current status.
curl https://infer.flow7.org/v1/responses \
-H "Authorization: Bearer $INFER_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: first-request" \
-d '{
"model": "infer/gpt-5.6-terra:low-cost",
"input": "Create a typed retry helper with exponential backoff.",
"max_output_tokens": 1200,
"relay": {
"session_id": "repo-acme-42",
"privacy": "standard"
}
}'Price options
| Price option | Model suffix | Routing policy |
|---|---|---|
| Low cost | :low-cost | Lowest-cost qualified routing |
| Balanced | :balanced | Health-qualified routing, lowest landed cost first |
| Stable | :stable | Stricter health-qualified routing, lowest landed cost first |
| Official API | :official | Exact first-party provider route through Infer |
Infer publishes a price option only when at least one route currently passes that option’s model, health, privacy, capability, authorization, and margin checks. The live catalog is authoritative for which selectors are callable now.
Model IDs
A model ID combines the model name and price option. Exact model IDs stay within that model family. Dynamic model IDs return the model that actually ran in the response metadata.
| Model ID | Type | Price option |
|---|---|---|
infer/gpt-5.6-terra:low-cost | Exact family | Low cost |
infer/gpt-5.6-sol:balanced | Exact family | Balanced |
infer/gpt-5.6-sol:stable | Exact family | Stable |
infer/gpt-5.6-sol:official | Exact family | Official API |
API formats
Infer exposes three compatible text-inference request formats over one serving engine, plus native generation endpoints when those models appear in the live catalog. /v1/responses is the canonical Infer/OpenAI Responses surface, /v1/chat/completions accepts OpenAI Chat Completions clients, and /v1/messages accepts Anthropic Messages clients. All three use the same model selectors, wallet, spend limits, routing, privacy checks, prompt-cache accounting, media admission, receipts, and supplier privacy.
| Endpoint | Client format | Authentication |
|---|---|---|
/v1/responses | OpenAI Responses | Bearer Infer API key |
/v1/chat/completions | OpenAI Chat Completions | Bearer Infer API key |
/v1/messages | Anthropic Messages | x-api-key or Bearer Infer API key |
/v1/images/generations / /v1/images/edits | Native Images API | Bearer Infer API key |
/v1/videos | Native asynchronous Video API | Bearer Infer API key |
All three streaming modes forward upstream text, reasoning/thinking, and tool-call deltas as they arrive. Infer emits keepalives during quiet reasoning periods, does not impose a generation wall-clock timeout while the connection remains alive, and settles final usage, cache accounting, customer charge, and receipt metadata in the terminal stream event. Disconnecting the client cancels the upstream stream and reconciles the wallet reservation rather than leaving a long-running orphan request.
Responses API
Infer returns the normal model response plus the selected price option, final charge, token usage, and receipt under relay. The serving supply identity is never returned.
{
"id": "resp_…",
"object": "response",
"status": "completed",
"model": "infer/gpt-5.6-terra:low-cost",
"output_text": "…",
"usage": {
"input_tokens": 8240,
"input_tokens_details": {"cached_tokens": 7100, "cache_write_tokens": 240},
"output_tokens": 932
},
"relay": {
"receipt_id": "rcpt_…",
"price_version": "cloud-…",
"tier": "low-cost",
"resolved_model_class": "GPT-5.6 Terra",
"cache_status": "partial",
"environment": "live",
"provider_disclosed": false,
"customer_cost_usd": 0.004992
}
}Chat Completions compatibility
Existing OpenAI Chat Completions clients can point their base URL at https://infer.flow7.org/v1 and use an Infer model selector. Infer normalizes chat messages, function tools, tool results, structured-output requests, image parts, input audio, and file parts into its canonical request before routing.
curl https://infer.flow7.org/v1/chat/completions \
-H "Authorization: Bearer $INFER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "infer/gpt-5.6-terra:stable",
"messages": [{"role": "user", "content": "Return exactly CHAT_OK"}],
"max_completion_tokens": 32
}'stream: true emits chat.completion.chunk SSE and ends with data: [DONE]. stream_options.include_usage is honored. Infer currently supports one choice per request (n=1); audio output, prediction hints, web-search options, and returned logprobs are not part of this compatibility layer.
Anthropic Messages compatibility
The Anthropic SDK can use Infer by setting its base URL to https://infer.flow7.org, passing the Infer API key as x-api-key, and using an Infer model selector. The normal anthropic-version header is accepted. Text, images, documents/PDFs, custom function tools, tool_use/tool_result history, prompt-cache blocks, system prompts, and Anthropic-style streaming events are normalized through the same Infer engine.
curl https://infer.flow7.org/v1/messages \
-H "x-api-key: $INFER_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "infer/claude-haiku-4-5:official",
"max_tokens": 32,
"messages": [{"role": "user", "content": "Return exactly MESSAGES_OK"}]
}'stream: true emits Anthropic-compatible message_start, content_block_*, message_delta, and message_stop events. MCP-server and container request features are not currently exposed through Infer's Messages compatibility endpoint.
Native image and video generation
Generation-native models appear only when the live catalog has an authenticated, healthy route and a locked customer price. They currently use the Stable selector because the generation supply is admitted under Infer’s stricter Stable policy. Generation is billed by generated media—not by token rates.
| Model | Endpoint | Billing unit |
|---|---|---|
infer/gpt-image-2:stable | /v1/images/generations and /v1/images/edits | Per generated image, by quality and size |
infer/seedance-2.0:stable | /v1/videos | Per generated second, by resolution |
infer/seedance-2.5:stable | /v1/videos | Per generated second, by resolution |
Image generation is synchronous and returns either a short-lived Infer-signed URL or b64_json. Image edits accept JSON public-HTTPS references or one multipart PNG/JPEG/WebP upload up to 20 MiB. stream: true is intentionally rejected until Infer can preserve the upstream terminal event and billing contract without guessing.
curl https://infer.flow7.org/v1/images/generations \
-H "Authorization: Bearer $INFER_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: image-001" \
-d '{"model":"infer/gpt-image-2:stable","prompt":"A chrome flower on black glass","quality":"medium","size":"1024x1024"}'Video creation returns 202 with an Infer-owned task ID. Poll /v1/videos/{id}; after completion, download through /v1/videos/{id}/content. Infer also reconciles pending video tasks every five minutes, so terminal settlement does not depend on the client continuing to poll.
curl https://infer.flow7.org/v1/videos \
-H "Authorization: Bearer $INFER_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: video-001" \
-d '{"model":"infer/seedance-2.5:stable","prompt":"Slow cinematic dolly through mist","duration":6,"resolution":"720p"}'The full generation ceiling is reserved before dispatch. Infer never automatically retries a paid generation. If a post-dispatch transport outcome is ambiguous, the reservation remains locked rather than assuming the provider did or did not spend. Completed generation receipts expose the customer unit, variant, quantity, and final customer charge while keeping supplier identity and supplier cost private.
How requests are routed
- The API key type, model ID, routing option, and price are locked before the request starts.
- Infer checks model support, context size, privacy needs, current health, media capabilities, and price before choosing a route.
- When a client supplies
relay.session_id, the session stays on the same route while that route remains healthy. On supported native Responses routes, Infer derives a private organization-scoped sticky cache key from that session to maximize provider prompt-cache reuse without exposing the raw session ID. - Retries stay within the selected option’s price and reliability limits.
- The supplier and exact route remain private.
Prompt caching
Prompt cache reads and writes are first-class metered usage. usage.input_tokens_details.cached_tokens reports cache reads and cache_write_tokens reports tokens written to cache. Receipts preserve both buckets and their locked rates. For supported Claude Official routes, Infer enables ephemeral prompt caching automatically for standard and no-training requests; explicit cache_control and prompt_cache_key are also accepted when the selected route supports them. Zero-retention requests do not auto-create prompt-cache directives and reject explicit cache controls.
Multimodal input
Infer passes native Responses image, file/PDF, and audio objects through unchanged when the selected route advertises the corresponding modality. Media is never translated into a text-only protocol or silently discarded.
{
"model": "infer/gpt-5.6-sol:official",
"input": [{
"type": "message",
"role": "user",
"content": [
{"type": "input_text", "text": "What is in this image?"},
{"type": "input_image", "image_url": "https://provider.example/image.png"}
]
}],
"relay": {"session_id": "vision-session-1", "privacy": "standard"}
}The live catalog is authoritative: use a selector whose tier runtime reports vision, files, audio, or video for that input type. File/PDF inputs use input_file, audio uses input_audio, and video uses Infer's input_video extension. Chat Completions clients may send video_url content. Infer chooses the verified upstream wire format per capability, so video-capable Official routes may use Chat Completions internally while image/file/audio remain on native Responses. Anthropic Messages does not define a video content block and therefore does not expose video input.
Billing and receipts
Customers fund a prepaid Infer service-credit wallet through secure checkout with no separate Infer checkout fee. The first funding minimum is $20 and later reloads are $50. Applicable tax may be added separately. Wallet credit becomes available only after payment is confirmed. Before a request starts, Infer temporarily holds the maximum estimated cost. After it finishes, Infer charges the actual amount and immediately returns the rest to the wallet. Inspect a completed receipt locally.
| Field | Meaning |
|---|---|
input_tokens | Total input processed, including cache read/write buckets |
cached_input_tokens | Input served from prompt cache |
cache_write_tokens | Input written into prompt cache |
fresh_input_tokens | Input remaining after cache read/write buckets |
image_inputs | Image input parts observed in the admitted request |
output_tokens | Generated output |
price_version | Recorded pricing code used by the request |
customer_cost_usd | Final charge |
cache_status distinguishes miss, write, partial, hit, and an explicitly enabled whole-response response_hit. Prompt caching does not imply whole-response caching; Infer does not globally enable whole-response caching because that would change normal stochastic response semantics.
Errors
| Code | Meaning |
|---|---|
invalid_api_key | Missing, revoked, or invalid key |
insufficient_credits | The wallet cannot cover the estimated request cost |
capacity_unavailable | No working route supports the selected option |
request_in_progress | The idempotency key is already active |
request_failed_use_new_idempotency_key | The prior operation failed; retry the operation with a new idempotency key |
upstream_unavailable | Available routes failed before the request completed |