Checking Infer routes The comparison is dated; current Infer availability refreshes separately. Read current status

Decision guide · checked 11 August 2026

Choose the operating model before the gateway.

OpenRouter, Vercel AI Gateway, LiteLLM, Together AI, and Infer solve different infrastructure jobs. This sheet compares current first-party documentation and standard online terms without assigning a universal winner.

products compared
5
source policy
first-party only
research date
2026-08-11
Infer now
checking status…

Method

Six questions, asked in the same order.

Vendor documentation describes published behavior; only applicable terms or negotiated agreements create legal commitments. Dynamic counts and prices are snapshots dated 11 August 2026; unverified features are omitted.

01Who operates it?Hosted marketplace, hosted multi-provider gateway, self-hosted proxy, direct model platform, or hosted paid API.
02Which API surface?Responses, Chat Completions, Messages, multimodal endpoints, or a translated provider surface.
03Who controls routing?Managed provider selection, application-owned fallback, or an operator-configured router.
04When is spend enforced?After-accounted budgets, in-flight caps, prepaid balance, or a reservation before provider dispatch.
05What evidence returns?Usage metadata, reports, invoices, routing traces, or a completed-response final-charge record.
06Where does data go?Serving-provider disclosure, retention, training, upstream terms, and operator-controlled logging.

Product × requirement

The decision matrix.

“Not documented” means the reviewed first-party sources did not establish the behavior. It is not proof that the product can never provide it.

Scroll horizontally to compare every column.

AI gateway and direct model-platform decision matrix checked 11 August 2026
ProductOperating modelPublished API emphasisRouting and fallbackSpend and request evidenceData and provider identityCommercial entry
OpenRouterBroad hosted model and provider marketplace HostedManaged gatewayShared credits, optional BYOK, provider controls, and a large changing catalog. Several API shapesChat Completions, stateless Responses, Anthropic Messages, embeddings, and model-dependent media endpoints. Provider and model controlsOrder, restrict, sort, cap price, require parameters, and use pre-output provider or cross-model fallback. Usage cost + generation recordResponses include usage cost, while generation lookups add provider and model metadata. Key limits, guardrails, and activity exports are documented; workspace budgets are Enterprise-only, and already-dispatched requests can cause a slight overage. Configurable provider policyPrompt logging is opt-in at OpenRouter, while selected providers have their own terms. Routing can require no data collection or ZDR; opt-in metadata can identify providers and attempts. Free models or PAYG25+ free models at a daily request limit; paid credit purchases start at $5 and include a 5.5% platform fee.
Vercel AI GatewayHosted multi-provider gateway with Vercel and AI SDK integrations HostedManaged gatewayUsable from Vercel deployments or external runtimes through API keys, with Vercel OIDC, team billing, a public model catalog, and paid-tier BYOK. AI SDK + compatible APIsAI SDK, OpenAI Chat Completions, OpenAI Responses with documented store and previous_response_id, Anthropic Messages, OpenResponses, and Cohere Rerank. Capability varies by surface, model, and provider. Managed and policy-controlled routingProvider ordering or restriction; input-price, TTFT, or TPS sorting; ordered model fallbacks; and beta team routing rules. Custom provider timeouts are beta and BYOK-only; unsupported cancellation can still incur provider charges. Soft budgets + per-request evidenceTeam, project, and API-key budgets are checked before requests, but the crossing request completes and can slightly exceed the limit. BYOK provider spend is excluded. Per-request cost/fallback metadata, dashboards, alerts, and metered Custom Reporting are documented; a price-version-locked settlement record is not. Gateway deletion + provider filtersVercel documents deleting prompts and outputs after completion and retaining request details for 30 days; provider ZDR and no-training filters separately govern upstream routing. The Responses reference also exposes store for later retrieval, but the reviewed docs do not clarify its storage owner or retention period. $5/month free tierCredit begins with the first Gateway request and ends once the team purchases credits. Free credit covers eligible models at lower limits. Paid tokens use provider list rates; payment processing and metered controls can add fees.
LiteLLMOperator-controlled open-source proxy and commercial Enterprise layer Self-hostedYou operate the data planeOSS is MIT-licensed outside the enterprise directory; infrastructure and upstream charges remain yours. Translated multi-provider surfaceOpenAI-compatible proxy routes plus documented provider-native and pass-through endpoints. Support varies by provider. Operator-configured routerWeighted, rate-limit-aware, latency-based, least-busy, custom, and asynchronous lowest-cost strategies, with retries inside a model group and ordered fallbacks to other configured groups. Pre-dispatch checks on accounted spendPostgres-backed key, team, user, and customer budgets are checked before routing; per-call spend is recorded asynchronously after the response. Fail-closed enforcement is configurable, but an atomic monetary reservation for in-flight work is not documented. Operator-controlled, not provider-freeSelf-hosting sends no telemetry to LiteLLM-operated servers, but prompts still reach configured upstreams and enabled callbacks can export content or metadata. Database retention is the operator’s decision. OSS license fee: $0You cover infrastructure and any upstream-provider charges. Enterprise is an annual commercial license priced by gateway request capacity, deployment architecture, and support needs—not per token—with a documented 30-day trial.
Together AINamed-model infrastructure; opt-in passthrough is a separate data path Direct platformServerless or reserved infrastructureTogether-hosted serverless, batch, dedicated endpoints, and custom model deployments; separately enabled passthrough models forward traffic to an upstream provider. OpenAI-compatible inference APIsChat Completions, completions, embeddings, images, and audio endpoints are documented. The reviewed compatibility matrix does not expose /v1/responses. Application-owned cross-model routingFor serverless, the caller names a model and implements any cross-model selection or fallback. Dedicated model inference can split, A/B test, or shadow traffic across configured deployments. Per-call tokens; aggregate cost evidenceChat responses include a response ID, returned model, and token usage. Project keys support attribution, expiration, and revocation but not individual spend/rate caps. Cost analytics and invoices are documented; a per-call dollar settlement receipt is not. Default ZDR; opt-in passthroughTogether says inputs and outputs are not stored by default, though temporary caching may be used, and training sharing is opt-in. Passthrough requires enabling prompt/response storage and then follows the upstream provider’s data policy. Non-passthrough third-party-author models stay on Together infrastructure; serverless offers no region selection. $5 prepaid entryNo current free trial; platform access requires a minimum $5 purchased credit and a positive balance. Dedicated endpoints accrue per-minute charges while running.
Infer by Flow7Public self-service paid access Hosted APIManaged Responses gatewayTwenty-two model entries each expose Low Cost, Balanced, and Stable policy selectors in this dated snapshot. These are routing options, not distinct models; “Stable” is not an SLA. Responses onlyInfer’s public OpenAPI document exposes GET /v1/models and POST /v1/responses. Chat Completions and full OpenAI parity are not claimed. Model family + price optionInfer fixes the selected family, policy, and price version before dispatch, then applies private internal routing. Serving supplier identity is not disclosed. Reserve, then settleA conservative ceiling counts against the key before provider dispatch. Completed responses return measured usage, receipt ID, locked price version, and final customer charge. Private supplier, standard modeCurrently callable selectors advertise standard privacy only. Serving supplier identity remains private, standard responses may be retained briefly for replay, and no current ZDR or no-training claim is made. $20 first fundingLater reloads start at $50. Wallet credit funds metered live API requests.

Counts, availability, and pricing are snapshots, not guarantees. OpenRouter and Vercel catalog totals do not establish identical capability across every model. Infer’s Low Cost, Balanced, and Stable choices are routing policies; 17 separate Official API records were unavailable in the dated snapshot.

Choose by fit

Use the product whose operating model matches the job.

Each recommendation below deliberately sends some buyers elsewhere. The point is to shorten the decision, not stretch Infer into every workload.

OPENROUTER

Choose breadth and provider-level routing.

Use OpenRouter when model/provider switching, BYOK, provider policy filters, multimodal endpoints, and marketplace breadth matter—and stateless Responses is sufficient.

  • Responses is documented as stateless.
  • OpenRouter returns per-response usage cost and supports generation-record lookups. The reviewed documentation does not describe a price-version-locked or signed settlement receipt.
  • Standard Terms restrict competing resale; Infer does not present OpenRouter as an authorized upstream.
Read provider routing ↗ · Read data policy ↗ · Read pricing ↗
VERCEL

Choose a hosted multi-provider gateway.

Use Vercel AI Gateway when a managed service callable from any runtime, native AI SDK and Vercel integration, several compatible API shapes, provider and model failover, team policy controls, and consolidated observability and billing reduce operational work.

  • OpenAI Responses is currently labeled beta, with documented store and previous_response_id; support still varies by model and provider.
  • BYOK is paid-tier only: Vercel tries customer credentials first, documents fallback to system credentials on failure, requires purchased Gateway credits, and excludes BYOK provider spend from Gateway budgets.
  • Budgets are soft caps checked before requests; the crossing request completes and can produce a small overrun.
  • The AI Product Terms exclude AI Products and Services from Vercel’s SLA, restrict sensitive personal information in inputs, and make provider terms relevant. Public API resale or service-bureau use needs contract-specific review under the base Terms.
Read SDKs and APIs ↗ · Read BYOK ↗ · Read pricing ↗
LITELLM

Choose control and accept operations.

Use LiteLLM when your team can operate the gateway, Postgres, credentials, routing policy, and telemetry, and wants broad provider translation under its own control.

  • Documented routing strategies and fallbacks are operator-configured.
  • Budgets are checked before routing and can fail closed against authoritative recorded spend, but the docs do not describe a per-request monetary reservation for in-flight work.
  • A standalone hosted LiteLLM Cloud offer was not verifiable in the current public material.
Read routing ↗ · Read budget controls ↗ · Read pricing ↗
TOGETHER

Choose direct named-model infrastructure.

Use Together AI when Chat Completions compatibility is sufficient and direct named-model serverless, batch, dedicated, or custom deployment is the requirement.

  • For serverless, the application chooses the model and implements any cross-model routing or fallback.
  • Project keys do not have individual spend or rate caps.
  • The May 19, 2026 standard Terms prohibit using or accessing the Services to develop a competing product, conduct competitive analysis or benchmarking, or resell or offer the Services standalone; confirm any gateway-supply use under the applicable agreement.
Read inference options ↗ · Read credits ↗ · Read Terms ↗
INFER

Choose reserve-before-dispatch spend control.

Use Infer when a Responses-only integration is acceptable and the buying requirement is model-family selection, an explicit price option, a conservative reservation against the key before provider dispatch, and a final-charge record in every completed response.

  • Public availability is dynamic and is not an SLA or uptime guarantee.
  • Serving suppliers remain private; a receipt does not attest supplier identity or model weights.
  • SSE is currently assembled after the upstream response completes, so first-token streaming latency is not claimed.
Read the public API description → · Read the bounded reservation record →

Evidence boundary

What this guide does not prove.

The rows summarize current first-party documentation and Infer’s own product evidence. They are not a benchmark, legal opinion, endorsement, or future-availability promise.

No speed rankingNo controlled cross-product latency or throughput test was run. Status-page history and vendor fleet reports are not comparable benchmarks.
No lowest-price claimProvider prices, platform fees, credits, infrastructure, premium controls, cache behavior, and fallback outcomes make one headline price incomplete.
No universal parityA compatible endpoint does not establish every parameter, hosted tool, stateful workflow, streaming behavior, or provider-native feature.
No supplier inferenceA model creator, catalog label, or family name does not necessarily identify the serving provider. Read each product’s routing metadata and terms.
No availability overclaimCurrent availability is point-in-time and is not an SLA or uptime guarantee.
Selected primary sourcesOfficial pages accessed 11 August 2026. These are selected references, not a complete source ledger. Legal summaries concern standard online terms; separately negotiated agreements may differ.

Need prepaid enforcement and per-completed-response charge evidence?

Check the selected route, inspect the public API description, and start with one capped request. Choose another product when breadth, self-hosting, or direct model infrastructure is the requirement.