| OpenRouterBroad hosted model and provider marketplace |
HostedManaged gatewayShared credits, optional BYOK, provider controls, and a large changing catalog. |
Several API shapesChat Completions, stateless Responses, Anthropic Messages, embeddings, and model-dependent media endpoints. |
Provider and model controlsOrder, restrict, sort, cap price, require parameters, and use pre-output provider or cross-model fallback. |
Usage cost + generation recordResponses include usage cost, while generation lookups add provider and model metadata. Key limits, guardrails, and activity exports are documented; workspace budgets are Enterprise-only, and already-dispatched requests can cause a slight overage. |
Configurable provider policyPrompt logging is opt-in at OpenRouter, while selected providers have their own terms. Routing can require no data collection or ZDR; opt-in metadata can identify providers and attempts. |
Free models or PAYG25+ free models at a daily request limit; paid credit purchases start at $5 and include a 5.5% platform fee. |
| Vercel AI GatewayHosted multi-provider gateway with Vercel and AI SDK integrations |
HostedManaged gatewayUsable from Vercel deployments or external runtimes through API keys, with Vercel OIDC, team billing, a public model catalog, and paid-tier BYOK. |
AI SDK + compatible APIsAI SDK, OpenAI Chat Completions, OpenAI Responses with documented store and previous_response_id, Anthropic Messages, OpenResponses, and Cohere Rerank. Capability varies by surface, model, and provider. |
Managed and policy-controlled routingProvider ordering or restriction; input-price, TTFT, or TPS sorting; ordered model fallbacks; and beta team routing rules. Custom provider timeouts are beta and BYOK-only; unsupported cancellation can still incur provider charges. |
Soft budgets + per-request evidenceTeam, project, and API-key budgets are checked before requests, but the crossing request completes and can slightly exceed the limit. BYOK provider spend is excluded. Per-request cost/fallback metadata, dashboards, alerts, and metered Custom Reporting are documented; a price-version-locked settlement record is not. |
Gateway deletion + provider filtersVercel documents deleting prompts and outputs after completion and retaining request details for 30 days; provider ZDR and no-training filters separately govern upstream routing. The Responses reference also exposes store for later retrieval, but the reviewed docs do not clarify its storage owner or retention period. |
$5/month free tierCredit begins with the first Gateway request and ends once the team purchases credits. Free credit covers eligible models at lower limits. Paid tokens use provider list rates; payment processing and metered controls can add fees. |
| LiteLLMOperator-controlled open-source proxy and commercial Enterprise layer |
Self-hostedYou operate the data planeOSS is MIT-licensed outside the enterprise directory; infrastructure and upstream charges remain yours. |
Translated multi-provider surfaceOpenAI-compatible proxy routes plus documented provider-native and pass-through endpoints. Support varies by provider. |
Operator-configured routerWeighted, rate-limit-aware, latency-based, least-busy, custom, and asynchronous lowest-cost strategies, with retries inside a model group and ordered fallbacks to other configured groups. |
Pre-dispatch checks on accounted spendPostgres-backed key, team, user, and customer budgets are checked before routing; per-call spend is recorded asynchronously after the response. Fail-closed enforcement is configurable, but an atomic monetary reservation for in-flight work is not documented. |
Operator-controlled, not provider-freeSelf-hosting sends no telemetry to LiteLLM-operated servers, but prompts still reach configured upstreams and enabled callbacks can export content or metadata. Database retention is the operator’s decision. |
OSS license fee: $0You cover infrastructure and any upstream-provider charges. Enterprise is an annual commercial license priced by gateway request capacity, deployment architecture, and support needs—not per token—with a documented 30-day trial. |
| Together AINamed-model infrastructure; opt-in passthrough is a separate data path |
Direct platformServerless or reserved infrastructureTogether-hosted serverless, batch, dedicated endpoints, and custom model deployments; separately enabled passthrough models forward traffic to an upstream provider. |
OpenAI-compatible inference APIsChat Completions, completions, embeddings, images, and audio endpoints are documented. The reviewed compatibility matrix does not expose /v1/responses. |
Application-owned cross-model routingFor serverless, the caller names a model and implements any cross-model selection or fallback. Dedicated model inference can split, A/B test, or shadow traffic across configured deployments. |
Per-call tokens; aggregate cost evidenceChat responses include a response ID, returned model, and token usage. Project keys support attribution, expiration, and revocation but not individual spend/rate caps. Cost analytics and invoices are documented; a per-call dollar settlement receipt is not. |
Default ZDR; opt-in passthroughTogether says inputs and outputs are not stored by default, though temporary caching may be used, and training sharing is opt-in. Passthrough requires enabling prompt/response storage and then follows the upstream provider’s data policy. Non-passthrough third-party-author models stay on Together infrastructure; serverless offers no region selection. |
$5 prepaid entryNo current free trial; platform access requires a minimum $5 purchased credit and a positive balance. Dedicated endpoints accrue per-minute charges while running. |
| Infer by Flow7Public self-service paid access |
Hosted APIManaged multi-format inference gatewayThese are routing options, not distinct models. Infer publishes only the price options that are currently callable for each model, so availability can differ by model and over time. “Stable” is not an SLA; it names a routing option and does not promise uptime. |
Three request formatsInfer exposes /v1/responses, OpenAI-compatible /v1/chat/completions, and Anthropic-compatible /v1/messages over one routing, wallet, cache, media, privacy, and receipt engine. Compatibility does not claim every optional feature of either upstream API. |
Model family + price optionInfer fixes the selected family, policy, and recorded price before dispatch, then applies private internal routing. Serving supplier identity is not disclosed. |
Reserve, then settleA conservative ceiling counts against the key before provider dispatch. Completed responses return measured usage, receipt ID, locked pricing record, and final customer charge. |
Private supplier, standard modeCurrently callable selectors advertise standard privacy only. Serving supplier identity remains private, standard responses may be retained briefly for replay, and no current ZDR or no-training claim is made. |
$20 first fundingLater reloads start at $50. Wallet credit funds metered live API requests. |