| OpenRouterBroad hosted model and provider marketplace |
HostedManaged gatewayShared credits, optional BYOK, provider controls, and a large changing catalog. |
Several API shapesChat Completions, stateless Responses, Anthropic Messages, embeddings, and model-dependent media endpoints. |
Provider and model controlsOrder, restrict, sort, cap price, require parameters, and use pre-output provider or cross-model fallback. |
Usage cost + generation recordResponses include usage cost, while generation lookups add provider and model metadata. Key limits, guardrails, and activity exports are documented; workspace budgets are Enterprise-only, and already-dispatched requests can cause a slight overage. |
Configurable provider policyPrompt logging is opt-in at OpenRouter, while selected providers have their own terms. Routing can require no data collection or ZDR; opt-in metadata can identify providers and attempts. |
Free models or PAYG25+ free models at a daily request limit; paid credit purchases start at $5 and include a 5.5% platform fee. |
| Vercel AI GatewayHosted multi-provider gateway with Vercel and AI SDK integrations |
HostedManaged gatewayUsable from Vercel deployments or external runtimes through API keys, with Vercel OIDC, team billing, a public model catalog, and paid-tier BYOK. |
AI SDK + compatible APIsAI SDK, OpenAI Chat Completions, OpenAI Responses with documented store and previous_response_id, Anthropic Messages, OpenResponses, and Cohere Rerank. Capability varies by surface, model, and provider. |
Managed and policy-controlled routingProvider ordering or restriction; input-price, TTFT, or TPS sorting; ordered model fallbacks; and beta team routing rules. Custom provider timeouts are beta and BYOK-only; unsupported cancellation can still incur provider charges. |
Soft budgets + per-request evidenceTeam, project, and API-key budgets are checked before requests, but the crossing request completes and can slightly exceed the limit. BYOK provider spend is excluded. Per-request cost/fallback metadata, dashboards, alerts, and metered Custom Reporting are documented; a price-version-locked settlement record is not. |
Gateway deletion + provider filtersVercel documents deleting prompts and outputs after completion and retaining request details for 30 days; provider ZDR and no-training filters separately govern upstream routing. The Responses reference also exposes store for later retrieval, but the reviewed docs do not clarify its storage owner or retention period. |
$5/month free tierCredit begins with the first Gateway request and ends once the team purchases credits. Free credit covers eligible models at lower limits. Paid tokens use provider list rates; payment processing and metered controls can add fees. |
| LiteLLMOperator-controlled open-source proxy and commercial Enterprise layer |
Self-hostedYou operate the data planeOSS is MIT-licensed outside the enterprise directory; infrastructure and upstream charges remain yours. |
Translated multi-provider surfaceOpenAI-compatible proxy routes plus documented provider-native and pass-through endpoints. Support varies by provider. |
Operator-configured routerWeighted, rate-limit-aware, latency-based, least-busy, custom, and asynchronous lowest-cost strategies, with retries inside a model group and ordered fallbacks to other configured groups. |
Pre-dispatch checks on accounted spendPostgres-backed key, team, user, and customer budgets are checked before routing; per-call spend is recorded asynchronously after the response. Fail-closed enforcement is configurable, but an atomic monetary reservation for in-flight work is not documented. |
Operator-controlled, not provider-freeSelf-hosting sends no telemetry to LiteLLM-operated servers, but prompts still reach configured upstreams and enabled callbacks can export content or metadata. Database retention is the operator’s decision. |
OSS license fee: $0You cover infrastructure and any upstream-provider charges. Enterprise is an annual commercial license priced by gateway request capacity, deployment architecture, and support needs—not per token—with a documented 30-day trial. |
| Together AINamed-model infrastructure; opt-in passthrough is a separate data path |
Direct platformServerless or reserved infrastructureTogether-hosted serverless, batch, dedicated endpoints, and custom model deployments; separately enabled passthrough models forward traffic to an upstream provider. |
OpenAI-compatible inference APIsChat Completions, completions, embeddings, images, and audio endpoints are documented. The reviewed compatibility matrix does not expose /v1/responses. |
Application-owned cross-model routingFor serverless, the caller names a model and implements any cross-model selection or fallback. Dedicated model inference can split, A/B test, or shadow traffic across configured deployments. |
Per-call tokens; aggregate cost evidenceChat responses include a response ID, returned model, and token usage. Project keys support attribution, expiration, and revocation but not individual spend/rate caps. Cost analytics and invoices are documented; a per-call dollar settlement receipt is not. |
Default ZDR; opt-in passthroughTogether says inputs and outputs are not stored by default, though temporary caching may be used, and training sharing is opt-in. Passthrough requires enabling prompt/response storage and then follows the upstream provider’s data policy. Non-passthrough third-party-author models stay on Together infrastructure; serverless offers no region selection. |
$5 prepaid entryNo current free trial; platform access requires a minimum $5 purchased credit and a positive balance. Dedicated endpoints accrue per-minute charges while running. |
| Infer by Flow7Public self-service paid access |
Hosted APIManaged Responses gatewayTwenty-two model entries each expose Low Cost, Balanced, and Stable policy selectors in this dated snapshot. These are routing options, not distinct models; “Stable” is not an SLA. |
Responses onlyInfer’s public OpenAPI document exposes GET /v1/models and POST /v1/responses. Chat Completions and full OpenAI parity are not claimed. |
Model family + price optionInfer fixes the selected family, policy, and price version before dispatch, then applies private internal routing. Serving supplier identity is not disclosed. |
Reserve, then settleA conservative ceiling counts against the key before provider dispatch. Completed responses return measured usage, receipt ID, locked price version, and final customer charge. |
Private supplier, standard modeCurrently callable selectors advertise standard privacy only. Serving supplier identity remains private, standard responses may be retained briefly for replay, and no current ZDR or no-training claim is made. |
$20 first fundingLater reloads start at $50. Wallet credit funds metered live API requests. |