Back to vendors
L

LiteLLM

Also known as: BerriAI LiteLLM, LiteLLM Proxy, LiteLLM AI Gateway

Visit site
Entry priceOpen source free to self-host; Enterprise annual, quoted by salesFull pricing detail

Open source AI gateway and Python SDK: one OpenAI compatible API to 140+ model providers, plus MCP and agent gateways, keys, budgets, routing and self-hosting.

LiteLLM, built by BerriAI, is an open source AI gateway that puts a team's models, MCP servers and agents behind one OpenAI compatible API and one login. It ships in two forms that share a name: a Python SDK that runs inside application code, and a proxy server, the gateway, deployed as a central service for a team or organization. The core is MIT licensed and free to self-host.

The gateway reaches more than 140 providers and 1,800 models, and the model behind a request is set in the customer's configuration, so teams swap models without changing application code. Around that sit virtual keys, per key, team and project budgets that stop requests at the cap, rate limits, load balancing, fallbacks, caching, spend tracking for chargeback, and request logging that routes to tools such as Langfuse, LangSmith, Arize and Prometheus.

The MCP Gateway gives agents one endpoint for every registered tool server with per team access and OAuth, the A2A Agent Gateway registers and proxies agents built on Azure AI Foundry, Vertex AI Agent Engine, Bedrock AgentCore, LangGraph and Pydantic AI, and newer endpoints add a per user and per team memory store and a document ingestion pipeline into provider vector stores. The Auto Router, labeled an add-on, sends each prompt to a model tier and can be tested against live traffic with shadow evaluations before a switch.

Because it is self-hosted, LiteLLM runs in the customer's own cloud, on premises or fully air gapped, and its data and keys stay in that environment. A commercial Enterprise license, priced through sales on annual request capacity rather than per token, adds SSO and SCIM (SSO is free for up to five users), JWT authentication, audit logs, secret managers, organization and team admins, multi-region deployment and support SLAs, and LiteLLM states SOC 2 Type II with the report in its Trust Center. The tradeoff is ownership: the software is free, while running the database, cache and upgrades, which ship roughly weekly, sits with the adopting team.

Vendor details

Canonical URL

https://www.litellm.ai

Category

Agent infrastructure

Subcategory

LLM gateway and routing

Funding status

Independent, built by BerriAI (legal entity Berrie AI Incorporated). MIT licensed open source core with a commercial Enterprise license sold alongside it.

Company status

independent

Use cases & customers

Primary use cases

model access and routingspend tracking and budgetsself hosted gatewayprovider failoverMCP and agent gateway

Target customers

developersplatform teamsenterprise

Deployment options

self-hostedVPCon-premair-gapped

Integrations

One OpenAI compatible interface to more than 140 providers as an SDK or proxy. MCP Gateway for registered tool servers with per key and team access and OAuth; A2A Agent Gateway for Azure AI Foundry, Vertex AI Agent Engine, Bedrock AgentCore, LangGraph and Pydantic AI agents. Logs route to Langfuse, LangSmith, Arize, Prometheus and GCS or Azure Blob. Procurement through AWS Marketplace and authorized resellers.

In practice

Your platform team wants one gateway across every provider, fully in your own cloud. You deploy the MIT licensed LiteLLM proxy, issue virtual keys per squad, and enforce per team budgets that hard stop when exceeded.

You have air gapped or strict residency requirements that rule out SaaS. LiteLLM self hosts entirely in your environment, so provider keys and request logs never leave infrastructure you control.

Your CISO needs single sign on, role based access, and audit logs across LLM traffic. You add the LiteLLM Enterprise license on top of the open source proxy to unlock governance without changing the deployment model.

Agentic Index coverage score

9.5 / 14 capabilities · 68%

Integrations & Tool Calling Full

LiteLLM's MCP Gateway gives agents one fixed endpoint for every registered MCP server, with access controlled per key and team and server authentication documented for OAuth, OAuth passthrough, on-behalf-of, AWS SigV4 and JWT signer schemes. On the Responses and Chat Completions endpoints the proxy fetches the MCP server's tools, and when configured to, executes the returned tool calls itself. That is an authenticated action framework, so agents take authenticated action in outside systems through the gateway rather than handing a suggestion back.

Sourcedocs.litellm.ai/docs/mcpread 2026-09-22

Workflow Orchestration Partial

With an MCP tool's require_approval set to never, the proxy executes the model's tool calls and feeds the results back into the model before returning the answer, so the gateway runs a multi-step agent loop itself, and the A2A Agent Gateway routes calls to registered agents. No workflow definition, deterministic node or versioned flow is documented, so sequencing, branching and retries that mix deterministic steps with agent steps are not covered.

Sourcedocs.litellm.ai/docs/mcpread 2026-09-22

Knowledge Grounding & RAG Full

The /rag/ingest endpoint is an all in one ingestion pipeline (upload, chunk, embed, write to a vector store) into OpenAI vector stores, Bedrock Knowledge Bases, Vertex AI RAG Engine, Gemini or AWS S3 Vectors, and /rag/query searches the ingested content and generates a response from it; vector store access is permissioned per team in the gateway. New customer knowledge enters a persistent index without retraining, so the retrieval structure over the customer's knowledge is maintained, persists, scales past the context window and stays queryable.

Sourcedocs.litellm.ai/docs/rag_ingestread 2026-09-22

Human Oversight & Guardrails Partial

Guardrails run before, during or after a call and can be scoped per key and team (PII masking, prompt injection and secret detection, content moderation and third-party guardrail providers), and the MCP Gateway adds per-key and per-team tool permissions and guardrails on tool results. These are constraints the customer controls on what an agent may do.

When an MCP tool's require_approval is set to anything other than never, the proxy returns the tool calls to the client so they can be reviewed and executed manually; that review happens in the customer's own client. LiteLLM itself documents no surface where a person reviews and approves an agent action before it commits.

Sourcedocs.litellm.ai/docs/mcpread 2026-09-22

Security, Identity & Governance Full

Access control is documented on the Enterprise page: SSO for the Admin UI through Okta, Azure AD, Google Workspace or any OIDC or SAML provider (free for up to five users, an Enterprise license beyond that), SCIM, JWT authentication against the customer's own identity provider, role-based access control across organizations, teams and user roles, IP allowlists, public and private route controls, key rotation and external secret managers.

LiteLLM's Data Privacy and Security page (docs.litellm.ai/docs/data_security) states SOC 2 Type II, with the current report available through the LiteLLM Trust Center. That pairs a named access model with identity integration and an attestation, and audit logs with retention policies add to it. Most controls sit in the Enterprise tier.

Sourcedocs.litellm.ai/docs/enterpriseread 2026-09-22

Observability & Auditability Full

The gateway records what each caller's agent did through it: request and response content for model and agent calls with user, key and team attribution, latency and cost (the A2A Agent Gateway page shows this in the Logs tab for invoked agents), and Prometheus metrics in the open source core.

The Enterprise page adds per-key or per-team log routing to Langfuse, LangSmith, Arize and other callbacks, audit logs of admin actions with retention policies kept apart from request logs, and log export to GCS or Azure Blob. Traffic that passes through the gateway can be inspected step by step, with audit logs kept apart from runtime traces.

Sourcedocs.litellm.ai/docs/a2aread 2026-09-22

Memory & State Persistence Partial

The proxy's Memory Management API (/v1/memory, LiteLLM v1.83.10 or later with PostgreSQL connected) keeps entries across sessions, scoped per user and team with role-based read and write rules, and supports create, read, update, list by key prefix and delete, so a buyer can say where memory lives (the gateway's own Postgres) and delete one user's entries without touching the rest. The application reads entries and places them in the prompt. The scope is stated, but no lifetime is.

Sourcedocs.litellm.ai/docs/proxy/memoryread 2026-09-22

Deployment & Data Residency Full

Deployment is self hosted. The homepage documents official Docker images, a Helm chart and a Terraform module, running on the customer's own Postgres and Redis, one-click deploy into AWS, GCP or Azure, and fully air-gapped installation. The pricing page FAQ (litellm.ai/pricing) confirms air-gapped deployment on Enterprise and states that the customer's data and keys never leave its own infrastructure, and the Enterprise docs page adds multi-region deployment under one license with an admin and worker split. The customer can run it in its own cloud, on premises or fully air gapped.

Sourcelitellm.airead 2026-09-22

Prebuilt Agents, Templates & Packs Not documented

No agents of LiteLLM's own are offered for a buyer to adopt. The AI Hub on the Enterprise page publishes a page of the models, agents, MCP servers and skills the customer has registered, which is a directory of the customer's own assets, and the Google AI Studio managed agents LiteLLM supports live entirely on Google's side, with LiteLLM as the auth and routing layer. Nothing there is a ready made workflow, template or role specific agent; a model or tool catalog is integration, not a pack.

Sourcedocs.litellm.ai/docs/enterpriseread 2026-09-22

Triggers & Channel Coverage Not documented

Work reaches agents through LiteLLM when a caller sends a request: model, MCP and A2A calls are all invoked by the client. No schedule, event, webhook or inbound queue that starts an agent run is documented; budget alerts notify people rather than wake an agent. A gateway that only answers callers has no way to wake an agent on its own.

Sourcedocs.litellm.ai/docs/a2aread 2026-09-22

Model Flexibility & Routing Full

The customer chooses. The homepage documents one OpenAI-compatible API to 140+ providers and 1,800+ models, with models set in the customer's own configuration and swapped without changing application code, the customer's internal, fine-tuned and self-hosted models behind the same key, load balancing across providers, regions and keys, lowest-cost routing, and an Auto Router that sends prompts to model tiers the customer configures.

Sourcelitellm.airead 2026-09-22

APIs, SDKs & MCP Extensibility Full

A Python SDK and a proxy with a documented OpenAI compatible REST API ship with LiteLLM, and the gateway's own features are callable from outside. The A2A Agent Gateway serves a proxied agent card for each registered agent, pinned to A2A 0.3 or 1.0, and agents are invoked through the A2A SDK or the OpenAI SDK; the Memory Management and /rag/ingest pages document REST endpoints (/v1/memory, /v1/rag/ingest) with curl and Python examples.

Sourcedocs.litellm.ai/docs/a2aread 2026-09-22

Testing, Debugging & Optimization Full

Shadow evaluations sample a key's, team's or user's live traffic, send each sampled request through a candidate router configuration without returning that answer to the client, and have an LLM judge compare it blind against the answer the current model served; a job runs up to 30 days and can compare several configurations on the same traffic before anything changes, and after a switch each request carries its routing decision and savings. A change is evaluated against the customer's own traffic, with a judge verdict.

Sourcedocs.litellm.ai/docs/auto_router/evaluateread 2026-09-22

Browser & Computer Use Not documented

The product is a gateway for models, MCP servers and agents; no browser, desktop or remote computer session that LiteLLM runs for an agent is documented. The sandboxes the Managed Agents Platform announcement describes, and the blog's swap of OpenAI's Code Interpreter for E2B or OpenSandbox, are code execution, not control of a real interface that an agent drives itself.

Sourcelitellm.airead 2026-09-22

The Agentic Index coverage score grades every vendor Full, Partial or Not documented against the same 14 buyer facing capabilities, from public evidence only. Each capability links to how all vendors in the index score on it. How this evidence is graded

Recent platform changes

2026-10-01·Observability / auditabilityPartially Verified

Lens, new from LiteLLM, uses AI agents to review the traces passing through the LiteLLM gateway, group recurring failures and link each finding to the exact step where it happened. Investigations run on demand or on a schedule, and the traces stay in a store the customer runs beside the gateway.

Bears on: Observability / auditability

View source
2026-09-25·MCP / tool calling / APIVerified

LiteLLM launched the public preview of LiteAgents, an SDK that supports Deep Agents, Pydantic AI, Claude Agent SDK, Codex and OpenCode through a common query interface. Profiles select the runtime while retaining shared tools, MCP connections and model configuration; native runtime options remain available. Optional Temporal integration adds recovery from recorded operations and checkpoints, while simple local agents require neither Temporal nor PostgreSQL. Changing runtimes starts new runs and does not migrate an active native session.

Bears on: Workflow orchestration

View source
2026-09-14·Workflow orchestrationVerified

LiteLLM v1.101.0 adds a heuristic auto router, semantic search over MCP tools and support for off peak pricing. It also restricts agent and vector store listings for non admin users to the resources they have been granted.

Bears on: Workflow orchestration

View source
View all 9 changes for LiteLLM →Tracked since Jul 2026 · Verified from public vendor sources

Pricing

Open source free to self-host; Enterprise annual, quoted by sales

license and self hosted infrastructure

Free tierTrial available

Included quota

The MIT licensed open source proxy and SDK are free with no usage fees, log limits, or request quotas imposed by LiteLLM; you fund only infrastructure and pay providers directly. LiteLLM Enterprise is a self hosted commercial license adding governance features.

What is public

The open source tier is published at $0, self-hosted. Enterprise has no published price: the pricing page says it is sized to annual gateway request capacity, deployment architecture and support needs, never per token, with volume discounts.

Billing mechanics

Open source: no license fee; the proxy runs on the customer's infrastructure and providers bill the customer directly. Enterprise: a license key unlocks governance features on the same self-hosted deployment, bought direct, through AWS Marketplace or authorized resellers.

Cost watchouts

The customer runs and pays for the infrastructure (Postgres, Redis, the proxy fleet) and pays model providers directly. LiteLLM ships a new minor line roughly weekly and supports only the four most recent stable lines, so staying supported means a regular upgrade cadence. Standard Enterprise support has no guaranteed response time; 24/7 SLAs cost extra.

Variable cost rationale

The software license is fixed or zero, but total cost is dominated by self hosted infrastructure, engineering operations, and direct provider spend, all of which scale with usage.

Overage / add-ons

No usage metering on the open source core. Enterprise is sold by annual request capacity with volume discount tiers; provider and infrastructure costs are the customer's own.

Sales call required

Mixed (some tiers require a call)

Free / trial

MIT open source core free to self-host; 30-day Enterprise trial key, no credit card

Lowest paid plan

Enterprise, annual, priced through sales by gateway request capacity, deployment architecture and support needs; no published figure

Commercial notes

Independent, built by BerriAI (Y Combinator). More than 40,000 GitHub stars. Widely adopted as the default open source LLM gateway.

Key ambiguities

Enterprise figures are quoted per deployment; the Auto Router add-on's price is not published.

Support SLA / resale

Standard support included with Enterprise (weekday Slack or Teams channel, no guaranteed response time); 24/7 SLAs for an additional fee

Agentic Index verified 2026-09-22

Alternatives to LiteLLM

The closest documented capability profiles to LiteLLM among agent infrastructure platforms tracked by Agentic Index, ordered by similarity on the same 14 point evidence the rankings use. No vendor pays for placement.

  • Kong8.5 / 14A lighter documented profile than LiteLLM
  • Arcade7.5 / 14A lighter documented profile than LiteLLM
  • Cartesia10.5 / 14Adds documented Prebuilt Agents, Templates & Packs and Triggers & Channel Coverage
  • Inworld AI8.5 / 14Adds documented Triggers & Channel Coverage
  • Portkey7.5 / 14A lighter documented profile than LiteLLMLiteLLM vs Portkey →
  • Haystack12.0 / 14Adds documented Prebuilt Agents, Templates & Packs and Triggers & Channel Coverage

Similarity is computed from each vendor's Agentic Index coverage score evidence, axis by axis, not from the totals. How this evidence is graded

Contact us

Found a vendor we missed? Have feedback on the index? We'd love to hear from you.