
Who Litellm is for#
Central LLM access for an engineering organization
Run the proxy once, issue virtual keys per team with budgets, and let every developer use their usual OpenAI client against one endpoint.
Skip if:
You are a single developer calling one provider, where the provider's own SDK is simpler.
Multi-provider apps with fallbacks
Use the SDK and Router to spread traffic across deployments and fail over automatically when a provider rate limits or errors.
Skip if:
Your application will only ever use a single model from a single provider.
Routing coding agents and MCP tools
Point agents and editors at the gateway so model calls and MCP tool access share one set of keys, logs, and budgets.
Skip if:
You do not want to operate any infrastructure and prefer a managed service.
The problem it solves#
Calling several LLM providers means juggling different SDKs, authentication schemes, request formats, and error types. Once more than one team shares those models, the harder questions arrive: who holds the provider keys, who is allowed to call which model, what each team spent, and what happens when a provider rate limits or goes down. Hosted gateways answer those questions by sitting between your traffic and the providers. If you would rather keep keys, logs, and routing rules on your own infrastructure, you need a gateway you can run yourself.
How it solves it#
One OpenAI-format interface for 100+ providers
Call OpenAI, Anthropic, Vertex AI, Bedrock, Azure, Ollama, and many more through the same completion() call. Responses follow the OpenAI Chat Completions format, and provider errors map to OpenAI exception types.
Router with retries, fallbacks, and load balancing
Spread requests across multiple deployments and set automatic fallbacks, so one rate-limited or failing deployment does not take your application down.
Self-hosted proxy with virtual keys and budgets
The proxy is an OpenAI-compatible gateway. Issue virtual keys with per-key, team, and user budgets and rate limits, and track spend across every provider in one place.
Guardrails, caching, and logging
Add content filtering, PII masking, and safety checks at the gateway, cache responses, and send input and output to observability tools such as Langfuse, MLflow, and Helicone.
MCP and A2A gateway
Expose MCP servers through a central endpoint with per-key access control, and add A2A agents to the same gateway, so models, tools, and agents share one entry point.
Strengths and trade-offs#
Strengths
- Provider switching without code rewritesBecause every response uses the OpenAI format, changing the model string is usually the whole migration, and existing OpenAI clients work against the proxy unchanged.
- Cost and access control in one placeVirtual keys, budgets, and spend tracking per key, team, and user give platform teams a single place to see and limit LLM usage.
- Two ways to adopt itStart with the Python SDK inside one application, then move to the shared proxy when a team needs central keys and logging. A leaner litellm-core package exists for SDK-only use.
- Stated gateway overheadThe README reports 8ms P95 latency at 1k requests per second and links benchmarks, so you can check the figure against your own load.
Trade-offs
- -Not plain MIT across the whole repositoryThe non-enterprise code is MIT-licensed, and you can use, modify, deploy, and redistribute it freely, including commercially. Content under the enterprise/ directory is governed by a separate license that is not included in the repository, so check those terms before relying on anything from that folder. GitHub reports no SPDX license for the repo.
- -You operate the gatewaySelf-hosting means you run, secure, upgrade, and monitor the proxy and its supporting services. A hosted gateway removes that work; this one does not.
- -Large and fast-moving projectThe repository lists over 5,000 open issues, and the docs show packaging changes in flight, such as the separate litellm-core distribution that cannot be installed alongside litellm. Pin versions and test upgrades.
- -Some features sit in the enterprise tierThe docs list SSO/SAML, audit logs, and advanced security under an Enterprise offering, so teams with those needs should confirm what the open source tier covers.
Litellm vs alternatives#
LiteLLM vs Portkey
Portkey is a commercial AI gateway. LiteLLM covers similar ground in an open source package that you host: a unified OpenAI-format API, virtual keys with budgets, guardrails, caching, load balancing, and observability callbacks. The practical difference is where it runs. With LiteLLM, the proxy and its logs live on your infrastructure and you carry the operational work. Check Portkey's current feature list and pricing directly before deciding, since this page only describes LiteLLM from its own docs.
LiteLLM vs OpenRouter
OpenRouter is a commercial service for reaching many models through one API. LiteLLM gives you the same single-interface idea, but you bring your own provider accounts and keys, and the gateway runs where you put it. That means you pay providers directly and keep routing rules under your control, in exchange for running the service yourself. LiteLLM also adds team budgets, an admin UI, and MCP and A2A gateways.
Which to pick
Choose LiteLLM when key custody, self-hosting, and per-team spend control matter more than avoiding operations work. Choose a hosted gateway when you want someone else to run the infrastructure.
Quick start#
Self-hosting the proxy needs Python with uv, or Docker plus a litellm_config.yaml for your models.
```bash
uv tool install 'litellm[proxy]'
litellm --model gpt-4o
docker run -v $(pwd)/litellm_config.yaml:/app/config.yaml -p 4000:4000 docker.litellm.ai/berriai/litellm:latest --config /app/config.yaml
```What it's built on#
- Languages
- GoPythonRustTypeScript
- Frameworks
- LangChainNext.jsReact
- Tooling
- esbuild
FAQ#
What is LiteLLM?
LiteLLM is an open source AI gateway that lets you call 100+ LLM providers using the OpenAI format. You can use it as a Python SDK inside your code or deploy the proxy server as a shared gateway for a team.
Is LiteLLM free to self-host?
Yes. The non-enterprise portions are MIT-licensed, so you can use, modify, deploy, and redistribute them for any purpose, including commercial production use. Content under the enterprise/ directory is governed by a separate license that is not included in the repository, so review it before using that code.
How do I start the proxy?
Install it with uv tool install 'litellm[proxy]' and run litellm with a model, or run the Docker image with a litellm_config.yaml mounted. The proxy listens on port 4000 and accepts any OpenAI client.
Can it handle MCP servers and agents?
Yes. The gateway can host a central MCP endpoint with per-key access control and can add A2A agents, alongside the 100+ LLM providers.
Similar open-source tools#
OmniRoute
352-provider AI gateway with automatic fallback and token compression
freellmapi
One OpenAI-compatible key for 635 free LLM endpoints
sub2api
One API gateway for Claude, OpenAI, Gemini, and Grok subscriptions
Switchyard
LLM proxy with API translation and multi-backend routing
rtk
CLI proxy that removes 89% of agent context noise on average
9Router
Smart AI Router with 3-Tier Fallback

