LiteLLM proxy routes between Ollama, vLLM, and cloud APIs from a single endpoint
LiteLLM is an open-source proxy that exposes a single OpenAI-compatible API endpoint and routes requests to 100+ model providers, including local Ollama and vLLM instances. It handles automatic fallback, load balancing, and per-key budget controls.
LiteLLM is an open-source proxy that sits between your applications and your model providers, exposing a single OpenAI-compatible API endpoint that routes to over 100 backends. The project has become a standard part of the self-hosted AI stack for anyone running multiple model providers.
The architecture is simple. Your apps, agents, scripts, and tools all point to http://litellm:4000/v1. LiteLLM translates the requests to the correct format for each backend: Ollama, vLLM, Anthropic, OpenAI, Azure, Google, or any other supported provider. Your applications never need to know which provider is handling any given request.
The most powerful feature is conditional routing. You configure multiple models under the same alias with different priorities in a YAML file. Simple requests go to a fast local model. Complex requests go to a larger cloud model. If one backend is down, LiteLLM falls back to the next one automatically. This gives you resilience without any application-level retry logic.
Budget controls let you set per-key spending limits. This is useful for teams where different projects or users should have different cost ceilings. The proxy tracks token usage across all providers and enforces the limits in real time. You can see exactly how much each key has consumed and set alerts when budgets approach their limits.
The setup is a single YAML configuration file. You define your models, API keys, routing rules, and fallback chains. A Docker image is available for quick deployment, or you can install it directly with pip. Most setups are running within 15 minutes of starting the configuration.
One critical caveat: CVE-2026-42208, a SQL injection vulnerability with a CVSS score of 9.3, was disclosed in April. The flaw allows unauthenticated attackers to read and modify the proxy database through crafted Authorization headers. This affects any LiteLLM proxy exposed to a network. Make sure you are running a patched version before deploying, and never expose the proxy admin interface to the public internet without additional authentication.
For anyone managing multiple model backends, LiteLLM solves a real operational problem. Instead of configuring each application to know about each provider's API format, endpoint, and authentication, you configure it once in LiteLLM and everything downstream just works.