Both sit in the model provider layer with similar commercial terms, so the practical split is language: LiteLLM targets Python and vLLM targets Python.
Self-hosted proxy that speaks one API to 100+ providers, with keys and budgets.
Give every agent its own virtual key so you can attribute and cap spend per surface.
High-throughput open-weight serving on your own GPUs.
| LiteLLM | vLLM | |
|---|---|---|
| Licence | Open source | Open source |
| Pricing | open source | open source |
| Self-hostable | Yes | Yes |
| Languages | Python, TypeScript | Python |
| Install | pip install 'litellm[proxy]' | pip install vllm |
| Keys required | None | None |
Both sit in the model provider layer with similar commercial terms, so the practical split is language: LiteLLM targets Python and vLLM targets Python. LiteLLM: Self-hosted proxy that speaks one API to 100+ providers, with keys and budgets. vLLM: High-throughput open-weight serving on your own GPUs.
LiteLLM can run on your own infrastructure. vLLM can run on your own infrastructure.
LiteLLM is open source (open source). vLLM is open source (open source).
Every page here answers to Accept: text/markdown and returns the same content at roughly a tenth the tokens. No separate site, no toggle — same URL.
curl -s -H "Accept: text/markdown" https://newagent.build/compare/litellm-vs-vllm