The deciding difference is operational: vLLM can run on infrastructure you control, Google Vertex AI cannot. If that constraint is real for you, it settles the choice before anything else is considered.
Claude and Gemini under GCP billing, IAM and regional controls.
High-throughput open-weight serving on your own GPUs.
| Google Vertex AI | vLLM | |
|---|---|---|
| Licence | Proprietary | Open source |
| Pricing | usage based | open source |
| Self-hostable | No | Yes |
| Languages | TypeScript, Python | Python |
| Install | — | pip install vllm |
| Keys required | GOOGLE_CLOUD_PROJECT, GOOGLE_APPLICATION_CREDENTIALS | None |
The deciding difference is operational: vLLM can run on infrastructure you control, Google Vertex AI cannot. If that constraint is real for you, it settles the choice before anything else is considered. Google Vertex AI: Claude and Gemini under GCP billing, IAM and regional controls. vLLM: High-throughput open-weight serving on your own GPUs.
Google Vertex AI is a hosted service only. vLLM can run on your own infrastructure.
Google Vertex AI is proprietary (usage based). vLLM is open source (open source).
Every page here answers to Accept: text/markdown and returns the same content at roughly a tenth the tokens. No separate site, no toggle — same URL.
curl -s -H "Accept: text/markdown" https://newagent.build/compare/vertex-vs-vllm