Free
Free
Open-source engine; homepage invites anyone to deploy and serve models locally or on their hardware
- Access
- Open-source engine; homepage invites anyone to deploy and serve models locally or on their hardware
LLM inference engine · Vendor site 2026-10-04
vLLM is an open-source inference and serving engine for large language models focused on high throughput and memory-efficient deployment. The project ships with an OpenAI-compatible API, PagedAttention scheduling, and documentation for CUDA, ROCm, CPU, and Docker installs.
Understand the total cost
Free
Free
Open-source engine; homepage invites anyone to deploy and serve models locally or on their hardware
No retained pricing changes yet. A current price alone does not establish a historical trend.
Reviewed vendor source
Read original source ↗Facts apply to the named version and check date. Send a sourced correction if something changed.
Yes. vLLM has an ongoing free plan. Open-source engine; homepage invites anyone to deploy and serve models locally or on their hardware
vLLM is listed as free. No paid plan with a published price was found.
API or developer access is mentioned. Documentation highlights a drop-in OpenAI-compatible API for integration