toolcompass.

Find your next AI tool

Search by product name or task. Press Escape to close.

vLLM pricing

● Identity checked

LLM inference engine · Vendor site 2026-10-04

vLLM is an open-source inference and serving engine for large language models focused on high throughput and memory-efficient deployment. The project ships with an OpenAI-compatible API, PagedAttention scheduling, and documentation for CUDA, ROCm, CPU, and Docker installs.

Updated 2026-10-04View sources
Official website
Category
LLM inference engine
Free access
Open-source engine; homepage invites anyone to deploy and serve models locally or on their hardware
API access
Documentation highlights a drop-in OpenAI-compatible API for integration

Understand the total cost

vLLM pricing & plans

Free

Free

Open-source engine; homepage invites anyone to deploy and serve models locally or on their hardware

Access
Open-source engine; homepage invites anyone to deploy and serve models locally or on their hardware

Price history

No retained pricing changes yet. A current price alone does not establish a historical trend.

How this profile is supported

Facts apply to the named version and check date. Send a sourced correction if something changed.

vLLM pricing questions

Is vLLM free?

Yes. vLLM has an ongoing free plan. Open-source engine; homepage invites anyone to deploy and serve models locally or on their hardware

How much does vLLM cost?

vLLM is listed as free. No paid plan with a published price was found.

Does vLLM have an API?

API or developer access is mentioned. Documentation highlights a drop-in OpenAI-compatible API for integration