Free
Free
Open-source documentation and install guides are published without a paid tier on the project site.
- Access
- Open-source documentation and install guides are published without a paid tier on the project site.
Code · Vendor site 2026-10-04
vllm-mlx brings OpenAI- and Anthropic-compatible LLM inference to Apple Silicon using MLX. The project documents server deployment, continuous batching, multimodal models, embeddings, tool calling, and MCP tool support.
Developers running local or self-hosted LLM inference on Mac Apple Silicon who want vLLM-like serving features.
Performance and supported models depend on your Mac hardware and the MLX model registry you enable.
vllm-mlx documents how to install and run compatible inference servers and Python APIs optimized for Apple Silicon hardware.
What it can do
Unknown is different from unavailable. Each fact carries its own evidence.
| Capability | Value | Evidence | Checked |
|---|---|---|---|
| Apple Silicon inference | Homepage tagline promises OpenAI- and Anthropic-compatible inference on Apple Silicon. | Facts sourced | 2026-10-04 |
| Serving features | Navigation lists server mode, continuous batching, multimodal audio/video, embeddings, and tool calling. | Facts sourced | 2026-10-04 |
Understand the total cost
Free
Free
Open-source documentation and install guides are published without a paid tier on the project site.
It supports local and cloud LLMs, codebase indexing on paid tiers, and a free Basic plan with your own API keys.
Explore toolBolt.new is a browser-based AI app builder where users describe a web or mobile app in chat and get a running project with hosting, databases and custom domains on paid tiers.
Explore toolCleric is an AI SRE agent platform that follows production changes, detects regressions, investigates incidents, and proposes fixes while change context is still fresh. Teams interact in Slack or the web app and pay in credits for completed investigative work.
Explore toolThe practical questions
vllm-mlx brings OpenAI- and Anthropic-compatible LLM inference to Apple Silicon using MLX. The project documents server deployment, continuous batching, multimodal models, embeddings, tool calling, and MCP tool support.
Open-source documentation and install guides are published without a paid tier on the project site.. This record lists ongoing free access; check the plan limits before starting.
No paid monthly price is listed; this record treats the product as free to start. See the plan cards for entitlements, billing commitments and seat minimums.
Docs describe a Python API, CLI, and HTTP server with OpenAI- and Anthropic-compatible endpoints.. API access and subscription access may have different terms; consult the linked sources.
Performance and supported models depend on your Mac hardware and the MLX model registry you enable.
No retained pricing changes yet. A current price alone does not establish a historical trend.
Reviewed vendor source
Read original source ↗Facts apply to the named version and check date. Send a sourced correction if something changed.