Pricing
Not listed
Official monthly price not confirmed
- API
- Not confirmed
Assistant · Vendor site 2026-10-04
LLMKube is an open-source Kubernetes operator for self-hosted LLM inference across NVIDIA, Apple Silicon, and AMD hardware, supporting runtimes such as vLLM, llama.cpp, and mlx-server with CLI deployment, autoscaling, GPU sharding, and Grafana metrics.
Platform and ML teams that want Kubernetes-native orchestration for local or private-cloud LLM inference instead of per-token cloud APIs.
You supply clusters, GPUs, and operational expertise; LLMKube does not publish a hosted SaaS price list.
LLMKube targets teams scaling beyond single-machine Docker setups by adding HPA autoscaling, inference metrics, and pluggable LLM runtimes under Kubernetes control.
What it can do
Unknown is different from unavailable. Each fact carries its own evidence.
| Capability | Value | Evidence | Checked |
|---|---|---|---|
| Kubernetes operator | Homepage markets LLMKube as a Kubernetes operator for self-hosted inference with vLLM, llama.cpp, and related runtimes. | Facts sourced | 2026-10-04 |
| CLI deploy | Site shows llmkube CLI commands to deploy named models with selectable GPU runtimes. | Facts sourced | 2026-10-04 |
| Hardware support | Header badges list Kubernetes native support for NVIDIA, Apple Silicon, and AMD accelerators. | Facts sourced | 2026-10-04 |
Understand the total cost
Pricing
Not listed
Official monthly price not confirmed
Stakpak is an open-source autonomous agent that runs on your machines to keep applications healthy, respond to incidents, and report through chat channels. Paid cloud tiers add managed LLM credits, multi-app autopilot, and team billing.
Explore toolRiver is an AI desktop for project work where teams collaborate with AI across documents, notes, sheets, slides, files, and browser connections. It combines native workspace apps with model access and usage-based AI credits instead of chat-only interfaces.
Explore toolOrkas lets solo founders and small teams direct a roster of AI agents by chat to produce research, documents, slides, images, videos, websites, and operational workflows. The desktop app ships free and open source for Mac and Windows, with optional membership credits for managed frontier models, collaboration, and cloud sync.
Explore toolThe practical questions
LLMKube is an open-source Kubernetes operator for self-hosted LLM inference across NVIDIA, Apple Silicon, and AMD hardware, supporting runtimes such as vLLM, llama.cpp, and mlx-server with CLI deployment, autoscaling, GPU sharding, and Grafana metrics.
Project is described as open source on the homepage with public documentation and GitHub access.. This record does not confirm an ongoing free plan.
A listed monthly price has not been confirmed. See the plan cards for entitlements, billing commitments and seat minimums.
Not confirmed. API access and subscription access may have different terms; consult the linked sources.
You supply clusters, GPUs, and operational expertise; LLMKube does not publish a hosted SaaS price list.
No retained pricing changes yet. A current price alone does not establish a historical trend.
Reviewed vendor source
Read original source ↗Facts apply to the named version and check date. Send a sourced correction if something changed.