Pricing
Not listed
Official monthly price not confirmed
- API
- Not confirmed
Code · Vendor site 2026-10-04
OneTriangle builds a KV cache-aware inference engine that prefills on a smaller model and decodes on a larger one. By transferring key-value cache between models, it targets lower prefill cost and faster first responses on long inputs without sacrificing decode quality when pairs pass quality gates.
Teams running multi-model LLM inference who want to cut redundant prefill work on long prompts.
Cache transfer only applies to vetted model pairs; ordinary prefill runs when a pair fails quality or latency gates.
OneTriangle optimizes inference by running prefill on a compact model, transferring KV cache into a larger decoder, and falling back to standard prefill when a pair does not meet held-out quality checks.
What it can do
Unknown is different from unavailable. Each fact carries its own evidence.
| Capability | Value | Evidence | Checked |
|---|---|---|---|
| KV cache transfer | Homepage explains projecting small-model K/V tensors into the target model attention space before decode. | Facts sourced | 2026-10-04 |
| Cost focus | Marketing claims roughly 20% inference cost and time savings on long inputs via smaller prefill models. | Facts sourced | 2026-10-04 |
| Open-weight models | Site references serviced open-weight models for the inference engine. | Facts sourced | 2026-10-04 |
Understand the total cost
Pricing
Not listed
Official monthly price not confirmed
It targets agencies, small businesses, and creators who want deployable AI tools without long custom development cycles.
Explore toolBolt.new is a browser-based AI app builder where users describe a web or mobile app in chat and get a running project with hosting, databases and custom domains on paid tiers.
Explore toolDualite is a vibe-coding platform that turns prompts into web and mobile apps, dashboards, and AI agents without traditional coding. It supports Figma-to-code, GitHub import, backend databases, authentication, and downloads with models such as GPT 5.1, Claude Sonnet 4.5, and Gemini 3 Pro.
Explore toolThe practical questions
OneTriangle builds a KV cache-aware inference engine that prefills on a smaller model and decodes on a larger one. By transferring key-value cache between models, it targets lower prefill cost and faster first responses on long inputs without sacrificing decode quality when pairs pass quality gates.
Not confirmed. This record does not confirm an ongoing free plan.
A listed monthly price has not been confirmed. See the plan cards for entitlements, billing commitments and seat minimums.
Not confirmed. API access and subscription access may have different terms; consult the linked sources.
Cache transfer only applies to vetted model pairs; ordinary prefill runs when a pair fails quality or latency gates.
No retained pricing changes yet. A current price alone does not establish a historical trend.
Reviewed vendor source
Read original source ↗Facts apply to the named version and check date. Send a sourced correction if something changed.