Compare measured performance, explore strengths and find the right balance of capability and cost.
Published results from LiveBench ↗·66 configurations
BENCHMARK RELEASE
Imported snapshot
Collected Oct 5, 2026Results reflect this release, not a live API test.
Explore model performance
One release. Consistent conditions. Clear comparisons.
TURN THE NUMBERS INTO A DECISION
Find your model’s sweet spot.
Explore capability and cost together. Click any model to understand the trade-off.
Higher score. Lower cost.
Best score available at each costZoomed to 75–90 points · cost uses a log scale
12 of 65 priced configurations shown · 1 without usable cost excludedClick a logo to inspect · use + to compare
Logos identify model families and providers. Every point is a distinct model configuration from the same benchmark release. Scores are benchmark measurements, not a prediction of your application’s results.
READ THE RESULTS WITH CONTEXT
Useful evidence. Better decisions.
A leaderboard is one signal. Your workload, access requirements and actual API pricing still matter.
We calculate the mean of the category averages in this release. Each category is the mean of its task scores on a 0–100 scale. Incomplete configurations are excluded from the relevant ranking.
How is task cost calculated?
We divide the reported total task cost by the question count, then divide by the selected scope score and multiply by 100. This is a benchmark-specific efficiency measure in USD, not an API subscription price or a quote for your workload. Missing cost is never treated as free.
Are these live tests from ToolCompass?
No. These are dated imports of results published by LiveBench. The page checks our imported snapshots every five minutes while visible. Editors control source refreshes; a failed import preserves the last valid snapshot. ToolCompass has not independently evaluated these models.
Why do effort settings and releases matter?
Configurations can use different reasoning budgets and API settings. We preserve distinct model IDs and do not combine scores from different releases. Open weights labels come from source metadata; a missing label means unknown.
Source: LiveBench · release 2026-06-25 · 23 tasks across 7 categories · imported Oct 5, 2026. Scores are published benchmark measurements, not endorsements.