toolcompass.

Find your next AI tool

Search by product name or task. Press Escape to close.

← All model benchmarks

Anthropic · LIVEBENCH 2026-06-25

Claude 5 Opus Thinking Max Effort

Published benchmark results for this specific model configuration.

claude-opus-5-max-effort
Overall score80.1 / 100
Cost / successful task$0.699USD · benchmark workload
Weight accessNot reported

Where this configuration performs

Category averages from the same release, on a 0–100 scale.

Reasoning91.2
Coding81.4
Agentic Coding65.2
Mathematics95.7
Data Analysis74.6
Language88.7
Instruction Following63.8

Task-level results

theory of mind78.8zebra puzzle100.0spatial100.0logic with navigation86.0code generation80.3code completion82.6javascript77.3typescript43.3python75.0AMPS Hard99.0integrals with game97.0math comp94.1olympiad92.8consecutive events77.6tablejoin52.0tablereformat94.1connections99.3plot unscrambling74.7typos92.0paraphrase65.7simplify61.6story generation61.3summarize66.4

Source files: Scores ↗ Categories ↗ Costs ↗ . Missing cost is unknown, never free. Compare configurations within the same release.