toolcompass.

Find your next AI tool

Search by product name or task. Press Escape to close.

← All model benchmarks

Anthropic · LIVEBENCH 2026-06-25

Claude 4.5 Opus Thinking High Effort

Published benchmark results for this specific model configuration.

claude-opus-4-5-20251101-thinking-64k-high-effort
Overall score72.6 / 100
Cost / successful task$0.610USD · benchmark workload
Weight accessNot reported

Where this configuration performs

Category averages from the same release, on a 0–100 scale.

Reasoning80.1
Coding79.7
Agentic Coding39.7
Mathematics90.4
Data Analysis74.4
Language81.3
Instruction Following62.5

Task-level results

theory of mind78.8zebra puzzle77.5spatial96.0logic with navigation68.0code generation78.9code completion80.4javascript59.1typescript20.0python40.0AMPS Hard99.0integrals with game78.0math comp95.1olympiad89.5consecutive events79.4tablejoin45.9tablereformat98.0connections99.3plot unscrambling66.5typos78.0paraphrase65.7simplify55.0story generation65.8summarize63.7

Source files: Scores ↗ Categories ↗ Costs ↗ . Missing cost is unknown, never free. Compare configurations within the same release.