toolcompass.

Find your next AI tool

Search by product name or task. Press Escape to close.

← All model benchmarks

Anthropic · LIVEBENCH 2026-06-25

Claude 4.6 Opus Thinking High Effort

Published benchmark results for this specific model configuration.

claude-opus-4-6-thinking-auto-high-effort
Overall score74.5 / 100
Cost / successful task$0.404USD · benchmark workload
Weight accessNot reported

Where this configuration performs

Category averages from the same release, on a 0–100 scale.

Reasoning88.7
Coding78.2
Agentic Coding49.0
Mathematics89.3
Data Analysis69.9
Language83.3
Instruction Following63.3

Task-level results

theory of mind82.7zebra puzzle94.0spatial98.0logic with navigation80.0code generation80.3code completion76.1javascript63.6typescript33.3python50.0AMPS Hard97.0integrals with game73.0math comp95.1olympiad92.2consecutive events63.5tablejoin48.2tablereformat98.0connections99.3plot unscrambling66.5typos84.0paraphrase62.8simplify59.1story generation66.1summarize65.2

Source files: Scores ↗ Categories ↗ Costs ↗ . Missing cost is unknown, never free. Compare configurations within the same release.