toolcompass.

Find your next AI tool

Search by product name or task. Press Escape to close.

← All model benchmarks

Anthropic · LIVEBENCH 2026-06-25

Claude 5.5 Opus Thinking Max Effort

Published benchmark results for this specific model configuration.

claude-opus-5-5-max-effort
Overall score83.2 / 100
Cost / successful task$0.799USD · benchmark workload
Weight accessNot reported

Where this configuration performs

Category averages from the same release, on a 0–100 scale.

Reasoning92.2
Coding89.3
Agentic Coding71.7
Mathematics97.1
Data Analysis80.3
Language86.3
Instruction Following65.7

Task-level results

theory of mind84.6zebra puzzle100.0spatial98.0logic with navigation86.0code generation91.5code completion87.0javascript81.8typescript63.3python70.0AMPS Hard98.0integrals with game100.0math comp97.1olympiad93.3consecutive events90.6tablejoin52.3tablereformat98.0connections99.3plot unscrambling77.5typos82.0paraphrase64.3simplify58.7story generation70.9summarize69.1

Source files: Scores ↗ Categories ↗ Costs ↗ . Missing cost is unknown, never free. Compare configurations within the same release.