toolcompass.

Find your next AI tool

Search by product name or task. Press Escape to close.

← All model benchmarks

Anthropic · LIVEBENCH 2026-06-25

Claude 4.8 Opus Thinking Max Effort

Published benchmark results for this specific model configuration.

claude-opus-4-8-max-effort
Overall score76.2 / 100
Cost / successful task$0.983USD · benchmark workload
Weight accessNot reported

Where this configuration performs

Category averages from the same release, on a 0–100 scale.

Reasoning89.2
Coding81.8
Agentic Coding50.5
Mathematics94.3
Data Analysis66.0
Language79.7
Instruction Following72.0

Task-level results

theory of mind80.8zebra puzzle100.0spatial98.0logic with navigation78.0code generation78.9code completion84.8javascript68.2typescript33.3python50.0AMPS Hard98.0integrals with game89.0math comp98.0olympiad92.2consecutive events52.9tablejoin51.1tablereformat94.1connections99.3plot unscrambling61.6typos78.0paraphrase67.9simplify62.0story generation80.6summarize77.7

Source files: Scores ↗ Categories ↗ Costs ↗ . Missing cost is unknown, never free. Compare configurations within the same release.