toolcompass.

Find your next AI tool

Search by product name or task. Press Escape to close.

← All model benchmarks

Anthropic · LIVEBENCH 2026-06-25

Claude 4.6 Sonnet Thinking Medium Effort

Published benchmark results for this specific model configuration.

claude-sonnet-4-6-thinking-auto-medium-effort
Overall score73.0 / 100
Cost / successful task$0.306USD · benchmark workload
Weight accessNot reported

Where this configuration performs

Category averages from the same release, on a 0–100 scale.

Reasoning84.8
Coding79.3
Agentic Coding42.6
Mathematics87.0
Data Analysis77.9
Language76.1
Instruction Following63.2

Task-level results

theory of mind73.1zebra puzzle96.0spatial100.0logic with navigation70.0code generation80.3code completion78.3javascript54.5typescript23.3python50.0AMPS Hard76.0integrals with game90.0math comp94.1olympiad87.9consecutive events92.7tablejoin45.1tablereformat96.1connections99.3plot unscrambling57.0typos72.0paraphrase59.4simplify62.9story generation68.9summarize61.6

Source files: Scores ↗ Categories ↗ Costs ↗ . Missing cost is unknown, never free. Compare configurations within the same release.