toolcompass.

Find your next AI tool

Search by product name or task. Press Escape to close.

← All model benchmarks

Alibaba · LIVEBENCH 2026-06-25

Qwen 3.7 Max

Published benchmark results for this specific model configuration.

qwen3.7-max
Overall score73.1 / 100
Cost / successful task$0.182USD · benchmark workload
Weight accessNot reported

Where this configuration performs

Category averages from the same release, on a 0–100 scale.

Reasoning83.3
Coding74.2
Agentic Coding43.6
Mathematics85.2
Data Analysis71.8
Language79.7
Instruction Following74.0

Task-level results

theory of mind78.8zebra puzzle74.5spatial96.0logic with navigation84.0code generation78.9code completion69.6javascript59.1typescript26.7python45.0AMPS Hard98.0integrals with game59.0math comp97.1olympiad86.9consecutive events71.8tablejoin45.5tablereformat98.0connections96.5plot unscrambling58.7typos84.0paraphrase72.0simplify64.1story generation76.3summarize83.8

Source files: Scores ↗ Categories ↗ Costs ↗ . Missing cost is unknown, never free. Compare configurations within the same release.