Alibaba · LIVEBENCH 2026-06-25
Qwen 3.8 Max
Published benchmark results for this specific model configuration.
qwen3.8-maxOverall score78.5 / 100
Cost / successful task$0.275USD · benchmark workload
Weight accessOpen weights
Where this configuration performs
Category averages from the same release, on a 0–100 scale.
Task-level results
theory of mind78.8zebra puzzle100.0spatial100.0logic with navigation74.0code generation71.8code completion73.9javascript77.3typescript56.7python60.0AMPS Hard98.0integrals with game81.0math comp95.1olympiad91.2consecutive events87.1tablejoin50.1tablereformat98.0connections94.5plot unscrambling58.6typos86.0paraphrase73.3simplify67.3story generation76.5summarize79.2
Source files: Scores ↗ Categories ↗ Costs ↗ . Missing cost is unknown, never free. Compare configurations within the same release.