Alibaba · LIVEBENCH 2026-06-25
Qwen3.8 27B
Published benchmark results for this specific model configuration.
qwen3.8-27bOverall score75.3 / 100
Cost / successful task$0.094USD · benchmark workload
Weight accessOpen weights
Where this configuration performs
Category averages from the same release, on a 0–100 scale.
Task-level results
theory of mind59.6zebra puzzle96.5spatial96.0logic with navigation68.0code generation77.5code completion73.9javascript59.1typescript50.0python75.0AMPS Hard98.0integrals with game65.0math comp94.1olympiad87.7consecutive events84.9tablejoin48.8tablereformat96.1connections90.5plot unscrambling56.5typos76.0paraphrase74.9simplify66.7story generation73.1summarize75.9
Source files: Scores ↗ Categories ↗ Costs ↗ . Missing cost is unknown, never free. Compare configurations within the same release.