toolcompass.

Find your next AI tool

Search by product name or task. Press Escape to close.

← All model benchmarks

Alibaba · LIVEBENCH 2026-06-25

Qwen3.8 27B

Published benchmark results for this specific model configuration.

qwen3.8-27b
Overall score75.3 / 100
Cost / successful task$0.094USD · benchmark workload
Weight accessOpen weights

Where this configuration performs

Category averages from the same release, on a 0–100 scale.

Reasoning80.0
Coding75.7
Agentic Coding61.4
Mathematics86.2
Data Analysis76.6
Language74.3
Instruction Following72.7

Task-level results

theory of mind59.6zebra puzzle96.5spatial96.0logic with navigation68.0code generation77.5code completion73.9javascript59.1typescript50.0python75.0AMPS Hard98.0integrals with game65.0math comp94.1olympiad87.7consecutive events84.9tablejoin48.8tablereformat96.1connections90.5plot unscrambling56.5typos76.0paraphrase74.9simplify66.7story generation73.1summarize75.9

Source files: Scores ↗ Categories ↗ Costs ↗ . Missing cost is unknown, never free. Compare configurations within the same release.