toolcompass.

Find your next AI tool

Search by product name or task. Press Escape to close.

← All model benchmarks

Alibaba · LIVEBENCH 2026-06-25

Qwen 3.8 Flash Next

Published benchmark results for this specific model configuration.

qwen3.8-flash-next
Overall score76.2 / 100
Cost / successful task$0.042USD · benchmark workload
Weight accessOpen weights

Where this configuration performs

Category averages from the same release, on a 0–100 scale.

Reasoning87.4
Coding72.6
Agentic Coding61.6
Mathematics85.8
Data Analysis74.2
Language74.6
Instruction Following77.1

Task-level results

theory of mind80.8zebra puzzle98.8spatial100.0logic with navigation70.0code generation69.0code completion76.1javascript68.2typescript46.7python70.0AMPS Hard99.0integrals with game59.0math comp94.1olympiad91.2consecutive events75.6tablejoin51.0tablereformat96.1connections90.2plot unscrambling55.8typos78.0paraphrase79.5simplify70.3story generation81.4summarize77.3

Source files: Scores ↗ Categories ↗ Costs ↗ . Missing cost is unknown, never free. Compare configurations within the same release.