toolcompass.

Find your next AI tool

Search by product name or task. Press Escape to close.

← All model benchmarks

Moonshot AI · LIVEBENCH 2026-06-25

Kimi K2.6 Thinking

Published benchmark results for this specific model configuration.

kimi-k2.6-thinking
Overall score70.5 / 100
Cost / successful task$0.169USD · benchmark workload
Weight accessOpen weights

Where this configuration performs

Category averages from the same release, on a 0–100 scale.

Reasoning79.4
Coding78.6
Agentic Coding46.9
Mathematics84.3
Data Analysis65.1
Language75.1
Instruction Following64.4

Task-level results

theory of mind75.0zebra puzzle78.5spatial94.0logic with navigation70.0code generation78.9code completion78.3javascript59.1typescript26.7python55.0AMPS Hard97.0integrals with game54.0math comp96.1olympiad90.0consecutive events51.3tablejoin46.1tablereformat98.0connections89.3plot unscrambling58.1typos78.0paraphrase61.6simplify61.4story generation66.9summarize67.5

Source files: Scores ↗ Categories ↗ Costs ↗ . Missing cost is unknown, never free. Compare configurations within the same release.