Moonshot AI · LIVEBENCH 2026-06-25
Kimi K2.6 Thinking
Published benchmark results for this specific model configuration.
kimi-k2.6-thinkingOverall score70.5 / 100
Cost / successful task$0.169USD · benchmark workload
Weight accessOpen weights
Where this configuration performs
Category averages from the same release, on a 0–100 scale.
Task-level results
theory of mind75.0zebra puzzle78.5spatial94.0logic with navigation70.0code generation78.9code completion78.3javascript59.1typescript26.7python55.0AMPS Hard97.0integrals with game54.0math comp96.1olympiad90.0consecutive events51.3tablejoin46.1tablereformat98.0connections89.3plot unscrambling58.1typos78.0paraphrase61.6simplify61.4story generation66.9summarize67.5
Source files: Scores ↗ Categories ↗ Costs ↗ . Missing cost is unknown, never free. Compare configurations within the same release.