toolcompass.

Find your next AI tool

Search by product name or task. Press Escape to close.

← All model benchmarks

Moonshot AI · LIVEBENCH 2026-06-25

Kimi K3

Published benchmark results for this specific model configuration.

kimi-k3
Overall score79.2 / 100
Cost / successful task$0.348USD · benchmark workload
Weight accessOpen weights

Where this configuration performs

Category averages from the same release, on a 0–100 scale.

Reasoning90.7
Coding81.4
Agentic Coding62.2
Mathematics84.4
Data Analysis78.7
Language85.5
Instruction Following71.4

Task-level results

theory of mind82.7zebra puzzle100.0spatial100.0logic with navigation80.0code generation80.3code completion82.6javascript68.2typescript53.3python65.0AMPS Hard97.0integrals with game54.0math comp95.1olympiad91.6consecutive events89.8tablejoin48.4tablereformat98.0connections100.0plot unscrambling72.6typos84.0paraphrase74.4simplify66.4story generation75.4summarize69.3

Source files: Scores ↗ Categories ↗ Costs ↗ . Missing cost is unknown, never free. Compare configurations within the same release.