toolcompass.

Find your next AI tool

Search by product name or task. Press Escape to close.

← All model benchmarks

DeepSeek · LIVEBENCH 2026-06-25

DeepSeek V4 Flash 0731

Published benchmark results for this specific model configuration.

deepseek-v4-flash-0731
Overall score74.2 / 100
Cost / successful task$0.060USD · benchmark workload
Weight accessOpen weights

Where this configuration performs

Category averages from the same release, on a 0–100 scale.

Reasoning86.6
Coding75.0
Agentic Coding46.8
Mathematics86.8
Data Analysis79.3
Language79.2
Instruction Following65.5

Task-level results

theory of mind86.5zebra puzzle100.0spatial90.0logic with navigation70.0code generation76.1code completion73.9javascript63.6typescript26.7python50.0AMPS Hard97.0integrals with game65.0math comp96.1olympiad89.1consecutive events89.4tablejoin48.5tablereformat100.0connections97.3plot unscrambling58.2typos82.0paraphrase61.7simplify58.4story generation70.5summarize71.5

Source files: Scores ↗ Categories ↗ Costs ↗ . Missing cost is unknown, never free. Compare configurations within the same release.