toolcompass.

Find your next AI tool

Search by product name or task. Press Escape to close.

← All model benchmarks

Stealth · LIVEBENCH 2026-06-25

ox-alpha-max

Published benchmark results for this specific model configuration.

ox-alpha-max
Overall score69.2 / 100
Cost / successful taskNot reportedUSD · benchmark workload
Weight accessNot reported

Where this configuration performs

Category averages from the same release, on a 0–100 scale.

Reasoning76.6
Coding75.8
Agentic Coding52.6
Mathematics77.5
Data Analysis75.8
Language66.1
Instruction Following60.3

Task-level results

theory of mind78.8zebra puzzle79.5spatial96.0logic with navigation52.0code generation73.2code completion78.3javascript54.5typescript43.3python60.0AMPS Hard98.0integrals with game39.0math comp91.2olympiad82.0consecutive events80.4tablejoin46.9tablereformat100.0connections71.3plot unscrambling53.1typos74.0paraphrase57.5simplify54.5story generation70.9summarize58.2

Source files: Scores ↗ Categories ↗ Costs ↗ . Missing cost is unknown, never free. Compare configurations within the same release.