Free
Free
Benchmark pages and leaderboard access are free to use on the public EvalPlus site.
- Access
- Benchmark pages and leaderboard access are free to use on the public EvalPlus site.
Code · Vendor site 2026-10-04
EvalPlus publishes rigorous open benchmarks for large language model code performance, including HumanEval+, MBPP+, EvalPerf efficiency tests, and RepoQA long-context code understanding evaluators. The site hosts leaderboards and links to research artifacts maintained by the EvalPlus team.
ML researchers and engineers evaluating LLM coding ability who need the HumanEval+ and related benchmark suites.
EvalPlus is an evaluation toolkit and leaderboard hub, not a hosted coding assistant with commercial seat plans.
EvalPlus teams build high-quality, precise evaluators so researchers can compare LLM performance on code tasks beyond the original HumanEval and MBPP test counts.
What it can do
Unknown is different from unavailable. Each fact carries its own evidence.
| Capability | Value | Evidence | Checked |
|---|---|---|---|
| HumanEval+ / MBPP+ | Homepage states EvalPlus extended HumanEval and MBPP tests by 80× and 35× respectively for rigorous evaluation. | Facts sourced | 2026-10-04 |
| EvalPerf | Site describes EvalPerf for measuring efficiency of LLM-generated code using performance-exercising tasks. | Facts sourced | 2026-10-04 |
| RepoQA | Project section covers RepoQA evaluators for long-context repository understanding. | Facts sourced | 2026-10-04 |
Understand the total cost
Free
Free
Benchmark pages and leaderboard access are free to use on the public EvalPlus site.
ByteAsk is a terminal-based AI coding agent for C and C++ that edits repositories and verifies work with compilers, sanitizers, debuggers, and tests. It integrates LLVM, GCC, gdb, Valgrind, CMake, and related tooling with grounded standards and datasheet excerpts.
Explore toolDalus is an AI-native model-based systems engineering platform for hardware teams. It unifies requirements, architecture, analysis, and verification in a collaborative environment aimed at aerospace, defense, robotics, automotive, and energy programs.
Explore toolMaritime hosts persistent AI agents, each in its own Firecracker micro-VM with durable disk. Agents sleep when idle, wake in about a second, and scale from a free tier to paid plans with configurable RAM, SSD, and always-on options.
Explore toolThe practical questions
EvalPlus publishes rigorous open benchmarks for large language model code performance, including HumanEval+, MBPP+, EvalPerf efficiency tests, and RepoQA long-context code understanding evaluators. The site hosts leaderboards and links to research artifacts maintained by the EvalPlus team.
Benchmark pages and leaderboard access are free to use on the public EvalPlus site.. This record lists ongoing free access; check the plan limits before starting.
No paid monthly price is listed; this record treats the product as free to start. See the plan cards for entitlements, billing commitments and seat minimums.
Not confirmed. API access and subscription access may have different terms; consult the linked sources.
EvalPlus is an evaluation toolkit and leaderboard hub, not a hosted coding assistant with commercial seat plans.
No retained pricing changes yet. A current price alone does not establish a historical trend.
Reviewed vendor source
Read original source ↗Facts apply to the named version and check date. Send a sourced correction if something changed.