toolcompass.

Find your next AI tool

Search by product name or task. Press Escape to close.

SWE-bench

● Identity checked

Research · Vendor site 2026-10-04

SWE-bench is a family of open software-engineering benchmarks built from real GitHub issues, with public leaderboards comparing model and agent performance. The site tracks variants such as Verified, Lite, Multilingual, and Multimodal tasks and links related tools like mini-SWE-agent and SWE-agent.

Updated 2026-10-04View sources
Official website
Category
Research
Free access
Leaderboards and benchmark descriptions are published on the public site without a paid subscription.
API access
Not confirmed

Is SWE-bench right for you?

A good fit for

Researchers and teams evaluating coding agents and LLMs on realistic repository repair and software engineering tasks.

Before you choose

Verified Bash Only evaluation uses 500 curated instances rather than the full 2,294-instance SWE-bench set.

Software engineering evaluation

SWE-bench hosts official leaderboards and benchmark variants for measuring how well models and agents resolve real open-source issues.

What it can do

Features & capabilities

Unknown is different from unavailable. Each fact carries its own evidence.

CapabilityValueEvidenceChecked
Benchmark scopeSite describes SWE-bench as real GitHub issues from Python repositories with multiple subset benchmarks.Facts sourced2026-10-04
LeaderboardsPublic leaderboards report percent resolved across models with comparison and analysis views.Facts sourced2026-10-04

Understand the total cost

SWE-bench pricing & plans

Free

Free

Leaderboards and benchmark descriptions are published on the public site without a paid subscription.

Access
Leaderboards and benchmark descriptions are published on the public site without a paid subscription.
Explore pricing & history

Alternatives to SWE-bench

View all ↗

The practical questions

Frequently asked questions

What is SWE-bench used for?

SWE-bench is a family of open software-engineering benchmarks built from real GitHub issues, with public leaderboards comparing model and agent performance. The site tracks variants such as Verified, Lite, Multilingual, and Multimodal tasks and links related tools like mini-SWE-agent and SWE-agent.

Does SWE-bench have a free plan?

Leaderboards and benchmark descriptions are published on the public site without a paid subscription.. This record lists ongoing free access; check the plan limits before starting.

How much does SWE-bench cost?

No paid monthly price is listed; this record treats the product as free to start. See the plan cards for entitlements, billing commitments and seat minimums.

Can I use SWE-bench through an API?

Not confirmed. API access and subscription access may have different terms; consult the linked sources.

What should I check before choosing it?

Verified Bash Only evaluation uses 500 curated instances rather than the full 2,294-instance SWE-bench set.

Price history

No retained pricing changes yet. A current price alone does not establish a historical trend.

How this profile is supported

Facts apply to the named version and check date. Send a sourced correction if something changed.