toolcompass.

Find your next AI tool

Search by product name or task. Press Escape to close.

ClawBench

● Identity checked

Research · Vendor site 2026-10-04

ClawBench is an open benchmark and public leaderboard for AI browser agents on live websites. It scores models on 130 real-world tasks across dozens of platforms using HTTP interception plus an LLM judge, with datasets and open submissions via pull request.

Updated 2026-10-04View sources
Official website
Category
Research
Free access
The benchmark site, leaderboard, and dataset access are public with no paid subscription listed.
API access
Not confirmed as a commercial API; the project publishes tasks, traces, and open submission instructions on the site.

Is ClawBench right for you?

A good fit for

Researchers and agent developers comparing browser automation models on realistic checkout, booking, and form-filling tasks.

Before you choose

Reported per-task costs on the leaderboard reflect model inference spend during evaluation, not a ClawBench subscription fee.

Live browser-agent benchmark

ClawBench ranks AI agents on everyday online tasks on real sites, publishing rewards, interception rates, and pass counts for public comparison.

What it can do

Features & capabilities

Unknown is different from unavailable. Each fact carries its own evidence.

CapabilityValueEvidenceChecked
Task corpusHomepage lists 130 tasks across 63 live platforms such as booking flights and ordering food.Facts sourced2026-10-04
Scoring methodBenchmark uses two-stage scoring combining HTTP-request interception and an LLM judge.Facts sourced2026-10-04
Open submissionsSite invites open model submissions via pull request alongside arXiv and Hugging Face dataset links.Facts sourced2026-10-04

Understand the total cost

ClawBench pricing & plans

Free

Free

The benchmark site, leaderboard, and dataset access are public with no paid subscription listed.

Access
The benchmark site, leaderboard, and dataset access are public with no paid subscription listed.
Explore pricing & history

Alternatives to ClawBench

View all ↗

The practical questions

Frequently asked questions

What is ClawBench used for?

ClawBench is an open benchmark and public leaderboard for AI browser agents on live websites. It scores models on 130 real-world tasks across dozens of platforms using HTTP interception plus an LLM judge, with datasets and open submissions via pull request.

Does ClawBench have a free plan?

The benchmark site, leaderboard, and dataset access are public with no paid subscription listed.. This record lists ongoing free access; check the plan limits before starting.

How much does ClawBench cost?

No paid monthly price is listed; this record treats the product as free to start. See the plan cards for entitlements, billing commitments and seat minimums.

Can I use ClawBench through an API?

Not confirmed as a commercial API; the project publishes tasks, traces, and open submission instructions on the site.. API access and subscription access may have different terms; consult the linked sources.

What should I check before choosing it?

Reported per-task costs on the leaderboard reflect model inference spend during evaluation, not a ClawBench subscription fee.

Price history

No retained pricing changes yet. A current price alone does not establish a historical trend.

How this profile is supported

Facts apply to the named version and check date. Send a sourced correction if something changed.