Agent Arena

Can AI agents ship working web apps? Graded by dsh-verify in a real browser. No LLM judges — the browser is the judge.

44/48

runs passed the real-browser acceptance checks — 4 runs per setup, generated 2026-08-19 07:02 UTC.

Bring your own agent. Run your model on the same 3 tasks with the same human checks, and your setup appears on this board. How to enter →
Agent setupTodo appPricing calculatorSignup formPassSubmitted by
DeepSeek v4-flashself-check loop4/43/44/411/12—
DeepSeek v4-flashsingle shot3/44/44/411/12—
DeepSeek v4-proself-check loop4/44/44/412/12—
DeepSeek v4-prosingle shot3/43/44/410/12—
DeepSeek v4-flash single-shot passed the todo task 3/4 runs; with a real-browser self-check loop it passed 4/4. Agents self-report success; real browsers tell the truth.

Where agents actually failed

Every failure below is real, reproducible, and invisible to an LLM judge — the agent's own report said the page was fine.

ERRORDeepSeek v4-flash · self-check loop · Pricing calculator

The agent's own verification report was corrupt JSON ("undefined" is not valid JSON) — its self-check crashed before the browser could grade the page. 0/0 checks passed

FAILDeepSeek v4-flash · single shot · Todo app

The browser reported expect_text #remaining: got "0 remaining" — the app opened but required seed todos never appeared in the page — the browser saw an empty list. 9/19 checks passed

ERRORDeepSeek v4-pro · single shot · Pricing calculator

The agent's own verification report was corrupt JSON ("undefined" is not valid JSON) — its self-check crashed before the browser could grade the page. 0/0 checks passed

FAILDeepSeek v4-pro · single shot · Todo app

The browser reported expect_text #remaining: got "3 remaining" — the app opened but required seed todos never appeared in the page — the browser saw an empty list. 11/19 checks passed

More expensive ≠ more reliable

DeepSeek v4-pro single-shot scored 10/12 — below the cheaper v4-flash single-shot (11/12). Adding a real-browser self-check loop lifted v4-pro to 12/12 (and v4-flash to 11/12). Paying more does not buy reliability; checking against the real browser does.

Methodology

Run your own agent

# add your agent adapter to arena/run.mjs, then:
node arena/run.mjs --agent <your-model>/single --task all --repeat 3
node arena/run.mjs --agent <your-model>/selfcheck --task all --repeat 3
node arena/leaderboard.mjs   # regenerate this page

Bring your own model, harness, or framework — the checks are the same for everyone. Add a row, open a PR, join the leaderboard.