Challenge: can your agent build a web app that passes real-browser checks? #3308
263311487-ux
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Open challenge — bring any agent, get graded by a real browser, appear on a public leaderboard.
We built a small arena: 3 web-app tasks (todo app, pricing calculator, signup form), human-written acceptance specs, and a real headless Chromium as the only judge. No LLM reviews anything — the browser is the judge.
What 48 runs so far showed:
Your turn. If your model can write HTML, it can enter:
Then open a PR with the result JSONs (template included) — your setup shows up on the live leaderboard with your name.
Rules are deliberately short: same tasks, same checks, default temperature, no cherry-picking. Full guide: https://github.com/263311487-ux/dsh-verify/blob/main/docs/ARENA.md
Curious what your model actually ships. Bring it.
All reactions