Skip to content

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

ux-agent

Ask your coding agent to use your web app the way a person would, then hand you the evidence:

"Record me a video of that broken dialog — on mobile, like the user reported."

"Make a user guide for adding a new user to a team."

"Honestly, how hard is it to move a job from one team to another? Try it and tell me where it hurts."

ux-agent is a set of Claude Code skills and a sub-agent built on playwright-cli. Three modes, one theme: the agent drives the real app in a real browser, and what you get back is something a human can actually use. A video, a documented walkthrough, or a first-person usability report.

What's in the box

record-bug-video — reproduces a flow under the conditions the developer describes (viewport, browser, logged-in role, app state) and records it raw. No staging, no overlays — the mess is the point. Pairs the video with the console and network evidence a video can't show. "Couldn't reproduce" is a reported finding, not a failure.

user-guide-video — produces end-user documentation: short, human-paced video clips (one per step, typed at human speed, with highlight rings on the control that matters), stitched together by prose in a doc.yaml, rendered to a preview HTML page. A replay script ships alongside each guide, so when the UI changes you re-run it and the videos re-film themselves.

ux-journey (plus the ux-journey-probe agent) — task-driven UX discovery. A context-free sub-agent is given only a URL, a login, and a task in the words a manager would use. It attempts the task cold, keeping an in-the-moment think-aloud log (expectation, action, result, written down before each click resolves), with screenshots at every "...now what?" moment. The session then compiles that log into a journey report: a first-person account of friction, wrong turns, and dead ends, with a stats footer ("4 clicks became 18 across 7 pages"). The probe is tool-restricted and hook-guarded so it can't peek at your codebase: its ignorance is the instrument.

The judgement stays with you. The probe reports experience; whether a rough edge is a bug or a deliberate trade-off is a decision for the developer with the code and the context in front of them.

Show me

The examples/ directory holds real output from the first runs of each skill, against a Laravel/Livewire app:

Install

You'll need Claude Code, @playwright/cli (npm install -g @playwright/cli), ffmpeg, and uv (the user-guide render CLI is a single-file uv-run Python script).

git clone git@github.com:ohnotnow/ux-agent.git
cd ux-agent
./install.sh                        # global: ~/.claude/
./install.sh ~/code/some-project    # or per-project: <path>/.claude/

That's the whole install — the script just copies claude/skills/* and claude/agents/* into place, and it's safe to re-run after a git pull. The skills don't reference the clone, so nothing else to wire up.

A word about house conventions

These skills encode our habits: Laravel + Livewire apps, Lando local dev (*.lndo.site, self-signed certs), a seeded admin2x/secret login in every app, Laravel Debugbar to hide or ignore. The mechanics are stack-agnostic; the conventions are not. Read the skills through and swap in your own before adopting. Each one keeps its conventions in a clearly marked section. That includes the look of the rendered user guides: the house stylesheet sits in one marked block at the top of claude/skills/user-guide-video/render.py, ready to be swapped for your own brand.

Not using Claude Code?

The skills are plain markdown: point your harness at claude/skills/ (via AGENTS.md or whatever your tool reads) and you have most of the value. The ux-journey-probe agent's isolation and its no-peeking guard hook are Claude Code-specific; if your harness supports restricted sub-agents, the probe's markdown tells you exactly what to recreate.

Status

Early days, shared as-is. The preview.html files in the examples are genuine output from the render CLI (claude/skills/user-guide-video/render.py). Expect the odd rename while things settle.

You might also be interested in

  • a11y-agent — the same idea pointed at accessibility: a deterministic CLI (axe scan, keyboard tab-order walk, screen-reader transcript) your agent can run against any local page, plus a ready-made skill and background checker agent. The audit you were never going to get round to by hand.
  • agentic-stuff — the wider collection of Claude Code skills and agents these tools grew up alongside (Laravel/Livewire-leaning), including a quality-gate review stack that can call on both of these as part of a post-feature check.

Licence

MIT.

About

Agent-driven bug videos, user-guide videos, and cold usability probes for web apps

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages