A portable, LLM-driven browser-automation & web-scraping MCP server — one binary, three engines, any OS.
Webrain exposes ~45 browser/scraping tools over the Model Context Protocol. Install it on any system, point any LLM client (GitHub Copilot, Claude, Codex, Cursor, …) at it, and the model decides everything — search, crawl, scrape, navigate, interact — from a plain-language prompt. No hardcoded intent detection, no daemon, no Node.js.
Web automation shouldn't mean wiring up a driver, a browser download, and your own tool wrappers before an LLM can touch a page. Webrain collapses that into one binary + one install command and speaks MCP, so any LLM client drives it directly.
| webrain | Raw Playwright / Puppeteer | |
|---|---|---|
| Setup | one binary + webrain install |
runtime + driver + browser download + your own wrappers |
| Browsers | Chrome + lightpanda + obscura through one CDP backend | one engine, usually Chromium |
| LLM-ready | MCP server, ~45 tools, built-in decision guide | you hand-roll tool functions |
| Anti-bot | challenge detection + stealth sidecar | manual |
| Extraction | autoschema → JSON / regex / table / spider / batch | you write selectors |
| Runs on | any OS, any LLM client, or plain CLI | tied to your stack |
- Scrape at scale — batch pagination + spider with auto-throttle and checkpoint/resume;
webrain_sitemap/webrain_scanto map a site first. - Structured data without hand-written selectors —
webrain_autoschemaprobes the DOM, then JSON / regex / table extractors read container-level structure. - Stealth login — real-Chrome profiles with an encrypted local credential vault (AES-256-GCM + optional TOTP); transfer cookies across engines.
- Get past challenges — reads the
challengefield on every navigate and solves Cloudflare/Turnstile via a real-Chrome stealth sidecar. - See the page — a11y / semantic tree, snapshots, and vision tiles (screenshot → vector store) for tables and charts.
- Read anything — PDFs (extract + render), JSON-LD, media, plus
fetch_httpfor static pages 10–100× faster than a browser.
- Language: Rust (edition 2024, MSRV 1.85)
- Workspace:
webrain-core(CDP client, engines, vault, launch, install) ·webrain-mcp(MCP server + tool schemas) ·webrain-cli(single binary) - Protocol: Model Context Protocol (stdio + HTTP transports)
- Browser automation: Chrome DevTools Protocol over
tokio-tungstenite(raw WebSocket) - Engines: Chrome for Testing · Lightpanda (Zig) · Obscura (Rust, V8) ·
fetch_http - Crypto: AES-256-GCM vault, SHA-256/HMAC/SHA-1 (TOTP)
- PDF: lopdf + pdf-inspector (pure Rust)
- Deployment: Docker (multi-arch), GitHub Actions CI/CD
- OS: Windows, macOS, or Linux (x86_64 / arm64).
- A browser engine — run
webrain installonce (downloads Chrome for Testing). Obscura and Lightpanda are optional extra engines. - Linux: system libraries for Chrome (list below) before first run.
- Python 3 — only for the stealth sidecar (
scripts/stealth_solve.py) to bypass Cloudflare/Turnstile challenges. - Docker — only to run the obscura / lightpanda engines in containers.
- An MCP-capable client (VS Code + Copilot, Claude, Codex, Cursor, …) — optional; the CLI works standalone.
Install model mirrors vercel-labs/agent-browser: a native binary plus an
installcommand that downloads engines into a cache dir.
One command installs webrain on your PATH:
Linux / macOS:
curl -fsSL https://raw.githubusercontent.com/prokopis3/webrain/main/scripts/install.sh | bashWindows (PowerShell):
irm https://raw.githubusercontent.com/prokopis3/webrain/main/scripts/install.ps1 | iexThen, in any terminal:
webrain install # Download Chrome for Testing (first time only)
webrain mcp --http 9223 # start the MCP serverPrefer a manual download? The per-OS commands below do the same without a script.
Windows (PowerShell):
$dir = "$env:LOCALAPPDATA\Programs\webrain"; New-Item -ItemType Directory -Force -Path $dir | Out-Null
Invoke-WebRequest "https://github.com/prokopis3/webrain/releases/latest/download/webrain-windows.exe" -OutFile "$dir\webrain.exe"
[Environment]::SetEnvironmentVariable("Path", "$([Environment]::GetEnvironmentVariable('Path','User'));$dir", "User")
# open a new terminal, then:
webrain install # Download Chrome for Testing (first time only)
webrain mcp --http 9223 # start the MCP serverLinux / macOS:
curl -L -o ~/.local/bin/webrain https://github.com/prokopis3/webrain/releases/latest/download/webrain-linux
chmod +x ~/.local/bin/webrain
webrain install # Download Chrome for Testing (first time only)
webrain mcp --http 9223 # start the MCP servermacOS: swap
webrain-linuxforwebrain-macosin the curl line.
Pin a version as a local dependency:
cargo add webrain --git https://github.com/prokopis3/webrain
# or build from this repo (see From Source)cargo install --git https://github.com/prokopis3/webrain webrain-cli
webrain install # Download Chrome (first time only)brew tap prokopis3/webrain
brew install webrain
webrain install # Download Chrome (first time only)From the official extras bucket (after PR #18455 merges):
scoop install extras/webrainOr from the project's own bucket:
scoop bucket add webrain https://github.com/prokopis3/scoop-webrain
scoop install webrain
webrain install # Download Chrome (first time only)On Linux, Chrome needs system libraries. Install them with your package manager before first run:
# Debian/Ubuntu
sudo apt-get install -y libnss3 libnspr4 libxkbcommon0 libatk1.0-0 \
libatk-bridge2.0-0 libxcomposite1 libxdamage1 libxrandr2 libxfixes3 \
libxcursor1 libxi6 libxtst6 libxss1 libxext6 fonts-liberation
# Fedora
sudo dnf install -y nss nspr libxkbcommon atk at-spi2-atk at-spi2-core \
libXcomposite libXdamage libXrandr libXfixes libXcursor libXi libXtst \
libXScrnSaver libXextIf you see "shared library" errors when running Chrome, that's the missing-deps symptom — install the list above.
git clone https://github.com/prokopis3/webrain
cd webrain
cargo build --release --bin webrain
./target/release/webrain install # Download Chrome (first time only)
./target/release/webrain mcp --http 9223Upgrade to the latest version:
webrain upgradeDetects your installation method (Homebrew, Scoop, or a manual install) and
updates automatically — brew upgrade webrain, scoop update webrain, or
self-updates the binary in place.
If installed via cargo:
cargo install --git https://github.com/prokopis3/webrain webrain-cli --forceRe-download engines after an update (cache dir stays, versions are additive):
webrain install
webrain install --engine obscuraYes — the codebase targets all three desktop OSes. Engines per OS:
| Engine | Windows | macOS | Linux |
|---|---|---|---|
Chrome for Testing (webrain install) |
✅ | ✅ | ✅ |
Obscura (webrain install --engine obscura) |
✅ x86_64 | ✅ x86_64 / arm64 | ✅ x86_64 / arm64 |
Lightpanda (webrain lightpanda, binary needed) |
✅ | ✅ | |
Docker (docker/Dockerfile) |
via Docker Desktop | via Docker Desktop | ✅ |
Binary discovery is automatic on every OS: env override → PATH → ~/.lightpanda, ~/.obscura, ~/.local/bin → the webrain engine cache.
# 1. Get a browser engine
webrain install
# 2. Start the MCP server (stdio for VS Code/Copilot, or HTTP)
webrain mcp # stdio
webrain mcp --http 9223 # HTTP transport on 127.0.0.1:9223
# 3. Point an LLM at it. Register the MCP server in your client (see "Marketplace / IDE Plugins"):
# stdio: {"servers": {"webrain": {"command": "webrain", "args": ["mcp"]}}}
# HTTP: webrain mcp --http 9223 → {"servers": {"webrain": {"type": "http", "url": "http://127.0.0.1:9223/mcp"}}}
# 4. Ask the LLM to do something, e.g. "scrape all product titles + prices from URL X"Prefer to drive it by hand? Every engine + tool has a CLI twin:
webrain launch scrapingcourse demo "https://example.com/login" --port 9222 # headed Chrome + login
webrain doctor # full diagnosis: engines, MCP, CDP, vault, sidecar
webrain fetch <url> # attach to CDP_URL and fetch
webrain screenshot <url>
webrain eval "document.title"The MCP tools take plain CSS selectors directly (e.g. webrain_click with
selector: "#submit"), and the CLI works on snapshot refs — the same elements
snapshot reports, 1-indexed:
webrain launch https://example.com
webrain snapshot # prints interactive elements, 1-indexed
webrain click 3 # click element #3
webrain type 5 "text" # type into element #5
webrain eval 'document.querySelector("#email").value'webrain mcp [--http <port>] # MCP server (stdio, or HTTP on a port)
webrain doctor # diagnose the install (engines, MCP, CDP, vault)
webrain install [--engine chrome|obscura] [--stealth] # download a browser engine
webrain upgrade # update to the latest release
webrain launch <service> <profile> [url] # stealth Chrome, persistent profile
webrain login <service> <profile> [url] # interactive login into a profile
webrain vault set|list|user|rm # encrypted credential vault (AES-256-GCM + TOTP)
webrain cookies / setcookies <file> # export / import cookies
webrain fetch <url> # attach to CDP_URL and fetch
webrain screenshot <url> # screenshot (single or full page)
webrain spider <url> [--depth N --pages N --respect-robots] # crawl
webrain click <i> / type <i> <text> / eval <js> # drive the CDP_URL backend
webrain obscura / lightpanda [--port N] # spawn a CDP serverThe full ~45-tool MCP surface is discovered dynamically — webrain_guide lists
it for the LLM (see MCP Tools).
| Need | Engine | How to get it |
|---|---|---|
| Material / interactive SPA (Google Flights, calendars, dropdowns), real screenshots, Cloudflare/Turnstile | real Chrome | webrain install |
| Fast scraping of non-challenged JS pages, parallel tabs | obscura (stealth) | webrain install --engine obscura [--stealth], then webrain obscura |
| Fastest/lightest, real a11y + semantic tree, no rendering | lightpanda | install the binary, then webrain lightpanda |
| Static HTML, no JS/auth | fetch_http (no browser) | built-in |
Key rules (see Agent Decision Guide):
- Never use obscura/lightpanda for Material/SPA interaction or screenshots — they have no layout/paint engine. Obscura errors loudly on screenshots; lightpanda returns a fake placeholder PNG. Route interactive SPAs to real Chrome via
cdp_urls:["http://127.0.0.1:9222"]. - Read the
challengefield after everywebrain_navigate. If it's non-null, the page is gated — use the real-Chrome stealth sidecar. - Extract from container/card-level DOM, not bare
$text nodes (Google Flights renders a spurious price grid).
All tools are discovered dynamically (webrain_guide lists them for the LLM). Highlights:
| Category | Tools |
|---|---|
| Navigate / observe | webrain_navigate, webrain_snapshot, webrain_a11y, webrain_semantic_tree, webrain_get_html, webrain_console |
| Interact | webrain_click, webrain_type, webrain_press, webrain_scroll, webrain_nav, webrain_tab, webrain_dismiss_overlays |
| Extract | webrain_autoschema, webrain_extract_json, webrain_extract_regex, webrain_table, webrain_get_jsonld, webrain_pdf_extract, webrain_pdf_images |
| Crawl | webrain_batch, webrain_spider (AutoThrottle + checkpoint/resume), webrain_sitemap, webrain_scan |
| Vision | webrain_pixel, webrain_vision_index, webrain_vision_retrieve, webrain_screenshot |
| Auth / state | webrain_login, webrain_profiles, webrain_cookies, webrain_setcookies, webrain_open_session, webrain_close_session |
| Utility | webrain_fetch_http, webrain_download, webrain_search, webrain_validate_urls, webrain_clean, webrain_media, webrain_get_images, webrain_guide, webrain_eval |
Full per-tool reference: docs/ARCHITECTURE.md. Tool/browser/challenge decisions: docs/AGENT_DECISION_GUIDE.md.
| Command | Description |
|---|---|
webrain mcp [--http <port>] |
Start the MCP server (stdio, or HTTP on a port) |
webrain install [--force] [--engine chrome|obscura] [--stealth] |
Download a browser engine into the cache |
webrain obscura [--port N] |
Spawn the Obscura CDP server (default 9224) |
webrain lightpanda [--port N] |
Spawn the Lightpanda CDP server (default 9225) |
webrain launch <service> <profile> [url] [--headless] [--port N] |
Spawn a stealth Chrome with a persistent per-account profile |
webrain login <service> <profile> [url] [--port N] |
Launch + attach for interactive login |
webrain cookies [--port N] [--out file] / webrain setcookies <file> |
Export / import cookies |
webrain fetch <url> · webrain screenshot <url> · webrain spider <url> · webrain click <i> · webrain type <i> <text> · webrain eval <js> |
Drive the CDP_URL backend |
webrain vault set|list|user|rm |
Manage encrypted credentials (hidden prompts) |
webrain doctor |
Full install diagnosis — version, MCP server, CDP ports (9222/9224/9225), engine discovery, vault, Python sidecar, recommend. --doctor alias |
| Variable | Description | Default |
|---|---|---|
CDP_URL |
CDP endpoint to attach to | http://127.0.0.1:9222 |
WEBRAIN_CHROME |
Explicit Chrome binary path | auto-discovered |
WEBRAIN_BROWSERS_DIR |
Engine cache dir (Chrome for Testing, Obscura) | %LOCALAPPDATA%\webrain\browsers / ~/.cache/webrain/browsers |
WEBRAIN_LIGHTPANDA / WEBRAIN_OBSCURA |
Explicit lightpanda/obscura binary paths | auto-discovered |
WEBRAIN_PROFILES_DIR |
Per-account browser profile root | %APPDATA%\webrain\profiles / ~/.config/webrain/profiles |
WEBRAIN_VAULT_DIR |
Encrypted credential vault dir | %APPDATA%\webrain / ~/.config/webrain |
WEBRAIN_USER / WEBRAIN_PASS |
Login fallback credentials (env channel) | — |
RUST_LOG |
Log verbosity (webrain=info,tungstenite=warn) |
as above |
Webrain is an MCP server, so it plugs into any MCP-capable IDE. There is no separate extension to publish — you register the server, and the ~45 tools appear.
Both transports work in every client: stdio (webrain mcp) or HTTP
(webrain mcp --http 9223, endpoint http://127.0.0.1:9223/mcp).
VS Code (GitHub Copilot)
Add to your user settings.json under "mcp":
Claude Desktop — ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"webrain": { "command": "webrain", "args": ["mcp"] }
}
}Cursor — project .mcp.json:
{
"mcpServers": {
"webrain": { "command": "webrain", "args": ["mcp"] }
}
}Agent skills (optional) — this repo ships agent guidance so the LLM drives it well:
skills/webrain/SKILL.md— browser/challenge/extraction decision guide + auth-cookie transfer procedures..github/skills/— repo-local skills (credentials,update-changelog,git-cleanup).
webrain (single binary)
├── webrain-core CDP client + engines + vault + launch + install
│ ├── backends/cdp.rs BrowserBackend over CDP (WS): navigate/eval/click/a11y/… (+ STEALTH_JS)
│ ├── engines.rs Tile (vision tiles), Spider (BFS/DFS/best-first + AutoThrottle), extract, BM25
│ ├── install.rs engine download/discovery (Chrome for Testing, Obscura, Lightpanda)
│ ├── launch.rs spawn Chrome/lightpanda/obscura, wait for CDP
│ ├── login.rs / vault.rs encrypted credential vault + CDP login injection (TOTP)
│ └── vision.rs screenshot-tile embedding + vector store
├── webrain-mcp MCP server: list_tools / call_tool, ~45 tool schemas
└── webrain-cli subcommand dispatch (mcp | install | launch | login | …)
How it works:
CdpBackendconnects to a CDP endpoint (Chrome, lightpanda, or obscura — all speak CDP) over a raw WebSocket.- On attach it applies stealth hardening (UA override,
Emulation.setAutomationOverride, JS patches) — so it can log into real sites. - The MCP layer exposes every action as a tool. An LLM picks tools by intent;
webrain_guide+AGENT_DECISION_GUIDE.mdencode the which-browser / which-tool decisions so the LLM never guesses. - Extraction is generic — autoschema probes the DOM, JSON/regex/table extractors read container-level structure, spider/batch/sitemap crawl at scale, vision tiles give the model "eyes" for tables/charts.
Data flow (LLM → browser):
LLM prompt → tool call (webrain_navigate/extract/batch/…) → CdpBackend → CDP → engine (Chrome/lightpanda/obscura)
← JSON result (PageState, extracted rows, stats) ←
Read docs/AGENT_DECISION_GUIDE.md before browser tasks. It encodes, with live-verified results:
- Browser selection — Material/SPA → real Chrome via
cdp_urls; Cloudflare/Turnstile → Chrome + stealth sidecar; fast JS scraping → obscura; lightpanda for real a11y with minimal footprint; static →fetch_http. - Challenge handling — read
challengeafter every navigate; obscura/lightpanda cannot pass interactive challenges. - Extraction matrix — autoschema → extract_json → batch → regex → table → spider; never guess selectors.
- a11y — Google/Material widgets are
combobox/option/tab, notbutton;filtermatches name/value/css_path; ifrole=<x>returns[], drop the role and filter by label.
# run the self-check tests in webrain-core
cargo test --package webrain-core
# run everything
cargo test --workspace
# lint
cargo clippy --workspace --all-targets
cargo fmt --checkThe repo favors one runnable self-check per non-trivial module (assert-based, no test framework sprawl) — see the #[cfg(test)] blocks in webrain-core/src/engines.rs and install.rs.
GitHub Actions (.github/workflows/):
ci.yml—cargo fmt --check+clippy+teston every push/PR tomain.release.yml— on av*tag push: builds the binary natively on Linux/Windows/macOS, generates the CHANGELOG entry for that version (if missing), and creates the GitHub Release with the changelog section as the body.pr-lint.yml— enforce conventional-commit PR titles (<type>(<scope>): <description>).changelog-enforce.yml— require the changelog to be updated when source changes (PRs and pushes tomain).
Commit convention: <type>(<scope>): <description> — types feat|fix|docs|style|refactor|perf|test|build|ci|chore|revert, scopes core|api|cli|handlers|tools|plugin|memory|cache|pipeline|extraction|login|navigation|antibot|llm|token|browser|serp|search|config|deps|docs|tests|ci.
A self-contained image with Chromium bundled lives in docker/ with a docker-compose.yml:
# build (context is the repo root, so the workspace crates are visible)
docker build -f docker/Dockerfile -t webrain .
# run (HTTP MCP on 9223)
docker run -p 9223:9223 webrain mcp --http 9223
# or via compose — also mounts persistent volumes for the vault + profiles + engine cache
docker compose -f docker/docker-compose.yml up -d
# multi-arch
docker buildx build --platform=linux/amd64,linux/arm64 -t ghcr.io/prokopis3/webrain -f docker/Dockerfile .failed to spawn Chrome → Chrome isn't installed or WEBRAIN_CHROME is wrong. Run webrain install to download Chrome for Testing.
port X already has a CDP endpoint → a browser is already running there. Stop it, or pick another port (--port N).
Screenshots fail / blank on obscura or lightpanda → they have no paint engine. Obscura errors loudly; lightpanda returns a fake placeholder PNG. Use real Chrome.
Empty a11y tree → the page likely never rendered (consent/challenge page) or the control needs interaction. Check webrain_navigate's challenge; try real Chrome; drop the role filter and use filter on the label.
lightpanda not found / obscura not found → install the binary (see Browser Engines) or set WEBRAIN_LIGHTPANDA / WEBRAIN_OBSCURA.
webrain doctor shows MCP down → start webrain mcp --http 9223; if CDP ports are down, webrain install then start an engine.
Cloudflare/Turnstile blocks the scrape → read the challenge field and use the real-Chrome stealth sidecar (python scripts/stealth_solve.py <url> --cdp-port 9222 --headed), then re-attach webrain to that CDP port.
See CHANGELOG.md. The changelog follows Keep a Changelog + SemVer: unreleased changes live under ## [Unreleased], grouped by ### Added / Changed / Fixed / …, with per-entry **scope**: prefixes matching the commit scopes. CI enforces it (changelog-enforce.yml). Releases are cut by the release workflow, which versions from conventional commits and regenerates the changelog entry.
- Read
AGENTS.mdanddocs/AGENT_DECISION_GUIDE.mdfirst. - For source changes, add a
CHANGELOG.mdentry under[Unreleased]and use conventional commits. - Keep the ponytail contract: delete over add, stdlib before deps, one self-check per non-trivial module. Run
cargo test --workspaceandcargo clippybefore pushing. - Open a PR —
pr-lint.ymlvalidates the title,ci.ymlruns lint + tests.
{ "mcp": { "servers": { "webrain": { "type": "stdio", "command": "webrain", "args": ["mcp"], "env": { "CDP_URL": "http://127.0.0.1:9222" } } } } }