Repository navigation
Releases: vernikr/browser-agent-mcp
Release list
v1.0.10
Added
-
use_proxyonbrowser_action. The daemon has honoured the field on/action
for its whole life, but it was in neither the zod schema nor the advertised
inputSchema, so an agent could only say "all fetches via WARP" and never "this one
goto via WARP". It is now a realbooleanon both surfaces, forwarded verbatim.The description carries the two consequences that are otherwise silent, because an
agent that learns them from a failed turn has already wasted one: switching the route
restarts Chrome and drops cookies and page state, so it belongs once per session
rather than per click; andgotoreturns{ok, url, title}with noblockedfield,
so a forced route that degraded to direct looks like a success.
Changed
browser_actionnow omitsuse_proxyinstead of sendingfalse.false
forbids WARP and drags a proxied session back to direct; omitting it leaves the
session on the route it is on.browser_fetchstill sendsnullon purpose, because
therenullis what enables the daemon's auto-retry — the two tools want different
defaults for the same field, and a test pins each.
v1.0.9
Added
-
The HTTP contract is now checked, not assumed. The
/statusfields this bridge
reads were declared nowhere, so a daemon-side rename failed a live agent with a
confusing 500 instead of this repo's CI.test/contract-fields.mjsdeclares the fields;test/contract.test.mjsasserts
the mock serves them and starts the real daemon and checks its live/status
against the same list. CI installs the daemon for that, which is whyastral-sh/setup-uv
now appears here — the daemon is not a runtime dependency of this repo.The honest limit, documented in the test and in
AGENTS.md: CI installs the daemon
from PyPI, so a rename is caught once it is published, not at the moment it is
merged. When the daemon is absent the live half is skipped, because a missing
binary is an environment problem rather than a contract violation — and a check
that fails for the wrong reason gets disabled.
v1.0.8
Added
-
doctornow hands the agent a runnable repair instead of a report. Every failing
check carries afixfield with the exact command that clears it, and the
plain-text output prints it indented under the check it belongs to. A passing check
carries no fix, because a repair for something that works is noise. The daemon's
doctor does the same, so the two stay symmetric.The
hintcheck that appears when the daemon is unreachable now carries the tunnel
command as a machine-readable field rather than burying it in prose, so an agent can
run it rather than retype it.
v1.0.7
Fixed
-
The bridge told agents to read
journalctlfor their own mistakes. The daemon
reports every failure as a 500, and the bridge classified by status code alone, so
a missingurlparameter and a Chrome that would not start looked identical and
both carried the advice to check the host's logs. That advice is wrong for a bad
argument, and following it for one trains the agent to ignore the real 500s.The daemon now labels its failures (
error_kind: validation | chrome | internal)
and the bridge reads that label instead of guessing from the status. A
validationfailure gets a hint about the argument names and types, is not
retried — the identical request cannot succeed — and no longer holds the
daemon's lock for the length of three doomed attempts. Achromefailure gets the
retry-once-then-check-the-logs advice it deserves.A label the bridge does not recognise is ignored rather than guessed at, so a
future daemon label degrades to the status-based classification instead of
becoming an unknown kind.
v1.0.6
Fixed
-
tunnel stopcould kill processes that have nothing to do with the tunnel.
It ranpkill -f "127.0.0.1:8766:127.0.0.1:8765", which matches any command line
containing that string: a shell that happens to be running the same command,
an editor with the forward spec open, a log tail. Stopping a tunnel could take
an unrelated session down with it.The tunnel is now started as a detached child whose PID is recorded in the
state directory, andstopkills exactly that process — after confirming from
its command line that it really is an ssh carrying this forward, because a PID
is only unique among live processes and whatever recycled it is innocent. The
anchored pattern match remains as the fallback for a tunnel started by an older
version or by hand.ssh -fwas the reason the PID could not be recorded before: it double-forks,
so the child's PID is gone before it could be written down. Detaching the
spawn instead keeps the child alive past this CLI process while leaving its PID
ours to record.
Changed
- The daemon's
/agentnow reportsok: falsewhen the vision loop runs out of
steps or out of its wall-clock budget instead of returning a confident success
containing no answer.BROWSER_AGENT_VISION_BUDGET_SECsets the budget.
v1.0.5
Fixed
browser_agentreturned its whole JSON payload uncapped while every other tool
capped at 40k, so a single run could put tens of kB into the caller's context — the
one thing the "context hygiene is a feature" invariant exists to prevent. Steps are
now projected to the reasoning trail ({step, action, thought}); the coordinates,
keystrokes and typed text were consumed when the action was issued.browser_action
gets the same cap for symmetry.truncate()kept only the head of a long page. The 40k limit already cuts a large
page, and it cut the end — the last rows of a table, the conclusion of an article.
It now keeps a head and a tail with the elision stated between them, and
browser_fetchgainstext_offsetso a long page can be paged deliberately instead
of guessed at.save_text_tostill writes the whole page: paging applies to the
inline path only./fetchasked the daemon for the full rendered HTML on every call and then dropped
it. The daemon has to read the DOM either way, so the default bought nothing while
costing a full-page transfer and a large transient JS heap on a 1 GB host. HTML is
now requested only whensave_html_tois used;html_lengthis still reported
either way.- The
gitleaksjob getsGITHUB_TOKEN, so it no longer fails on every pull request
while passing on pushes tomain.
Changed
info/debuglog lines are buffered and flushed onsetImmediateinstead of one
synchronousappendFileSyncper line — two per tool call plus one per HTTP attempt,
each blocking the stdio server's event loop.warnanderrorstay synchronous:
they are rare, they are what you need after a failure, and a crash is exactly when a
buffer would be lost.- The log file rotates to a single
.1generation past 8 MB instead of growing
without bound into a directory on a 20 GB SSD shared with a live VPN. browser_action's coordinate description claimed a "1280x720 viewport" when the
space is the whole X11 window, page area included. It now describes the real space
and states no offset number, because that gap varies and the daemon measures it.
v1.0.4
Fixed
- A bridge timeout is now reported as
timeoutrather than being indistinguishable
from a client-side cancellation. The two share an exit status but mean opposite
things — one is worth retrying after a cold start, the other means the caller has
already gone away. Timeouts are classifiedtimeoutand are not retried. use_proxyis sent to the daemon asnullwhen unset instead offalse. The
daemon reads the three states differently — unset lets it decide (WARP auto-retry
included),falseforces the proxy off — so the old value silently disabled the
daemon's own fallback. The tool and JSON-schema descriptions now state the
tri-state explicitly.- A missing
sshbinary no longer crashes the bridge with an uncaughtENOENT.
Both spawn sites route through onesshRun()helper whoseerrorhandler prints
a message and exits 1;sshAvailablealso requires a cleanstatus === 0result,
so a non-zero exit no longer counts as "available". ci.yml: thegitleaksjob getsGITHUB_TOKEN. Without it the check fails on
everypull_requestevent while still passing on pushes tomain, so a red PR
said nothing about its own diff.
Changed
- Upstream retry backoff starts at 1 s instead of 3 s, capped at 8 s, so a cold
start is not penalised by a fixed 3 s wait before the first retry.
v1.0.3
Changed
- CI:
ci.ymlgainedconcurrencycancellation and per-jobtimeout-minutes;
the redundanttypecheckstep was dropped —pnpm run testalready compiles
the project withtsc(declarations included) before runningnode:test, so
pnpm run verifywas compiling the same project twice. - CLI numeric flags (
--port,--local,--remote-port,--vnc-port) now
reject a missing, non-numeric or out-of-range value with a usage message and
exit code 2, instead of reachinglisten(NaN)and failing with an opaque
error. Validation lives in the pureparsePorthelper so it is unit-testable
without spawning a process.
Fixed
test/upstream.test.mjs: the abort test's mock upstream waited 5 s although
the abort fires at 100 ms; the wait is now 1.5 s. Same discriminating power,
~3.5 s less per CI run.
Added
.github/dependabot.yml— weekly grouped updates for npm and github-actions.test/cli.test.mjs— coversparsePortend to end.
v1.0.2
Changed
init --runner npxnow writesnpx -y @vernikr/browser-agent-mcp@latest:
npx revalidates the dist-tag against the registry every ~5 minutes, so
bridge updates reach MCP clients within minutes — instead of pnpm dlx's
24-hour TTL (--runner pnpmunchanged; pin an exact version for
reproducibility).
Fixed (docs)
- "Getting bridge updates" previously claimed
npxcaches likepnpm dlx—
wrong; it doesn't (registry packumentCache-Control: max-age=300only).
v1.0.1
Fixed
- Test script works on Node 20 again: the quoted
test/*.test.mjsglob only
expands on Node 22+; the shell now expands it (unquoted) on all versions. - Release workflow upgrades npm CLI before publishing — OIDC trusted
publishing requires npm >= 11.5.1 (Node 22 ships 10.x).