Releases: ChenneyZhuang/laya-browser-agent
Release list
v0.3.2 — Empty-submit guard (a real finding from the v32b default switch)
Found by testing the new default checkpoint against real pages. Asked to "Search products for 'kettle'", v32b clicks the Search button at p=0.74 with the field still empty. The confidence gate can't catch this — the proposal is confident. (The older v10s default only escaped because its confidence on the same state was 0.06: safety by accident.)
Added
- Empty-submit guard in the loop: a click on a submit-like control (search / submit / send / sign-in / checkout / …) while the still-empty field it plainly pairs with is refused — the model is asked again with the fresh observation. Two refusals in one run stop with a clear error instead of burning the step budget. Filling the field first makes the same click go through.
- 5 contract tests covering the guard: refusal, two-strike stop, filled-field pass-through, submit-word-without-empty-field stands down, non-submit words never blocked.
Changed
- The search-flow and reveal-flow live tests accept both guard refusals (confidence gate or empty-submit guard) — the safety outcome either way, as the tests always intended.
- READMEs (en/zh/ja/es): the guard documented alongside the measured failure.
Verified
- 89 fast tests green; ruff clean.
- Live suite 34/34 on v32b AND 34/34 on browser-legacy (no regression).
- CI green.
v0.3.1 — v32b is now the default checkpoint
The fine-tuned v32b checkpoint is now what model="browser" loads. Verified end-to-end on Apple Silicon: download, load, and a real decision through the MLX path.
Changed
- Default model switch:
LayaTorchBackend()/LayaMLXBackend()now loadichenney/laya-browser-v32b. The upstream cklxx checkpoint stays available asmodel="browser-legacy". - README overhaul (en/zh/ja/es): install moved up; honest three-way comparison tables; 1.3 GB default download noted; historical head-to-heads labeled
v10s era — kept for provenance. - Reference table: cklxx/laya-browser documented as the fine-tuning base of the default, not the default.
Fixed
- Version drift: pyproject stayed at 0.2.3 through the v0.3.0 release; now 0.3.1 and CHANGELOG covers 0.3.0 + 0.3.1.
- Nested standalone repos (laya-training-log/, profile-readme/) excluded from the harness repo.
Full training story: laya-training-log
v0.3.0 — Fine-tuned v32b checkpoint: beats the official browser model on 6/8 benchmarks
What's new
The v32b checkpoint — ichenney/laya-browser-v32b
A frozen-encoder, head-only fine-tune of cklxx/laya-browser (Apache-2.0 throughout).
Trained via 15 iterations (v17→v32) of targeted data injection + low-weight head blending:
| Benchmark | v32b | official td | hosted Jev API |
|---|---|---|---|
| recovery2-holdout (240) | 0.7125 | 0.425 | — |
| MiniWoB-116 | 0.9138 | 0.6638 | — |
| browser-suite v4 (70) | 0.5143 | 0.500 | — |
| browser-suite v5 (110) | 0.5636 | 0.5818 | — |
| JevBench hard (111) | 0.4144 | 0.243 | 0.7207 |
| JevBench easy (48) | 0.8542 | 0.979 | 1.0000 |
| decision latency (p50) | 27 ms | — | 854 ms |
Key finding: the upstream training pipeline never produced noul (statement-holds)
training items — probe accuracy for negative judgments was 37%. A purpose-built 8k-item
noul corpus raised it to 100% and lifted done_judgment from 0.615 → 0.769 (ensemble).
Value vs hosted Jev: not absolute accuracy (their 322M+-class cloud model wins
JevBench mean 0.862 vs 0.533) — but $0, fully private, offline, 31× faster, and it
beats Jev on score questions (0.667 vs 0.333) and temporal_numeric (0.33 vs 0.20).
Docs
reports/v20/MULTIDIM_COMPARISON.md— full recipe
lineage v17→v32, per-family breakdowns, the noul root-cause analysisreports/v20/JEV_COMPARISON.md— fresh 231-item
head-to-head against the hosted TypeSafe Jev API (re-measured 2026-09-25 with a live key)
Try it
from localdecide.backends import LayaTorchBackend
backend = LayaTorchBackend(model="ichenney/laya-browser-v32b", subfolder="v32b")Full changes: README benchmark section, curated eval reports, gitignore curation.
v0.2.3 — CORS preflight fix, version unification, HTTP+MCP test suites
What's new
Fixed
- CORS preflight was broken in v0.2.0–v0.2.2.
do_OPTIONSwas defined twice inserve.py— the second definition silently overrode the CORS-correct first one, soOPTIONSresponses carried noAccess-Control-*headers. Browser-based clients (extensions, local web consoles) would have failed at the preflight stage. Now a single handler routing through_send. This was caught by the new HTTP test suite: hand-verification had only checked the 204 status code, never the headers. - Version drift: three different versions existed in the tree (pyproject
0.2.2,mcp_server.SERVER_INFO0.2.0,__init__.__version__0.1.0). All now read from installed package metadata — one source of truth. goal_tokenswas public API by usage but missing from__all__; now exported.
Added
tests/test_serve.py(14 tests): CORS on every response, preflight headers, routing, body limits, error codes — with an injected fake backend, no checkpoint needed.tests/test_mcp.py(13 tests): MCP spec version negotiation (current / legacy / unknown), tools list and call, JSON-RPC error paths.ruffwired intopyproject.tomland CI (correctness rules only — dead imports, redefinitions, loop-variable closures).- Coverage moved from 50% → 65% overall; the HTTP and MCP layers went from 0% to 83% / 81%.
Verification
- 130 tests pass locally (contract + tough + live + serve + mcp + jev_compat)
- CI green on all 7 jobs: 6 contract platforms + the real-checkpoint
apple-silicon-livejob
v0.2.2 — Repo-root portability sweep
What's new
Fixed
- Repo-root portability sweep — after PR #1 exposed one hardcoded path in CI, this release removes the entire class of "only works on the author's machine" assumptions:
scripts/run_live_batched.shresolves the repo root from its own location (fixes theapple-silicon-liveCI job on GitHub runners)- all 17 diagnostic scripts in
examples/diagnostics/resolve paths from__file__instead of/Volumes/SSD/localdecide— they now run from any clone tests/test_live.pydropdown test follows the PR #1 contract (select_option_<index>on multi-dropdown pages)
.gitignorenow explicitly covers.envand localreports/(neither was ever tracked — verified in git history — but the rules are now explicit)
Added
CHANGELOG.md(Keep-a-Changelog format, all releases since 0.1.0 documented)
Verification
- Fresh Python 3.12 venv → editable install with all extras → 103 tests pass (contract + tough + live)
test_jev_compatpasses with a local serve on 8791- CI green on all 7 jobs (contract × 6 platforms + apple-silicon-live with real checkpoint)
v0.2.1 — Literature-backed docs & upstream calibration baselines
What's new
Research-grounded documentation (4 languages)
- The performance table now carries upstream reference points alongside local numbers: Laya's published p50 of 38 ms per question and post-calibration ECE 0.030 across 13 task families — so every local benchmark has a baseline to be judged against.
- New Literature section (EN / 中文 / 日本語 / ES):
- arXiv 2609.23959 — Open-Jev Judgments on CallScreenBench (Sep 2026). Independent evidence that the typed-decision readout reaches AUROC .974 with calibration error .052 at 64.5 ms/decision on a consumer GPU when trained for the task. This grounds the README's phishing finding: v10s misses scams (p≈0.14-0.23) because of training data, not architecture.
- arXiv 2402.09769 — Learning Using a Single Forward Pass: the non-autoregressive decision lineage this model class descends from.
- The Research & framing table gained both papers with explicit "what it contributed" statements.
Also in this release series (v0.2.0)
- Head-to-head batteries vs hosted Jev (single-step, multi-step flows, text classification, browser edge cases) — all reproducible from
examples/diagnostics/ - SPA-aware
page_changedfix in both drivers - MCP protocol-version negotiation per the 2025-06-18 spec
- TypeSafe-compatible
GET /v1/models, CORS + OPTIONS + HEAD on the HTTP service - Credibility pass: every number traced to a results file, every external link verified (13/13 live), star counts and badges current
v0.2.0 — Head-to-head vs hosted Jev + SPA page_changed fix
What's new
Head-to-head vs hosted Jev — measured, not claimed
Two new reproducible batteries in examples/diagnostics/:
jev_head_to_head.py— 12 single-step element-table decisions, same pages, same question contract, local v10s vs TypeSafe productionjev-1.13.0jev_flow_h2h.py— six full task flows through the real browser loop (history, scoping, guards) with either engine answering
Key numbers (single-step, 6 languages): hosted Jev 8/12 strict element hits vs local 4/12; cross-lingual 5/6 vs 1/6; latency parity (618 vs 716 ms median); local confidence skews overconfident (0.90 vs 0.81). Full results and honest interpretation in the README (EN/中文/日本語/ES).
Fixed: page_changed missed SPA-style progress
Both drivers computed page_changed from URL+title only. A click that reveals a dynamic panel without navigating (the single most common agent action on modern pages) reported "no change", so the loop's repeat-detector punished real progress as a stuck loop. The change signature now folds in a DOM fingerprint (visible interactive element count + body text length).
Fixed: batch script referenced removed test classes
scripts/run_live_batched.sh listed batches that no longer exist (TestShadowDOM, TestHiddenMenus, TestRTLAndExtendedScripts, TestGroundingMultiLanguage) and missed TestMemorySafety.
CI: all platforms green. Live suite: 33 passed, 1 skipped across 5 batches.
Follow-up (same release): text-classification and edge-case batteries
jev_text_h2h.py— 21 real business texts: pool-lead triage, Chinese SMS triage, robustness probes. Local v10s wins latency decisively (36 ms vs 738 ms median) but misses all 3 phishing cases (p=0.14-0.23); hosted Jev catches all 3 (p=0.96) and calibrates lead-quality scores correctly.jev_edge_h2h.py— 9 restraint goals: local 4/9 vs hosted 5/9, failing in opposite directions (local over-acts with ~0.9 confidence; hosted over-blocks). Harness guards remain the safety net for both.
Fine-tuning targets this data justifies: Chinese grounding, phishing detection, and goal-restraint — all documented in the four READMEs.
v0.1.1 — Verified against TypeSafe production Jev API
What's new
Verified against TypeSafe's production Jev API
The systemone wire dialect was validated end-to-end against the real hosted endpoint (api.typesafe.ai/v1/systemone, model jev-1.13.0):
criteriais required for every question type (noul / choice / score) — the API rejects questions without it- For
choice,criteriais a map of option → rubric description; forscore, an array of level names - The same payload pointed at the bundled
localdecide serveproduces the same answer shape from the local Laya checkpoint — hosted ↔ local switching is one base URL change
Docs
- New "Verified against the real Jev API" section with a working request/response example, in all four READMEs (EN / 中文 / 日本語 / ES)
v0.1.0 — First public release
v0.1.0 — First public release
Browser-agent decisions from a local, open-weight System 1 model. No cloud, no API key, no screenshots.
What works today
- Decision loop: observe → decide → act with confidence gating, a checkbox-undoing guard, loop detection, and a human-confirmation gate for destructive actions
- Two drivers: Playwright (own browser) and CDP (attach to your logged-in Chrome)
- Three wire dialects: TypeSafe-compatible
/v1/systemone, native/v1/decide,/v1/table, plus an MCP server over stdio (localdecide-mcp) for Claude Desktop and Cursor - Cross-language grounding: non-Latin goals get same-script option filtering (1/9 → 8/9 measured across 10 scripts)
- Agent skills: portable SKILL.md files +
python3 install_skills.pyfor Claude Code / Codex / Cursor / Hermes
Measured on an M4 (16 GB)
| Short-state decision | 10–30 ms |
| Scoped browser step (20 elements) | ~333 ms |
| Throughput | up to ~100 decisions/s |
| Multilingual (with grounding) | 8/9 goals hit exactly |
Install
pip install 'localdecide[mlx]' # Apple Silicon
pip install 'localdecide[torch]' # Linux / Windows / Intel Mac
localdecide doctor # hardware check + smoke testFull details, limitations, and per-device expectations in the README.