Skip to content

Releases: ChenneyZhuang/laya-browser-agent

v0.3.2 — Empty-submit guard (a real finding from the v32b default switch)

Choose a tag to compare

@ChenneyZhuang ChenneyZhuang released this 26 Sep 01:06

Found by testing the new default checkpoint against real pages. Asked to "Search products for 'kettle'", v32b clicks the Search button at p=0.74 with the field still empty. The confidence gate can't catch this — the proposal is confident. (The older v10s default only escaped because its confidence on the same state was 0.06: safety by accident.)

Added

  • Empty-submit guard in the loop: a click on a submit-like control (search / submit / send / sign-in / checkout / …) while the still-empty field it plainly pairs with is refused — the model is asked again with the fresh observation. Two refusals in one run stop with a clear error instead of burning the step budget. Filling the field first makes the same click go through.
  • 5 contract tests covering the guard: refusal, two-strike stop, filled-field pass-through, submit-word-without-empty-field stands down, non-submit words never blocked.

Changed

  • The search-flow and reveal-flow live tests accept both guard refusals (confidence gate or empty-submit guard) — the safety outcome either way, as the tests always intended.
  • READMEs (en/zh/ja/es): the guard documented alongside the measured failure.

Verified

  • 89 fast tests green; ruff clean.
  • Live suite 34/34 on v32b AND 34/34 on browser-legacy (no regression).
  • CI green.

v0.3.1 — v32b is now the default checkpoint

Choose a tag to compare

@ChenneyZhuang ChenneyZhuang released this 25 Sep 23:32

The fine-tuned v32b checkpoint is now what model="browser" loads. Verified end-to-end on Apple Silicon: download, load, and a real decision through the MLX path.

Changed

  • Default model switch: LayaTorchBackend() / LayaMLXBackend() now load ichenney/laya-browser-v32b. The upstream cklxx checkpoint stays available as model="browser-legacy".
  • README overhaul (en/zh/ja/es): install moved up; honest three-way comparison tables; 1.3 GB default download noted; historical head-to-heads labeled v10s era — kept for provenance.
  • Reference table: cklxx/laya-browser documented as the fine-tuning base of the default, not the default.

Fixed

  • Version drift: pyproject stayed at 0.2.3 through the v0.3.0 release; now 0.3.1 and CHANGELOG covers 0.3.0 + 0.3.1.
  • Nested standalone repos (laya-training-log/, profile-readme/) excluded from the harness repo.

Full training story: laya-training-log

v0.3.0 — Fine-tuned v32b checkpoint: beats the official browser model on 6/8 benchmarks

Choose a tag to compare

@ChenneyZhuang ChenneyZhuang released this 25 Sep 07:30

What's new

The v32b checkpoint — ichenney/laya-browser-v32b

A frozen-encoder, head-only fine-tune of cklxx/laya-browser (Apache-2.0 throughout).
Trained via 15 iterations (v17→v32) of targeted data injection + low-weight head blending:

Benchmark v32b official td hosted Jev API
recovery2-holdout (240) 0.7125 0.425 —
MiniWoB-116 0.9138 0.6638 —
browser-suite v4 (70) 0.5143 0.500 —
browser-suite v5 (110) 0.5636 0.5818 —
JevBench hard (111) 0.4144 0.243 0.7207
JevBench easy (48) 0.8542 0.979 1.0000
decision latency (p50) 27 ms — 854 ms

Key finding: the upstream training pipeline never produced noul (statement-holds)
training items — probe accuracy for negative judgments was 37%. A purpose-built 8k-item
noul corpus raised it to 100% and lifted done_judgment from 0.615 → 0.769 (ensemble).

Value vs hosted Jev: not absolute accuracy (their 322M+-class cloud model wins
JevBench mean 0.862 vs 0.533) — but $0, fully private, offline, 31× faster, and it
beats Jev on score questions (0.667 vs 0.333) and temporal_numeric (0.33 vs 0.20).

Docs

Try it

from localdecide.backends import LayaTorchBackend
backend = LayaTorchBackend(model="ichenney/laya-browser-v32b", subfolder="v32b")

Full changes: README benchmark section, curated eval reports, gitignore curation.

v0.2.3 — CORS preflight fix, version unification, HTTP+MCP test suites

Choose a tag to compare

@ChenneyZhuang ChenneyZhuang released this 24 Sep 14:14

What's new

Fixed

  • CORS preflight was broken in v0.2.0–v0.2.2. do_OPTIONS was defined twice in serve.py — the second definition silently overrode the CORS-correct first one, so OPTIONS responses carried no Access-Control-* headers. Browser-based clients (extensions, local web consoles) would have failed at the preflight stage. Now a single handler routing through _send. This was caught by the new HTTP test suite: hand-verification had only checked the 204 status code, never the headers.
  • Version drift: three different versions existed in the tree (pyproject 0.2.2, mcp_server.SERVER_INFO 0.2.0, __init__.__version__ 0.1.0). All now read from installed package metadata — one source of truth.
  • goal_tokens was public API by usage but missing from __all__; now exported.

Added

  • tests/test_serve.py (14 tests): CORS on every response, preflight headers, routing, body limits, error codes — with an injected fake backend, no checkpoint needed.
  • tests/test_mcp.py (13 tests): MCP spec version negotiation (current / legacy / unknown), tools list and call, JSON-RPC error paths.
  • ruff wired into pyproject.toml and CI (correctness rules only — dead imports, redefinitions, loop-variable closures).
  • Coverage moved from 50% → 65% overall; the HTTP and MCP layers went from 0% to 83% / 81%.

Verification

  • 130 tests pass locally (contract + tough + live + serve + mcp + jev_compat)
  • CI green on all 7 jobs: 6 contract platforms + the real-checkpoint apple-silicon-live job

v0.2.2 — Repo-root portability sweep

Choose a tag to compare

@ChenneyZhuang ChenneyZhuang released this 24 Sep 13:57

What's new

Fixed

  • Repo-root portability sweep — after PR #1 exposed one hardcoded path in CI, this release removes the entire class of "only works on the author's machine" assumptions:
    • scripts/run_live_batched.sh resolves the repo root from its own location (fixes the apple-silicon-live CI job on GitHub runners)
    • all 17 diagnostic scripts in examples/diagnostics/ resolve paths from __file__ instead of /Volumes/SSD/localdecide — they now run from any clone
    • tests/test_live.py dropdown test follows the PR #1 contract (select_option_<index> on multi-dropdown pages)
  • .gitignore now explicitly covers .env and local reports/ (neither was ever tracked — verified in git history — but the rules are now explicit)

Added

  • CHANGELOG.md (Keep-a-Changelog format, all releases since 0.1.0 documented)

Verification

  • Fresh Python 3.12 venv → editable install with all extras → 103 tests pass (contract + tough + live)
  • test_jev_compat passes with a local serve on 8791
  • CI green on all 7 jobs (contract × 6 platforms + apple-silicon-live with real checkpoint)

v0.2.1 — Literature-backed docs & upstream calibration baselines

Choose a tag to compare

@ChenneyZhuang ChenneyZhuang released this 23 Sep 14:02

What's new

Research-grounded documentation (4 languages)

  • The performance table now carries upstream reference points alongside local numbers: Laya's published p50 of 38 ms per question and post-calibration ECE 0.030 across 13 task families — so every local benchmark has a baseline to be judged against.
  • New Literature section (EN / 中文 / 日本語 / ES):
    • arXiv 2609.23959 — Open-Jev Judgments on CallScreenBench (Sep 2026). Independent evidence that the typed-decision readout reaches AUROC .974 with calibration error .052 at 64.5 ms/decision on a consumer GPU when trained for the task. This grounds the README's phishing finding: v10s misses scams (p≈0.14-0.23) because of training data, not architecture.
    • arXiv 2402.09769 — Learning Using a Single Forward Pass: the non-autoregressive decision lineage this model class descends from.
  • The Research & framing table gained both papers with explicit "what it contributed" statements.

Also in this release series (v0.2.0)

  • Head-to-head batteries vs hosted Jev (single-step, multi-step flows, text classification, browser edge cases) — all reproducible from examples/diagnostics/
  • SPA-aware page_changed fix in both drivers
  • MCP protocol-version negotiation per the 2025-06-18 spec
  • TypeSafe-compatible GET /v1/models, CORS + OPTIONS + HEAD on the HTTP service
  • Credibility pass: every number traced to a results file, every external link verified (13/13 live), star counts and badges current

v0.2.0 — Head-to-head vs hosted Jev + SPA page_changed fix

Choose a tag to compare

@ChenneyZhuang ChenneyZhuang released this 22 Sep 05:15

What's new

Head-to-head vs hosted Jev — measured, not claimed

Two new reproducible batteries in examples/diagnostics/:

  • jev_head_to_head.py — 12 single-step element-table decisions, same pages, same question contract, local v10s vs TypeSafe production jev-1.13.0
  • jev_flow_h2h.py — six full task flows through the real browser loop (history, scoping, guards) with either engine answering

Key numbers (single-step, 6 languages): hosted Jev 8/12 strict element hits vs local 4/12; cross-lingual 5/6 vs 1/6; latency parity (618 vs 716 ms median); local confidence skews overconfident (0.90 vs 0.81). Full results and honest interpretation in the README (EN/中文/日本語/ES).

Fixed: page_changed missed SPA-style progress

Both drivers computed page_changed from URL+title only. A click that reveals a dynamic panel without navigating (the single most common agent action on modern pages) reported "no change", so the loop's repeat-detector punished real progress as a stuck loop. The change signature now folds in a DOM fingerprint (visible interactive element count + body text length).

Fixed: batch script referenced removed test classes

scripts/run_live_batched.sh listed batches that no longer exist (TestShadowDOM, TestHiddenMenus, TestRTLAndExtendedScripts, TestGroundingMultiLanguage) and missed TestMemorySafety.

CI: all platforms green. Live suite: 33 passed, 1 skipped across 5 batches.

Follow-up (same release): text-classification and edge-case batteries

  • jev_text_h2h.py — 21 real business texts: pool-lead triage, Chinese SMS triage, robustness probes. Local v10s wins latency decisively (36 ms vs 738 ms median) but misses all 3 phishing cases (p=0.14-0.23); hosted Jev catches all 3 (p=0.96) and calibrates lead-quality scores correctly.
  • jev_edge_h2h.py — 9 restraint goals: local 4/9 vs hosted 5/9, failing in opposite directions (local over-acts with ~0.9 confidence; hosted over-blocks). Harness guards remain the safety net for both.

Fine-tuning targets this data justifies: Chinese grounding, phishing detection, and goal-restraint — all documented in the four READMEs.

v0.1.1 — Verified against TypeSafe production Jev API

Choose a tag to compare

@ChenneyZhuang ChenneyZhuang released this 22 Sep 03:45

What's new

Verified against TypeSafe's production Jev API

The systemone wire dialect was validated end-to-end against the real hosted endpoint (api.typesafe.ai/v1/systemone, model jev-1.13.0):

  • criteria is required for every question type (noul / choice / score) — the API rejects questions without it
  • For choice, criteria is a map of option → rubric description; for score, an array of level names
  • The same payload pointed at the bundled localdecide serve produces the same answer shape from the local Laya checkpoint — hosted ↔ local switching is one base URL change

Docs

  • New "Verified against the real Jev API" section with a working request/response example, in all four READMEs (EN / 中文 / 日本語 / ES)

v0.1.0 — First public release

Choose a tag to compare

@ChenneyZhuang ChenneyZhuang released this 21 Sep 13:00

v0.1.0 — First public release

Browser-agent decisions from a local, open-weight System 1 model. No cloud, no API key, no screenshots.

What works today

  • Decision loop: observe → decide → act with confidence gating, a checkbox-undoing guard, loop detection, and a human-confirmation gate for destructive actions
  • Two drivers: Playwright (own browser) and CDP (attach to your logged-in Chrome)
  • Three wire dialects: TypeSafe-compatible /v1/systemone, native /v1/decide, /v1/table, plus an MCP server over stdio (localdecide-mcp) for Claude Desktop and Cursor
  • Cross-language grounding: non-Latin goals get same-script option filtering (1/9 → 8/9 measured across 10 scripts)
  • Agent skills: portable SKILL.md files + python3 install_skills.py for Claude Code / Codex / Cursor / Hermes

Measured on an M4 (16 GB)

Short-state decision 10–30 ms
Scoped browser step (20 elements) ~333 ms
Throughput up to ~100 decisions/s
Multilingual (with grounding) 8/9 goals hit exactly

Install

pip install 'localdecide[mlx]'        # Apple Silicon
pip install 'localdecide[torch]'      # Linux / Windows / Intel Mac
localdecide doctor                    # hardware check + smoke test

Full details, limitations, and per-device expectations in the README.