Skip to content

Releases: coff33ninja/go-mcp-computer-use

v0.3.10

Choose a tag to compare

@github-actions github-actions released this 18 Sep 11:48

[0.3.10] - 2026-09-18

Post-0.3.9 computer-use reliability pass: focus/keylogger chain bugs found in live OpenCode testing, plus the ML outcome ledger (desktop-as-teacher). Includes a full local computer-use data reset for clean testing (datalog, training store, memory, transformer weights; ONNX models kept).

Fixed

  • focus_window no longer reverse-maximizes windows — FocusWindow called ShowWindow(SW_RESTORE) unconditionally, which un-maximizes a maximized window every time it is focused (OpenCode observed browsers/editors snap out of maximize). Restore is now applied only when IsIconic (minimized); otherwise SW_SHOW keeps the current maximized/normal placement.
  • Keylogger/replicate chain replay missing tool — keylogger_stop emitted "tool":"_focus" and eventsToSmartSteps emitted FocusWindow steps with an empty Tool. Chain dispatch then failed with unknown tool: _focus / unknown tool: (confirmed in chain_log). Fixes:
    • keylogger emits focus_window_by_title
    • smart steps carry Tool: focus_window_by_title + focus_window field
    • chain treats focus-only steps (empty Tool + focus field) as successful focus after auto-focus
    • toolDispatch gains focus_window_by_title, _focus (legacy alias), keylogger_start/stop/status, replicate
  • Local tests no longer require live APPDATA assumptions for ledger schema — ml_predictions is created on demand; supervised samples never include unknown/pending.

Added

  • ml_predictions outcome ledger (desktop as teacher) — predictions from ml_query / agent_suggest (transformer or statistical) are inserted as immutable rows (engine, query/ocr/window hashes, pred_tool, pred_x/y, confidence, model_version). Only resolution columns update: status, outcome_source, actual_*, verification_data, resolved_at.
  • Outcome resolver on click — ResolveMLPredictionsForAction scores pending preds after a real click:
    • hit — click succeeded within hit radius (~3% screen, min 80px) of a pending coord
    • miss — click failed near a pending pred (or same-tool retry nearby)
    • unknown — pending rows older than 10m (timeout); never treated as failure
    • recovered — a miss followed within ~45s by a successful nearby same-tool action
    • Successful clicks far from predictions leave rows pending (agent may have ignored ML) — no poison labels
  • ml_outcome_status tool + ml_status.outcome — hit/miss/unknown/recovered counts, hit_rate, recovery_rate, unknown_rate, last_hour_hit_rate, eligible_train.
  • Resolved-only supervised signal — SupervisedLedgerSamples returns only hit|miss|recovered; transformer train pool prefers training_pairs.success=1; unknown/pending never train.
  • Chain aliases for keylogger/replicate replay — focus_window_by_title, _focus, keylogger_start/stop/status, replicate (executes provided steps).

Changed

  • Clean-slate computer-use data for testing — local %APPDATA%\go-mcp-computer-use training/datalog/memory/ML weights archived then cleared so 0.3.10 testing starts from zero pairs (ONNX models in models\ retained). Agents must re-collect via normal use; agent_train / scripts/eval-ml.ps1 -Train rebuild from new data only.

Build / verification

  • scripts/build.ps1 — OK (version 0.3.10, Zig cc + CGO)
  • scripts/lint.ps1 — go vet clean + build OK
  • go test -short ./internal/actions — pass (includes focus/keylogger chain alias tests)

v0.3.9

Choose a tag to compare

@github-actions github-actions released this 18 Sep 10:39

[0.3.9] - 2026-09-18

Fixed

  • Transformer training never saw real click coordinates — production training_pairs.command_json stores args as a nested JSON string with mixed-case keys ({"args":"{\"X\":700,\"Y\":400}","tool":"click"}). ml/trainer.decodeCoords only unmarshaled top-level lowercase x/y, so on live datalogs every click/hover coord target was (0,0) (measured: 800/800 samples). Fix: ml/dataloader.NormalizeCommandJSON / ApplyNormalization unwrap string-or-object args, accept X/Y/from_x/to_x case-insensitively, and fill Sample.CoordX/Y + FromCoordX/Y. Trainer decodeCoords uses the same unwrap path; makeTargetFromSample prefers loader-normalized pixels. Unit tests cover the production nested-string shape.
  • Transformer unusable after restart (empty tokenizer) — MLEngine.LoadModel created a fresh unfitted tokenizer and never loaded vocab. tokenizer.Encode returns nil when unfitted, so Forward failed with token len 0 != maxLen 128 while IsReady() could still be true; predictions silently fell through to the statistical engine. Fix: training persists vocab.bin next to model.gob (tokenizer Save/Load now wired in production), plus ml_meta.json (vocab size, ArgDim/WindowDim/FromCoordDim, tool list, holdout metrics). LoadModel refuses ready-state without vocab and surfaces a clear last_error.
  • Vocab size vs embedding table mismatch — fitted vocabs on real OCR exceeded the hardcoded VocabSize=2000 (observed 2421), so high token IDs were dropped in embedding lookup. modelConfigFor() sizes the model from the fitted tokenizer (rounded up) and is shared by Train / LoadModel / online / finetune paths.
  • Spatial features were always zero — prepareBatch called encoder.Encode(0, 0) for every sample, and inference used an all-zero coord vector. Training now encodes the sample's decoded click/drag pixels; inference sets a context point from the live cursor via predict.Engine.SetContextPoint.
  • Online training destroyed the tokenizer — trainFromBuffer refit the vocab on a 32-sample replay batch then ran TrainEpoch over the full SQLite DB, scrambling token IDs against existing embeddings. Online updates now train on replay-derived samples without Fit, keep embedding IDs stable, and re-save vocab.bin + report real eval loss/accuracy (checkpoint accuracy is no longer hardcoded 0.0).
  • Honest tool-accuracy metric — trainer.Accuracy previously took argmax over toolStart (tools + coord + arg dims), mixing continuous outputs into classification and inflating scores. It now scores argmax over tool logits only, skips unknown tools, and Train logs click_acc / non_click_acc alongside overall holdout accuracy and the majority-class baseline.
  • agent_train only rebuilt statistical indexes — it now also trains the Go transformer (writes model.gob + vocab.bin + ml_meta.json) and returns ml_status in the tool payload.
  • scripts/lint.ps1 failed when launched from scripts/ — Get-Content VERSION resolved relative to the script directory. Now uses $PSScriptRoot\..\VERSION and builds ..\cmd\mcp-server.

Added

  • ml_status on agent ML tools — agent_train, agent_suggest, and chain_predict return transformer health: ready, model_loaded, vocab_loaded, vocab_size, paths, last_eval_loss, last_eval_accuracy, majority_baseline, last_error, and a source verdict (transformer | statistical_preferred | unavailable).
  • source field on PredictedAction — predictions are labeled transformer or statistical so agents know which engine produced a coordinate/command guess.
  • Honest transformer gate — MLEngine.Predict returns nil when outputs are near-uniform (no signal) or when last holdout tool accuracy is below the majority-class baseline, so PredictActions falls back to the statistical adaptive engine instead of shipping junk coords.
  • cmd/ml-eval + scripts/eval-ml.ps1 — local eval harness: tool distribution, coord-decode sanity (with_coords %), majority baseline, optional -Train, ml_status dump, and smoke predict on real datalog OCR. Exit code 2 when model accuracy is below majority baseline.
  • Class balancing for click-heavy logs — balanceTools upsamples minority tools (cap 48) before augmentation so a 70%+ click datalog does not collapse the model to “always click”. Tool one-hot targets are amplified (2.0) so MSE is not drowned by zero coord/arg dims.
  • Production-shaped ML tests — ml/dataloader/normalize_test.go and ml/trainer/decode_coords_test.go lock the nested-args schema that unit tests previously missed (tests used {"x":...} while logs used {"args":"{\"X\":...}"}).

Changed

  • README ML section is truthful — the Go-native transformer is described as experimental with vocab/meta persistence, majority-baseline reporting, and an explicit preference for the statistical engine (ml_query/ml_teach) when the neural path underperforms. OCR/UIA/statistical priors remain the reliable locators.
  • agent_train tool description — documents dual training (adaptive indexes + transformer artifacts) and ml_status in the response.
  • docs/reference/tools.md regenerated after handler/description updates.
  • Live retrain on a real datalog (800 pairs) — after the data-path fix, with_coords went from 0% → 70.4% for transformer targets; model.gob + vocab.bin + ml_meta.json load correctly across process restarts. Holdout tool accuracy on noisy OCR remains at or near the majority baseline — the neural path is gated and labeled rather than silently preferred. Statistical ml_query/agent_suggest now actually learn token→coord averages from unwrapped production args.

Build / verification

  • scripts/build.ps1 — OK (mcp-server.exe, Zig cc + CGO, version 0.3.9)
  • scripts/lint.ps1 — go vet clean + build OK
  • go test — internal/actions short suite + ml/dataloader, ml/trainer, ml/predict, ml/tokenizer, ml/transformer pass

v0.3.8

Choose a tag to compare

@github-actions github-actions released this 16 Sep 03:38

[0.3.8] - 2026-08-27

Fixed

  • YOLO UI detection now letterboxes non-square inputs instead of aspect-distorting them - preprocessYOLO previously rescaled X and Y independently to a 640x640 square, so a non-square input (e.g. the 3200x1980 virtual desktop) was squashed by 5x on X and 3.09x on Y, destroying every element's geometry. The detector then emitted floor-confidence (~0.5) sigma-degenerate 0x0 boxes everywhere. Fix: aspect-preserving letterbox (one uniform scale min(640/w,640/h), centered gray padding) shared between the encoder and decoder via a yoloLetterbox struct, and parseYOLOOutput now inverse-maps boxes with (out - pad) / scale. Verified live: full-screen detection now yields real, varied boxes and click points with confidence 0.5-1.0 instead of thousands of 0x0 floor boxes.
  • elements_query / onnx_detect no longer return garbage off-screen boxes from the letterbox padding - ONNXDetect inverse-mapped detections whose centers sat inside the 640x640 letterbox's gray padding band back into negative or oversized source coordinates (observed: {948,842,839,236} -> click y:960 on an 860-tall window). Fix: a new clipElementToImage gate filters every detection against the captured bitmap bounds - boxes whose center lands outside the image are dropped (padding false positives), and partially off-screen boxes are clipped to the image so screen_box/click_point are always actionable. Verified by unit tests covering all reported garbage shapes.
  • ocr_window / screenshot_element now report + key the captured window, not the foreground window - annotateCaptureOpts stamped window_title from getActiveWindowTitle() (the FOREGROUND window) regardless of the handle being captured, so handle-scoped captures mislabeled the wrong window and keyed per-window ML-priors from that wrong title. Fix: annotateCaptureOpts takes a windowTitle param and AnnotateWindow resolves getWindowTitle(handle) for the actual captured handle. Verified live: ocr_window(132082) (Mozilla Firefox) now reports window_title="Mozilla Firefox" (was "OpenCode").
  • Capture responses no longer dump a giant element array on every ocr/screenshot/onnx_detect call - the fused YOLO+MobileNet annotation (from 0.3.5/0.3.6) embedded every detected box into every capture response with no cap, so a busy desktop returned thousands of degenerate icon boxes (~1751 observed) and blew the payload to several hundred KB even with include_image:false. The AI was forced to grep a monolithic JSON blob. Fix: annotateCaptureOpts now runs FilterAndCapElements before returning - it drops boxes whose fused combined_confidence is below a floor (default 0.40) and caps the surviving list (default 50, highest-confidence first). Detection/ML still runs at full fidelity internally; only what's embedded on the wire is bounded. Verified by unit tests; a 1751-element dump is now at most 50 high-trust rows.

Added

  • include_image opt-out for the base64 image bloat in capture-tool responses - every ocr/screenshot/screenshot_element/ocr_window/ocr_active_window/onnx_detect/onnx_classify response embeds the base64 screenshot (~0.7-1.5MB), which wastes model context and forces truncation for text-only AIs. This adds a config flag include_image (default true, preserving current behavior) that globally strips the returned image_b64 copy while keeping OCR text, element boxes/confidence, and all metadata intact - the internal ML detection pipeline still runs on the frame; only the returned copy is removed. A per-call include_image override (true/false) on any capture tool wins over the global config, giving callers a tri-state. Exposed through set_config/get_config and persisted. Verified live: with include_image:false an ocr_window returns metadata only (no image_b64); a per-call include_image:true restores it; onnx_detect(include_image:false) still returns 1677 elements / 350 clickable with the image stripped.
  • exclude_elements opt-out for text-only captures - a config flag exclude_elements (default false, preserving current behavior) plus a per-call exclude_elements override on ocr, ocr_window, ocr_active_window, screenshot, screenshot_element, and onnx_detect empties the returned element array so a text-only client never drags the detection boxes onto the wire. Like include_image, it flows through the single safeHandler choke point (stripImageIfExcluded), is persisted via set_config/get_config, and the per-call override wins over the global value.
  • elements_query tool - a grep-style query over the fused capture for text-based LLMs - instead of dumping the full annotated JSON, the AI asks "where is the submit button" and gets a compact flat projection: filter by source (screen/window/region), YOLO class, MobileNet label, an x,y,w,h region, min_confidence, clickable_only, and OCR text (returns matching words with screen coords). Each returned row is just {index, class, label, confidence, clickable, screen_box, image_box, click_point} - the coordinates to click - with no base64 and no nested classifier bloat. Registration mirrors the existing onnx_* tool family.
  • scripts/gen-icons.ps1 no longer writes a UTF-8 BOM into winres/winres.json - Set-Content -Encoding UTF8 (PowerShell 7) prepends a BOM, which go-winres rejects with invalid character 'ï' looking for beginning of value, breaking the resource-generation step and hence the whole build. Fix: write with -Encoding utf8NoBOM so the generated JSON parses cleanly.
  • key_press key names expanded to the full Windows virtual-key surface + case-insensitive lookup + generic modifier combos - vkModMap now exposes the explicit left/right modifier variants (LCTRL/LCONTROL, RCTRL/RCONTROL, LALT, RALT, LSHIFT/RSHIFT, WIN/META/LWIN/SUPER/CMD/COMMAND, RWIN), and vkSpecialMap gained the numpad block (NUMPAD0-NUMPAD9, NUMPAD_MULTIPLY/ADD/SEPARATOR/SUBTRACT/DECIMAL/DIVIDE), the media/volume keys (VOLUME_MUTE/VOLUME_DOWN/VOLUME_UP, MEDIA_NEXT_TRACK/PREV_TRACK/STOP/PLAY_PAUSE), plus SNAPSHOT, APPS/CONTEXT, CLEAR, EXECUTE, and SLEEP. keyNameToVK now upper-cases its input so lookups are case-insensitive, and single-char punctuation resolves through the charToVK table. KeyPress generalized the old CTRL-only MOD+key prefix into any modmap modifier followed by a single letter or digit (CTRL+A, WIN+R, ALT+F4, SHIFT+1, ...) and emits the requested modifier's real VK instead of hardcoding Ctrl. Covered by new internal/actions/keyboard_test.go unit tests.

v0.3.7

Choose a tag to compare

@github-actions github-actions released this 27 Aug 14:55

[0.3.7] - 2026-08-27

Added

  • --version, --license, and --help CLI flags on the server binary - mcp-server.exe --version prints the build version (go-mcp-computer-use 0.3.7), --license prints the full Apache-2.0 text plus the NOTICE, and --help lists the available subcommands/flags. LICENSE and NOTICE are embedded into the binary via go:embed (kept in sync by scripts/build.ps1), so the license text is always retrievable from the shipped executable with no files alongside.

Changed

  • Project relicensed from MIT to Apache-2.0 - LICENSE replaced with the full Apache-2.0 text, (c) 2026 coff33ninja. Added a NOTICE file wiring the Apache-2.0 attribution and clarifying that the bundled gpa_gui_detector.onnx (a converted Salesforce GPA-GUI-Detector) remains under the upstream MIT license, not Apache-2.0. README gained a license badge, a License section, and updated model-attribution wording.
  • Version-info, copyright, and admin manifest now embedded in the executable - resource generation switched from akavel/rsrc (icon only, no manifest/version) to go-winres (scripts/gen-icons.ps1 + a committed winres/winres.json.template). The exe now carries a VS_VERSION_INFO block visible in Windows Explorer Properties -> Details (CompanyName coff33ninja, LegalCopyright "Copyright (c) 2026 coff33ninja. Licensed under the Apache License, Version 2.0.", ProductName, File/Product version from VERSION) and an application manifest declaring requireAdministrator (admin is required for the tool's UIA/UIPI automation to fully work) plus per-monitor-v2 DPI awareness. Previously the exe had no embedded manifest at all (as-invoker, no DPI declaration), so this is an explicit behavioral improvement that enforces the project's real admin requirement at the OS level.
  • License decisions: LICENSE + NOTICE are shipped with each release - .github/workflows/release.yml now attaches LICENSE and NOTICE to every versioned release alongside mcp-server.exe, so the license and third-party model attribution always accompany the binary.

v0.3.6

Choose a tag to compare

@github-actions github-actions released this 27 Aug 13:04

[0.3.6] - 2026-08-27

Added

  • onnx_classify tool — MobileNetV3 GUI element classifier — a new advisory tier that classifies UI content using mobilenetv3_small.onnx across 15 classes (button, checkbox, container, dropdown, icon_button, image, label, link, menu_item, scrollbar, slider, tab, text_input, toggle, unknown). Supported source values: screen (full screen), window (active window), region (x,y,w,h), elements (classify every YOLO-detected element crop), crop (raw base64 image). Returns label + confidence for top-N (default 3) per target. Advisory only — it returns an additional signal and never hard-blocks an action.
  • MobileNet as a chain verification tier — chain verify steps accept an optional classify config ({enabled, top_n, source}). When enabled, the executed step's result carries a classify advisory payload with the top classifications, independent of the pass/fail decision.
  • MobileNet as a watcher advisory tier — the background watcher now classifies each YOLO-detected element crop and attaches classifications to its cached detection, giving the AI type+confidence context per element without affecting detection or training. Best-effort and skipped entirely when the model file is absent.
  • AnnotatedCapture / AnnotateScreen / AnnotateRegion / AnnotateWindow — the fused, best-effort annotated-capture production layer. All classifier signals run on the SAME captured frame so geometry stays aligned. Any engine that is unavailable yields an empty sub-block (noted in errors) without failing the capture.
  • scripts/gpa-gui-export/ — redoable GPA-GUI→ONNX conversion — a uv-based project (pyproject.toml + export.py + README.md) that downloads Salesforce/GPA-GUI-Detector model.pt via huggingface_hub, validates the single icon class layout, exports to ONNX (1,3,640,640)→(1,5,8400) opset 12, and writes gpa_gui_detector.onnx. uv sync + uv run export.py regenerates the artifact from source; it is what CI uses to produce the release asset.
  • Main README "Models" section — documents the three runtime models, auto-download behavior, and how to regenerate the detector, plus a "UI-aware element detection" feature bullet.

Changed

  • Every click now returns an advisory validated block — click and find_text_and_click results carry a post-click validation combined into MobileNet classification of the click target (top-N label + confidence), the element-priors DB (sample count, prior-adjusted confidence, known-location flag, learned frequency/position), and an ML-memory cross-reference (whether the model has seen this click context before). Hooked at the central Click() choke point so every click source gets it. The validated block is merged at the top level of the tool result so the AI sees a consistent shape whether or not OCR auto-verify ran. Best-effort and non-blocking: if the model or capture fails, the click still succeeds and the validation block simply carries minimal/no data.
  • Internal refactor — ClassifyElements now delegates to a shared classifyElementList helper (also used by the watcher), keeping crop capture + model inference in one place.
  • Annotated capture on all AI-facing capture tools — screenshot, screenshot_element, ocr, ocr_window, and ocr_active_window now always-on return a fused AnnotatedCapture (OCR text + YOLO element boxes + MobileNet per-element classification + element-priors + ML-memory) with each element exposed in BOTH bitmap-image space and virtual-screen space (screen_box), so text-bound AIs know exactly what is on screen and where to click regardless of whether a vision model is available. Screenshot tools keep the raw b64 as text content for vision AIs while the annotated map rides alongside as structured result. screenshot/screenshot_element gained an optional language param passed to the OCR signal.
  • onnx_detect and onnx_classify now return the same fused annotated capture — instead of separate raw detection/classification payloads, onnx_detect and onnx_classify (screen/window/region) now produce the identical AnnotatedCapture shape as every other capture tool, so the classifier signals are aligned on the same frame with dual image_box/screen_box coords, click_point, combined_confidence, and the clickable gate — no redundant double-inference. onnx_detect honors its threshold/iou_threshold args; onnx_classify elements/crop sources return the classification wrapped in the same shape via a shim (no live capture-derived screen boxes in those cases).
  • The clickable gate and confidence fusion now trust the MobileNet UI label as the authoritative signal — the YOLO proposal tier emits general COCO object classes (person, vase, ...) that are essentially never interactive, so gating on the YOLO class alone always returned clickable=false even when MobileNet had correctly identified a real control (button, text_input, link, ...). ElementIsClickable now accepts an element when either the MobileNet top-1 classified label OR the YOLO class names an interactive type, and CombineElementConfidence weights an interactive MobileNet label more heavily regardless of YOLO agreement. This makes the fused capture actually recommend clickable controls in the live pipeline.
  • Coordinate safety + anti-over-click assurance — each annotated element carries a screen_box in virtual-screen physical pixels (the exact space click consumes, no DPI double-scaling), a click_point centroid to press, a fused combined_confidence, and a clickable gate (interactive-class whitelist + min confidence) so the AI avoids false-positive clicks on low-trust or non-interactive boxes. The capture also exposes width/height, dpi_scale, and the full virtual_screen bounds so the AI can reason about screen size and multi-monitor layout. The ml_teach/priors feedback loop weights previously-seen UI higher, so unfamiliar or moved elements are flagged rather than silently clicked.
  • GUI element detector swapped from generic COCO YOLO to Salesforce GPA-GUI-Detector (single-class icon) — the previous yolo11n.onnx is a COCO-80 general object detector whose proposals (person, car, bicycle, ...) flooded the watcher/priors loop with meaningless, never-interactive regions. The detector is now a UI-native, single-class (icon) ONNX export of Salesforce's MIT-licensed GPA-GUI-Detector (fine-tuned from OmniParser), which proposes actual interactive-element boxes. The authoritative per-control UI type remains the 15-class MobileNet tier.
  • Watcher/priors novelty gate now keys on the MobileNet UI label instead of the detector class — because the detector is now single-class (icon), the watcher dedup and element-priors DB would otherwise bucket every control under one key. The watcher now runs MobileNet classification BEFORE persisting crops and threads each element's top UI label (button, text_input, ...) into ElementKnownConfidently/saveElementRegionSamples, so the priors learn per-control-type locations. DetectedElement gained a mobile_net_label field carrying the authoritative class.
  • Detector model auto-downloads when missing — ONNXDetect now fetches gpa_gui_detector.onnx from https://github.com/coff33ninja/go-mcp-computer-use/releases/latest/download/gpa_gui_detector.onnx on first use if absent (best-effort, once per process, serialized against concurrent watcher/tool paths). GitHub redirects that URL to the newest non-draft release's asset, so it follows every version bump with no hardcoded tag.
  • Detector model ships with releases — .github/workflows/release.yml rebuilds gpa_gui_detector.onnx in CI (via the new export project) and attaches it to each versioned release alongside mcp-server.exe.

v0.3.5

Choose a tag to compare

@github-actions github-actions released this 27 Aug 10:10

[0.3.5] - 2026-08-27

Changed

  • Action-triggered training snapshots are now cropped around the target — SaveSnapshotAfterAction (used by click, type, drag, hover, find_text_and_click, type_and_submit, select_all_and_type, and click_menu_item) now captures a 400×400 region centered on the action target instead of saving a full-screen screenshot. This gives the ML model a focused view of the element being acted on, dramatically reducing the size of training data. Older call sites fall back to full-screen capture automatically.
  • Watcher training samples are now cropped around detected elements — The background watcher (which screenshots the screen every cycle to feed the priors/ML systems) now saves a small padded crop around each detected UI element instead of a full-screen 3200×1980 PNG every cycle. When no elements are detected in a frame, nothing is saved. This prevents the tens of gigabytes of near-identical full-screen images that previously accumulated.
  • Watcher is AI-gated with an achievement lock — The background watcher now starts locked (watcher_locked: true by default). While locked it still runs ONNX detection and caches results for reference, but never persists training crops. The AI must first prove competence through real interactive actions (click/type/record) and then explicitly call watcher_unlock to grant the watcher permission to save crops. The unlocked state persists across reboots, so the watcher stays unlocked once earned.
  • Watcher confidentiality / novelty dedup — The watcher now consults the priors system before saving a crop via a new ElementKnownConfidently check: an element is considered known when its (class, window) pair has enough samples AND its current normalized location falls within tolerance of the learned position. Familiar elements are skipped; only new/moved/uncertain elements are saved. Snapshots therefore decay as the ML gains confidence, replacing the old fragile count-only gate.
  • Wallpaper guard — The watcher will never save training crops when the foreground window is the desktop/shell (e.g. "Program Manager", empty title, Shell_*, Windows Shell Experience), regardless of the lock state. This prevents the wallpaper from ever being mined as a training signal.
  • Training DB pruning actually reclaims disk — PruneOldSamples now runs VACUUM after deleting rows, and a new PruneOrphanedSamples pass (run alongside the periodic retention pruner) removes rows whose image file no longer exists on disk. This reclaims the space that stayed locked inside samples.db after the large full-screen PNG cleanup.

Added

  • watcher_lock_status tool — reports whether the watcher's training-crop capture is achievement-locked, plus the count of real-action training samples gathered (action_signal) as guidance on when unlocking is warranted.
  • watcher_unlock tool — grants the watcher permission to persist training crops after the AI has gathered confident data from real interactive actions. The unlock persists across reboots.
  • set_config ... watcher_locked / get_config watcher_locked — the watcher lock can be inspected and toggled through the standard config surface alongside the dedicated tools.

v0.3.4

Choose a tag to compare

@github-actions github-actions released this 26 Aug 16:33

[0.3.4] - 2026-08-26

Fixed

  • scripts/gen-icons.ps1 — rewritten to handle encoding issues and missing dependencies gracefully. No longer fails with PowerShell parse errors when rsrc is not installed or icon path contains spaces. Build script (scripts/build.ps1) now works end-to-end again.

v0.3.3

Choose a tag to compare

@github-actions github-actions released this 26 Aug 13:09

[0.3.3] - 2026-08-26

Added

  • ml_query — ask the ML engine "where is X on this screen?" Pass a query (what you're looking for) plus current OCR text. Returns coordinate predictions ranked by confidence, matched OCR keywords, and related commands the ML has seen. Searches both coordIndex (per-tool coordinate distributions) and wordToCmds (command frequency). Query tokens get priority matching, context tokens add breadth.
  • ml_teach — feed confirmed correct answers back to the ML after every action. Pass what was being looked for, the screen OCR, which tool was used, coordinates, and success/fail. Updates coordIndex and wordToCmds directly with both query and context tokens. The learning loop: ml_query → AI acts → ml_teach reinforces. Each cycle strengthens token→coordinate associations.

Changed

  • The ML feedback loop is now complete: query → predict → act → teach. Whether the AI follows an ML prediction or discovers the correct answer itself, ml_teach ensures the ML learns from every outcome — including when the user shows the AI the right answer.
  • CI workflows updated from v0.2.x to v0.3.x (ci.yml, auto-tag.yml, jekyll-gh-pages.yml, mod-maintenance.yml).
  • README status block rewritten — v0.2.x framed as testing/iteration ground, v0.3.x as current stable. Recording & replication and ML feedback loop documented in features section.
  • docs/architecture.md — added record_replicate.go to code map, ML Loop layer in agent stack diagram, updated adaptive.go and chain.go descriptions.
  • docs/ci-cd-pipeline.md — updated branch references and diagram to v0.3.x.
  • _config.yml logo URL updated to v0.3.x branch.
  • scripts/gen-tools-doc.go — added ml_query + ml_teach to Adaptive Agent category (now 5 tools).

Fixed

  • scripts/gen-icons.ps1 — rewritten to handle encoding issues and missing dependencies gracefully. No longer fails with PowerShell parse errors when rsrc is not installed or icon path contains spaces. Build script (scripts/build.ps1) now works end-to-end again.

v0.3.2

Choose a tag to compare

@github-actions github-actions released this 26 Aug 11:01

[0.3.2] - 2026-08-26

Added

  • Timed recording no longer blocks MCP — record(duration_secs=N) now starts a background goroutine for the sleep+auto-stop, returning immediately with a confirmation. Previously time.Sleep() inside the handler caused the MCP SDK to kill the request as timed out. Both manual (duration_secs=0) and timed modes now work reliably.
  • Recording feeds ML on stop — RecordStop() now calls LogEnrichPatternsFromSession() asynchronously, feeding OCR, UIA, and ML enrichment payloads back to the adaptive engine immediately when recording ends. No manual wiring required — the full loop is: record → stop → enrich → ML learns.

Verified

  • 30-second timed recording test: started, returned immediately, auto-stopped after 10 seconds, keylogger inactive — no MCP timeout.
  • Manual recording + record_stop: full session returned with enrichment. OCR snapshots jumped 1668→1817 (+149 from enrichment logging). ML engine received 692 command sequences, 535 click patterns, 26 double-clicks, 26 long-presses, 49 type patterns. Top sequences: "reminder"→click (260 samples), "delete"→click (154), "schedule"→click (152).
  • AI replication of recorded sessions not yet tested — next validation step.

v0.3.1

Choose a tag to compare

@github-actions github-actions released this 26 Aug 10:23

See docs/meta/CHANGELOG.md for details.