Repository navigation
Releases: coff33ninja/go-mcp-computer-use
Releases · coff33ninja/go-mcp-computer-use
Release list
v0.3.10
[0.3.10] - 2026-09-18
Post-0.3.9 computer-use reliability pass: focus/keylogger chain bugs found in live OpenCode testing, plus the ML outcome ledger (desktop-as-teacher). Includes a full local computer-use data reset for clean testing (datalog, training store, memory, transformer weights; ONNX models kept).
Fixed
focus_windowno longer reverse-maximizes windows —FocusWindowcalledShowWindow(SW_RESTORE)unconditionally, which un-maximizes a maximized window every time it is focused (OpenCode observed browsers/editors snap out of maximize). Restore is now applied only whenIsIconic(minimized); otherwiseSW_SHOWkeeps the current maximized/normal placement.- Keylogger/replicate chain replay missing tool —
keylogger_stopemitted"tool":"_focus"andeventsToSmartStepsemittedFocusWindowsteps with an emptyTool. Chain dispatch then failed withunknown tool: _focus/unknown tool:(confirmed inchain_log). Fixes:- keylogger emits
focus_window_by_title - smart steps carry
Tool: focus_window_by_title+focus_windowfield - chain treats focus-only steps (empty Tool + focus field) as successful focus after auto-focus
toolDispatchgainsfocus_window_by_title,_focus(legacy alias),keylogger_start/stop/status,replicate
- keylogger emits
- Local tests no longer require live APPDATA assumptions for ledger schema —
ml_predictionsis created on demand; supervised samples never includeunknown/pending.
Added
ml_predictionsoutcome ledger (desktop as teacher) — predictions fromml_query/agent_suggest(transformer or statistical) are inserted as immutable rows (engine,query/ocr/windowhashes,pred_tool,pred_x/y,confidence,model_version). Only resolution columns update:status,outcome_source,actual_*,verification_data,resolved_at.- Outcome resolver on
click—ResolveMLPredictionsForActionscores pending preds after a real click:- hit — click succeeded within hit radius (~3% screen, min 80px) of a pending coord
- miss — click failed near a pending pred (or same-tool retry nearby)
- unknown — pending rows older than 10m (timeout); never treated as failure
- recovered — a miss followed within ~45s by a successful nearby same-tool action
- Successful clicks far from predictions leave rows pending (agent may have ignored ML) — no poison labels
ml_outcome_statustool +ml_status.outcome— hit/miss/unknown/recovered counts,hit_rate,recovery_rate,unknown_rate,last_hour_hit_rate,eligible_train.- Resolved-only supervised signal —
SupervisedLedgerSamplesreturns onlyhit|miss|recovered; transformer train pool preferstraining_pairs.success=1; unknown/pending never train. - Chain aliases for keylogger/replicate replay —
focus_window_by_title,_focus,keylogger_start/stop/status,replicate(executes provided steps).
Changed
- Clean-slate computer-use data for testing — local
%APPDATA%\go-mcp-computer-usetraining/datalog/memory/ML weights archived then cleared so 0.3.10 testing starts from zero pairs (ONNX models inmodels\retained). Agents must re-collect via normal use;agent_train/scripts/eval-ml.ps1 -Trainrebuild from new data only.
Build / verification
scripts/build.ps1— OK (version 0.3.10, Zig cc + CGO)scripts/lint.ps1—go vetclean + build OKgo test -short ./internal/actions— pass (includes focus/keylogger chain alias tests)
v0.3.9
[0.3.9] - 2026-09-18
Fixed
- Transformer training never saw real click coordinates — production
training_pairs.command_jsonstores args as a nested JSON string with mixed-case keys ({"args":"{\"X\":700,\"Y\":400}","tool":"click"}).ml/trainer.decodeCoordsonly unmarshaled top-level lowercasex/y, so on live datalogs every click/hover coord target was(0,0)(measured: 800/800 samples). Fix:ml/dataloader.NormalizeCommandJSON/ApplyNormalizationunwrap string-or-objectargs, acceptX/Y/from_x/to_xcase-insensitively, and fillSample.CoordX/Y+FromCoordX/Y. TrainerdecodeCoordsuses the same unwrap path;makeTargetFromSampleprefers loader-normalized pixels. Unit tests cover the production nested-string shape. - Transformer unusable after restart (empty tokenizer) —
MLEngine.LoadModelcreated a fresh unfitted tokenizer and never loaded vocab.tokenizer.Encodereturnsnilwhen unfitted, soForwardfailed withtoken len 0 != maxLen 128whileIsReady()could still be true; predictions silently fell through to the statistical engine. Fix: training persistsvocab.binnext tomodel.gob(tokenizer Save/Load now wired in production), plusml_meta.json(vocab size, ArgDim/WindowDim/FromCoordDim, tool list, holdout metrics).LoadModelrefuses ready-state without vocab and surfaces a clearlast_error. - Vocab size vs embedding table mismatch — fitted vocabs on real OCR exceeded the hardcoded
VocabSize=2000(observed 2421), so high token IDs were dropped in embedding lookup.modelConfigFor()sizes the model from the fitted tokenizer (rounded up) and is shared by Train / LoadModel / online / finetune paths. - Spatial features were always zero —
prepareBatchcalledencoder.Encode(0, 0)for every sample, and inference used an all-zero coord vector. Training now encodes the sample's decoded click/drag pixels; inference sets a context point from the live cursor viapredict.Engine.SetContextPoint. - Online training destroyed the tokenizer —
trainFromBufferrefit the vocab on a 32-sample replay batch then ranTrainEpochover the full SQLite DB, scrambling token IDs against existing embeddings. Online updates now train on replay-derived samples withoutFit, keep embedding IDs stable, and re-savevocab.bin+ report real eval loss/accuracy (checkpoint accuracy is no longer hardcoded0.0). - Honest tool-accuracy metric —
trainer.Accuracypreviously took argmax overtoolStart(tools + coord + arg dims), mixing continuous outputs into classification and inflating scores. It now scores argmax over tool logits only, skips unknown tools, and Train logsclick_acc/non_click_accalongside overall holdout accuracy and the majority-class baseline. agent_trainonly rebuilt statistical indexes — it now also trains the Go transformer (writesmodel.gob+vocab.bin+ml_meta.json) and returnsml_statusin the tool payload.scripts/lint.ps1failed when launched fromscripts/—Get-Content VERSIONresolved relative to the script directory. Now uses$PSScriptRoot\..\VERSIONand builds..\cmd\mcp-server.
Added
ml_statuson agent ML tools —agent_train,agent_suggest, andchain_predictreturn transformer health:ready,model_loaded,vocab_loaded,vocab_size, paths,last_eval_loss,last_eval_accuracy,majority_baseline,last_error, and asourceverdict (transformer|statistical_preferred|unavailable).sourcefield onPredictedAction— predictions are labeledtransformerorstatisticalso agents know which engine produced a coordinate/command guess.- Honest transformer gate —
MLEngine.Predictreturnsnilwhen outputs are near-uniform (no signal) or when last holdout tool accuracy is below the majority-class baseline, soPredictActionsfalls back to the statistical adaptive engine instead of shipping junk coords. cmd/ml-eval+scripts/eval-ml.ps1— local eval harness: tool distribution, coord-decode sanity (with_coords%), majority baseline, optional-Train,ml_statusdump, and smoke predict on real datalog OCR. Exit code2when model accuracy is below majority baseline.- Class balancing for click-heavy logs —
balanceToolsupsamples minority tools (cap 48) before augmentation so a 70%+clickdatalog does not collapse the model to “always click”. Tool one-hot targets are amplified (2.0) so MSE is not drowned by zero coord/arg dims. - Production-shaped ML tests —
ml/dataloader/normalize_test.goandml/trainer/decode_coords_test.golock the nested-args schema that unit tests previously missed (tests used{"x":...}while logs used{"args":"{\"X\":...}"}).
Changed
- README ML section is truthful — the Go-native transformer is described as experimental with vocab/meta persistence, majority-baseline reporting, and an explicit preference for the statistical engine (
ml_query/ml_teach) when the neural path underperforms. OCR/UIA/statistical priors remain the reliable locators. agent_traintool description — documents dual training (adaptive indexes + transformer artifacts) andml_statusin the response.docs/reference/tools.mdregenerated after handler/description updates.- Live retrain on a real datalog (800 pairs) — after the data-path fix,
with_coordswent from 0% → 70.4% for transformer targets;model.gob+vocab.bin+ml_meta.jsonload correctly across process restarts. Holdout tool accuracy on noisy OCR remains at or near the majority baseline — the neural path is gated and labeled rather than silently preferred. Statisticalml_query/agent_suggestnow actually learn token→coord averages from unwrapped production args.
Build / verification
scripts/build.ps1— OK (mcp-server.exe, Zig cc + CGO, version 0.3.9)scripts/lint.ps1—go vetclean + build OKgo test—internal/actionsshort suite +ml/dataloader,ml/trainer,ml/predict,ml/tokenizer,ml/transformerpass
v0.3.8
[0.3.8] - 2026-08-27
Fixed
- YOLO UI detection now letterboxes non-square inputs instead of aspect-distorting them -
preprocessYOLOpreviously rescaled X and Y independently to a 640x640 square, so a non-square input (e.g. the 3200x1980 virtual desktop) was squashed by 5x on X and 3.09x on Y, destroying every element's geometry. The detector then emitted floor-confidence (~0.5) sigma-degenerate 0x0 boxes everywhere. Fix: aspect-preserving letterbox (one uniform scalemin(640/w,640/h), centered gray padding) shared between the encoder and decoder via ayoloLetterboxstruct, andparseYOLOOutputnow inverse-maps boxes with(out - pad) / scale. Verified live: full-screen detection now yields real, varied boxes and click points with confidence 0.5-1.0 instead of thousands of 0x0 floor boxes. elements_query/onnx_detectno longer return garbage off-screen boxes from the letterbox padding -ONNXDetectinverse-mapped detections whose centers sat inside the 640x640 letterbox's gray padding band back into negative or oversized source coordinates (observed:{948,842,839,236}-> click y:960 on an 860-tall window). Fix: a newclipElementToImagegate filters every detection against the captured bitmap bounds - boxes whose center lands outside the image are dropped (padding false positives), and partially off-screen boxes are clipped to the image soscreen_box/click_pointare always actionable. Verified by unit tests covering all reported garbage shapes.ocr_window/screenshot_elementnow report + key the captured window, not the foreground window -annotateCaptureOptsstampedwindow_titlefromgetActiveWindowTitle()(the FOREGROUND window) regardless of the handle being captured, so handle-scoped captures mislabeled the wrong window and keyed per-window ML-priors from that wrong title. Fix:annotateCaptureOptstakes awindowTitleparam andAnnotateWindowresolvesgetWindowTitle(handle)for the actual captured handle. Verified live:ocr_window(132082)(Mozilla Firefox) now reportswindow_title="Mozilla Firefox"(was "OpenCode").- Capture responses no longer dump a giant element array on every
ocr/screenshot/onnx_detectcall - the fused YOLO+MobileNet annotation (from 0.3.5/0.3.6) embedded every detected box into every capture response with no cap, so a busy desktop returned thousands of degenerateiconboxes (~1751 observed) and blew the payload to several hundred KB even withinclude_image:false. The AI was forced to grep a monolithic JSON blob. Fix:annotateCaptureOptsnow runsFilterAndCapElementsbefore returning - it drops boxes whose fusedcombined_confidenceis below a floor (default0.40) and caps the surviving list (default50, highest-confidence first). Detection/ML still runs at full fidelity internally; only what's embedded on the wire is bounded. Verified by unit tests; a 1751-element dump is now at most 50 high-trust rows.
Added
include_imageopt-out for the base64 image bloat in capture-tool responses - everyocr/screenshot/screenshot_element/ocr_window/ocr_active_window/onnx_detect/onnx_classifyresponse embeds the base64 screenshot (~0.7-1.5MB), which wastes model context and forces truncation for text-only AIs. This adds a config flaginclude_image(defaulttrue, preserving current behavior) that globally strips the returnedimage_b64copy while keeping OCR text, element boxes/confidence, and all metadata intact - the internal ML detection pipeline still runs on the frame; only the returned copy is removed. A per-callinclude_imageoverride (true/false) on any capture tool wins over the global config, giving callers a tri-state. Exposed throughset_config/get_configand persisted. Verified live: withinclude_image:falseanocr_windowreturns metadata only (noimage_b64); a per-callinclude_image:truerestores it;onnx_detect(include_image:false)still returns 1677 elements / 350 clickable with the image stripped.exclude_elementsopt-out for text-only captures - a config flagexclude_elements(defaultfalse, preserving current behavior) plus a per-callexclude_elementsoverride onocr,ocr_window,ocr_active_window,screenshot,screenshot_element, andonnx_detectempties the returned element array so a text-only client never drags the detection boxes onto the wire. Likeinclude_image, it flows through the singlesafeHandlerchoke point (stripImageIfExcluded), is persisted viaset_config/get_config, and the per-call override wins over the global value.elements_querytool - a grep-style query over the fused capture for text-based LLMs - instead of dumping the full annotated JSON, the AI asks "where is the submit button" and gets a compact flat projection: filter by source (screen/window/region), YOLOclass, MobileNetlabel, anx,y,w,hregion,min_confidence,clickable_only, and OCRtext(returns matching words with screen coords). Each returned row is just{index, class, label, confidence, clickable, screen_box, image_box, click_point}- the coordinates to click - with no base64 and no nested classifier bloat. Registration mirrors the existingonnx_*tool family.scripts/gen-icons.ps1no longer writes a UTF-8 BOM intowinres/winres.json-Set-Content -Encoding UTF8(PowerShell 7) prepends a BOM, whichgo-winresrejects withinvalid character 'ï' looking for beginning of value, breaking the resource-generation step and hence the whole build. Fix: write with-Encoding utf8NoBOMso the generated JSON parses cleanly.key_presskey names expanded to the full Windows virtual-key surface + case-insensitive lookup + generic modifier combos -vkModMapnow exposes the explicit left/right modifier variants (LCTRL/LCONTROL,RCTRL/RCONTROL,LALT,RALT,LSHIFT/RSHIFT,WIN/META/LWIN/SUPER/CMD/COMMAND,RWIN), andvkSpecialMapgained the numpad block (NUMPAD0-NUMPAD9,NUMPAD_MULTIPLY/ADD/SEPARATOR/SUBTRACT/DECIMAL/DIVIDE), the media/volume keys (VOLUME_MUTE/VOLUME_DOWN/VOLUME_UP,MEDIA_NEXT_TRACK/PREV_TRACK/STOP/PLAY_PAUSE), plusSNAPSHOT,APPS/CONTEXT,CLEAR,EXECUTE, andSLEEP.keyNameToVKnow upper-cases its input so lookups are case-insensitive, and single-char punctuation resolves through thecharToVKtable.KeyPressgeneralized the old CTRL-onlyMOD+keyprefix into any modmap modifier followed by a single letter or digit (CTRL+A,WIN+R,ALT+F4,SHIFT+1, ...) and emits the requested modifier's real VK instead of hardcoding Ctrl. Covered by newinternal/actions/keyboard_test.gounit tests.
v0.3.7
[0.3.7] - 2026-08-27
Added
--version,--license, and--helpCLI flags on the server binary -mcp-server.exe --versionprints the build version (go-mcp-computer-use 0.3.7),--licenseprints the full Apache-2.0 text plus theNOTICE, and--helplists the available subcommands/flags. LICENSE and NOTICE are embedded into the binary viago:embed(kept in sync byscripts/build.ps1), so the license text is always retrievable from the shipped executable with no files alongside.
Changed
- Project relicensed from MIT to Apache-2.0 -
LICENSEreplaced with the full Apache-2.0 text, (c) 2026 coff33ninja. Added aNOTICEfile wiring the Apache-2.0 attribution and clarifying that the bundledgpa_gui_detector.onnx(a converted Salesforce GPA-GUI-Detector) remains under the upstream MIT license, not Apache-2.0. README gained a license badge, aLicensesection, and updated model-attribution wording. - Version-info, copyright, and admin manifest now embedded in the executable - resource generation switched from
akavel/rsrc(icon only, no manifest/version) togo-winres(scripts/gen-icons.ps1+ a committedwinres/winres.json.template). The exe now carries aVS_VERSION_INFOblock visible in Windows Explorer Properties -> Details (CompanyNamecoff33ninja, LegalCopyright "Copyright (c) 2026 coff33ninja. Licensed under the Apache License, Version 2.0.", ProductName, File/Product version fromVERSION) and an application manifest declaringrequireAdministrator(admin is required for the tool's UIA/UIPI automation to fully work) plus per-monitor-v2 DPI awareness. Previously the exe had no embedded manifest at all (as-invoker, no DPI declaration), so this is an explicit behavioral improvement that enforces the project's real admin requirement at the OS level. - License decisions: LICENSE + NOTICE are shipped with each release -
.github/workflows/release.ymlnow attachesLICENSEandNOTICEto every versioned release alongsidemcp-server.exe, so the license and third-party model attribution always accompany the binary.
v0.3.6
[0.3.6] - 2026-08-27
Added
onnx_classifytool — MobileNetV3 GUI element classifier — a new advisory tier that classifies UI content usingmobilenetv3_small.onnxacross 15 classes (button, checkbox, container, dropdown, icon_button, image, label, link, menu_item, scrollbar, slider, tab, text_input, toggle, unknown). Supportedsourcevalues:screen(full screen),window(active window),region(x,y,w,h),elements(classify every YOLO-detected element crop),crop(raw base64 image). Returns label + confidence for top-N (default 3) per target. Advisory only — it returns an additional signal and never hard-blocks an action.- MobileNet as a chain verification tier — chain
verifysteps accept an optionalclassifyconfig ({enabled, top_n, source}). When enabled, the executed step's result carries aclassifyadvisory payload with the top classifications, independent of the pass/fail decision. - MobileNet as a watcher advisory tier — the background watcher now classifies each YOLO-detected element crop and attaches
classificationsto its cached detection, giving the AI type+confidence context per element without affecting detection or training. Best-effort and skipped entirely when the model file is absent. AnnotatedCapture/AnnotateScreen/AnnotateRegion/AnnotateWindow— the fused, best-effort annotated-capture production layer. All classifier signals run on the SAME captured frame so geometry stays aligned. Any engine that is unavailable yields an empty sub-block (noted inerrors) without failing the capture.scripts/gpa-gui-export/— redoable GPA-GUI→ONNX conversion — a uv-based project (pyproject.toml+export.py+README.md) that downloadsSalesforce/GPA-GUI-Detectormodel.ptviahuggingface_hub, validates the singleiconclass layout, exports to ONNX(1,3,640,640)→(1,5,8400)opset 12, and writesgpa_gui_detector.onnx.uv sync+uv run export.pyregenerates the artifact from source; it is what CI uses to produce the release asset.- Main README "Models" section — documents the three runtime models, auto-download behavior, and how to regenerate the detector, plus a "UI-aware element detection" feature bullet.
Changed
- Every click now returns an advisory
validatedblock —clickandfind_text_and_clickresults carry a post-click validation combined into MobileNet classification of the click target (top-N label + confidence), the element-priors DB (sample count, prior-adjusted confidence, known-location flag, learned frequency/position), and an ML-memory cross-reference (whether the model has seen this click context before). Hooked at the centralClick()choke point so every click source gets it. Thevalidatedblock is merged at the top level of the tool result so the AI sees a consistent shape whether or not OCR auto-verify ran. Best-effort and non-blocking: if the model or capture fails, the click still succeeds and the validation block simply carries minimal/no data. - Internal refactor —
ClassifyElementsnow delegates to a sharedclassifyElementListhelper (also used by the watcher), keeping crop capture + model inference in one place. - Annotated capture on all AI-facing capture tools —
screenshot,screenshot_element,ocr,ocr_window, andocr_active_windownow always-on return a fusedAnnotatedCapture(OCR text + YOLO element boxes + MobileNet per-element classification + element-priors + ML-memory) with each element exposed in BOTH bitmap-image space and virtual-screen space (screen_box), so text-bound AIs know exactly what is on screen and where to click regardless of whether a vision model is available. Screenshot tools keep the raw b64 as text content for vision AIs while the annotated map rides alongside as structured result.screenshot/screenshot_elementgained an optionallanguageparam passed to the OCR signal. onnx_detectandonnx_classifynow return the same fused annotated capture — instead of separate raw detection/classification payloads,onnx_detectandonnx_classify(screen/window/region) now produce the identicalAnnotatedCaptureshape as every other capture tool, so the classifier signals are aligned on the same frame with dualimage_box/screen_boxcoords,click_point,combined_confidence, and theclickablegate — no redundant double-inference.onnx_detecthonors itsthreshold/iou_thresholdargs;onnx_classifyelements/cropsources return the classification wrapped in the same shape via a shim (no live capture-derived screen boxes in those cases).- The
clickablegate and confidence fusion now trust the MobileNet UI label as the authoritative signal — the YOLO proposal tier emits general COCO object classes (person,vase, ...) that are essentially never interactive, so gating on the YOLO class alone always returnedclickable=falseeven when MobileNet had correctly identified a real control (button,text_input,link, ...).ElementIsClickablenow accepts an element when either the MobileNet top-1 classified label OR the YOLO class names an interactive type, andCombineElementConfidenceweights an interactive MobileNet label more heavily regardless of YOLO agreement. This makes the fused capture actually recommend clickable controls in the live pipeline. - Coordinate safety + anti-over-click assurance — each annotated element carries a
screen_boxin virtual-screen physical pixels (the exact spaceclickconsumes, no DPI double-scaling), aclick_pointcentroid to press, a fusedcombined_confidence, and aclickablegate (interactive-class whitelist + min confidence) so the AI avoids false-positive clicks on low-trust or non-interactive boxes. The capture also exposeswidth/height,dpi_scale, and the fullvirtual_screenbounds so the AI can reason about screen size and multi-monitor layout. Theml_teach/priors feedback loop weights previously-seen UI higher, so unfamiliar or moved elements are flagged rather than silently clicked. - GUI element detector swapped from generic COCO YOLO to Salesforce GPA-GUI-Detector (single-class
icon) — the previousyolo11n.onnxis a COCO-80 general object detector whose proposals (person,car,bicycle, ...) flooded the watcher/priors loop with meaningless, never-interactive regions. The detector is now a UI-native, single-class (icon) ONNX export of Salesforce's MIT-licensed GPA-GUI-Detector (fine-tuned from OmniParser), which proposes actual interactive-element boxes. The authoritative per-control UI type remains the 15-class MobileNet tier. - Watcher/priors novelty gate now keys on the MobileNet UI label instead of the detector class — because the detector is now single-class (
icon), the watcher dedup and element-priors DB would otherwise bucket every control under one key. The watcher now runs MobileNet classification BEFORE persisting crops and threads each element's top UI label (button,text_input, ...) intoElementKnownConfidently/saveElementRegionSamples, so the priors learn per-control-type locations.DetectedElementgained amobile_net_labelfield carrying the authoritative class. - Detector model auto-downloads when missing —
ONNXDetectnow fetchesgpa_gui_detector.onnxfromhttps://github.com/coff33ninja/go-mcp-computer-use/releases/latest/download/gpa_gui_detector.onnxon first use if absent (best-effort, once per process, serialized against concurrent watcher/tool paths). GitHub redirects that URL to the newest non-draft release's asset, so it follows every version bump with no hardcoded tag. - Detector model ships with releases —
.github/workflows/release.ymlrebuildsgpa_gui_detector.onnxin CI (via the new export project) and attaches it to each versioned release alongsidemcp-server.exe.
v0.3.5
[0.3.5] - 2026-08-27
Changed
- Action-triggered training snapshots are now cropped around the target —
SaveSnapshotAfterAction(used by click, type, drag, hover,find_text_and_click,type_and_submit,select_all_and_type, andclick_menu_item) now captures a 400×400 region centered on the action target instead of saving a full-screen screenshot. This gives the ML model a focused view of the element being acted on, dramatically reducing the size of training data. Older call sites fall back to full-screen capture automatically. - Watcher training samples are now cropped around detected elements — The background watcher (which screenshots the screen every cycle to feed the priors/ML systems) now saves a small padded crop around each detected UI element instead of a full-screen 3200×1980 PNG every cycle. When no elements are detected in a frame, nothing is saved. This prevents the tens of gigabytes of near-identical full-screen images that previously accumulated.
- Watcher is AI-gated with an achievement lock — The background watcher now starts locked (
watcher_locked: trueby default). While locked it still runs ONNX detection and caches results for reference, but never persists training crops. The AI must first prove competence through real interactive actions (click/type/record) and then explicitly callwatcher_unlockto grant the watcher permission to save crops. The unlocked state persists across reboots, so the watcher stays unlocked once earned. - Watcher confidentiality / novelty dedup — The watcher now consults the priors system before saving a crop via a new
ElementKnownConfidentlycheck: an element is considered known when its (class, window) pair has enough samples AND its current normalized location falls within tolerance of the learned position. Familiar elements are skipped; only new/moved/uncertain elements are saved. Snapshots therefore decay as the ML gains confidence, replacing the old fragile count-only gate. - Wallpaper guard — The watcher will never save training crops when the foreground window is the desktop/shell (e.g. "Program Manager", empty title,
Shell_*, Windows Shell Experience), regardless of the lock state. This prevents the wallpaper from ever being mined as a training signal. - Training DB pruning actually reclaims disk —
PruneOldSamplesnow runsVACUUMafter deleting rows, and a newPruneOrphanedSamplespass (run alongside the periodic retention pruner) removes rows whose image file no longer exists on disk. This reclaims the space that stayed locked insidesamples.dbafter the large full-screen PNG cleanup.
Added
watcher_lock_statustool — reports whether the watcher's training-crop capture is achievement-locked, plus the count of real-action training samples gathered (action_signal) as guidance on when unlocking is warranted.watcher_unlocktool — grants the watcher permission to persist training crops after the AI has gathered confident data from real interactive actions. The unlock persists across reboots.set_config ... watcher_locked/get_configwatcher_locked— the watcher lock can be inspected and toggled through the standard config surface alongside the dedicated tools.
v0.3.4
[0.3.4] - 2026-08-26
Fixed
scripts/gen-icons.ps1— rewritten to handle encoding issues and missing dependencies gracefully. No longer fails with PowerShell parse errors whenrsrcis not installed or icon path contains spaces. Build script (scripts/build.ps1) now works end-to-end again.
v0.3.3
[0.3.3] - 2026-08-26
Added
ml_query— ask the ML engine "where is X on this screen?" Pass a query (what you're looking for) plus current OCR text. Returns coordinate predictions ranked by confidence, matched OCR keywords, and related commands the ML has seen. Searches bothcoordIndex(per-tool coordinate distributions) andwordToCmds(command frequency). Query tokens get priority matching, context tokens add breadth.ml_teach— feed confirmed correct answers back to the ML after every action. Pass what was being looked for, the screen OCR, which tool was used, coordinates, and success/fail. UpdatescoordIndexandwordToCmdsdirectly with both query and context tokens. The learning loop:ml_query→ AI acts →ml_teachreinforces. Each cycle strengthens token→coordinate associations.
Changed
- The ML feedback loop is now complete: query → predict → act → teach. Whether the AI follows an ML prediction or discovers the correct answer itself,
ml_teachensures the ML learns from every outcome — including when the user shows the AI the right answer. - CI workflows updated from
v0.2.xtov0.3.x(ci.yml, auto-tag.yml, jekyll-gh-pages.yml, mod-maintenance.yml). - README status block rewritten — v0.2.x framed as testing/iteration ground, v0.3.x as current stable. Recording & replication and ML feedback loop documented in features section.
docs/architecture.md— addedrecord_replicate.goto code map, ML Loop layer in agent stack diagram, updatedadaptive.goandchain.godescriptions.docs/ci-cd-pipeline.md— updated branch references and diagram to v0.3.x._config.ymllogo URL updated to v0.3.x branch.scripts/gen-tools-doc.go— addedml_query+ml_teachto Adaptive Agent category (now 5 tools).
Fixed
scripts/gen-icons.ps1— rewritten to handle encoding issues and missing dependencies gracefully. No longer fails with PowerShell parse errors whenrsrcis not installed or icon path contains spaces. Build script (scripts/build.ps1) now works end-to-end again.
v0.3.2
[0.3.2] - 2026-08-26
Added
- Timed recording no longer blocks MCP —
record(duration_secs=N)now starts a background goroutine for the sleep+auto-stop, returning immediately with a confirmation. Previouslytime.Sleep()inside the handler caused the MCP SDK to kill the request as timed out. Both manual (duration_secs=0) and timed modes now work reliably. - Recording feeds ML on stop —
RecordStop()now callsLogEnrichPatternsFromSession()asynchronously, feeding OCR, UIA, and ML enrichment payloads back to the adaptive engine immediately when recording ends. No manual wiring required — the full loop is: record → stop → enrich → ML learns.
Verified
- 30-second timed recording test: started, returned immediately, auto-stopped after 10 seconds, keylogger inactive — no MCP timeout.
- Manual recording +
record_stop: full session returned with enrichment. OCR snapshots jumped 1668→1817 (+149 from enrichment logging). ML engine received 692 command sequences, 535 click patterns, 26 double-clicks, 26 long-presses, 49 type patterns. Top sequences:"reminder"→click(260 samples),"delete"→click(154),"schedule"→click(152). - AI replication of recorded sessions not yet tested — next validation step.
v0.3.1
See docs/meta/CHANGELOG.md for details.