Skip to content

Releases: wolf0x/FoxIR

FoxIR v1.0.17

Choose a tag to compare

@wolf0x wolf0x released this 29 Sep 01:03

FoxIR v1.0.17

Two documentation commits on top of v1.0.16 and no code changes at all — runtime behavior is
what v1.0.16 already was. If you came here looking for a fix, there is none in this release.

What landed

docs/changes/pluggable-memory-provider/ — an SDD for a pluggable MemoryProvider front door:
FoxIR's core would know no specific memory service, and wiring mem0 / MemoryHub / anything later
would be one config entry or one adapter instead of an embedded code path. Status is draft, not
implemented
— the existing deep_memory / SQLite behavior is untouched, and the baseline
recorded in the docs is b542a35.

Why the design body lives at specs/…-spec.md

.gitignore:54 carries the bare filename pattern SPEC.md ("local-only documentation"). A bare
pattern matches at any depth, and on this OS it is case-insensitive — so every
docs/changes/*/spec.md is silently un-addable. The first commit of this release contained
tasks.md and log.md and not the design body; the body was renamed to
specs/pluggable-memory-provider-spec.md, the same path style that let 72d62b8 publish a spec.

Three other design bodies are still under that rule and remain local: multi-agent-orchestration,
memory-hub-sync, adopt-agentskills-skill-spec. The rule itself was not changed — publishing
those was nobody's request.

Verification

cargo test --lib → 379 passed / 0 failed / 18 ignored, finished in 1.13s.

That is two more than v1.0.16's note recorded (377 passed, 2 known failures in cases that read the
real Prefetch / Recent files). Those two passed this round because they depend on machine state,
not because anything improved.

The 18 #[ignore] real-machine browser cases were not re-run: there is no code change for them
to guard. This is not a green-light claim about the browser path.

Release mechanics

Cargo.toml and Cargo.lock are bumped in this one commit. v1.0.15 split them across two commits,
which left the tag's lock file at 1.0.14 — building from that tag dirties the working tree.

FoxIR v1.0.16

Choose a tag to compare

@wolf0x wolf0x released this 26 Sep 02:39

FoxIR v1.0.16

One commit on top of v1.0.15, all of it in the built-in browser (browser_cdp).

Snapshot indices now die when the page moves

snapshot numbers the clickable elements of the page; click / type_text can address them by
that number instead of guessing a CSS selector (CSS cannot express "the link whose text is 设置").
Numbers were only good until the page changed — and an out-of-date number often still resolves,
it just points at a different element. Measured on the real machine: inserting two buttons moved
the DOM from 9 to 11 nodes, and the old index 0 clicked "新按钮 0" and reported success.

Now every index-based action compares the live DOM node count against the stamp taken at
snapshot time and refuses with an explicit instruction to re-snapshot. An explicit CSS handle
([data-fx="0"]) still works after invalidation — that is a selector you wrote down, not a number
you assume is still valid.

The browser left behind by a launcher handoff is adopted in visible-window mode

msedge.exe hands its work off to an already-running browser and exits 0. The surviving browser
stays alive, keeps the profile locked, and answers CDP — but chromiumoxide's pipe transport only
saw the exited launcher, so from then on every launch failed the same way. That is the state
behind "内置浏览器起不来" on machines that use the visible-window mode for logging in.

The guard for adopting such a leftover used to be "it has zero pages". That is wrong: the handoff
launch itself wakes the empty shell up (measured: 147 ms → 0 pages, 479 ms → edge://newtab/,
584 ms → plus a sync-confirmation dialog that closes itself after ~2.6 s), so counting pages read
our own leftover as "a human is using this". The condition is now the page URL: browser startup
pages are adoptable, any page with real content is left alone.

Verification

  • 18 real-browser regression tests (cargo test --lib -- --ignored) pass against installed Edge.
  • cargo test --lib: 377 passed; the 2 failures are the pre-existing tests that assert against
    this machine's real C:\Windows\Prefetch and Recent folders.

Asset: FoxIR-v1.0.16-win64.zip, containing a single FoxIR.exe (the UI is embedded at compile
time; no other files are needed). Windows x64.

FoxIR v1.0.15

Choose a tag to compare

@wolf0x wolf0x released this 25 Sep 08:13

FoxIR v1.0.15 — The Built-in Browser Gets Reliable, and Elements Get Addressed by Number

Heads-up on scope: this is the first published release since v1.0.8. Tags v1.0.9 … v1.0.14
were cut but never packaged, so this zip carries everything up to 3fb33cc — 51 commits of
history since the download you may have now. The short version of that gap, by release commit:
P0 security work + multi-session (v1.0.10), per-session mode/STOP (v1.0.11), main-chat
protection and language rules (v1.0.12–13), the Tools-page capability center and the first
browser launch fix (v1.0.14).

This round: browser_cdp reliability

Every item below was reproduced on the real machine (Edge 154) before it was changed, and each
has a test that fails on the old behaviour.

  • Browser state is sharded per agent — one page per invocation inside one shared browser, so
    a finished run can no longer close the browser another session is using. In visible-window
    mode it is one page per chat session, so a login page stops multiplying into tabs.
  • A browser that outlived its own launcher is now adopted, not mourned. Edge's launcher
    process hands the work to the real browser process and exits 0; the pipe transport then reports
    "browser exited before the websocket URL could be resolved", while a live browser keeps owning
    the profile — every later launch failed the same way. The port is recovered from
    <profile>/DevToolsActivePort, verified against /json/version (same browser family, headless
    UA) and connected.
  • fetch_targets() is gone. In chromiumoxide 0.9.1 it replays on_target_created, which
    replaces every registered target and silently kills all held page handles — the source of the
    channel disconnected errors. Listing goes through a plain CDP command now.
  • Shutdown waits. Every path that drops a browser now waits for the process to exit before
    re-launching, because Chromium's single-instance mutex is per user-data-dir.
  • Memory can no longer declare a tool dead. The fact curator refuses "the built-in browser is
    broken, use Edge headless instead", and the prompt states that memory is authoritative about the
    past, not about whether a tool works right now. (A stale fact had been making the model skip the
    tool entirely, then improvise with shell_exec and write evidence outside the case directory.)
  • Launch failures name the actual cause, and the profile stays inside the case directory.

Three action-level bugs, with their measurements

  • Default max_chars collapsed to 1 character. unwrap_or(0).min(budget).max(1) turned "not
    provided" into "give me one character", so get_text / get_html / execute_js returned ".
    A field test: 30 characters came back as 1; now the context budget is used and
    max_chars: 1 still truncates with next_offset.
  • A page that never finishes loading was reported as a failed navigation. On a local server
    that stalls mid-body, Page.navigate answers after 30.0s with Request timed out while the
    page had already arrived — the caller retried four times and burned two minutes. A late reply
    now returns loaded: false plus requested_url, and url only ever says what the browser
    itself will admit to.
  • A rejected selector was reported as a missing element after an 8s poll (real text:
    Error -32000: DOM Error while querying at 4.7ms). Bad syntax now fails fast and says so;
    genuinely absent elements keep the old message. navigate also stopped claiming the browser's
    stale about:blank as the destination.

New: snapshot + index addressing

CSS cannot match visible text, and that is the thing the agent keeps needing: on four real pages,
0 of 4 text targets could be written as a legal CSS selector, and only 15–16% had an id to
hang one on — so the model wrote a:contains('设置'), failed, and re-scanned the DOM by hand.

  • snapshot numbers the visible interactive elements and returns [N]<button 第一个按钮/>
    lines (measured: 25–27 elements ≈ 1.0–1.7 KB; 300 elements ≈ 10 KB; the pass itself is
    7–48 ms).
  • click / type_text accept index, resolved through the existing element path.
  • Stale indices say "re-snapshot, don't guess another number" and return in ~2.2s instead of
    polling for 8s.
  • Known cost: numbering writes a data-fx attribute into the page (attributes only), which can
    trip a site's own MutationObserver; iframe and shadow-DOM contents are not numbered.

Verification

Real machine (Edge 154): 15 #[ignore] browser tests pass, including the handoff-adoption,
visible-window switch and bounded-click cases. cargo test --lib 376 passed; the 2 failures are
pre-existing environment tests that read live Prefetch/Recent folders.

Not covered / known gaps

  • The Tools and Settings pages have not been re-clicked by a human after this version bump; the
    visible-window login path is verified by tests and by confirming a real window appears, not
    end-to-end by a person.
  • Element::click() on a link that navigates still never gets a reply from Chromium — actions are
    bounded around it and report "delivered, result unknown", which is honest but not the same as a
    confirmed click.
  • Iframes, shadow DOM and file uploads remain unimplemented in browser_cdp.
  • Windows only for the browser path; Linux browser discovery is still a stub.

FoxIR v1.0.8

Choose a tag to compare

@wolf0x wolf0x released this 14 Sep 14:51

FoxIR v1.0.8 — Parallel Multi-Session, Cross-Session Isolation, and Memory-Layer Cleanup

24 commits since v1.0.7. This release delivers the full multi-session story: you can create up to 5 isolated sessions (Settings-hot), switch between them from the sidebar, and run several concurrently — each with its own conversation, STOP control, tool view, and TASKS/TODO. Also lands the memory-layer simplification (shallow memory removed, deep memory projected back to MEMORY.md) and the parallel-session isolation fixes.

New Features

Parallel multi-session (P1/P2 落地)

  • Multi-session navigation index + /api/sessions + history filtered per session; sidebar panel with per-session snapshot/switch/rename/delete
  • Max sessions = 5 (Settings hot-reload); per-session cancel wired into Instant runs; executor slot registry so concurrent runs don’t collide
  • 片② demux/spawn: each Instant run is drained by its own spawned task; Expert stays inline. Same-session follow-ups queue behind the running task; interject messages render on execution-entry as a user bubble (no tool calls / session isolation)
  • per-session STOP is isolated end-to-end; switching sessions no longer stops another session's run, and each session keeps its own START/STOP state

Cross-session isolation (this fix)

  • drain_session_stream now injects the owning session into every forwarded WebSocket event; the frontend routes streaming by msg.session instead of a single global taskSessionId, so A and B running in parallel never cross-render (fixes "A's output appears in B")
  • TASKS/TODO are now per-session (todos-<session>.json) — each session's list is fully independent; /api/todos?session= serves the right list and the sidebar refreshes it on switch/new
  • A finishing background session no longer tears down the active session's live tool cards/text

Memory-layer cleanup

  • Removed the shallow-memory concept and all its artifacts (config, code paths, shallow_memories leftovers); deep facts + SQLite auto-summary remain
  • Deep memory is now projected to workspace MEMORY.md so you can review it in the original window; knowledge base is methodology-only (no auto-precipitation of session conclusions, threat-intel refs removed)

Verification

  • cargo check --lib clean; inline JS passes node --check; cargo build --release succeeds

FoxIR v1.0.7

Choose a tag to compare

@wolf0x wolf0x released this 13 Sep 08:08

FoxIR v1.0.7 — Sub-Agent Orchestration in Instant Mode, ir_process Inline Output, and ShimCache v3 Headerless Support

3 commits since v1.0.6. Focused release: makes Instant-mode sub-agent orchestration actually runnable/migrated from Expert, restores ir_process inline data (dropping the JSON dump), and adds parsing for the headerless AppCompatCache v3 blob that broke ShimCache on recent Windows 10/11.

New Features

Instant-mode sub-agent orchestration (Expert → Instant migration completed)

  • can_spawn now propagates so the Instant root can spawn workers; previously only the Expert executor path was gated
  • Fixed the P0 serde blocker: SubAgentSpec (from spawn_subagent) now applies #[serde(default)] to tools_allowlist / allow_write / allow_exec / skills, so a schema-compliant call with only role + prompt no longer fails with missing field. Unknown allowlist names now fail loudly instead of silently producing an empty worker (F8)
  • Added a per-Orchestrator concurrency ceiling so one round can’t spin up an unbounded number of workers (P2 cost/DoS guard)
  • Added an orchestration-gate diagnostic log line ([session:…] orchestration gate: prefilter=… -> orchard=…) so you can see at a glance why a run did/didn’t fan out (F7)

Typed sub-agent milestones + main-chat cards

  • Worker raw fragments (thinking/text/tool/progress) are no longer broadcast into the parent WebSocket stream — this collapses the wait-phase flood that could stall or drop the connection
  • Structured subagent_spawned / subagent_completed / subagent_failed / budget_update milestones are now emitted into the parent stream and rendered as grouped cards in the main conversation (plus the existing Expert drawer)
  • Frontend suppresses worker raw fragments by invocation_id/role, showing only the typed cards

Behavioral Changes

  • ir_process no longer writes output/ir_process-<timestamp>.json: the full classified process list is returned inline, consistent with the other ir_* tools (ir_scan, ir_artifacts, …). The old smart-truncation / file-dump path was removed
  • Wait-phase behavior: a closed consumer channel no longer tears down a run while sub-agents are still in flight — wait_subagent is allowed to collect worker results before the loop aborts normally

Bug Fixes

  • ShimCache parsing on headerless AppCompatCache v3 (fixes the v1.0.6 Known Issue): recent Windows 10/11 live registry blobs carry no 0x00000080 file header and start directly with the first v3 entry; parse_shimcache now recognizes the 0x30 / 0x34 entry-header leads and dispatches to a dedicated headerless parser (0x73743031 "10ts" marker walk). The machine-dependent test_shimcache failure should no longer occur on affected hosts

Verification

  • cargo test --release --no-fail-fast — lib suite passes including the new test_win11_headerless_parses_entry and subagent_spec_partial_json_uses_safe_defaults; prior known-issue test_shimcache live test is resolved by the headerless parser
  • cargo build --release — success; FoxIR.exe 38.6 MB

Known Issue

  • None blocking. Headless run still has no sub-agent UI events (a documented limitation for later work); Instant orchestration now has both backend and frontend milestone wiring.

FoxIR v1.0.6

Choose a tag to compare

@wolf0x wolf0x released this 13 Sep 03:09

FoxIR v1.0.6 — Expert Multi-Agent Orchestration, Dual-Mode Routing, and ir_process Returns Data Again

25 commits since v1.0.5. Absorbs upstream v1.0.5 and the shell Recycle-Bin change (0914b81).

New Features

Expert mode multi-agent orchestration

  • Manager–Executor–Auditor role separation: Manager plans each round, Executor runs the tools, Auditor verifies claims before a task may complete
  • TaskContract state persisted to SQLite — a crashed or stopped run resumes instead of restarting
  • Bidirectional sync between TaskContract and update_plan, so the Expert task tree reflects real execution
  • Auditor reports are surfaced back to Manager/Executor as cross-round memory; a Done claim with zero verified findings is rejected and fed back as a synthetic audit note
  • Expert multi-agent UI view: sub-agent status, token budget and plan updates streamed over WebSocket
  • Two-layer capability routing prompts

Behavioral Changes

  • Orchestration gate flipped Expert → Instant. The Instant root now orchestrates workers; Expert runs serially (Manager → Executor → Auditor)
  • can_spawn=false is enforced for the Expert executor (no nested sub-agents from an executor)
  • Removed the parallel collect path from the managed runner, plus parallel.rs, orch_selftest and the parallel_subtasks field on ManagerPlan/PendingPlan; dead orchestration_template config deleted
  • Orchestration delivery and construction are now gated on the rule-based orchestration_prefilter signal

Bug Fixes

  • ir_process had never returned any data — since v1.0.0. A stray } in the embedded PowerShell collection script made the whole script block fail to parse, so the child produced zero stdout and the tool answered {"status":"ok","raw":"","classified":[]}: success-shaped, empty inside, and indistinguishable from "no suspicious processes found" unless you looked at the key names. The parse failure also silenced itself ($ErrorActionPreference='SilentlyContinue' + stderr discarded). Now fixed and verified end to end. Additionally, the payload is no longer dumped whole into the context: the full list is written to output/ir_process-<timestamp>.json (full_output_path) and only non-safe processes are returned inline.
    Note for archaeology: this code fix rides in commit c9274d3, whose message does not mention ir_process.
  • Expert rounds are STOP-interruptible and hang-proof (parallel collection, executor start)
  • Manager planning, the Executor stream and the Auditor now all honor STOP
  • Cancellation propagates to workers on STOP — no more orphaned sub-agent processes
  • Repaired artifact-path auditing, cmd_format false positives, the anti-stagnation gate, and skill-injection truncation

Configuration

  • [orchestration.*] keys move to [modes.instant.*]. The legacy section still loads through a fallback and is normalized on validation; existing config.toml files keep working unchanged
  • max_tokens_per_run is clamped to max_total_tokens when it exceeds it, with a warning

Upstream changes included in this release

  • feat(shell): literal shell deletes in shell_exec are redirected to the OS Recycle Bin (recoverable deletion)

Verification (measured on the build host, Windows 11 23H2)

  • cargo build --release — success in 9m44s, warnings only; FoxIR.exe 38.72 MB
  • cargo test --release --no-fail-fast — lib 279 passed / 1 failed, bin 304 passed / 1 failed, integration suites and doc-tests green. The single failure is the pre-existing, machine-dependent forensics::live_tests::test_shimcache (see Known Issue); src/forensics/ is byte-identical to v1.0.5, so this is not a regression from this release
  • ir_process smoke test — the embedded enumeration script parses with 0 syntax errors, exits 0, emits 254 520 bytes and 460 process records ("pid": fields), versus zero output before the fix

Known Issue

  • ShimCache parsing fails on Windows 11 23H2: parse_shimcache only accepts the 0x00000080 (Win10/11) and 0xBADC0FEE (Win7) headers, but this build's AppCompatCache blob starts directly with per-entry data (first u32 = 0x00000034), so the parser returns Unknown ShimCache signature: 0x00000034. The test_shimcache live test unwraps that error and therefore fails on affected hosts. Tracked for a follow-up release; the unit-level parsers for other artifacts are unaffected.

FoxIR v1.0.5

Choose a tag to compare

@wolf0x wolf0x released this 11 Sep 01:58

v1.0.5

  • feat(shell): redirect literal shell deletes in local shell_exec to the OS Recycle Bin (safe, recoverable delete for direct-path del/Remove-Item/rmdir).

FoxIR v1.0.4

Choose a tag to compare

@wolf0x wolf0x released this 10 Sep 14:16

FoxIR v1.0.4

Removal: sub-agent orchestration (net-negative)

  • Removed the predefined sub-agent subsystem entirely (run_agent / run_skill / fetch_agent_result / wait_agents tools, /agent WS path, /api/agents handlers, agents.json store, run_sub_agent + SubAgentRunParams, sub-agent norms injection, and the whole Agents UI + slash-command dispatch): low dispatch success and lost skill-orchestration context.
  • Kept the skill hot-injection main logic intact (contract_block / step_contract / install_skill / list_skills / remove_skill).
  • Purged all residual sub-agent references (rabbit-hole trim whitelist, agents.json migration, sub- session prefix, docs).

Preserved FoxIR features (baseline commit)

  • WinRM integration, evidence ledger, session recall, recycle-bin-safe file delete, heartbeat gating, deep-memory curator.

Asset

  • Windows x64 binary packaged as FoxIR-v1.0.4-win-x64.zip

FoxIR v1.0.3

Choose a tag to compare

@wolf0x wolf0x released this 10 Sep 03:26

FoxIR v1.0.3

i18n (console logs)

  • Translated all remaining Chinese console/log output to English:
    • Deep/shallow memory injection + finite-brain arbitration + recall-touch logs
    • Two-tier shallow memory storage failure warning
    • SOP invalid-JSON parse warning
    • Shallow memory recall-rate test output
    • Tool-call envelope parse-empty warning

Misc

  • Version bump to 1.0.3

FoxIR v1.0.2

Choose a tag to compare

@wolf0x wolf0x released this 10 Sep 02:54

FoxIR v1.0.2

UI / Navigation

  • Replaced all sidebar emoji with consistent outlined SVG icons
  • Active state: soft background + left gradient indicator bar (replaces full-lighter fill)
  • Grouped 9 nav items into Core / Capabilities / Automation / System
  • Mobile (<700px): hide labels/group headers, center icons

i18n

  • Dashboard Context Budget: English labels (Category / Measured Tokens)
  • Empty-state message localized to English

Misc

  • Version bump to 1.0.2