Skip to content

Releases: qf-studio/navigator

Navigator v8.2.1

Choose a tag to compare

@github-actions github-actions released this 05 Oct 10:37

Navigator v8.2.1 Release Notes

Release Date: 2026-10-05
Type: Patch — the reject line fits the reads card (TASK-88 follow-up)

Fixed

  • The reads card names the op. v8.2.0 ended the card with
    rejects today 1 · last 12:04 read_guard; at the card's 26% width the first live screenshot
    showed rejects today 1 · … — the op name, the one part worth reading, was the part cut off.
    The line is now 1 reject · 12:04 read_guard (no rejects when the log is empty), 27
    characters, so the op fits. l still opens the last eight full lines.

Tests: the pane-model test asserts the new form and the plural. Docs: CLAUDE.md, the archived
TASK-88 doc and the docs-site pages (nav-pane, reject-log) say the new line.

Navigator v8.2.0

Choose a tag to compare

@github-actions github-actions released this 05 Oct 09:20

Navigator v8.2.0 Release Notes

Release Date: 2026-10-05
Type: Minor — the reject log: every refusal the runtime makes, one line, greppable (TASK-88)

Added

  • Reject log. Navigator already refused deterministically at three points — the read guard
    denies the fifth repeated .agent/ Read, the prompt gate drops a loop prompt that follows a
    skipped WORKFLOW CHECK, the stop gate forces one continuation on an unfinished mutating turn —
    but kept only tallies: no reason, no time. The 2026-10-03 stop-gate over-fire (six blocks on
    read-only turns) was diagnosed from the transcript. Each refusing op now attaches
    reject: {reason, evidence} to its blocking result; the runtime strips the key before the
    merge and appends one compact JSON line to .agent/.nav-rejects.jsonl from a single point
    per runtime (runtime._log_reject, runner.ts logReject). The line:
    {ts, session, event, op, tool?, reason, evidence, suppressed?}. Evidence per op: read guard
    {path, count, threshold}; stop gate {met, unmet, mutating_tools}, so a Bash-only
    over-fire reads "mutating_tools":["Bash"] in one grep; prompt gate {trigger}.
  • Pilot writes it too. Under PILOT_EXECUTOR the merge belt strips every block; the log
    line survives with suppressed: true — the refusal was computed, not applied. The autonomous
    case is the one that needs the record most.
  • Bounded, on by default. The file keeps its newest 500 lines (rewritten past 600). It
    observes and never blocks, so reject_log.enabled ships on; nav-features disable reject_log
    turns it off. Gitignored here and by nav-init.
  • In the pane. With d, the reads card ends with rejects today N · last HH:MM <op>
    (warning color when today is non-zero); l opens the last eight lines. The log never enters
    the model's context — the pane renders it, stderr stays sentinel-redacted.

Parity

Python and the mod emit byte-identical lines: a generated corpus
(scripts/mod_fixtures/rejects.py → hooks/mod/tests/fixtures/rejects.gen.ts) is asserted in
rejects.test.ts; the op fixtures carry the new reject key on both sides.

Tests: 146 kit tests (+7: line parity, one line per refusal, strip, off switch, Pilot
suppression, bounded append, pane model); Python runtime tests (+7) and the three op suites
assert the summaries. Config: reject_log.enabled (default true).
Design: .agent/tasks/archive/TASK-88-reject-log.md.

Navigator v8.1.2

Choose a tag to compare

@github-actions github-actions released this 05 Oct 08:23

Navigator v8.1.2 Release Notes

Release Date: 2026-10-05
Type: Patch — the /nav pane follows every kind of doc edit (TASK-89); the session-start drift check is hook-safe and the auto-update docs are honest (TASK-81)

Fixed

  • The pane follows Bash edits, git moves and agent runs. The task list, marker, memories
    and the next card reloaded only after a turn in which Edit or Write touched .agent/. A doc
    changed through a Bash heredoc or sed -i, a task archived with git mv, a commit, or a
    subagent's edits left the pane stale until r. All of them now set the reload; and each edit
    of the active task doc re-reads its checklist at once, so the next card advances while the
    turn is still running. The full reload still waits for the turn to end.
  • The hook drift check never runs the Claude CLI. auto_updater.py --check-drift, run by
    both session-start ops, resolved the plugin version with claude plugin list — a 10 s
    subprocess inside a 4 s hook budget, and the source of the wrong "behind plugin (v7.0.0)" line
    in September. It now reads the plugin's own manifest, then the cache tree. 0.08 s live.
  • Auto-update docs say what happens. The nav-features, nav-upgrade and nav-sync-claude
    skills, the feature table, and eight docs-site pages no longer claim Navigator updates itself
    on session start. It shows the notice; claude plugin update navigator@navigator-marketplace
    or nav-start Step 1.5 applies it. auto_update.enabled: false means no check and no notice.

Tests: 139 kit tests (+9 across update and register suites); Python drift tests with the CLI
patched to fail plus a static guard; the golden fixture drops the version-dependent drift
section (tests/golden/README.md, deviations).
Design: .agent/tasks/archive/TASK-89-pane-follows-bash-and-agent-edits.md,
.agent/tasks/archive/TASK-81-auto-update-truth.md.

Navigator v8.1.1

Choose a tag to compare

@github-actions github-actions released this 04 Oct 11:07

Navigator v8.1.1 Release Notes

Release Date: 2026-10-04
Type: Patch — the /nav reads card sees reads made through Bash (TASK-87)

Fixed

  • Reads through Bash now count. The reads card and its fan-out verdict counted the Read
    tool only, so a session that read docs with cat / sed -n / head showed 0 0 docs.
    A read-only Bash command (cat, sed, head, tail, grep, rg, wc, diff) now counts
    every file it names — .agent/ paths as docs, the rest as code — exactly like a Read call.
    Not counted: ls, find, running a script, awk, anything that redirects to a file.
  • sed without an in-place flag is read-only in both runtimes' Bash allowlist
    (-i… / --in-place[=…] still mark the turn as mutating), so a sed -n inspection turn no
    longer trips the completion gate.

Tests: 131 kit tests; classifier cases in both suites; fixtures regenerated.
Design: .agent/tasks/archive/TASK-87-bash-reads-count.md.

Navigator v8.1.0

Choose a tag to compare

@github-actions github-actions released this 03 Oct 20:14

Navigator v8.1.0 Release Notes — "Show Me the Decisions"

Release Date: 2026-10-03
Type: Minor — the judge's decision trail and labeling in /nav (TASK-86); the plugin is now
published by QuantFlow Studio

Added

  • Judge trail behind j (/nav): under the session tally, the last eight judged prompts,
    one line each — 14:03 · task · substantial · unclear "make the onboarding better".
  • Label from the pane: while the latest decision is unlabeled, y confirms the judge's
    verdict as the label and x disputes it. Labels land in ~/.config/navigator/judge-labels.json
    (the personal config dir, like the ADHD switch) in the shape scripts/judge_label.py and
    scripts/judge_eval.py --fixture read: a confirmed verdict scores immediately; a disputed one
    has tier: null, so judge_label.py label --fixture ~/.config/navigator/judge-labels.json
    walks it for the real label later. Entries are deduped by prompt text.

Changed

  • Navigator is a QuantFlow Studio plugin: plugin.json author and marketplace.json owner,
    README and security contact, and every live install path and link now say
    qf-studio/navigator and quantflow.studio. The old alekspetrov/navigator path redirects;
    existing installs keep updating.
  • The release-check URL, auto_updater.py, version_detector.py, plugin_updater.py,
    claude_updater.py and check-version.sh read the qf-studio repo directly.

Tests: 128 kit tests; Python suites unchanged except fixtures that name the repo.
Design: .agent/tasks/TASK-86-judge-trail-labels.md.

Navigator v8.0.1

Choose a tag to compare

@github-actions github-actions released this 03 Oct 18:29

Navigator v8.0.1 Release Notes

Release Date: 2026-10-03
Type: Patch — the completion gate stops over-firing on read-only turns (TASK-85)

Fixed

  • stop_completion forced continuations on read-only turns when two sessions shared a repo.
    The previous Stop's working-tree digest lived in the session-scoped completion section; any
    event from another session in the same repo reset it, so the tree-evidence rule could not
    fire and the Bash allowlist alone judged the turn. The digest now lives per session in
    tree.digests (bounded to the 8 most recent sessions), outside the scoped section, in both
    runtimes. completion.tree_digest is still written.
  • Read-only allowlist: lsof, pgrep, nproc, sw_vers; curl is read-only unless it
    names an output (-o, -O, -sSo…, --output…, --remote-name…).
  • Mod only: a Bash-only turn whose every call Claude Code held isReadOnly is never
    mutating, whatever the allowlist says. Python keeps the allowlist; parity fixtures replay with
    no such flag.

Tests: Python op suite 65 (4 new), mod kit 124 (classifier, flag, end-to-end gate pair),
fixtures regenerated. Verified live on 2026-10-03 against the turn shape that over-fired.

Full design: .agent/tasks/TASK-85-stop-gate-shared-state.md.

Navigator v8.0.0

Choose a tag to compare

@github-actions github-actions released this 03 Oct 12:51

Navigator v8.0.0 Release Notes — "In Process"

Release Date: 2026-10-03
Type: Major — the hook runtime becomes a Claude Code mod (TASK-84); the Python runtime stays as
the fallback

Summary

Navigator's workflow used to run as Python processes that Claude Code started on every hook
event and talked to through stdin and stdout. In v8 the same ops run inside Claude Code as a
mod (Claude Code 2.1.287 or newer): no process per event, real state, and a user interface.
Every op was ported with byte parity against the Python version, measured on generated
corpora rather than spot checks. Python stays registered and takes over automatically on older
Claude Code, where an organization blocks mods, or for any op that crashes three times in a
session. Nothing to configure.

What's New

The Navigator mod (TASK-84)

  • hooks/hooks.json names hooks/mod/register.tsx; the classic hooks in plugin.json are
    unchanged. On Claude Code ≥ 2.1.287 the mod runs all 16 ops; it announces them in
    NAVIGATOR_MOD_OWNS and the Python dispatcher skips them (fast exit when an event is fully
    owned).
  • One shared state file, .agent/.nav-runtime-state.json (schema 2), so an op handed between
    the runtimes mid-session sees the same state.
  • Parity: scripts/gen_mod_data.py runs the real Python ops over generated corpora (852
    scorer prompts, 60 recorded judge responses, ~3,600 prompt-op cases, ~560 tool-op cases,
    ~2,100 Stop cases, plus lifecycle cases) and writes fixtures the mod's kit tests must match
    byte for byte. make test-all runs both runtimes' suites.

/nav: one screen, cards

What you see by default is a surprise or a press; the readouts wait behind d:

  • context: fill and a bar, one verdict (compact safe / good moment to compact /
    compact due), and the fill over the last turns as a braille area in a subdued gradient
    (grom's stat texture, github.com/qf-studio/grom); a flat trend draws dim;
  • session: with the Prometheus of .agent/grafana/ (port 9092) answering, today's cost,
    tokens and cache hit rate, the 7-day cost and commit count, and tokens/min over two hours
    as the same braille area;
    without it, Claude Code's own cost and rate-limit window, the phase, and the graph size.
    Read on /nav, on r, and after every completed turn, never under Pilot; one probe first,
    so a stopped stack costs one refused connection per turn. The task list, marker and memories
    reload after a turn that wrote under .agent/, and after /clear, /resume or /compact. dashboard.enabled: false turns it off; dashboard.prometheus_url points
    it at another Prometheus on this machine (loopback URLs only; read with curl, so it works
    with CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC set and nothing leaves the host);
  • reads (behind d): Read calls this session and how many were docs; use an Agent
    when a turn read three or more code files;
  • judge (only after a judged prompt): the typed judge's verdict in words
    (task · substantial · unclear), what Navigator did with it (→ brief shown, task mode,
    loop mode, direct), and the axes where it overrode the keyword rule; j opens the
    session tally per axis. With CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC set, Claude Code
    refuses the mod's fetch; the judge then goes through curl (key and body via environment and
    stdin), so enabling the judge and providing a key is what decides, as in the Python runtime;
  • next: the destination (the goal Claude states in a brief, else the active task), then the
    current leg as a button (● 9/14 Docs: …, turns spent on it): n or Enter submits
    Do the next leg of TASK-84: … as your prompt. Under it the leg after, with an ETA from the
    pace so far (turns per finished leg × seconds per turn × legs left), and the newest marker.
    Legs come from the task's checklist, else the numbered steps of its plan section
    (## Work breakdown, ## Implementation plan, …) with ### Step n — … ✅ progress headings
    marking them done, else research → impl → verify → complete;
  • off route (only while drifting): two prompts in a row that share nothing with the
    destination open a warning with park (writes a parked task stub), back, or switch;
  • memories: up to three recalled for the open tasks, one line each; ▸ pins one into the
    next prompt;
  • tasks (behind d): up to five marked in progress, the destination marked.
  • A one-line band above the prompt that leads with the leg:
    nav · TASK-84 · ● 3/5 verify · 3 legs left, low fuel 72% · compact after this leg · …,
    off route: … · /nav to park or go back, or nothing.
  • A Pilot custom theme (/theme → Pilot) shipped through the plugin manifest.

Truthful update notice (TASK-81)

At session start the mod compares its own version with the latest GitHub release (at most
every auto_update.check_interval_hours, never under Pilot) and shows the update command.
Navigator never updates itself from a hook. The nav-start skill's Step 1.5 now runs
auto_updater.py once and reports its JSON instead of a template the model could echo.

Behavior changes

  • read_guard's warning reaches the model as context at the warn threshold (v7 printed it to
    stderr only). The block at the escalate threshold is unchanged.
  • config_guard and setup notices appear as toasts (mod classic results carry no
    systemMessage).
  • Judge key: the mod reads TYPESAFE_API_KEY or ~/.config/typesafe/api_key; a custom
    judge.api_key_env name is honored only by the Python fallback.
  • Matchers: MultiEdit is no longer a tool on Claude Code 2.1.287; the mod watches Edit,
    Write and NotebookEdit.

Removed

  • The v6 scorer shims skills/nav-start/functions/workflow_detector.py and
    skills/nav-brief/functions/ambiguity_scorer.py, and scoring.v6_exports. The nav-workflow
    CLIs (complexity_detector.py, skill_detector.py) stay as thin entry points.

Requirements and fallback

  • Mods need Claude Code 2.1.287+. Older versions load the classic hooks exactly as v7 did
    (verified on 2.1.284: hooks/hooks.json with modules is tolerated).
  • allowManagedModsOnly, disableAllHooks or a refused mod leave the Python runtime in charge.
  • Conformance re-driven on 2.1.287 (tests/harness-conformance/results/cc-2.1.287.json); probe
    S2 now measures delivery, not obedience: decision:block still forces the continuation.

Upgrade

claude plugin update navigator@navigator-marketplace

Restart Claude Code afterwards. Then try /nav, and /theme → Pilot.

Navigator v7.9.0

Choose a tag to compare

@github-actions github-actions released this 01 Oct 12:41

Navigator v7.9.0 Release Notes — "Say the Word"

Release Date: 2026-10-01
Type: Minor — ADHD mode (TASK-82) and the removal of the deprecated multi-Claude orchestration

Summary

Two things. A per-person switch that reshapes replies for an ADHD reader and flips on or
off mid-session by saying so, with zero model turn. And the end of the multi-Claude shell
orchestration that was deprecated in v6.15: native Claude Code Workflows, the Agent tool
and Pilot own that job now, so the skills, scripts, templates and config block are gone.

What's New

ADHD mode (TASK-82)

Say adhd mode on, adhd mode off or adhd mode at any prompt. The new prompt_adhd op
answers the phrase itself through the same decision: block channel as Tier-1 (no model
turn) and writes a personal switch to ~/.config/navigator/adhd-mode.json
(NAVIGATOR_CONFIG_HOME overrides the directory). The switch follows you across repos and
takes effect on the next prompt, no restart.

While on, every prompt carries a short declarative rule block as injected context: one next
action first, time-critical items first with the deadline in bold, bullets over prose, lists
capped at five, numbered steps with "step k of n", flat tone for errors, no preamble, one
sub-two-minute closing action. Error output, test results, explanations you asked for and
warnings before destructive actions are never shortened. While off, the rules exist nowhere
in the context. Subagents and the Pilot executor never see the block.

  • Config block adhd_mode: enabled (seeds true, only makes the machinery available)
    and on (true/false pins it for a repo in .nav-config.json or the .local
    override; null defers to the person). Resolution: repo pin > personal switch > off.
  • nav-features: new adhd_mode row. enable adhd_mode writes the personal file, not
    the repo; --local pins adhd_mode.on in .nav-config.local.json.
  • Session start prints ADHD mode: on (personal switch ...) when a switch is set.
  • Lib: nav_hook_lib.adhd (phrases, resolution, rule block) and nav_hook_lib.personal
    (per-person JSON files under the config home).
  • Design and reference review: .agent/tasks/TASK-82-adhd-mode.md. The rule set builds
    on the user's own rules plus ayghri/i-have-adhd.

Research agents on Sonnet (documented)

navigator-research, task-planner and the deep-research fetcher already declare
model: sonnet in their frontmatter; on Claude Code 2.1.284 that resolves to Sonnet 5.5.
CLAUDE.md now says so. No behaviour change.

Breaking Changes

Multi-Claude orchestration removed (deprecated since v6.15.0, TASK-25). Deleted:
nav-multi and nav-install-multi-claude skills, scripts/navigator-multi-claude*.sh
and their helpers (sub-claude-monitor.sh, resume-workflow.sh,
multi-claude-dashboard.sh, install-multi-claude.sh, simple-poc.sh, POC-LEARNINGS.md),
templates/multi-claude/, both multi-Claude SOPs, memory mem-020, and the multi_agent
config block (DEFAULTS, migrator, tier-1 feature list, nav-features table). A stale
multi_agent block in an existing .nav-config.json is ignored. Use the Workflow and Agent
tools, or Pilot, for parallel and multi-phase work.

Tests

  • 34 new tests: test_personal.py, test_adhd.py, test_prompt_adhd.py (incl. the full
    dispatcher subprocess path: toggle on, block injected, toggle off), session-start notice,
    nav-features personal-switch CLI.
  • Registry, config-defaults, migrator and golden-fixture suites updated for the new block
    and the removed one. make test green.

Getting Started

claude plugin update navigator@navigator-marketplace   # then restart Claude Code

Then, in any Navigator project: adhd mode on.

Navigator v7.8.0

Choose a tag to compare

@github-actions github-actions released this 28 Sep 19:51

Navigator v7.8.0 Release Notes — "One Repo, Many People"

Release Date: 2026-09-28
Type: Minor — team-repo features and fixes from the 2026-09-28 issue batch (GH-30 … GH-34)

Summary

Everything in Navigator's .agent/ used to assume one person per checkout. This release
makes the shared parts shareable and the personal parts personal: a per-contributor config
override, onboarding state outside the repo, GitHub-allocated task IDs, and the gitignore
entries nav-init should have written all along. One deep-research fix rides along.

What's New

Personal config override (GH-30)

.agent/.nav-config.json stays committed and shared. A new .agent/.nav-config.local.json
(gitignored) merges over it last — DEFAULTS < shared < local — in the hook runtime
(nav_hook_lib.config.load) and in nav-features. A contributor without a TypeSafe key
turns the judge off for themselves only:

python3 skills/nav-features/functions/feature_manager.py disable judge --local

show marks rows the local file decides with L and prints a legend; info names the
override; toggling the shared value while a local override shadows it prints a warning.
config_guard validates the local file the same way it validates the shared one.

GitHub-issue task IDs (GH-32)

task_id_source: github in the config makes nav-task create the GitHub issue first and
name the doc GH-<n>-<slug>.md. Issue numbers are allocated by GitHub, so two contributors
branching at once can never mint the same ID, and Pilot already addresses tasks as
GH-<n>. A failing gh (no auth, no network) is an error, never a silent local number.
The default stays local (sequential TASK-NN). Index updates, graph sync and the
TaskCreated/TaskCompleted lifecycle events now accept any PREFIX-<n> filename.

python3 skills/nav-task/functions/task_id_generator.py --title "Add OAuth" --label pilot --json
# {"id": "GH-57", "source": "github", "url": "https://github.com/o/r/issues/57"}

Bug Fixes

  • nav-onboard state is per person (GH-31). Progress, the personal workflow guide and
    the .completed marker live under ~/.config/navigator/onboarding/<repo-id>/
    (NAVIGATOR_ONBOARDING_HOME overrides the base; <repo-id> is the directory name plus
    an 8-character hash of the checkout path). Nothing is written inside the repo, so one
    contributor's finished onboarding no longer makes everyone else's skip. A pre-existing
    .agent/onboarding/ in the repo is ignored. When the project has its own setup skill,
    nav-onboard points to it first.
  • nav-init gitignore (GH-33) now appends .agent/.nav-runtime-state.json, its .lock,
    .agent/.nav-config.local.json and .agent/onboarding/, each only if absent, so a
    fresh repo's first git add -A no longer commits hook session state.
  • Deep-research fallback notes (GH-34). source_store write used to dedup on the
    canonical URL regardless of status, so the WebFetch fallback after a blocked raw fetch
    was silently dropped and the 403 stub won. An ok write now replaces a blocked or
    skipped stub under the stub's id (superseded: true), a retry of a stub is allowed,
    and ok notes still dedup as before. The fetcher agent is told not to delete stubs by hand.

Breaking Changes

None. Existing TASK-NN docs, the shared config file and every default are unchanged.
Onboarding runs completed before 7.8.0 are not migrated; re-run nav-onboard once if you
want the personal workflow guide back.

Tests

make test green. Two new suites are wired into the Makefile: skills/nav-features/functions
(personal override layering) and skills/nav-task/functions (local and GitHub ID sources,
fake gh on PATH). Config, config_guard, graph_sync, task_to_graph, source_store and
nav-onboard suites gained cases for the new behavior.

Getting Started

claude plugin update navigator@navigator-marketplace   # then restart Claude Code

Existing projects need no config change. Team repos: run the nav-init gitignore step once
(the six lines in skills/nav-init/SKILL.md §7), and set "task_id_source": "github" if
you want issue-numbered task docs.

Navigator v7.7.1

Choose a tag to compare

@github-actions github-actions released this 23 Sep 11:53

Navigator v7.7.1 Release Notes — "Count What the Judge Does"

Release Date: 2026-09-23
Type: Patch — telemetry and eval tooling for the typed judge; no gating changes

Summary

v7.7.0 put a typed judge behind the keyword scorers and measured it on 60 invented
prompts. This patch makes the judge measurable in use (TASK-80, workstreams A and B).

Override telemetry

Runtime state gains a judge section (30-day TTL, cumulative like tier1): calls,
failures, last and max latency, and per axis — loop, complexity, task-shaped, ambiguity —
whether the judge overrode the keyword decision, agreed with it, or stayed
undecided inside the band. For the score axes "agreed" compares the decision side of
the threshold, not the raw number. Counters only; no prompt text ever enters state.

nav stats shows two lines:

judge: 86 calls / 1 failed · 499 ms last, 1271 ms max
judge axes: 4 overridden / 240 agreed / 15 undecided

The nav-stats skill report carries the same as a row.

Real-session eval tooling

scripts/judge_label.py builds a labeled set from prompts you actually typed:

  • extract walks local transcripts, keeps genuine user prompts (no tool results, hook
    feedback, slash-command echoes, or subagent transcripts), redacts key-shaped tokens,
    caps and deduplicates, samples across projects.
  • label is one keystroke per prompt (n not a task, d small task, t substantial,
    a substantial and needs a brief, l loop). sheet / import do the same through a
    markdown file. status counts.
  • The output is gitignored. It holds real prompts from private projects and must never be
    committed; only aggregate numbers belong in the repo. scripts/judge_eval.py skips
    unlabeled rows, so a partial set already produces numbers.

First live read

Four days, 92 judged prompts across two projects: loop 91 agreed / 0 overridden,
complexity 81 agreed / 10 undecided, task 81 agreed / 4 overridden / 6 undecided. The
keyword heuristics are right almost every time on ordinary prompts; the judge's value is
in the tails, and no tail case occurred in the window. It stays opt-in and off by default.

Also in this release

  • TASK-79 marked released; docs site synced to 7.7.0 on 2026-09-19.
  • Repository note: GitHub reports the repo moved to qf-studio/navigator; the old path
    redirects.