Skip to content

GitPulse v0.0.5

Choose a tag to compare

@github-actions github-actions released this 05 Sep 06:23

Added

  • Fleet: one surface for every open repository, and every recent one. A workspace of two dozen tabs could only be inspected one tab at a time — "is anything unsaved anywhere", "which of these has an agent running", "what is all this costing on disk" were unanswerable without visiting each in turn. Fleet (Shift+F10, the leftmost chip in the repository strip, or the command palette) puts them on one grid: changes, sync, conflicts, stash, parked operations, worktrees, agent sessions and last activity, plus lines of code, disk usage, dependency audits and coverage on demand. It is deliberately not a sixteenth view — a ViewTab is stored on the active repository's session and its pane lives inside {#key currentPath}, so a view would be scoped to the wrong thing and rebuilt on every repository switch. It sits beside the repository pane instead, and the two swap by hiding rather than unmounting, because that subtree holds the live terminal PTY.
  • Every Fleet cell is a value, not scanned, or could not read — never a reassuring zero. The failure a workspace dashboard invites is reporting a fleet of clean, empty, vulnerability-free repositories that nobody ever scanned. So the three states are distinct in the model, in the markup and in the totals: a repository with no audit shows "not scanned", an audit that ran but could not finish is marked as a floor, a ledger that could not be opened fails its four cells rather than emptying them, and a total says "1.50 GB — counted across 14 of 21, 1 failed, 6 not scanned" rather than a bare number implying the whole workspace. A repository whose sweep failed is reported unknown, never clean — but never at the cost of downgrading one that has real conflicts.
  • Expensive scans stay opt-in, on the same posture as automatic coverage and the release check. Storage walks up to 250,000 files behind a 20-second deadline and the dependency audit spawns npm audit / cargo audit with a 90-second timeout; neither ever runs from an effect. The cheap tier costs nothing at all — changes, sync and conflicts are already hydrated in memory — and the middle tier is a new cmd_fleet_snapshot, two git spawns per repository in one rayon-parallel round trip, deliberately narrower than cmd_insights_snapshot (which probes every worktree and cross-scans up to 16 for collisions: correct for one repository on screen, several hundred subprocesses across a workspace). Family sweeps run two at a time for the expensive families and report successes, failures and skips separately.
  • Scan results are cached in each repository's own ledger, with their age. A new additive fleet_metrics table records one row per repository, every family carrying its own value and its own timestamp, so "never scanned" is NULL rather than 0 and a displayed number can always be dated. Families are independent: a storage scan cannot blank last week's audit. Reads go through a read-only, no-create SQLite open, so rendering a row for a repository that is merely in the recents list never writes a .devcouncil/ directory into a repository the user has not opened.
  • Fleet repository removal with accessible action button and keyboard navigation (Delete / Backspace). Rows in the Fleet dashboard provide a direct removal action via an action button or keyboard shortcut (Delete / Backspace): removing an open repository closes its tab and drops it from the workspace, while removing a recent repository purges it from workspace history and last-closed tabs so stale paths stay clean without lingering.
  • The package is now installable and connectable as a Claude Code plugin. The Agent Plugins 1.0 package described the server, but no client could find it: Claude Code discovers a package through a .claude-plugin/marketplace.json at the repo root and then reads <source>/.claude-plugin/plugin.json and <source>/.mcp.json, none of which existed. Those files now sit alongside the Agent Plugins and Codex manifests in the one canonical package, sharing its single skills/ tree, and the marketplace points at it. Verified end to end: claude plugin install records 0.0.5, the component inventory resolves both skills and the MCP server, and claude mcp list reports the server connected.
  • GitPulse is now a native Codex plugin, not only an MCP binary that happened to exist. plugins/gitpulse/.codex-plugin/plugin.json declares the native surface, .agents/plugins/marketplace.json makes it installable from the repository, and the same canonical package is bundled into Contents/Resources/plugin for Settings discovery. The portable manifest still launches gitpulse-mcp from PATH; this host's Codex MCP registration uses the absolute installed path so desktop, CLI, and IDE launches do not depend on shell initialization. Verified with a fresh Codex process that loaded gitpulse@gitpulse, invoked gitpulse_insights, and received the live branch and change count.
  • npm run mcp:install and npm run mcp:doctor. The MCP manifests spawn the bare token gitpulse-mcp off PATH — correct for a published plugin, but it means the server a client actually connects to is whatever is on PATH, which no build step owned. mcp:install puts it there through cargo install, so the binary is tracked and refreshable. mcp:doctor completes a real handshake and compares the version the server reports against this tree's, keeping absent, unresponsive and stale distinct from matching — the first of those is the one a naive check would report as a silent pass. Deliberately not in ci:local: CI has no reason to install the server, and a check that cannot run there must not be made to look like one that passed.
  • A repository that pins its own toolchain is told how to honour that pin. Go, Swift, .NET and Dart still refuse when their runtime is absent — installing a language runtime is a host-wide change with no bounded, reversible command, and a coverage panel does not get to make that trade for you. But "install Go, then rescan" is a poor answer for a checkout that has already written down which Go it wants. When mise.toml, .mise/config.toml or .tool-versions is present, the refusal names the pin and the one command that honours it, and it names only a manager actually on PATH — a mise-only config never suggests asdf, which cannot read it. GitPulse still does not run it: the command writes outside the repository, so it stays on your side of the line.
  • Line count, coverage and storage now track the repository instead of the moment a panel was opened. All three were one-shot fetches keyed on the repository path changing and nothing else, so a day of editing left the headline LOC number, the coverage report and the disk usage exactly as they were when the tab was first shown — with nothing on screen saying so. Meanwhile the backend already emitted repo-changed on every settled write and only repoStore listened. A new metric layer (src/lib/metrics/freshness.ts) owns freshness for all three: one measurement serves every panel, the watcher revalidates it, and each metric carries its own debounce and minimum interval derived from what its command actually costs — 20 s for LOC, 30 s for coverage, 120 s for the storage walk, so a build writing continuously into target/ cannot pin a 20-second tree scan. Change count needed none of this: cmd_branch_stats was already refetched by handleRepoChanged, and re-measuring it here would have run the same git subprocesses twice per change.
  • A value that could not be refreshed no longer looks refreshed. A failed revalidation keeps the last good report — panels must not blank out on a transient error — but the snapshot carries value and stale independently, and every panel reads both. Storage shows an "out of date" badge, the Pulse LOC tile degrades to partial, and a truncated or stale reading is never recorded as a point on the line-count trend.
  • C/C++ coverage, via an out-of-tree instrumented CMake build. This family used to be a flat dead end on the grounds that GitPulse cannot add coverage flags to a project's build files. The premise was right and the conclusion too strong: CMake takes compiler and linker flags on the configure command line, into a separate build directory, so cmake -S . -B build-gitpulse-coverage -DCMAKE_C_FLAGS=--coverage → cmake --build → ctest --test-dir → gcovr --lcov produces coverage without touching CMakeLists.txt or the developer's own build/. gcovr installs into a project-local virtualenv, exactly as pytest already did. Make-only projects stay a dead end and say why: make CFLAGS=--coverage replaces a project's own flags rather than adding to them.
  • Storage now reports a reclaim audit, not just sizes. The report published raw numbers and left every judgement to the reader — gc_recommended was a bare boolean, prunable worktree admin was not detected at all, and nothing said how many bytes were actually recoverable. Each row now carries the action that reclaims it, whether the bytes were measured or estimated, and whether a human has to decide. The headline counts measured and safe bytes only: estimates (a repack's saving depends on the content) and items needing review (committed build output, the reflog, a large file that may be someone's dataset) are carried as separate figures rather than folded into a total the repository cannot deliver. New detection: .git/worktrees/<name> left behind by a worktree deleted with rm -rf, which the old report counted as space belonging to a live worktree.
  • Manvi Control Plane & Operator Panel:
    • ManviHarnessPane, ManviOpsPanel, and HarnessBadge integration for real-time agent session monitoring, task attribution, and operation gating.
    • Dedicated harness store (harnessStore.ts) with lease tracking, task status synchronization, and worktree association.
  • Durable WAL SQLite Ledger & Redaction:
    • Embedded WAL SQLite database for durable append-only event logging across repositories.
    • Sensitive credential and secret redaction (ledger/redact.rs) before persisting mutation records.
    • Idempotent catch-up replay of agent transcripts and reflogs for automated attribution even when the app is closed via gitpulsed.
  • Diagnostics & Health Reporting:
    • In-app Diagnostics modal displaying real-time system metrics, IPC command performance, crash logs, open file descriptor usage, and environment context.
    • Machine-readable and exportable diagnostics report generator for troubleshooting.
  • Enhanced Multi-Tab File Viewer & Editing:
    • Multi-file editor tab management (editorTabs.ts) with persistent draft preservation across workspace switches (editorDraftRegistry.ts).
    • Serial file saving queue (serialSave.ts) preventing write collisions and index locking.
    • Interactive merge conflict editor (ConflictEditor.svelte) with visual diff resolution and instant staging.
    • MarkDev viewer and rich media viewers for markdown and media asset inspection.
  • LivePulse Analytics Dashboard:
    • Dynamic interactive pulse dashboard (LivePulseDashboard.svelte) and export modal (PulseExportModal.svelte).
  • IPC & Wire-Type Safety:
    • Expanded IPC surface to 136 verified Rust cmd_* handlers with zero untracked orphans.
    • 762 wire-type contract fields strictly synchronized between Rust Serde structs and TypeScript interfaces across 48 contracts.
  • Automatic macOS glass styling for window chrome, sidebar, menus, and dialogs; fluid view-selection and dialog transitions with reduced-motion, reduced-transparency, and contrast fallbacks. In-app glass preserves existing distribution options.

Changed

  • Fifteen views are four. The header had grown a button per panel — four tabs plus an Inspect and a More dropdown — and most of what sat behind them was not a destination at all. Diff carried its own commit picker so you would not have to walk back to Graph for the commit you had just selected; Blame carried its own explorer rail and its own path box so you would not have to walk back to Files for the file you already had open. Those pickers are the tell: a tab that has to rebuild the previous tab's context is a lens on a shared subject, not a place. Each subject now owns one view, and its lenses are sections switched by a segmented control inside it, so the subject survives the switch. Work (Overview · Resolve · Remote · Stack · Policy) shares the worktree row; Code (Explorer · Blame) shares selectedFilePath; History (Graph · Diff · Reflog) shares selectedCommitId; Insights (Pulse · Coverage · Health · Storage) shares the repository. The terminal became a dock beneath whichever view is on screen — the PTY always had to outlive a view switch, so it was already mounted once and hidden thereafter, which is a dock wearing a tab's clothes. Fleet stays workspace-scoped and is deliberately not a view.
  • A retired view lands where its content went, not at the start of the app. Sessions persist a viewTab, and eleven of those ids no longer exist. migrateViewTab used to send every unrecognised id to Work, which would have stranded anyone whose last session was on Diff, Coverage or Blame. Retirements are a map now (RETIRED_VIEWS), each naming the view and the section its content became, so a session written by 0.0.5 reopens on Code showing Blame, not on Code's default pane — landing on the right view showing the wrong thing reads to the user as the content having been deleted. The same map drives the tests: the check that every retired view kept a command-palette door is derived from it rather than hand-listed, so retiring a view without giving it a door fails the suite instead of silently removing the only route a keyboard user had.
  • The header's dropdown machinery is deleted, not left standing beside the tabs. Menus existed because fifteen views could not fit a title bar; four can, so the last group emptied and the dropdown branch became unreachable. Rather than keep a grouping layer with one group, the portal, the capture-phase dismissal, the viewNav group layer and menuGroup on every registration are gone — four regression tests about dismissing that menu were replaced by one asserting it stayed deleted, since a reintroduced menu would bring those regressions back with it. View-hiding in Settings is unaffected and now renders one flat list instead of headings describing a tabs-versus-menu distinction that no longer exists. Digit accelerators close up to ⌘1–⌘3 with no gap, which the native menu asserts.
  • The language strip became a status-bar segment. It was a 32px full-width bar that re-fetched cmd_get_language_stats on its own, next to a comment in the metric registry saying the headline LOC number and the language bar are two readings of one scan "and deriving them separately is how they came to disagree". It reads locMetric now, so the two cannot drift, its failures land in diagnostics instead of blanking silently, and the breakdown moved into a popover — reclaiming a strip that cost every session and paid occasionally.
  • Stack draws the chain as a chain. A stack is its shape, and a flat list with "based on X" on every row made the reader rebuild it in their head. Rows are a tree now, each joined to what the branch list already knows: commits ahead of its parent, commits behind the default branch, tracking state (↑/↓, upstream gone, or untracked — never 0↑ 0↓, which is what a pushed and current branch looks like), and when it last moved and by whom. A branch the progressive stats pass has not measured contributes nothing rather than a confident zero.
  • Updating a stack carries the stack. Rebasing one branch moves every branch above it off the commit it was cut from — and, because the hierarchy is tip-anchored, moves them out of the tree at the same time, so a single restack stranded the rest of the stack invisibly. The plan is computed from the tree on screen before the first rewrite, which is the last moment those fork points exist; the confirmation names every branch it will touch, because "restack 4 branches" does not say which four and this rewrites commits on all of them. Steps run parent-before-child, each an independently gated and independently rolled-back cmd_restack, and a cascade that stops names what moved and what did not.
  • Work's Overview says where you are standing. It described every worktree in flight and never the one checked out in front of the reader — not the branch, not whether it had drifted from its remote, not what was uncommitted in it. A strip above the rows carries the branch, its tracking and base comparison, and the working tree split into staged, unstaged and conflicted, with the parked operation getting its own line into Resolve. A branch the branch list has not reached yet reads "sync not measured yet" rather than as parity with a remote nobody has asked about, and an operation probe that could not run says so instead of leaving the same empty space as an idle worktree.
  • Overview's counts became doors. "3 blocked" over a list sorted by weight was a number the reader then had to go find. Each tile selects exactly the rows it counted — strip and list share one predicate, so a tile saying three can never sit above a list showing none of them — and a filter box matches on branch, path, task and pull request. Rows carry the age of the most recent commit on their branches, because a worktree nobody has touched in three weeks and one from ten minutes ago read identically without it and want opposite remedies. A filter matching nothing says so and offers to clear itself, rather than borrowing the wording of a repository with nothing in flight; narrowing is dropped on a repository switch.
  • Remote is two columns, not one ragged grid. Five listings shared a grid whose rows are as tall as their tallest cell, so twenty open pull requests left a screen of white space beside a three-line releases card — and Workflows sat a grid row away from the runs it produces. Pull requests and issues take the wide column because they are what a reader acts on; workflows, runs and releases sit together in a CI rail. The queue's own counts are the way into it (All / Awaiting review / Failing / Drafts, each filtering the list it counted, plus a search across number, title and both refs), issues carry when they were last updated and are searchable, runs carry their age and narrow to the checked-out branch, and the header stamps how long ago the context was fetched — a listing with no timestamp reads as current however long it has been sitting there. Failing means a red verdict only: a run still going and a repository whose checks never start are neither passing nor failing, and neither is folded into the other. The CI:local report is a card the reader folds up or dismisses instead of a full-width banner that pushed every listing down and could not be put away.
  • The diff page is rebuilt around the diff itself, not around the last thing you clicked. The header printed selectedFilePath whatever the body held, so a commit diff was labelled with one file's name, one file's language icon and the whole commit's line count — .github/workflows/codeql-analysis.yml, 2,405 lines, above a body showing server.go. Identity is parsed from the diff now (diff/outline.ts): the title is the file's path when the diff covers exactly one and N files when it does not, the churn is +19,250 −2,217 read off the diff's own lines, and the language badge is dropped entirely for a multi-file diff rather than borrowed from whichever file happened to be selected. Below it a context strip names the file, hunk and enclosing function the reader is currently inside, updated as the list scrolls, so a 200-file commit stops being 25,000 anonymous lines. The file stepper reads 4/200 instead of –/26.
  • Split view puts a replaced line beside its replacement. The old builder walked the line list holding one pending deletion, so a block of D deletions and A additions came out as D−1 rows with an empty right column, one paired row, then A−1 rows with an empty left — the two sides offset by the size of the block, which is exactly the case split view exists for. Pairing is now a shared model (diff/rowModel.ts) that both layouts read: row k of a replacement block holds del[k] beside add[k], the longer side spills into rows with one empty cell, context is the same object on both sides, and file headers and hunk headers span both columns instead of leaving a column-height hole. The same model backs the unified list, so the two views can no longer disagree about which line replaced which — the intra-line word-diff highlight used to depend on which view you had opened first.
  • One horizontal scrollbar for the diff, not one per row. Every row was its own scroll container, so a wide diff drew a grey bar under each line, the rows scrolled independently of one another, and the line-number gutter scrolled away with the code. The rows now share the list's single scroller: the gutter is position: sticky with an opaque backing (a translucent tint let code slide visibly underneath it), and VirtualList grew an opt-in contentWidth so a row's background tint spans the full content width rather than stopping at the viewport edge when scrolled right.
  • Diff rows are syntax-highlighted, in each file's own language. The file viewer beside it has had a tokenizer the whole time; the diff printed one colour per line type. Three overlapping layers — syntax, the word diff's changed spans, and the search hit — do not nest, so they are flattened into one span list per line (diff/highlight.ts), cached per line object and skipped above 2,000 characters. Crucially the language is resolved per section: a commit diff is a stack of files in different languages, and one global language coloured a 200-file commit's Rust with the JSON tokenizer, so // comments came out as plain text and the commas were confidently picked out as punctuation.
  • The minimap maps the list that is on screen. It built its ticks from the unified line list and then wrote the resulting offset into whichever pane was showing, so in split view — a different list of a different length — every click landed off by the ratio between the two. It also mapped a click to ratio × contentHeight, which top-aligns: clicking the last tick scrolled past the end, clamped, and left the thing you aimed at off screen. It now takes the tones of the list actually being drawn and centres the target, and it draws the viewport band and the file boundaries that turn a decoration into a map. A bucket holding both additions and deletions is a modification, not growth.
  • The file rail tells its files apart. Two hundred rows of manuscript_citation_hy… and manuscript_citation_m… name nothing. Rows are disambiguated against each other (the shared disambiguateLabels), so each shows the shortest directory prefix that makes it unique; there is a filter with a live 33 of 200 files count, a list/tree toggle that reuses the repository file-tree builder, a drag- and keyboard-resizable width (a real ARIA window splitter: arrows step, Home/End jump to the bounds), and virtualization past 60 rows so a 200-file commit renders 44 buttons rather than 200.
  • Find in a diff is the file viewer's search, hardened. Both had their own loop; they are now one module (text/lineSearch.ts). It refuses a pattern whose nesting can backtrack catastrophically instead of running it — (a+)+c against 28 characters took 111 seconds in the old path, and a JavaScript regex cannot be interrupted once started — bounds the scan with a deadline checked every 256 lines, caps the match list, and reports 2 of 5+ rather than pretending the cap is the total. ⌘F now belongs to the diff when the diff is what is on screen, instead of always opening the commit filter.
  • The frame survives an empty diff, an image, and a fetch in flight. Each of those used to replace the whole pane, so a clean merge or a .png left no way to reach the next file except by leaving for the Graph. The rail, toolbars and file stepper are now outside the branch. A diff being fetched is a first-class state (selectedDiffPending) rather than the previous file's rows sitting there looking current, and the wait is announced — the region reports aria-busy and the message is a live status, because a spinner is a picture of nothing to a screen reader.

Fixed

  • Release notes for this tag would have failed to publish. The 0.0.5 changelog section is 54 KB, past GitHub Actions' 48 KB cap on an environment-file variable. The workflow had captured it into GITHUB_ENV and handed it to tauri-action as releaseBody, which would have failed after every platform had already built. Preflight now asserts the section exists; the verify job writes it with gh release edit --notes-file. The v1 updater-json input is uploadUpdaterJson (the v0 name was ignored, so the default remained on).
  • Language mix was ordered by category, not by share. The backend sorted programming languages ahead of higher-volume markup and data files, and the status-bar picker kept that order after preferring them for display slots, so a repository of HTML 40% / TypeScript 35% / CSS 15% / Rust 10% was labelled TypeScript and drawn as 35% → 10% → 40% → 15%. Stats now sort by code lines (language name as a tie-break). The picker still keeps programming languages on the bar when lockfiles and markup would otherwise crowd them off, then sorts the shown list by percentage with Other last, so the dominant label is the language with the highest share among what is drawn.
  • The commit graph drew branches it had no name for. The history walk asked git for --all — every namespace under refs/ — while the ref listing read only refs/heads, refs/remotes and refs/tags. Anything else opened a lane nothing in the UI could label. That is not a corner case: agent harnesses keep a ref per turn (refs/cmux/last-turn/*, refs/codex/turn-diffs/*), git maintenance writes refs/prefetch/remotes/*, CI mirrors write refs/pull/*. On one repository here, 18 turn-checkpoint refs — each a stash-shaped pair forked from the same base commit — turned 65 commits of straight history into 101 rows and 35 lanes: 34 anonymous rails descending thirty-odd rows to a shared parent far below the fold, and at the default graph width 21 of those rows drew their node off-canvas with nothing on screen to say so. The lane solver was not at fault and is unchanged; its output was internally consistent throughout. Which refs the graph is about is now one decision (graph/ref_scope.rs) answering both the walk and the labels, with a contract test pinning that they agree — the drift that made unnameable lanes possible cannot recur. The default scope is branches, remote-tracking branches, tags and HEAD; All refs in Settings → Graph restores the old reach and labels those refs by their full path instead of leaving them blank. Same repository, default scope: 65 rows, 8 lanes, every node on screen.
  • History the graph does not draw is now reported rather than simply absent. A narrower walk that says nothing is the same failure as a check that could not run reporting a pass. Each load counts the commits reachable only from refs outside the walked set and names the namespaces holding them — "36 commit(s) reachable only from refs outside branches, remotes and tags are not drawn. Those refs live in: refs/archive/* (1 ref(s)), refs/cmux/* (18 ref(s))" — as a diagnostics breadcrumb, not a banner. It stays silent when nothing is missing, including the common case of a custom ref that merely points at a commit already on a branch: that commit is drawn, so there is nothing to report. Both probes are bounded, and the ref scan runs only when the commit probe has already found something.
  • Hardening pass over the ref scope, driven by tests written to break it. Six defects, each now covered by a test that fails against the code as it stood. (1) The named-ref test matched a string prefix, not path components, so refs/headsfoo/x, refs/tags-archive/x and friends were classified as branches git does not walk — their commits were neither drawn nor reported, the worst of both. The classifier is now checked against git itself over a table of adversarial names rather than against our belief about --branches. (2) That misclassification could leave the report claiming hidden commits with no namespace behind them, printing a sentence that trailed off after "live in:". (3) The bounded namespace list was ordered alphabetically, so a repository with ten tiny namespaces and one holding ten thousand refs named the ten and hid the one that explained the graph; it ranks by size now. (4) refs/stash was described as refs/stash/*, a directory it does not have. (5) The all-refs scope produced one decoration per ref with no ceiling — a CI mirror's refs/pull/* would have shipped six figures of chips over IPC on every graph load. It is capped, and the cap is reported, which also closes a pre-existing hole: tags have been silently truncated at 200 since the cap was written, so a repository with 250 of them showed 200 chips and said nothing. (6) A 209-character agent-checkpoint ref path drawn whole pushed the commit summary off the row; chips fold to the namespace and keep the full path in the title.
  • The "Refs drawn" setting did nothing. Changing it updated the preference and stopped there: the graph's fetch scheduler keys requests on path, revision and query, so the key never moved, the scheduler saw a request it had already served, and nothing reloaded until an unrelated event happened to refresh the pane. The scope is part of the request identity now. The scheduler's own 200-iteration fuzz harness could not have caught this — its no-loss invariant compared the settled load's key against the latest key, both from the same key function, so a key that fails to distinguish two different requests satisfied it trivially. It now compares the request itself, and reproduces the bug at seed 1.
  • A degradation that appeared while history stood still was never reported. loadGraph short-circuits on a structurally identical payload to keep canvas caches alive, and that return came before warnings were handled. Warnings are not part of the rendered-history signature, so a ref listing that started failing, or a namespace that started hiding commits, was silently swallowed on every quiet repository. Whether history changed and whether the load degraded are two different questions, and they are asked in that order now. Warnings are also logged when the SET changes rather than on every load: the watcher reloads on every settled write, the diagnostics ring only coalesces consecutive repeats, and two persistent warnings alternate — filling the ring with the same pair forever and burying everything else.
  • The real-repository smoke check labelled under a different scope than it walked. It walked under the scope being tested and then called list_ref_decorations with RefScope::Named hard-coded, reproducing the original defect inside the harness meant to detect it: an all-refs dump of MarkDev showed 35 lanes and 7 labels. Same scope on both sides now — that dump carries 26 refs, and every lane-opening tip has one.
  • Pulse metrics counted machine-written commits as work. The analytics walk was the same --all, so on a repository with agent-harness namespaces every contributor and churn figure included commits no person wrote — a metric measuring the tooling. It now walks the same named set as the graph.
  • The real-repository smoke check ran a walk the app does not perform. real_repo_smoke.rs shelled out to its own git log --all, making it a third independent copy of "which refs the graph is about" — so it could not have caught this, and its committed fixtures described a pipeline that no longer existed. It goes through GitReader now.
  • parse().unwrap_or(0) made three more counts read as zero when they were enormous. Same class as the --shortstat bug and found by sweeping for it: git rev-list --left-right --count (a wildly diverged branch reporting as in sync), git count-objects -v (the largest repositories reporting as holding no git data), and the three --numstat parse sites. All now saturate through one shared helper that keeps the other half of the rule intact — --numstat writes - for a binary file, and there zero really is the right answer, so text that is not a number stays absent rather than saturating to the maximum. (cvss_to_severity was checked too and is already correct: an unparseable score falls back to high, not to zero.)
  • A metric subscription could attach to a cell that had just been evicted. Opening one more repository than the tracking bound, while every existing entry was being watched, evicted the cell that had just been created — the only one with no listeners yet, because subscribe registers its listener after the cell exists. The subscriber then held a reference to something no longer in the map: it never fired again, and the panel read idle forever. The bound is now soft, and the cell being created is never a candidate. Found by a randomized soak test, not by a hand-written case.
  • Two watcher tests carried their own 8-second budget instead of the suite's shared one. panicking_on_change_cannot_leak_the_watch_slot failed once under cargo test --workspace and passed 3/3 in isolation: FSEvents delivery plus the 400 ms debounce stretches far past its idle latency when the whole workspace runs at once, which the neighbouring test already documents and allows 20 s for. Both now use PRIME_DEADLINE, raised to match. Costs nothing on a passing run — every user is a loop that breaks on success — and only lengthens how long a genuinely broken backend takes to be declared broken.
  • The repository's headline line count was wrong for every language with block comments. LocCounter classified a line by one test — does it start with this language's single comment prefix — so a four-line /* … */ counted as four lines of code, a Python module docstring counted as code, and a CSS or HTML comment counted exactly one comment line (the opener) with the rest as code. The _ => "//" catch-all was also wrong for PowerShell, CMake, Git Config, GDScript, OCaml, PureScript and WebAssembly, and a UTF-8 BOM hid the first comment of any file an editor had written one into. Measured on GitPulse's own source: 10,557 of 188,004 "code" lines — 5.6%, across 339 of 719 files — were comments. Replaced with a per-language scanner that tracks block-comment nesting, docstrings and string literals, so a "/*" inside a string can no longer open a comment that swallows the rest of the file. The bound that makes it safe: a single-line string cannot cross a newline, so one unbalanced apostrophe in YAML prose can no longer reclassify every comment below it. language.rs no longer keeps a second comment table that could disagree with the first — it did, returning Some("/*") for CSS, a block opener applied as a line prefix.
  • Go, JavaScript and PHP coverage claimed a tool was ready without ever probing it. Go was the plain case: it was the only ready-producing planner in the file that probed nothing at all, so a go.mod on a machine without the Go toolchain published a Run button that failed at spawn — a check that never ran, reported exactly like a check that ran and passed. JavaScript emitted npm run … and PHP emitted vendor/bin/phpunit on the same unexamined assumption. All three now probe, and a shared test asserts no planner may return a plan that is simultaneously "ready" and unrunnable.
  • A git diff --shortstat count too large for usize parsed as zero. unwrap_or(0) meant the largest possible diff reported as no change — a reading a caller cannot tell apart from a genuinely empty diff. Counts now saturate.
  • Credentials named by a JSON object key reached the ledger, the durable log and the diagnostics report in full. Both redactors — ledger/redact.rs (write path for ledger rows, every logging.rs line, and gitpulsed) and diagnostics.ts (the copyable report) — traversed JSON objects by value and discarded the key. An opaque token therefore matched no pattern and was written through unchanged, so {"access_token": "…"}, the shape of every OAuth response and most API error bodies, was stored and displayed intact while the identical secret one syntax away (access_token=…, --access-token …) was redacted. Object keys are now consulted against the same credential-name table the CLI-flag path already used, normalized across client_secret, clientSecret, X-Api-Key and ACCESS_TOKEN, with compound names such as github_token covered by suffix and non-credential names such as public_key and cache_key deliberately left alone. Non-string values under a credential key fail closed rather than being recursed into.
  • A credential used as a JSON object key was never scanned at all. The other half of the same blindness: a vendor-shaped token in key position — a cache or rate-limit map keyed by the credential — was emitted verbatim while the same token in value position was redacted. Keys are now scanned through the same boundary, and a rename that would collide with an existing key is disambiguated rather than allowed to overwrite, so no entry is silently dropped.
  • The two credential-name tables had nothing binding them together. SECRET_FIELD_NAMES exists once in Rust and once in TypeScript, and redact.rs's own comment warns that drifting tables are "discovered by the leak" — but no contract test compared them. scripts/diagnostics-contract.test.ts now derives both lists from the source that owns them and fails on any divergence, on a missing key-consult or key-scan on either side, and on the bare word key entering the suffix table (which would redact public_key and cache_key and gut the report).
  • Discarding an untracked file was gated and recorded as a modification, not a deletion. cmd_discard_changes declared op = "modify" unconditionally, but GitWriter::discard_changes runs both git restore and git clean -f against the path: for a tracked file the restore acts, for an untracked one the clean acts and the file is removed. op is sent to the policy sidecar as policy.check.file's op and recorded in the ledger as file.<op>, so every discard of an untracked file asked the gate to judge — and the durable record to state — a gentler act than the one that ran. The operation is now resolved from the index via a new GitReader::is_tracked (:(literal) pathspec, so a path of * cannot be answered for by some other file), and an unreadable index fails closed to delete rather than to the gentler claim.
  • Two new contracts bind the gate to the command that actually runs. command-policy-contract.test.ts proved a gate was present; nothing proved it was told the truth. file-gate-fidelity-contract.test.ts derives the destructive writers from source and fails when a command that may delete declares a fixed modify/write/create. command-gate-fidelity-contract.test.ts compares, for 17 guarded commands, every literal flag handed to guard() against the flags its GitWriter can actually pass — following private helpers such as commit_inner transitively — and requires every command it cannot compare to be documented as derived, so a new command can neither drift nor fall silently out of the contract. Audited all 38 guarded commands: no argv drift existed beyond the discard case above.
  • The five local-AI features were exercised by no test in CI. generate_commit_message, explain_commit, suggest_branch_name, fix_health and coverage_report were reachable only through local_ai_live.rs, which is opt-in behind GITPULSE_LIVE_AI=1 and needs a real model server — so 482 lines of ai/mod.rs, including prompt assembly, the token budget and the reply path, ran in no automated check. They are now driven end to end against a loopback model server over a real socket, with the MANVI binary forced absent so the assertions land on the degraded path — the ordinary case for a user who has never installed the harness. The pipeline must still produce a result AND say the harness never parsed the reply: degrading is correct, degrading silently is not. The first version of this test was itself machine-dependent — a real manvi on the developer's PATH answered chat.prepare/chat.settle, so it silently exercised the harness-present path and asserted nothing about the degraded one; the binary override makes the outcome the same on every machine.
  • Every codeintel read surface reported success on a repository whose code map was never built. search, impact, dependencies, dead_symbols and trace_between opened the devmap database, found no indexed generation, and returned available: true with an empty list — so a repository mid-index answered "no callers", "no dependencies" and "no dead symbols", which reads as verified-clean rather than not-yet-known. status() reported the same repository as unavailable at the same moment. All five now go through a new open_indexed_store that requires a generation; status deliberately keeps the looser gate so it can still tell "no database", "no generation yet" and "database unreadable" apart, since those call for three different actions.
  • Nothing stopped two native menu items from claiming the same accelerator. A collision is silent — muda binds one item and the other's shortcut simply never fires — and build_native_menu cannot be unit tested to catch it: muda::MenuChild can only be constructed on the main thread, so calling it from a test harness panics in the platform layer (verified by running it, not assumed from the comment). view-menu-contract.test.ts now derives every accelerator from the menu source, resolving &str constants and skipping PredefinedMenuItem title arguments, and fails on any duplicate. It is cfg-aware, because Settings binds CmdOrCtrl+, twice on purpose — once under cfg(target_os = "macos") and once under cfg(not(...)) — and those never coexist in one binary.
  • Redaction cost is now linear in the number of object keys. Disambiguating a redacted key that collides with another probed for a free name linearly from #2, which is quadratic on precisely the input that is cheapest to construct: N distinct tokens of equal length all redact to identical text, so key n paid n probes. Measured at 20,000 such keys: 31s before, 277ms after, with all 20,000 entries preserved. Guarded in both languages by a budget that always runs.
  • A complexity check silently skipped itself on fast machines and failed on loaded ones. markDevParser's quadratic-growth test compared wall-clock timings behind if (small > 5), so when the small case rounded under 5 ms the assertion never ran and the test reported green having asserted nothing; when the machine was loaded the same tiny denominator made the ratio explode and the build failed on machine speed rather than on complexity (observed 13.7 against a threshold of 12, with the correct implementation). The always-running absolute budget is the real guard — measured, the unbounded pattern misses it by 2.5x — and complexity is now asserted structurally: every asymmetric bracket scan ([^\]], [^)]) must carry an explicit bound, while the symmetric delimiters that fail fast are correctly left alone.
  • Timing-based tests failed on a busy machine, for reasons that had nothing to do with the code. Ten suites asserted absolute wall-clock budgets, and seven more relied on Vitest's 5-second default timeout; both encode the speed of the machine they were written on. Measured under a concurrent build (load average 45 on 18 cores): ten test files failed, GraphRendering.rails taking 31,420 ms where it takes 2,198 ms on a quiet machine — a 14x spread on an identical tree, and every failure a false alarm. Raising the numbers would have bought quiet by weakening the guard on a fast machine, so budgets are now expressed as multiples of work the machine can do right now: src/lib/__tests__/perfBudget.ts times a fixed-instruction reference loop immediately after the measured work, and the ratio cancels the machine out of the assertion, leaving the algorithmic claim the test actually means. The unit counts were read off a GITPULSE_PERF_REPORT=1 run rather than guessed — the tightest case had only 3.6x headroom and is now at 11x. Stress and fuzz cases whose assertions are invariants rather than speed carry an explicit STRESS_TIMEOUT_MS instead, which weakens no check: a real non-termination still fails, just later. Verified by saturating the machine to load average 117 with 54 CPU hogs on 18 cores and running the whole suite: 3,355 of 3,355 passed.
  • build_native_menu is now reachable from a test, and the recent-repositories branch with it. muda::MenuChild can only be constructed on the process main thread, and a libtest harness runs every case on a worker — so the largest function in the desktop module, carrying every menu entry and accelerator, ran in no test. The constraint is on the thread, not on testability: src-tauri/tests/native_menu_main_thread.rs is declared harness = false, giving it a fn main that is the main thread. Reaching the populated branch also meant extracting set_recent_menu from cmd_set_recent_menu, whose #[tauri::command] signature had bound the cap, the state write and the menu rebuild to the concrete Wry handle no test can construct; the command is now the thin wrapper it should have been, and the cap, its ordering, degenerate paths and the return to the empty placeholder are all covered.
  • The plugin manifest advertised a version the binary did not have, and the release gate could not see it. plugin/plugin.json said 0.0.4 while package.json, package-lock.json, tauri.conf.json, Cargo.toml and Cargo.lock all said 0.0.5; check:release gated exactly those five files and printed OK: all version sources agree on 0.0.5 while the drift sat one directory away. The MCP binary answers initialize with the crate version, so a client would have recorded an 0.0.4 install against an 0.0.5 server with nothing in the handshake to reveal it — confirmed by the install, which now records 0.0.5 only because the manifest was corrected first.
  • The release gate now discovers plugin manifests instead of naming them. One package ships a manifest per agent client (plugin.json, .claude-plugin/plugin.json, .codex-plugin/plugin.json) and that set grows whenever a client is added, so a hand-maintained list stops covering the newest manifest — the one most likely to be wrong. The gate walks plugins/<name>/, requires plugin.json in each package so a deletion cannot read as "nothing to check", fails closed when no package exists at all, and checks every optional client manifest it finds. It went from 6 version sources to 9 without a path being added by hand, the Codex manifest among them.
  • gitpulse-mcp on PATH was an untracked copy that would have gone stale in silence. The binary at ~/.cargo/bin/gitpulse-mcp had been placed by hand: byte-identical to the release build at the time, but absent from Cargo's install records, so nothing tied it to the repo and the next rebuild would have left the connected server serving an older tree while still answering handshakes normally. cargo install reported it as replacing package unknown, confirming the missing provenance. It is now a tracked install, refreshable with npm run mcp:install and checkable with npm run mcp:doctor.
  • A Claude Code manifest under plugin/ could never have been committed. .gitignore's broad .claude* rule matches .claude-plugin/ at any depth, and only the root marketplace and the canonical package carry negations — so a .claude-plugin/plugin.json placed beside the legacy package was invisible to git and would have been absent from a fresh clone, failing there as "missing manifest" rather than "ignored file", far from the cause. The package converged on the one canonical directory, and plugin-contract.test.ts now asserts that every file under a marketplace source survives git check-ignore.
  • The contract-test table's exemption list is derived rather than hand-kept. documented-counts-contract.test.ts held a literal set of "tests that belong to a script", so every new script test had to be added by hand or the table would demand a contract row for a plain unit test. A foo.test.ts beside a foo.mjs is now recognized as that script's unit test, with vite-config — whose subject is vite.config.ts — the one documented exception.
  • Blame had two ways to name a file, and one of them could disagree. It carried a path box because it was reachable with nothing selected; the store's selectedFilePath was the other. As a section of Code — with Explorer one click away and sharing that selection — the box is gone. Its one irreplaceable job came back explicitly: retrying a failed blame used to mean re-typing the path and pressing Enter, and is now a Retry button in the error state.
  • A latent startup panic in the native View menu. The menu was built from nine hand-unrolled indexes into VIEW_TAB_BINDINGS, so shortening that list — exactly what consolidation does — would have panicked at launch, and lengthening it would have dropped the extra view with nothing to say so. It is built from the list itself now.
  • A stale-session race in applyToSession. Callers spread a session captured before an await and wrote it back after, so a concurrent update landing in between was silently overwritten. It accepts a patch function receiving the live session, and the three call sites that switch to History's diff section use it.
  • Exclude iPhones, iPads, and desktop-mode iPads from Mac-specific window chrome and appearance.
  • The Stack page could only ever show stacks that needed nothing done to them. build_stack_hierarchy records a parent only while a branch's first-parent walk lands on another branch's literal current tip, so one extra commit on feat-a made feat-b report main as its base and the stack silently became two siblings. The button was the wrong way round in both directions: Restack was offered exactly when it was a no-op — the branch already sat on its parent's tip — and hidden exactly when it was needed. This is not a gap in the backend to close by inference. Git stores no "cut from" edge, and once the parent moves nothing on disk distinguishes a drifted child from an unrelated branch sharing history; the two are structurally symmetric, and picking a direction would be fabricated hierarchy. So the page states the limit instead, on every populated stack rather than only on the empty one, and lists the local branches the walk placed nowhere under On no stack — a stack that has fallen apart must not read as a repository that never had one.
  • A cascading restack replays each child from the parent tip the stack was read at, not from a recomputed one. Once a parent is rebased, merge-base(parent, child) has collapsed back to the trunk, so replaying from it tells git to re-apply the parent's own commits on top of the parent. merge-base --fork-point normally rescues that from the parent's reflog — and does not in a fresh clone, in a bare repository (core.logAllRefUpdates defaults off there), or once gc.reflogExpire has run, which is the ordinary state of a long-lived stack. cmd_restack takes an optional fork_point, refuses one that is not an ancestor of the branch rather than quietly widening the rewrite the caller planned, and is covered by a test that builds exactly that repository: with the reflog expired and the parent's commits revised during the update, the computed plan drags the parent's stale pre-image into the child and conflicts on a file the child never touched.
  • A cascade that stopped part-way no longer erases its own report. The failure path set a banner naming which branches had been rebased and which were still on their old base, then reloaded the tree — and the reload cleared the banner on the way past, leaving a half-rebased repository with a screen that said nothing had happened. Clearing belongs to the next attempt, which does it before its first await; a watcher tick is not an attempt.
  • gp-pill was written by a dozen call sites across the Remote, MANVI and Stack panes and defined nowhere. Every one of them rendered as bare inline text, with the !bg-…/!text-… overrides beside them landing on nothing. Defined once in app.css rather than replaced at each call site.
  • CI verdicts carried one theme's shade. text-green-400 / text-red-400 / text-amber-400 are tuned for the dark theme and sit near 2:1 against the light theme's near-white card — on precisely the labels a reader opens the page to check. Every verdict carries both shades now, in the panel and in ciStepClass.
  • Three copies of the git path parser disagreed. patchBuilder and wordDiff each carried their own unquoting and prefix-stripping, and octal escapes in a quoted path were decoded as characters rather than as bytes, so "sp ace/\303\251.ts" came out mangled instead of as sp ace/é.ts. There is one parser now (diff/gitPaths.ts), decoding octal escapes as UTF-8 bytes and splitting a bare/bare diff --git header at the b/ where both sides name the same file.
  • A source file with a raw NUL byte is invisible to search. ripgrep and grep sniff for NUL, classify the whole file as binary and skip it in silence, so a search for a symbol that is in the file reports that the file does not use it. DiffViewer.svelte carried four (cache keys joined on a literal NUL), and fileRail.ts, FleetView.svelte and operation.test.ts one each. All five are written as escapes now, and a new greppable-source-contract walks the tree so the next one fails the suite instead of quietly shrinking every future search.
  • The shortcuts modal still listed ⌘1–9 for nine views that no longer exist. It reads ⌘1–3 … Code, History, Insights, matching what the native menu actually binds, and gained the diff's own chord list.