Releases: CodeHalwell/codypendent
Release list
codypendent v0.14.0 (build 146)
v0.14.0
Everything since v0.13.0. This release combines the council and TUI work that
landed after that tag with an eleven-part adversarial review of the execution,
recovery, control-plane, client and release boundaries.
The minor bump reflects new durable council behavior, repository-scoped
control-plane synchronization, and the first complete remote-runner policy
boundary. The rest of the release is intentionally dominated by correctness:
work that succeeds must not be repeated, authority supplied by a remote job
must not outrank local policy, and a client must not claim a server capability
that does not exist.
Councils that preserve and review the work
Council deliberation now has an explicit board shared across rounds, and every
round reaches a chair decision instead of merely producing a final summary.
Members have concrete role obligations; a member that exhausts its time keeps
and receives credit for its partial work; an independent reviewer reads the
synthesis before handoff; and citation verification checks the cited material
instead of accepting claims that it was checked.
A runner with a real local trust boundary
A claimed job can now narrow, but never expand, an immutable local runner
policy. Host mounts, working directory, environment access, resource bounds
and data classification are validated locally and bound into deterministic job
and attestation hashes.
Container execution keeps a stable daemon-owned identity and kills and reaps
the real container on every terminal path. Process execution polls host
cancellation and reaps the process group. Combined output is bounded while it
is read rather than after it has consumed memory.
Artifact paths are resolved beneath the attempt directory with no-follow
filesystem operations, required output failures are terminal, and upload
progress is journaled so a completed non-idempotent job is not re-executed just
because registration was interrupted. Cleanup handles mode-000 trees without
following symlinks and quarantines an attempt it cannot securely remove.
Claims use absolute expiry plus bounded renew/finalize calls; a failed or
non-monotonic renewal cancels the workload.
Recovery that does not repeat effects
Writing runs require launch and turn checkpoints before destructive work.
Paused runs reserve one recovery owner, restore their own durable assistant and
tool turns, reattach the existing worktree lease, retain approvals and
cancellation, and fail closed when safe replay cannot be established.
Terminal transitions are compare-and-set, so a late executor error cannot
overwrite cancellation, pause or completion. Repository and model provenance
is resolved at the originating durable sequence; corrupt provenance never
falls back to the daemon's current directory. A forced worktree release is
refused if its safety patch cannot capture the tree.
The daemon now starts a real control-plane synchronization service using the
generated protocol, repository-scoped cursors, durable backoff and owner-only
OS credential storage. Provider connections have explicit connect/idle
timeouts and bounded error bodies.
Control-plane authentication and tenant isolation
Refresh, pairing, daemon and WebSocket credentials use OS-generated 256-bit
secret material. JWT validation fixes algorithm, issuer and audience and
enforces issued-at, expiry and maximum-lifetime bounds. Refresh rotation is an
atomic compare-and-set; replay revokes the stolen descendant family rather
than every session belonging to the user.
Authorization rechecks active users, memberships and organizations and carries
credential purpose and audience into each route decision. Pairing completion
locks and validates the challenge, membership, scope, daemon and credential in
one transaction in both the memory and PostgreSQL stores. The live PostgreSQL
suite covers concurrent completion and rollback.
Caller-asserted identity linking is disabled with an explicit 501 until a
verified provider flow exists. Object uploads recheck both organization and
daemon policy before writing, direct arbitrary PUT is disabled, and downloads
derive their key only from authorized metadata. WebSocket clients use
short-lived one-use scoped tickets; subscription begins before complete,
repository-scoped replay and event-id deduplication close the replay/live race.
Synchronization that converges under crashes and policy changes
Authoritative session, run, artifact, approval, fork and graph writes now feed
the control-plane outbox in their production transactions, with startup repair
for legacy rows. Federated and legacy local repository identities resolve only
through the authenticated repository catalog and a hash-verified consent
manifest; catalog responses expose the live organization∩repository policy,
not a stale wider repository row.
Organization policy is drained to a stable cursor before the first outbound
byte. Artifact classifications and every graph fact are checked at the daemon
and again by the control plane. Policy-blocked rows can recover after a later
widening, while malformed and terminally rejected rows move to a durable
dead-letter state so they cannot starve the queue.
Policy repair appends a fresh sanitized occurrence instead of rewriting a
possibly committed sequence, and it preserves per-subject chronology. Session
and graph deletions supersede ambiguous queued publications; session
tombstones also dominate late summaries in a repository-and-daemon-scoped
ledger, preventing resurrection without letting another repository suppress a
same-named session. Revisited run states carry a durable sync revision so
Running → Paused → Running remains three distinct occurrences.
Cancelled runs can no longer leave child sessions permanently uncloseable when
cancellation lands during assembly-owned startup. After a short grace period, a
state- and event-guarded watchdog idempotently supplies missing RunCompleted
evidence; council cleanup reuses its attached connection and leaves a write-free
backoff between close-barrier polls. The real-daemon lifecycle race passed six
consecutive focused runs in addition to the affected package suites.
Clients that agree on what is live
Desktop session attachment is generation-correlated. The old session remains
authoritative until the replacement is accepted, attach-time frames are
buffered, stale history is rejected, snapshots set sequence watermarks and gap
repair remains scoped to the correct session through reconnects.
Web routes have an authentication guard and stale async responses cannot
replace newer state. React consumers share a multicast control-plane stream
with listener isolation and reference-counted teardown.
Bearer credentials now keep protected routes available even though the server
does not yet expose current-user lookup. Stream subscriptions require an
explicit supported stream in both the core and React APIs, so an omitted scope
cannot be reported as an all-stream connection while the server silently
narrows its ticket to sync events.
The control-plane SDK now maps to the actual Axum routes and generated wire
types. API-compatible methods for capabilities the server has not implemented
reject locally with UnsupportedControlPlaneCapabilityError instead of
issuing fabricated requests.
The TUI moves new/switch/fork/reconnect preparation off its sole input loop,
rejects superseded completions by generation, bounds deferred input and
deduplicates reconnect catch-up. Accessible mode now has command parity,
incremental streaming output, cold blackboard loading and visible writer
failure. The composer also edits like an editor and its surfaces answer the
keys they advertise.
Release integrity
Every third-party workflow action is pinned to a full commit SHA. Jobs default
to read-only permissions, checkouts do not persist credentials, and only the
publisher receives release-write authority. The PostgreSQL service image is
digest-pinned.
Both Rust workspaces, both lockfiles, generated protocol families, browser
clients, PostgreSQL integration tests and dependency policies are release
gates. Root and control-plane migration checksum manifests are immutable even
in shallow clones, with a regression that tampers with both SQL and manifest
from a depth-one checkout. A structural workflow test prevents these controls
from silently drifting.
Malformed durable sync rows now fail closed as unknown events instead of
inventing an empty subject and presenting a partial projection as a valid
delta. New sync writes reject blank, oversized, and control-character subject
identifiers before persistence.
Validation
The pre-final-review root workspace baseline completed 3,991 tests with 9
expected ignores and no failures. The final changed-package gates then passed
406 daemon unit tests, 14 daemon synchronization integrations, 294 composition
root unit tests, the real-router synchronization end-to-end test and 129
control-plane tests. Strict all-target/all-feature Clippy, both
generated-protocol checks and formatting passed. The separate Tauri workspace
passed format, locked all-target compilation, strict Clippy, 33 unit tests and
doc tests.
The focused client suites passed 158 desktop tests, 460 protocol SDK tests, 46
control-plane SDK tests, 10 React binding tests and 16 web tests. Both
Cargo-deny policies, both migration manifests and release workflow security
checks passed.
The full review and the deliberately retained architectural follow-ups are in
docs/reviews/2026-08-29-massive-review.md.
v0.14.0
v0.14.0
Everything since v0.13.0. This release combines the council and TUI work that
landed after that tag with an eleven-part adversarial review of the execution,
recovery, control-plane, client and release boundaries.
The minor bump reflects new durable council behavior, repository-scoped
control-plane synchronization, and the first complete remote-runner policy
boundary. The rest of the release is intentionally dominated by correctness:
work that succeeds must not be repeated, authority supplied by a remote job
must not outrank local policy, and a client must not claim a server capability
that does not exist.
Councils that preserve and review the work
Council deliberation now has an explicit board shared across rounds, and every
round reaches a chair decision instead of merely producing a final summary.
Members have concrete role obligations; a member that exhausts its time keeps
and receives credit for its partial work; an independent reviewer reads the
synthesis before handoff; and citation verification checks the cited material
instead of accepting claims that it was checked.
A runner with a real local trust boundary
A claimed job can now narrow, but never expand, an immutable local runner
policy. Host mounts, working directory, environment access, resource bounds
and data classification are validated locally and bound into deterministic job
and attestation hashes.
Container execution keeps a stable daemon-owned identity and kills and reaps
the real container on every terminal path. Process execution polls host
cancellation and reaps the process group. Combined output is bounded while it
is read rather than after it has consumed memory.
Artifact paths are resolved beneath the attempt directory with no-follow
filesystem operations, required output failures are terminal, and upload
progress is journaled so a completed non-idempotent job is not re-executed just
because registration was interrupted. Cleanup handles mode-000 trees without
following symlinks and quarantines an attempt it cannot securely remove.
Claims use absolute expiry plus bounded renew/finalize calls; a failed or
non-monotonic renewal cancels the workload.
Recovery that does not repeat effects
Writing runs require launch and turn checkpoints before destructive work.
Paused runs reserve one recovery owner, restore their own durable assistant and
tool turns, reattach the existing worktree lease, retain approvals and
cancellation, and fail closed when safe replay cannot be established.
Terminal transitions are compare-and-set, so a late executor error cannot
overwrite cancellation, pause or completion. Repository and model provenance
is resolved at the originating durable sequence; corrupt provenance never
falls back to the daemon's current directory. A forced worktree release is
refused if its safety patch cannot capture the tree.
The daemon now starts a real control-plane synchronization service using the
generated protocol, repository-scoped cursors, durable backoff and owner-only
OS credential storage. Provider connections have explicit connect/idle
timeouts and bounded error bodies.
Control-plane authentication and tenant isolation
Refresh, pairing, daemon and WebSocket credentials use OS-generated 256-bit
secret material. JWT validation fixes algorithm, issuer and audience and
enforces issued-at, expiry and maximum-lifetime bounds. Refresh rotation is an
atomic compare-and-set; replay revokes the stolen descendant family rather
than every session belonging to the user.
Authorization rechecks active users, memberships and organizations and carries
credential purpose and audience into each route decision. Pairing completion
locks and validates the challenge, membership, scope, daemon and credential in
one transaction in both the memory and PostgreSQL stores. The live PostgreSQL
suite covers concurrent completion and rollback.
Caller-asserted identity linking is disabled with an explicit 501 until a
verified provider flow exists. Object uploads recheck both organization and
daemon policy before writing, direct arbitrary PUT is disabled, and downloads
derive their key only from authorized metadata. WebSocket clients use
short-lived one-use scoped tickets; subscription begins before complete,
repository-scoped replay and event-id deduplication close the replay/live race.
Synchronization that converges under crashes and policy changes
Authoritative session, run, artifact, approval, fork and graph writes now feed
the control-plane outbox in their production transactions, with startup repair
for legacy rows. Federated and legacy local repository identities resolve only
through the authenticated repository catalog and a hash-verified consent
manifest; catalog responses expose the live organization∩repository policy,
not a stale wider repository row.
Organization policy is drained to a stable cursor before the first outbound
byte. Artifact classifications and every graph fact are checked at the daemon
and again by the control plane. Policy-blocked rows can recover after a later
widening, while malformed and terminally rejected rows move to a durable
dead-letter state so they cannot starve the queue.
Policy repair appends a fresh sanitized occurrence instead of rewriting a
possibly committed sequence, and it preserves per-subject chronology. Session
and graph deletions supersede ambiguous queued publications; session
tombstones also dominate late summaries in a repository-and-daemon-scoped
ledger, preventing resurrection without letting another repository suppress a
same-named session. Revisited run states carry a durable sync revision so
Running → Paused → Running remains three distinct occurrences.
Cancelled runs can no longer leave child sessions permanently uncloseable when
cancellation lands during assembly-owned startup. After a short grace period, a
state- and event-guarded watchdog idempotently supplies missing RunCompleted
evidence; council cleanup reuses its attached connection and leaves a write-free
backoff between close-barrier polls. The real-daemon lifecycle race passed six
consecutive focused runs in addition to the affected package suites.
Clients that agree on what is live
Desktop session attachment is generation-correlated. The old session remains
authoritative until the replacement is accepted, attach-time frames are
buffered, stale history is rejected, snapshots set sequence watermarks and gap
repair remains scoped to the correct session through reconnects.
Web routes have an authentication guard and stale async responses cannot
replace newer state. React consumers share a multicast control-plane stream
with listener isolation and reference-counted teardown.
Bearer credentials now keep protected routes available even though the server
does not yet expose current-user lookup. Stream subscriptions require an
explicit supported stream in both the core and React APIs, so an omitted scope
cannot be reported as an all-stream connection while the server silently
narrows its ticket to sync events.
The control-plane SDK now maps to the actual Axum routes and generated wire
types. API-compatible methods for capabilities the server has not implemented
reject locally with UnsupportedControlPlaneCapabilityError instead of
issuing fabricated requests.
The TUI moves new/switch/fork/reconnect preparation off its sole input loop,
rejects superseded completions by generation, bounds deferred input and
deduplicates reconnect catch-up. Accessible mode now has command parity,
incremental streaming output, cold blackboard loading and visible writer
failure. The composer also edits like an editor and its surfaces answer the
keys they advertise.
Release integrity
Every third-party workflow action is pinned to a full commit SHA. Jobs default
to read-only permissions, checkouts do not persist credentials, and only the
publisher receives release-write authority. The PostgreSQL service image is
digest-pinned.
Both Rust workspaces, both lockfiles, generated protocol families, browser
clients, PostgreSQL integration tests and dependency policies are release
gates. Root and control-plane migration checksum manifests are immutable even
in shallow clones, with a regression that tampers with both SQL and manifest
from a depth-one checkout. A structural workflow test prevents these controls
from silently drifting.
Malformed durable sync rows now fail closed as unknown events instead of
inventing an empty subject and presenting a partial projection as a valid
delta. New sync writes reject blank, oversized, and control-character subject
identifiers before persistence.
Validation
The pre-final-review root workspace baseline completed 3,991 tests with 9
expected ignores and no failures. The final changed-package gates then passed
406 daemon unit tests, 14 daemon synchronization integrations, 294 composition
root unit tests, the real-router synchronization end-to-end test and 129
control-plane tests. Strict all-target/all-feature Clippy, both
generated-protocol checks and formatting passed. The separate Tauri workspace
passed format, locked all-target compilation, strict Clippy, 33 unit tests and
doc tests.
The focused client suites passed 158 desktop tests, 460 protocol SDK tests, 46
control-plane SDK tests, 10 React binding tests and 16 web tests. Both
Cargo-deny policies, both migration manifests and release workflow security
checks passed.
The full review and the deliberately retained architectural follow-ups are in
docs/reviews/2026-08-29-massive-review.md.
codypendent v0.14.0 (build 144)
v0.14.0
Everything since v0.13.0. This release combines the council and TUI work that
landed after that tag with an eleven-part adversarial review of the execution,
recovery, control-plane, client and release boundaries.
The minor bump reflects new durable council behavior, repository-scoped
control-plane synchronization, and the first complete remote-runner policy
boundary. The rest of the release is intentionally dominated by correctness:
work that succeeds must not be repeated, authority supplied by a remote job
must not outrank local policy, and a client must not claim a server capability
that does not exist.
Councils that preserve and review the work
Council deliberation now has an explicit board shared across rounds, and every
round reaches a chair decision instead of merely producing a final summary.
Members have concrete role obligations; a member that exhausts its time keeps
and receives credit for its partial work; an independent reviewer reads the
synthesis before handoff; and citation verification checks the cited material
instead of accepting claims that it was checked.
A runner with a real local trust boundary
A claimed job can now narrow, but never expand, an immutable local runner
policy. Host mounts, working directory, environment access, resource bounds
and data classification are validated locally and bound into deterministic job
and attestation hashes.
Container execution keeps a stable daemon-owned identity and kills and reaps
the real container on every terminal path. Process execution polls host
cancellation and reaps the process group. Combined output is bounded while it
is read rather than after it has consumed memory.
Artifact paths are resolved beneath the attempt directory with no-follow
filesystem operations, required output failures are terminal, and upload
progress is journaled so a completed non-idempotent job is not re-executed just
because registration was interrupted. Cleanup handles mode-000 trees without
following symlinks and quarantines an attempt it cannot securely remove.
Claims use absolute expiry plus bounded renew/finalize calls; a failed or
non-monotonic renewal cancels the workload.
Recovery that does not repeat effects
Writing runs require launch and turn checkpoints before destructive work.
Paused runs reserve one recovery owner, restore their own durable assistant and
tool turns, reattach the existing worktree lease, retain approvals and
cancellation, and fail closed when safe replay cannot be established.
Terminal transitions are compare-and-set, so a late executor error cannot
overwrite cancellation, pause or completion. Repository and model provenance
is resolved at the originating durable sequence; corrupt provenance never
falls back to the daemon's current directory. A forced worktree release is
refused if its safety patch cannot capture the tree.
The daemon now starts a real control-plane synchronization service using the
generated protocol, repository-scoped cursors, durable backoff and owner-only
OS credential storage. Provider connections have explicit connect/idle
timeouts and bounded error bodies.
Control-plane authentication and tenant isolation
Refresh, pairing, daemon and WebSocket credentials use OS-generated 256-bit
secret material. JWT validation fixes algorithm, issuer and audience and
enforces issued-at, expiry and maximum-lifetime bounds. Refresh rotation is an
atomic compare-and-set; replay revokes the stolen descendant family rather
than every session belonging to the user.
Authorization rechecks active users, memberships and organizations and carries
credential purpose and audience into each route decision. Pairing completion
locks and validates the challenge, membership, scope, daemon and credential in
one transaction in both the memory and PostgreSQL stores. The live PostgreSQL
suite covers concurrent completion and rollback.
Caller-asserted identity linking is disabled with an explicit 501 until a
verified provider flow exists. Object uploads recheck both organization and
daemon policy before writing, direct arbitrary PUT is disabled, and downloads
derive their key only from authorized metadata. WebSocket clients use
short-lived one-use scoped tickets; subscription begins before complete,
repository-scoped replay and event-id deduplication close the replay/live race.
Synchronization that converges under crashes and policy changes
Authoritative session, run, artifact, approval, fork and graph writes now feed
the control-plane outbox in their production transactions, with startup repair
for legacy rows. Federated and legacy local repository identities resolve only
through the authenticated repository catalog and a hash-verified consent
manifest; catalog responses expose the live organization∩repository policy,
not a stale wider repository row.
Organization policy is drained to a stable cursor before the first outbound
byte. Artifact classifications and every graph fact are checked at the daemon
and again by the control plane. Policy-blocked rows can recover after a later
widening, while malformed and terminally rejected rows move to a durable
dead-letter state so they cannot starve the queue.
Policy repair appends a fresh sanitized occurrence instead of rewriting a
possibly committed sequence, and it preserves per-subject chronology. Session
and graph deletions supersede ambiguous queued publications; session
tombstones also dominate late summaries in a repository-and-daemon-scoped
ledger, preventing resurrection without letting another repository suppress a
same-named session. Revisited run states carry a durable sync revision so
Running → Paused → Running remains three distinct occurrences.
Cancelled runs can no longer leave child sessions permanently uncloseable when
cancellation lands during assembly-owned startup. After a short grace period, a
state- and event-guarded watchdog idempotently supplies missing RunCompleted
evidence; council cleanup reuses its attached connection and leaves a write-free
backoff between close-barrier polls. The real-daemon lifecycle race passed six
consecutive focused runs in addition to the affected package suites.
Clients that agree on what is live
Desktop session attachment is generation-correlated. The old session remains
authoritative until the replacement is accepted, attach-time frames are
buffered, stale history is rejected, snapshots set sequence watermarks and gap
repair remains scoped to the correct session through reconnects.
Web routes have an authentication guard and stale async responses cannot
replace newer state. React consumers share a multicast control-plane stream
with listener isolation and reference-counted teardown.
Bearer credentials now keep protected routes available even though the server
does not yet expose current-user lookup. Stream subscriptions require an
explicit supported stream in both the core and React APIs, so an omitted scope
cannot be reported as an all-stream connection while the server silently
narrows its ticket to sync events.
The control-plane SDK now maps to the actual Axum routes and generated wire
types. API-compatible methods for capabilities the server has not implemented
reject locally with UnsupportedControlPlaneCapabilityError instead of
issuing fabricated requests.
The TUI moves new/switch/fork/reconnect preparation off its sole input loop,
rejects superseded completions by generation, bounds deferred input and
deduplicates reconnect catch-up. Accessible mode now has command parity,
incremental streaming output, cold blackboard loading and visible writer
failure. The composer also edits like an editor and its surfaces answer the
keys they advertise.
Release integrity
Every third-party workflow action is pinned to a full commit SHA. Jobs default
to read-only permissions, checkouts do not persist credentials, and only the
publisher receives release-write authority. The PostgreSQL service image is
digest-pinned.
Both Rust workspaces, both lockfiles, generated protocol families, browser
clients, PostgreSQL integration tests and dependency policies are release
gates. Root and control-plane migration checksum manifests are immutable even
in shallow clones, with a regression that tampers with both SQL and manifest
from a depth-one checkout. A structural workflow test prevents these controls
from silently drifting.
Malformed durable sync rows now fail closed as unknown events instead of
inventing an empty subject and presenting a partial projection as a valid
delta. New sync writes reject blank, oversized, and control-character subject
identifiers before persistence.
Validation
The pre-final-review root workspace baseline completed 3,991 tests with 9
expected ignores and no failures. The final changed-package gates then passed
406 daemon unit tests, 14 daemon synchronization integrations, 294 composition
root unit tests, the real-router synchronization end-to-end test and 129
control-plane tests. Strict all-target/all-feature Clippy, both
generated-protocol checks and formatting passed. The separate Tauri workspace
passed format, locked all-target compilation, strict Clippy, 33 unit tests and
doc tests.
The focused client suites passed 158 desktop tests, 460 protocol SDK tests, 46
control-plane SDK tests, 10 React binding tests and 16 web tests. Both
Cargo-deny policies, both migration manifests and release workflow security
checks passed.
The full review and the deliberately retained architectural follow-ups are in
docs/reviews/2026-08-29-massive-review.md.
codypendent v0.13.0 (build 143)
v0.13.0
Everything since v0.12.4. Twenty commits, nearly all of them repairs to the two
clients an operator actually sits in front of — the desktop app and the TUI.
The prompt for most of it was blunt: the desktop app was reported as unusable,
with screenshots. What that turned out to mean is recorded below, and it was
not one bug. Pages flickered as fast as the machine allowed, the sidebar could
not be scrolled to the destinations it listed, chat arrived as raw Markdown,
the Providers page was inert, and panels that looked live had silently stopped
receiving anything. Two adversarial review rounds then found the rest.
The minor bump is for two additions: the Agent Skills SKILL.md format is
now a first-class package, and the code graph is drawn rather than counted.
Everything else is a fix.
Screens that could not be used
The councils and repository pages flickered continuously. Eight views take
their loader as a prop and run it from an effect keyed on that prop — and every
one of those loaders is an inline arrow in App.tsx, so it is a new function
on each render. Several also call setState on the app while they run, closing
the loop: render → effect → fetch → app setState → render, as fast as the
machine allows. Fixed with a useLoadOnMount hook that holds the callback in a
ref, rather than by hand-memoizing sixteen call sites, so a future call site
that forgets useCallback costs nothing instead of melting the view. Reverting
this does not merely fail the new test — it kills the vitest worker, which is
the same unbounded loop the screen was showing.
The sidebar could not be scrolled at all. Its nav list is a flex child with
no overflow and the default min-height: auto, which refuses to shrink below
its content — so with groups open it grew past the viewport, pushed the session
list off the bottom of a height: 100vh aside that has no overflow of its own,
and left destinations reachable neither by clicking nor by scrolling.
The chat showed raw Markdown, ## and ** as literal characters. It
renders now, through a small parser rather than a dependency, and the reason is
the security property: it emits React elements and never HTML, so model
output arriving over a socket cannot inject into the webview. A test asserts it
directly — <img src=x onerror=…> survives as visible text and produces no
element. Links render as their text plus the href rather than as a navigable
anchor, because a click target the model chose is not one the operator asked
for.
The Providers page was inert because it was rendered with no onSelect at
all — every click optional-chained into nothing. The credential and model form
already existed one view over, reachable only through Models → Add model; a
chosen provider now opens that flow already on it.
Long system notes fold behind a one-line summary, mirroring the TUI. Empty
lists offer to create the thing they are empty of. The Get Started page no
longer shouts.
Silently doing nothing
These are the ones with no error on screen, which is what made them expensive.
Memory extraction was disabled on every run, for any all-ACP setup. An ACP
entry is a full-agent executor, not a ChatClient, so the extractor refuses
it — and for a models.toml where every entry is ACP, which is the documented
and supported configuration, model selection could never succeed. Extraction
was therefore off since the machine was set up: zero rows across 4,941 events,
with the only evidence a once-per-run log line naming a protocol mismatch
rather than a consequence. Selection now falls back to any other configured
chat-capable model, and when nothing qualifies the warning says what is
actually true.
Loading skills from .claude and .agents registered nothing, and said
nothing about it. scan_skill_root filters on skill.toml before it looks
at a directory at all, and ecosystem skills carry SKILL.md with YAML
frontmatter and no skill.toml — so every one read as "not a package",
silently, because a package that is never tried produces no failure to report
either. The roots looked right and did nothing.
Analytics exports shipped one page of rows and called it complete. export
asked the query layer for max_rows + 1, but query clamps every caller to
its own 200-row page ceiling — so an export with the default 1,000-row budget
received 200 rows, and the truncation check compared that clamped 200 against
1,000 and was false for every export that could possibly have been truncated.
The artifact recorded 200 rows with truncated: false: a partial dataset
labelled whole, with nothing in it to say which rows were missing.
A gap in the desktop's live event stream was noticed and then forgotten. A
jump in sequence means events the client never received; the reducer detected
it, wrote a console warning and carried on, leaving the transcript permanently
short by that range with nothing marking where. It is now read back from the
durable log.
A stale teardown could kill the new connection, permanently.
daemon_disconnect took whatever connection was registered, with no notion of
which one the caller meant to close, so on a reconnect a deferred teardown
could shut down its replacement. Suppressing the Disconnected frame for a
deliberate disconnect — correct on its own terms — turned that race from a
visible glitch into silent death: the store kept reporting "connected" while
every command timed out.
Live watches did not survive a reconnect. A subscription belongs to the
connection that grew it, and a reconnect builds a new client, so afterwards the
daemon was streaming the open workflow run to nobody. The graph sat at its last
node transition and the blackboard at its last read, indistinguishable from a
run that had gone quiet.
Crashes, wedges and dead ends
A re-issued question crashed the TUI. QuestionAsked replaces a pending
question in place, but the card holding one answer slot per sub-question was
built only when no card existed — so a re-issue with more sub-questions left
the card sized for the previous shape and indexed past the end, panicking the
whole TUI while a question was blocking the operator. Daemon-triggerable.
A question from a closed session wedged the next one. begin_new_session
cleared runs, approvals and the composer but left pending_questions,
question_card_state and pending_prompts behind. The old question captured
the new session's composer, and answering it sent ResolveQuestion against the
new session id, which the daemon rejects — clearing nothing locally. A wedge
with no way out but restarting.
The transcript row counter saturated at 65,535 and hid the newest rows.
RunView::scroll, transcript_max_scroll and the measure pass's row counter
were u16 accumulated with saturating_add. That is reachable: a single model
entry may reach 256 KiB, roughly 3,300 wrapped rows on its own, so twenty of
them saturate the counter and a session holds many runs. Past that point follow
mode pinned to a bottom that was not the bottom and every row beyond became
unreachable — which on screen reads exactly like a hung run. Absolute offsets
are u32 now.
Switching or forking to a session skipped its history restore. Boot pages
the durable log when a catch-up arrives as a compact snapshot; SwitchSession
and ForkSession folded the snapshot alone. The same session opened blank when
reached from inside the TUI and complete when reached at boot, with nothing to
distinguish that from a session that never had a transcript.
An armed remote-UI confirmation outlived the notice announcing it. Arming
set the pending state and showed a notice for about two seconds; the notice
expired and the armed state did not. A stray Enter on the same control an hour
later executed a confirmed action with nothing on screen saying anything had
been armed.
Accessible mode named a command and then refused it. The controls line has
always read "yes or Enter confirms, no or Esc cancels" — but no was never in
the input map and fell through to "unrecognised accessible command". For a
screen-reader user that line is the interface for the dialog. Separately,
yes was reserved as a keypress in every mode, so answering an agent's
question with "yes" submitted the composer's draft instead — normally empty, so
nothing happened and nothing said why.
Authorisation surfaces
Terminal-escape injection reached the two places the operator decides
something. v0.12.0 fixed this for model prose and stopped there. The approval
modal and the question card render strings the model also chose — program,
arguments, environment, working directory, question text, option labels — and
those reached ratatui raw. Crossterm writes cell symbols verbatim, so a crafted
argument could emit OSC 52 to overwrite the clipboard, or reposition the cursor
and repaint the dialog to describe a command other than the one being approved.
That is the wrong half to have protected: the approval modal is the one place
in the app where the operator authorises something.
Journey stole the approval keys. Approvals now outrank it.
Data integrity
Concurrent preference saves could shred the repository selection. Every
save wrote through the same desktop.json.tmp, so two in flight put their bytes
into one file and both renamed it into place — and truncated JSON loads as "no
repository selected", silently discarding the operator's checkout. AuthStore
in the same tree already namespaces its temp file by pid; this sibling never
inherited it.
A failed key save left a model that looked configured. add_model loaded
auth.json before any write to catch a corrupt store — and then saved the key
after models.toml, so a save that failed on its own terms left the model
listed with no key behind it. The picker showed it as ready and the first
request failed, from an add that reported success. The key is saved...
codypendent v0.13.0 (build 142)
v0.13.0
Everything since v0.12.4. Twenty commits, nearly all of them repairs to the two
clients an operator actually sits in front of — the desktop app and the TUI.
The prompt for most of it was blunt: the desktop app was reported as unusable,
with screenshots. What that turned out to mean is recorded below, and it was
not one bug. Pages flickered as fast as the machine allowed, the sidebar could
not be scrolled to the destinations it listed, chat arrived as raw Markdown,
the Providers page was inert, and panels that looked live had silently stopped
receiving anything. Two adversarial review rounds then found the rest.
The minor bump is for two additions: the Agent Skills SKILL.md format is
now a first-class package, and the code graph is drawn rather than counted.
Everything else is a fix.
Screens that could not be used
The councils and repository pages flickered continuously. Eight views take
their loader as a prop and run it from an effect keyed on that prop — and every
one of those loaders is an inline arrow in App.tsx, so it is a new function
on each render. Several also call setState on the app while they run, closing
the loop: render → effect → fetch → app setState → render, as fast as the
machine allows. Fixed with a useLoadOnMount hook that holds the callback in a
ref, rather than by hand-memoizing sixteen call sites, so a future call site
that forgets useCallback costs nothing instead of melting the view. Reverting
this does not merely fail the new test — it kills the vitest worker, which is
the same unbounded loop the screen was showing.
The sidebar could not be scrolled at all. Its nav list is a flex child with
no overflow and the default min-height: auto, which refuses to shrink below
its content — so with groups open it grew past the viewport, pushed the session
list off the bottom of a height: 100vh aside that has no overflow of its own,
and left destinations reachable neither by clicking nor by scrolling.
The chat showed raw Markdown, ## and ** as literal characters. It
renders now, through a small parser rather than a dependency, and the reason is
the security property: it emits React elements and never HTML, so model
output arriving over a socket cannot inject into the webview. A test asserts it
directly — <img src=x onerror=…> survives as visible text and produces no
element. Links render as their text plus the href rather than as a navigable
anchor, because a click target the model chose is not one the operator asked
for.
The Providers page was inert because it was rendered with no onSelect at
all — every click optional-chained into nothing. The credential and model form
already existed one view over, reachable only through Models → Add model; a
chosen provider now opens that flow already on it.
Long system notes fold behind a one-line summary, mirroring the TUI. Empty
lists offer to create the thing they are empty of. The Get Started page no
longer shouts.
Silently doing nothing
These are the ones with no error on screen, which is what made them expensive.
Memory extraction was disabled on every run, for any all-ACP setup. An ACP
entry is a full-agent executor, not a ChatClient, so the extractor refuses
it — and for a models.toml where every entry is ACP, which is the documented
and supported configuration, model selection could never succeed. Extraction
was therefore off since the machine was set up: zero rows across 4,941 events,
with the only evidence a once-per-run log line naming a protocol mismatch
rather than a consequence. Selection now falls back to any other configured
chat-capable model, and when nothing qualifies the warning says what is
actually true.
Loading skills from .claude and .agents registered nothing, and said
nothing about it. scan_skill_root filters on skill.toml before it looks
at a directory at all, and ecosystem skills carry SKILL.md with YAML
frontmatter and no skill.toml — so every one read as "not a package",
silently, because a package that is never tried produces no failure to report
either. The roots looked right and did nothing.
Analytics exports shipped one page of rows and called it complete. export
asked the query layer for max_rows + 1, but query clamps every caller to
its own 200-row page ceiling — so an export with the default 1,000-row budget
received 200 rows, and the truncation check compared that clamped 200 against
1,000 and was false for every export that could possibly have been truncated.
The artifact recorded 200 rows with truncated: false: a partial dataset
labelled whole, with nothing in it to say which rows were missing.
A gap in the desktop's live event stream was noticed and then forgotten. A
jump in sequence means events the client never received; the reducer detected
it, wrote a console warning and carried on, leaving the transcript permanently
short by that range with nothing marking where. It is now read back from the
durable log.
A stale teardown could kill the new connection, permanently.
daemon_disconnect took whatever connection was registered, with no notion of
which one the caller meant to close, so on a reconnect a deferred teardown
could shut down its replacement. Suppressing the Disconnected frame for a
deliberate disconnect — correct on its own terms — turned that race from a
visible glitch into silent death: the store kept reporting "connected" while
every command timed out.
Live watches did not survive a reconnect. A subscription belongs to the
connection that grew it, and a reconnect builds a new client, so afterwards the
daemon was streaming the open workflow run to nobody. The graph sat at its last
node transition and the blackboard at its last read, indistinguishable from a
run that had gone quiet.
Crashes, wedges and dead ends
A re-issued question crashed the TUI. QuestionAsked replaces a pending
question in place, but the card holding one answer slot per sub-question was
built only when no card existed — so a re-issue with more sub-questions left
the card sized for the previous shape and indexed past the end, panicking the
whole TUI while a question was blocking the operator. Daemon-triggerable.
A question from a closed session wedged the next one. begin_new_session
cleared runs, approvals and the composer but left pending_questions,
question_card_state and pending_prompts behind. The old question captured
the new session's composer, and answering it sent ResolveQuestion against the
new session id, which the daemon rejects — clearing nothing locally. A wedge
with no way out but restarting.
The transcript row counter saturated at 65,535 and hid the newest rows.
RunView::scroll, transcript_max_scroll and the measure pass's row counter
were u16 accumulated with saturating_add. That is reachable: a single model
entry may reach 256 KiB, roughly 3,300 wrapped rows on its own, so twenty of
them saturate the counter and a session holds many runs. Past that point follow
mode pinned to a bottom that was not the bottom and every row beyond became
unreachable — which on screen reads exactly like a hung run. Absolute offsets
are u32 now.
Switching or forking to a session skipped its history restore. Boot pages
the durable log when a catch-up arrives as a compact snapshot; SwitchSession
and ForkSession folded the snapshot alone. The same session opened blank when
reached from inside the TUI and complete when reached at boot, with nothing to
distinguish that from a session that never had a transcript.
An armed remote-UI confirmation outlived the notice announcing it. Arming
set the pending state and showed a notice for about two seconds; the notice
expired and the armed state did not. A stray Enter on the same control an hour
later executed a confirmed action with nothing on screen saying anything had
been armed.
Accessible mode named a command and then refused it. The controls line has
always read "yes or Enter confirms, no or Esc cancels" — but no was never in
the input map and fell through to "unrecognised accessible command". For a
screen-reader user that line is the interface for the dialog. Separately,
yes was reserved as a keypress in every mode, so answering an agent's
question with "yes" submitted the composer's draft instead — normally empty, so
nothing happened and nothing said why.
Authorisation surfaces
Terminal-escape injection reached the two places the operator decides
something. v0.12.0 fixed this for model prose and stopped there. The approval
modal and the question card render strings the model also chose — program,
arguments, environment, working directory, question text, option labels — and
those reached ratatui raw. Crossterm writes cell symbols verbatim, so a crafted
argument could emit OSC 52 to overwrite the clipboard, or reposition the cursor
and repaint the dialog to describe a command other than the one being approved.
That is the wrong half to have protected: the approval modal is the one place
in the app where the operator authorises something.
Journey stole the approval keys. Approvals now outrank it.
Data integrity
Concurrent preference saves could shred the repository selection. Every
save wrote through the same desktop.json.tmp, so two in flight put their bytes
into one file and both renamed it into place — and truncated JSON loads as "no
repository selected", silently discarding the operator's checkout. AuthStore
in the same tree already namespaces its temp file by pid; this sibling never
inherited it.
A failed key save left a model that looked configured. add_model loaded
auth.json before any write to catch a corrupt store — and then saved the key
after models.toml, so a save that failed on its own terms left the model
listed with no key behind it. The picker showed it as ready and the first
request failed, from an add that reported success. The key is saved...
codypendent v0.13.0 (build 141)
v0.13.0
Everything since v0.12.4. Twenty commits, nearly all of them repairs to the two
clients an operator actually sits in front of — the desktop app and the TUI.
The prompt for most of it was blunt: the desktop app was reported as unusable,
with screenshots. What that turned out to mean is recorded below, and it was
not one bug. Pages flickered as fast as the machine allowed, the sidebar could
not be scrolled to the destinations it listed, chat arrived as raw Markdown,
the Providers page was inert, and panels that looked live had silently stopped
receiving anything. Two adversarial review rounds then found the rest.
The minor bump is for two additions: the Agent Skills SKILL.md format is
now a first-class package, and the code graph is drawn rather than counted.
Everything else is a fix.
Screens that could not be used
The councils and repository pages flickered continuously. Eight views take
their loader as a prop and run it from an effect keyed on that prop — and every
one of those loaders is an inline arrow in App.tsx, so it is a new function
on each render. Several also call setState on the app while they run, closing
the loop: render → effect → fetch → app setState → render, as fast as the
machine allows. Fixed with a useLoadOnMount hook that holds the callback in a
ref, rather than by hand-memoizing sixteen call sites, so a future call site
that forgets useCallback costs nothing instead of melting the view. Reverting
this does not merely fail the new test — it kills the vitest worker, which is
the same unbounded loop the screen was showing.
The sidebar could not be scrolled at all. Its nav list is a flex child with
no overflow and the default min-height: auto, which refuses to shrink below
its content — so with groups open it grew past the viewport, pushed the session
list off the bottom of a height: 100vh aside that has no overflow of its own,
and left destinations reachable neither by clicking nor by scrolling.
The chat showed raw Markdown, ## and ** as literal characters. It
renders now, through a small parser rather than a dependency, and the reason is
the security property: it emits React elements and never HTML, so model
output arriving over a socket cannot inject into the webview. A test asserts it
directly — <img src=x onerror=…> survives as visible text and produces no
element. Links render as their text plus the href rather than as a navigable
anchor, because a click target the model chose is not one the operator asked
for.
The Providers page was inert because it was rendered with no onSelect at
all — every click optional-chained into nothing. The credential and model form
already existed one view over, reachable only through Models → Add model; a
chosen provider now opens that flow already on it.
Long system notes fold behind a one-line summary, mirroring the TUI. Empty
lists offer to create the thing they are empty of. The Get Started page no
longer shouts.
Silently doing nothing
These are the ones with no error on screen, which is what made them expensive.
Memory extraction was disabled on every run, for any all-ACP setup. An ACP
entry is a full-agent executor, not a ChatClient, so the extractor refuses
it — and for a models.toml where every entry is ACP, which is the documented
and supported configuration, model selection could never succeed. Extraction
was therefore off since the machine was set up: zero rows across 4,941 events,
with the only evidence a once-per-run log line naming a protocol mismatch
rather than a consequence. Selection now falls back to any other configured
chat-capable model, and when nothing qualifies the warning says what is
actually true.
Loading skills from .claude and .agents registered nothing, and said
nothing about it. scan_skill_root filters on skill.toml before it looks
at a directory at all, and ecosystem skills carry SKILL.md with YAML
frontmatter and no skill.toml — so every one read as "not a package",
silently, because a package that is never tried produces no failure to report
either. The roots looked right and did nothing.
Analytics exports shipped one page of rows and called it complete. export
asked the query layer for max_rows + 1, but query clamps every caller to
its own 200-row page ceiling — so an export with the default 1,000-row budget
received 200 rows, and the truncation check compared that clamped 200 against
1,000 and was false for every export that could possibly have been truncated.
The artifact recorded 200 rows with truncated: false: a partial dataset
labelled whole, with nothing in it to say which rows were missing.
A gap in the desktop's live event stream was noticed and then forgotten. A
jump in sequence means events the client never received; the reducer detected
it, wrote a console warning and carried on, leaving the transcript permanently
short by that range with nothing marking where. It is now read back from the
durable log.
A stale teardown could kill the new connection, permanently.
daemon_disconnect took whatever connection was registered, with no notion of
which one the caller meant to close, so on a reconnect a deferred teardown
could shut down its replacement. Suppressing the Disconnected frame for a
deliberate disconnect — correct on its own terms — turned that race from a
visible glitch into silent death: the store kept reporting "connected" while
every command timed out.
Live watches did not survive a reconnect. A subscription belongs to the
connection that grew it, and a reconnect builds a new client, so afterwards the
daemon was streaming the open workflow run to nobody. The graph sat at its last
node transition and the blackboard at its last read, indistinguishable from a
run that had gone quiet.
Crashes, wedges and dead ends
A re-issued question crashed the TUI. QuestionAsked replaces a pending
question in place, but the card holding one answer slot per sub-question was
built only when no card existed — so a re-issue with more sub-questions left
the card sized for the previous shape and indexed past the end, panicking the
whole TUI while a question was blocking the operator. Daemon-triggerable.
A question from a closed session wedged the next one. begin_new_session
cleared runs, approvals and the composer but left pending_questions,
question_card_state and pending_prompts behind. The old question captured
the new session's composer, and answering it sent ResolveQuestion against the
new session id, which the daemon rejects — clearing nothing locally. A wedge
with no way out but restarting.
The transcript row counter saturated at 65,535 and hid the newest rows.
RunView::scroll, transcript_max_scroll and the measure pass's row counter
were u16 accumulated with saturating_add. That is reachable: a single model
entry may reach 256 KiB, roughly 3,300 wrapped rows on its own, so twenty of
them saturate the counter and a session holds many runs. Past that point follow
mode pinned to a bottom that was not the bottom and every row beyond became
unreachable — which on screen reads exactly like a hung run. Absolute offsets
are u32 now.
Switching or forking to a session skipped its history restore. Boot pages
the durable log when a catch-up arrives as a compact snapshot; SwitchSession
and ForkSession folded the snapshot alone. The same session opened blank when
reached from inside the TUI and complete when reached at boot, with nothing to
distinguish that from a session that never had a transcript.
An armed remote-UI confirmation outlived the notice announcing it. Arming
set the pending state and showed a notice for about two seconds; the notice
expired and the armed state did not. A stray Enter on the same control an hour
later executed a confirmed action with nothing on screen saying anything had
been armed.
Accessible mode named a command and then refused it. The controls line has
always read "yes or Enter confirms, no or Esc cancels" — but no was never in
the input map and fell through to "unrecognised accessible command". For a
screen-reader user that line is the interface for the dialog. Separately,
yes was reserved as a keypress in every mode, so answering an agent's
question with "yes" submitted the composer's draft instead — normally empty, so
nothing happened and nothing said why.
Authorisation surfaces
Terminal-escape injection reached the two places the operator decides
something. v0.12.0 fixed this for model prose and stopped there. The approval
modal and the question card render strings the model also chose — program,
arguments, environment, working directory, question text, option labels — and
those reached ratatui raw. Crossterm writes cell symbols verbatim, so a crafted
argument could emit OSC 52 to overwrite the clipboard, or reposition the cursor
and repaint the dialog to describe a command other than the one being approved.
That is the wrong half to have protected: the approval modal is the one place
in the app where the operator authorises something.
Journey stole the approval keys. Approvals now outrank it.
Data integrity
Concurrent preference saves could shred the repository selection. Every
save wrote through the same desktop.json.tmp, so two in flight put their bytes
into one file and both renamed it into place — and truncated JSON loads as "no
repository selected", silently discarding the operator's checkout. AuthStore
in the same tree already namespaces its temp file by pid; this sibling never
inherited it.
A failed key save left a model that looked configured. add_model loaded
auth.json before any write to catch a corrupt store — and then saved the key
after models.toml, so a save that failed on its own terms left the model
listed with no key behind it. The picker showed it as ready and the first
request failed, from an add that reported success. The key is saved...
codypendent v0.13.0 (build 140)
v0.13.0
Everything since v0.12.4. Twenty commits, nearly all of them repairs to the two
clients an operator actually sits in front of — the desktop app and the TUI.
The prompt for most of it was blunt: the desktop app was reported as unusable,
with screenshots. What that turned out to mean is recorded below, and it was
not one bug. Pages flickered as fast as the machine allowed, the sidebar could
not be scrolled to the destinations it listed, chat arrived as raw Markdown,
the Providers page was inert, and panels that looked live had silently stopped
receiving anything. Two adversarial review rounds then found the rest.
The minor bump is for two additions: the Agent Skills SKILL.md format is
now a first-class package, and the code graph is drawn rather than counted.
Everything else is a fix.
Screens that could not be used
The councils and repository pages flickered continuously. Eight views take
their loader as a prop and run it from an effect keyed on that prop — and every
one of those loaders is an inline arrow in App.tsx, so it is a new function
on each render. Several also call setState on the app while they run, closing
the loop: render → effect → fetch → app setState → render, as fast as the
machine allows. Fixed with a useLoadOnMount hook that holds the callback in a
ref, rather than by hand-memoizing sixteen call sites, so a future call site
that forgets useCallback costs nothing instead of melting the view. Reverting
this does not merely fail the new test — it kills the vitest worker, which is
the same unbounded loop the screen was showing.
The sidebar could not be scrolled at all. Its nav list is a flex child with
no overflow and the default min-height: auto, which refuses to shrink below
its content — so with groups open it grew past the viewport, pushed the session
list off the bottom of a height: 100vh aside that has no overflow of its own,
and left destinations reachable neither by clicking nor by scrolling.
The chat showed raw Markdown, ## and ** as literal characters. It
renders now, through a small parser rather than a dependency, and the reason is
the security property: it emits React elements and never HTML, so model
output arriving over a socket cannot inject into the webview. A test asserts it
directly — <img src=x onerror=…> survives as visible text and produces no
element. Links render as their text plus the href rather than as a navigable
anchor, because a click target the model chose is not one the operator asked
for.
The Providers page was inert because it was rendered with no onSelect at
all — every click optional-chained into nothing. The credential and model form
already existed one view over, reachable only through Models → Add model; a
chosen provider now opens that flow already on it.
Long system notes fold behind a one-line summary, mirroring the TUI. Empty
lists offer to create the thing they are empty of. The Get Started page no
longer shouts.
Silently doing nothing
These are the ones with no error on screen, which is what made them expensive.
Memory extraction was disabled on every run, for any all-ACP setup. An ACP
entry is a full-agent executor, not a ChatClient, so the extractor refuses
it — and for a models.toml where every entry is ACP, which is the documented
and supported configuration, model selection could never succeed. Extraction
was therefore off since the machine was set up: zero rows across 4,941 events,
with the only evidence a once-per-run log line naming a protocol mismatch
rather than a consequence. Selection now falls back to any other configured
chat-capable model, and when nothing qualifies the warning says what is
actually true.
Loading skills from .claude and .agents registered nothing, and said
nothing about it. scan_skill_root filters on skill.toml before it looks
at a directory at all, and ecosystem skills carry SKILL.md with YAML
frontmatter and no skill.toml — so every one read as "not a package",
silently, because a package that is never tried produces no failure to report
either. The roots looked right and did nothing.
Analytics exports shipped one page of rows and called it complete. export
asked the query layer for max_rows + 1, but query clamps every caller to
its own 200-row page ceiling — so an export with the default 1,000-row budget
received 200 rows, and the truncation check compared that clamped 200 against
1,000 and was false for every export that could possibly have been truncated.
The artifact recorded 200 rows with truncated: false: a partial dataset
labelled whole, with nothing in it to say which rows were missing.
A gap in the desktop's live event stream was noticed and then forgotten. A
jump in sequence means events the client never received; the reducer detected
it, wrote a console warning and carried on, leaving the transcript permanently
short by that range with nothing marking where. It is now read back from the
durable log.
A stale teardown could kill the new connection, permanently.
daemon_disconnect took whatever connection was registered, with no notion of
which one the caller meant to close, so on a reconnect a deferred teardown
could shut down its replacement. Suppressing the Disconnected frame for a
deliberate disconnect — correct on its own terms — turned that race from a
visible glitch into silent death: the store kept reporting "connected" while
every command timed out.
Live watches did not survive a reconnect. A subscription belongs to the
connection that grew it, and a reconnect builds a new client, so afterwards the
daemon was streaming the open workflow run to nobody. The graph sat at its last
node transition and the blackboard at its last read, indistinguishable from a
run that had gone quiet.
Crashes, wedges and dead ends
A re-issued question crashed the TUI. QuestionAsked replaces a pending
question in place, but the card holding one answer slot per sub-question was
built only when no card existed — so a re-issue with more sub-questions left
the card sized for the previous shape and indexed past the end, panicking the
whole TUI while a question was blocking the operator. Daemon-triggerable.
A question from a closed session wedged the next one. begin_new_session
cleared runs, approvals and the composer but left pending_questions,
question_card_state and pending_prompts behind. The old question captured
the new session's composer, and answering it sent ResolveQuestion against the
new session id, which the daemon rejects — clearing nothing locally. A wedge
with no way out but restarting.
The transcript row counter saturated at 65,535 and hid the newest rows.
RunView::scroll, transcript_max_scroll and the measure pass's row counter
were u16 accumulated with saturating_add. That is reachable: a single model
entry may reach 256 KiB, roughly 3,300 wrapped rows on its own, so twenty of
them saturate the counter and a session holds many runs. Past that point follow
mode pinned to a bottom that was not the bottom and every row beyond became
unreachable — which on screen reads exactly like a hung run. Absolute offsets
are u32 now.
Switching or forking to a session skipped its history restore. Boot pages
the durable log when a catch-up arrives as a compact snapshot; SwitchSession
and ForkSession folded the snapshot alone. The same session opened blank when
reached from inside the TUI and complete when reached at boot, with nothing to
distinguish that from a session that never had a transcript.
An armed remote-UI confirmation outlived the notice announcing it. Arming
set the pending state and showed a notice for about two seconds; the notice
expired and the armed state did not. A stray Enter on the same control an hour
later executed a confirmed action with nothing on screen saying anything had
been armed.
Accessible mode named a command and then refused it. The controls line has
always read "yes or Enter confirms, no or Esc cancels" — but no was never in
the input map and fell through to "unrecognised accessible command". For a
screen-reader user that line is the interface for the dialog. Separately,
yes was reserved as a keypress in every mode, so answering an agent's
question with "yes" submitted the composer's draft instead — normally empty, so
nothing happened and nothing said why.
Authorisation surfaces
Terminal-escape injection reached the two places the operator decides
something. v0.12.0 fixed this for model prose and stopped there. The approval
modal and the question card render strings the model also chose — program,
arguments, environment, working directory, question text, option labels — and
those reached ratatui raw. Crossterm writes cell symbols verbatim, so a crafted
argument could emit OSC 52 to overwrite the clipboard, or reposition the cursor
and repaint the dialog to describe a command other than the one being approved.
That is the wrong half to have protected: the approval modal is the one place
in the app where the operator authorises something.
Journey stole the approval keys. Approvals now outrank it.
Data integrity
Concurrent preference saves could shred the repository selection. Every
save wrote through the same desktop.json.tmp, so two in flight put their bytes
into one file and both renamed it into place — and truncated JSON loads as "no
repository selected", silently discarding the operator's checkout. AuthStore
in the same tree already namespaces its temp file by pid; this sibling never
inherited it.
A failed key save left a model that looked configured. add_model loaded
auth.json before any write to catch a corrupt store — and then saved the key
after models.toml, so a save that failed on its own terms left the model
listed with no key behind it. The picker showed it as ready and the first
request failed, from an add that reported success. The key is saved...
codypendent v0.13.0 (build 136)
v0.13.0
Everything since v0.12.4. Twenty commits, nearly all of them repairs to the two
clients an operator actually sits in front of — the desktop app and the TUI.
The prompt for most of it was blunt: the desktop app was reported as unusable,
with screenshots. What that turned out to mean is recorded below, and it was
not one bug. Pages flickered as fast as the machine allowed, the sidebar could
not be scrolled to the destinations it listed, chat arrived as raw Markdown,
the Providers page was inert, and panels that looked live had silently stopped
receiving anything. Two adversarial review rounds then found the rest.
The minor bump is for two additions: the Agent Skills SKILL.md format is
now a first-class package, and the code graph is drawn rather than counted.
Everything else is a fix.
Screens that could not be used
The councils and repository pages flickered continuously. Eight views take
their loader as a prop and run it from an effect keyed on that prop — and every
one of those loaders is an inline arrow in App.tsx, so it is a new function
on each render. Several also call setState on the app while they run, closing
the loop: render → effect → fetch → app setState → render, as fast as the
machine allows. Fixed with a useLoadOnMount hook that holds the callback in a
ref, rather than by hand-memoizing sixteen call sites, so a future call site
that forgets useCallback costs nothing instead of melting the view. Reverting
this does not merely fail the new test — it kills the vitest worker, which is
the same unbounded loop the screen was showing.
The sidebar could not be scrolled at all. Its nav list is a flex child with
no overflow and the default min-height: auto, which refuses to shrink below
its content — so with groups open it grew past the viewport, pushed the session
list off the bottom of a height: 100vh aside that has no overflow of its own,
and left destinations reachable neither by clicking nor by scrolling.
The chat showed raw Markdown, ## and ** as literal characters. It
renders now, through a small parser rather than a dependency, and the reason is
the security property: it emits React elements and never HTML, so model
output arriving over a socket cannot inject into the webview. A test asserts it
directly — <img src=x onerror=…> survives as visible text and produces no
element. Links render as their text plus the href rather than as a navigable
anchor, because a click target the model chose is not one the operator asked
for.
The Providers page was inert because it was rendered with no onSelect at
all — every click optional-chained into nothing. The credential and model form
already existed one view over, reachable only through Models → Add model; a
chosen provider now opens that flow already on it.
Long system notes fold behind a one-line summary, mirroring the TUI. Empty
lists offer to create the thing they are empty of. The Get Started page no
longer shouts.
Silently doing nothing
These are the ones with no error on screen, which is what made them expensive.
Memory extraction was disabled on every run, for any all-ACP setup. An ACP
entry is a full-agent executor, not a ChatClient, so the extractor refuses
it — and for a models.toml where every entry is ACP, which is the documented
and supported configuration, model selection could never succeed. Extraction
was therefore off since the machine was set up: zero rows across 4,941 events,
with the only evidence a once-per-run log line naming a protocol mismatch
rather than a consequence. Selection now falls back to any other configured
chat-capable model, and when nothing qualifies the warning says what is
actually true.
Loading skills from .claude and .agents registered nothing, and said
nothing about it. scan_skill_root filters on skill.toml before it looks
at a directory at all, and ecosystem skills carry SKILL.md with YAML
frontmatter and no skill.toml — so every one read as "not a package",
silently, because a package that is never tried produces no failure to report
either. The roots looked right and did nothing.
Analytics exports shipped one page of rows and called it complete. export
asked the query layer for max_rows + 1, but query clamps every caller to
its own 200-row page ceiling — so an export with the default 1,000-row budget
received 200 rows, and the truncation check compared that clamped 200 against
1,000 and was false for every export that could possibly have been truncated.
The artifact recorded 200 rows with truncated: false: a partial dataset
labelled whole, with nothing in it to say which rows were missing.
A gap in the desktop's live event stream was noticed and then forgotten. A
jump in sequence means events the client never received; the reducer detected
it, wrote a console warning and carried on, leaving the transcript permanently
short by that range with nothing marking where. It is now read back from the
durable log.
A stale teardown could kill the new connection, permanently.
daemon_disconnect took whatever connection was registered, with no notion of
which one the caller meant to close, so on a reconnect a deferred teardown
could shut down its replacement. Suppressing the Disconnected frame for a
deliberate disconnect — correct on its own terms — turned that race from a
visible glitch into silent death: the store kept reporting "connected" while
every command timed out.
Live watches did not survive a reconnect. A subscription belongs to the
connection that grew it, and a reconnect builds a new client, so afterwards the
daemon was streaming the open workflow run to nobody. The graph sat at its last
node transition and the blackboard at its last read, indistinguishable from a
run that had gone quiet.
Crashes, wedges and dead ends
A re-issued question crashed the TUI. QuestionAsked replaces a pending
question in place, but the card holding one answer slot per sub-question was
built only when no card existed — so a re-issue with more sub-questions left
the card sized for the previous shape and indexed past the end, panicking the
whole TUI while a question was blocking the operator. Daemon-triggerable.
A question from a closed session wedged the next one. begin_new_session
cleared runs, approvals and the composer but left pending_questions,
question_card_state and pending_prompts behind. The old question captured
the new session's composer, and answering it sent ResolveQuestion against the
new session id, which the daemon rejects — clearing nothing locally. A wedge
with no way out but restarting.
The transcript row counter saturated at 65,535 and hid the newest rows.
RunView::scroll, transcript_max_scroll and the measure pass's row counter
were u16 accumulated with saturating_add. That is reachable: a single model
entry may reach 256 KiB, roughly 3,300 wrapped rows on its own, so twenty of
them saturate the counter and a session holds many runs. Past that point follow
mode pinned to a bottom that was not the bottom and every row beyond became
unreachable — which on screen reads exactly like a hung run. Absolute offsets
are u32 now.
Switching or forking to a session skipped its history restore. Boot pages
the durable log when a catch-up arrives as a compact snapshot; SwitchSession
and ForkSession folded the snapshot alone. The same session opened blank when
reached from inside the TUI and complete when reached at boot, with nothing to
distinguish that from a session that never had a transcript.
An armed remote-UI confirmation outlived the notice announcing it. Arming
set the pending state and showed a notice for about two seconds; the notice
expired and the armed state did not. A stray Enter on the same control an hour
later executed a confirmed action with nothing on screen saying anything had
been armed.
Accessible mode named a command and then refused it. The controls line has
always read "yes or Enter confirms, no or Esc cancels" — but no was never in
the input map and fell through to "unrecognised accessible command". For a
screen-reader user that line is the interface for the dialog. Separately,
yes was reserved as a keypress in every mode, so answering an agent's
question with "yes" submitted the composer's draft instead — normally empty, so
nothing happened and nothing said why.
Authorisation surfaces
Terminal-escape injection reached the two places the operator decides
something. v0.12.0 fixed this for model prose and stopped there. The approval
modal and the question card render strings the model also chose — program,
arguments, environment, working directory, question text, option labels — and
those reached ratatui raw. Crossterm writes cell symbols verbatim, so a crafted
argument could emit OSC 52 to overwrite the clipboard, or reposition the cursor
and repaint the dialog to describe a command other than the one being approved.
That is the wrong half to have protected: the approval modal is the one place
in the app where the operator authorises something.
Journey stole the approval keys. Approvals now outrank it.
Data integrity
Concurrent preference saves could shred the repository selection. Every
save wrote through the same desktop.json.tmp, so two in flight put their bytes
into one file and both renamed it into place — and truncated JSON loads as "no
repository selected", silently discarding the operator's checkout. AuthStore
in the same tree already namespaces its temp file by pid; this sibling never
inherited it.
A failed key save left a model that looked configured. add_model loaded
auth.json before any write to catch a corrupt store — and then saved the key
after models.toml, so a save that failed on its own terms left the model
listed with no key behind it. The picker showed it as ready and the first
request failed, from an add that reported success. The key is saved...
v0.13.0
v0.13.0
Everything since v0.12.4. Twenty commits, nearly all of them repairs to the two
clients an operator actually sits in front of — the desktop app and the TUI.
The prompt for most of it was blunt: the desktop app was reported as unusable,
with screenshots. What that turned out to mean is recorded below, and it was
not one bug. Pages flickered as fast as the machine allowed, the sidebar could
not be scrolled to the destinations it listed, chat arrived as raw Markdown,
the Providers page was inert, and panels that looked live had silently stopped
receiving anything. Two adversarial review rounds then found the rest.
The minor bump is for two additions: the Agent Skills SKILL.md format is
now a first-class package, and the code graph is drawn rather than counted.
Everything else is a fix.
Screens that could not be used
The councils and repository pages flickered continuously. Eight views take
their loader as a prop and run it from an effect keyed on that prop — and every
one of those loaders is an inline arrow in App.tsx, so it is a new function
on each render. Several also call setState on the app while they run, closing
the loop: render → effect → fetch → app setState → render, as fast as the
machine allows. Fixed with a useLoadOnMount hook that holds the callback in a
ref, rather than by hand-memoizing sixteen call sites, so a future call site
that forgets useCallback costs nothing instead of melting the view. Reverting
this does not merely fail the new test — it kills the vitest worker, which is
the same unbounded loop the screen was showing.
The sidebar could not be scrolled at all. Its nav list is a flex child with
no overflow and the default min-height: auto, which refuses to shrink below
its content — so with groups open it grew past the viewport, pushed the session
list off the bottom of a height: 100vh aside that has no overflow of its own,
and left destinations reachable neither by clicking nor by scrolling.
The chat showed raw Markdown, ## and ** as literal characters. It
renders now, through a small parser rather than a dependency, and the reason is
the security property: it emits React elements and never HTML, so model
output arriving over a socket cannot inject into the webview. A test asserts it
directly — <img src=x onerror=…> survives as visible text and produces no
element. Links render as their text plus the href rather than as a navigable
anchor, because a click target the model chose is not one the operator asked
for.
The Providers page was inert because it was rendered with no onSelect at
all — every click optional-chained into nothing. The credential and model form
already existed one view over, reachable only through Models → Add model; a
chosen provider now opens that flow already on it.
Long system notes fold behind a one-line summary, mirroring the TUI. Empty
lists offer to create the thing they are empty of. The Get Started page no
longer shouts.
Silently doing nothing
These are the ones with no error on screen, which is what made them expensive.
Memory extraction was disabled on every run, for any all-ACP setup. An ACP
entry is a full-agent executor, not a ChatClient, so the extractor refuses
it — and for a models.toml where every entry is ACP, which is the documented
and supported configuration, model selection could never succeed. Extraction
was therefore off since the machine was set up: zero rows across 4,941 events,
with the only evidence a once-per-run log line naming a protocol mismatch
rather than a consequence. Selection now falls back to any other configured
chat-capable model, and when nothing qualifies the warning says what is
actually true.
Loading skills from .claude and .agents registered nothing, and said
nothing about it. scan_skill_root filters on skill.toml before it looks
at a directory at all, and ecosystem skills carry SKILL.md with YAML
frontmatter and no skill.toml — so every one read as "not a package",
silently, because a package that is never tried produces no failure to report
either. The roots looked right and did nothing.
Analytics exports shipped one page of rows and called it complete. export
asked the query layer for max_rows + 1, but query clamps every caller to
its own 200-row page ceiling — so an export with the default 1,000-row budget
received 200 rows, and the truncation check compared that clamped 200 against
1,000 and was false for every export that could possibly have been truncated.
The artifact recorded 200 rows with truncated: false: a partial dataset
labelled whole, with nothing in it to say which rows were missing.
A gap in the desktop's live event stream was noticed and then forgotten. A
jump in sequence means events the client never received; the reducer detected
it, wrote a console warning and carried on, leaving the transcript permanently
short by that range with nothing marking where. It is now read back from the
durable log.
A stale teardown could kill the new connection, permanently.
daemon_disconnect took whatever connection was registered, with no notion of
which one the caller meant to close, so on a reconnect a deferred teardown
could shut down its replacement. Suppressing the Disconnected frame for a
deliberate disconnect — correct on its own terms — turned that race from a
visible glitch into silent death: the store kept reporting "connected" while
every command timed out.
Live watches did not survive a reconnect. A subscription belongs to the
connection that grew it, and a reconnect builds a new client, so afterwards the
daemon was streaming the open workflow run to nobody. The graph sat at its last
node transition and the blackboard at its last read, indistinguishable from a
run that had gone quiet.
Crashes, wedges and dead ends
A re-issued question crashed the TUI. QuestionAsked replaces a pending
question in place, but the card holding one answer slot per sub-question was
built only when no card existed — so a re-issue with more sub-questions left
the card sized for the previous shape and indexed past the end, panicking the
whole TUI while a question was blocking the operator. Daemon-triggerable.
A question from a closed session wedged the next one. begin_new_session
cleared runs, approvals and the composer but left pending_questions,
question_card_state and pending_prompts behind. The old question captured
the new session's composer, and answering it sent ResolveQuestion against the
new session id, which the daemon rejects — clearing nothing locally. A wedge
with no way out but restarting.
The transcript row counter saturated at 65,535 and hid the newest rows.
RunView::scroll, transcript_max_scroll and the measure pass's row counter
were u16 accumulated with saturating_add. That is reachable: a single model
entry may reach 256 KiB, roughly 3,300 wrapped rows on its own, so twenty of
them saturate the counter and a session holds many runs. Past that point follow
mode pinned to a bottom that was not the bottom and every row beyond became
unreachable — which on screen reads exactly like a hung run. Absolute offsets
are u32 now.
Switching or forking to a session skipped its history restore. Boot pages
the durable log when a catch-up arrives as a compact snapshot; SwitchSession
and ForkSession folded the snapshot alone. The same session opened blank when
reached from inside the TUI and complete when reached at boot, with nothing to
distinguish that from a session that never had a transcript.
An armed remote-UI confirmation outlived the notice announcing it. Arming
set the pending state and showed a notice for about two seconds; the notice
expired and the armed state did not. A stray Enter on the same control an hour
later executed a confirmed action with nothing on screen saying anything had
been armed.
Accessible mode named a command and then refused it. The controls line has
always read "yes or Enter confirms, no or Esc cancels" — but no was never in
the input map and fell through to "unrecognised accessible command". For a
screen-reader user that line is the interface for the dialog. Separately,
yes was reserved as a keypress in every mode, so answering an agent's
question with "yes" submitted the composer's draft instead — normally empty, so
nothing happened and nothing said why.
Authorisation surfaces
Terminal-escape injection reached the two places the operator decides
something. v0.12.0 fixed this for model prose and stopped there. The approval
modal and the question card render strings the model also chose — program,
arguments, environment, working directory, question text, option labels — and
those reached ratatui raw. Crossterm writes cell symbols verbatim, so a crafted
argument could emit OSC 52 to overwrite the clipboard, or reposition the cursor
and repaint the dialog to describe a command other than the one being approved.
That is the wrong half to have protected: the approval modal is the one place
in the app where the operator authorises something.
Journey stole the approval keys. Approvals now outrank it.
Data integrity
Concurrent preference saves could shred the repository selection. Every
save wrote through the same desktop.json.tmp, so two in flight put their bytes
into one file and both renamed it into place — and truncated JSON loads as "no
repository selected", silently discarding the operator's checkout. AuthStore
in the same tree already namespaces its temp file by pid; this sibling never
inherited it.
A failed key save left a model that looked configured. add_model loaded
auth.json before any write to catch a corrupt store — and then saved the key
after models.toml, so a save that failed on its own terms left the model
listed with no key behind it. The picker showed it as ready and the first
request failed, from an add that reported success. The key is saved...
codypendent v0.13.0 (build 134)
v0.13.0
Everything since v0.12.4. Twenty commits, nearly all of them repairs to the two
clients an operator actually sits in front of — the desktop app and the TUI.
The prompt for most of it was blunt: the desktop app was reported as unusable,
with screenshots. What that turned out to mean is recorded below, and it was
not one bug. Pages flickered as fast as the machine allowed, the sidebar could
not be scrolled to the destinations it listed, chat arrived as raw Markdown,
the Providers page was inert, and panels that looked live had silently stopped
receiving anything. Two adversarial review rounds then found the rest.
The minor bump is for two additions: the Agent Skills SKILL.md format is
now a first-class package, and the code graph is drawn rather than counted.
Everything else is a fix.
Screens that could not be used
The councils and repository pages flickered continuously. Eight views take
their loader as a prop and run it from an effect keyed on that prop — and every
one of those loaders is an inline arrow in App.tsx, so it is a new function
on each render. Several also call setState on the app while they run, closing
the loop: render → effect → fetch → app setState → render, as fast as the
machine allows. Fixed with a useLoadOnMount hook that holds the callback in a
ref, rather than by hand-memoizing sixteen call sites, so a future call site
that forgets useCallback costs nothing instead of melting the view. Reverting
this does not merely fail the new test — it kills the vitest worker, which is
the same unbounded loop the screen was showing.
The sidebar could not be scrolled at all. Its nav list is a flex child with
no overflow and the default min-height: auto, which refuses to shrink below
its content — so with groups open it grew past the viewport, pushed the session
list off the bottom of a height: 100vh aside that has no overflow of its own,
and left destinations reachable neither by clicking nor by scrolling.
The chat showed raw Markdown, ## and ** as literal characters. It
renders now, through a small parser rather than a dependency, and the reason is
the security property: it emits React elements and never HTML, so model
output arriving over a socket cannot inject into the webview. A test asserts it
directly — <img src=x onerror=…> survives as visible text and produces no
element. Links render as their text plus the href rather than as a navigable
anchor, because a click target the model chose is not one the operator asked
for.
The Providers page was inert because it was rendered with no onSelect at
all — every click optional-chained into nothing. The credential and model form
already existed one view over, reachable only through Models → Add model; a
chosen provider now opens that flow already on it.
Long system notes fold behind a one-line summary, mirroring the TUI. Empty
lists offer to create the thing they are empty of. The Get Started page no
longer shouts.
Silently doing nothing
These are the ones with no error on screen, which is what made them expensive.
Memory extraction was disabled on every run, for any all-ACP setup. An ACP
entry is a full-agent executor, not a ChatClient, so the extractor refuses
it — and for a models.toml where every entry is ACP, which is the documented
and supported configuration, model selection could never succeed. Extraction
was therefore off since the machine was set up: zero rows across 4,941 events,
with the only evidence a once-per-run log line naming a protocol mismatch
rather than a consequence. Selection now falls back to any other configured
chat-capable model, and when nothing qualifies the warning says what is
actually true.
Loading skills from .claude and .agents registered nothing, and said
nothing about it. scan_skill_root filters on skill.toml before it looks
at a directory at all, and ecosystem skills carry SKILL.md with YAML
frontmatter and no skill.toml — so every one read as "not a package",
silently, because a package that is never tried produces no failure to report
either. The roots looked right and did nothing.
Analytics exports shipped one page of rows and called it complete. export
asked the query layer for max_rows + 1, but query clamps every caller to
its own 200-row page ceiling — so an export with the default 1,000-row budget
received 200 rows, and the truncation check compared that clamped 200 against
1,000 and was false for every export that could possibly have been truncated.
The artifact recorded 200 rows with truncated: false: a partial dataset
labelled whole, with nothing in it to say which rows were missing.
A gap in the desktop's live event stream was noticed and then forgotten. A
jump in sequence means events the client never received; the reducer detected
it, wrote a console warning and carried on, leaving the transcript permanently
short by that range with nothing marking where. It is now read back from the
durable log.
A stale teardown could kill the new connection, permanently.
daemon_disconnect took whatever connection was registered, with no notion of
which one the caller meant to close, so on a reconnect a deferred teardown
could shut down its replacement. Suppressing the Disconnected frame for a
deliberate disconnect — correct on its own terms — turned that race from a
visible glitch into silent death: the store kept reporting "connected" while
every command timed out.
Live watches did not survive a reconnect. A subscription belongs to the
connection that grew it, and a reconnect builds a new client, so afterwards the
daemon was streaming the open workflow run to nobody. The graph sat at its last
node transition and the blackboard at its last read, indistinguishable from a
run that had gone quiet.
Crashes, wedges and dead ends
A re-issued question crashed the TUI. QuestionAsked replaces a pending
question in place, but the card holding one answer slot per sub-question was
built only when no card existed — so a re-issue with more sub-questions left
the card sized for the previous shape and indexed past the end, panicking the
whole TUI while a question was blocking the operator. Daemon-triggerable.
A question from a closed session wedged the next one. begin_new_session
cleared runs, approvals and the composer but left pending_questions,
question_card_state and pending_prompts behind. The old question captured
the new session's composer, and answering it sent ResolveQuestion against the
new session id, which the daemon rejects — clearing nothing locally. A wedge
with no way out but restarting.
The transcript row counter saturated at 65,535 and hid the newest rows.
RunView::scroll, transcript_max_scroll and the measure pass's row counter
were u16 accumulated with saturating_add. That is reachable: a single model
entry may reach 256 KiB, roughly 3,300 wrapped rows on its own, so twenty of
them saturate the counter and a session holds many runs. Past that point follow
mode pinned to a bottom that was not the bottom and every row beyond became
unreachable — which on screen reads exactly like a hung run. Absolute offsets
are u32 now.
Switching or forking to a session skipped its history restore. Boot pages
the durable log when a catch-up arrives as a compact snapshot; SwitchSession
and ForkSession folded the snapshot alone. The same session opened blank when
reached from inside the TUI and complete when reached at boot, with nothing to
distinguish that from a session that never had a transcript.
An armed remote-UI confirmation outlived the notice announcing it. Arming
set the pending state and showed a notice for about two seconds; the notice
expired and the armed state did not. A stray Enter on the same control an hour
later executed a confirmed action with nothing on screen saying anything had
been armed.
Accessible mode named a command and then refused it. The controls line has
always read "yes or Enter confirms, no or Esc cancels" — but no was never in
the input map and fell through to "unrecognised accessible command". For a
screen-reader user that line is the interface for the dialog. Separately,
yes was reserved as a keypress in every mode, so answering an agent's
question with "yes" submitted the composer's draft instead — normally empty, so
nothing happened and nothing said why.
Authorisation surfaces
Terminal-escape injection reached the two places the operator decides
something. v0.12.0 fixed this for model prose and stopped there. The approval
modal and the question card render strings the model also chose — program,
arguments, environment, working directory, question text, option labels — and
those reached ratatui raw. Crossterm writes cell symbols verbatim, so a crafted
argument could emit OSC 52 to overwrite the clipboard, or reposition the cursor
and repaint the dialog to describe a command other than the one being approved.
That is the wrong half to have protected: the approval modal is the one place
in the app where the operator authorises something.
Journey stole the approval keys. Approvals now outrank it.
Data integrity
Concurrent preference saves could shred the repository selection. Every
save wrote through the same desktop.json.tmp, so two in flight put their bytes
into one file and both renamed it into place — and truncated JSON loads as "no
repository selected", silently discarding the operator's checkout. AuthStore
in the same tree already namespaces its temp file by pid; this sibling never
inherited it.
A failed key save left a model that looked configured. add_model loaded
auth.json before any write to catch a corrupt store — and then saved the key
after models.toml, so a save that failed on its own terms left the model
listed with no key behind it. The picker showed it as ready and the first
request failed, from an add that reported success. The key is saved...