Releases: MichelKerkmeester/skilled-agent-harness_spec-driven-loops
Release list
v4.0.0.2 — A Steadier Skill Advisor, Findable Changelogs and Leaner Goals
v4.0.0.2 is mostly a release for the skill advisor, whose prompt hook suggests a skill before each prompt reaches the model. The hook now holds up in every runtime that carries it.
The rest makes the framework easier to find your way around. Every changelog can be found by version, and a spec goal has one limit and one command that writes it. Doc validation gains off switches, and the README now says what the git hooks block.
Why This Release
The advisor's hook runs before every prompt, and running the advisor's own scenarios in five CLIs showed where it failed. It paid for routing work nobody read, a slow advisor left four runtimes with no suggestion and a turn without one looked the same as an outage.
Changelogs were hard to find. Only 52 of the 542 skill changelog entries declared a search phrase, so the index that Gate 1 and /speckit:search read held few of them.
Goals, doc validation and the git hooks each had a smaller gap. Parent goals ran over budget and shipped placeholders, turning validation off trapped every commit that staged a spec doc and nothing a new user reads said the git hooks block commits.
What's New at a Glance
- The advisor holds up under load. A slow advisor no longer empties the turn, and a second launcher no longer shuts a healthy one down.
- The hook does less and says more. Casual prompts never reach the advisor, and a turn without a suggestion names the reason.
- Lost suggestions are back. Prompts that missed out because of the old
.opencodepath, a two-letter executor name or their length now route. - Any changelog entry is one lookup away. Ask for a component and version, and that entry comes back first.
- Release notes have their own folder.
.skilled/changelog/skilled/holds one entry per Skilled release, and older histories share one format. - A goal has one limit, one command and one send rule. Up to 4,000 characters passes, and
/create:goalwrites and checks it. - Goals behave the same on every runtime. Goal commands fail loudly, and live runs fixed Hermes, Pi and OpenCode.
- Doc validation has safe off switches. A switched-off run reports as skipped, and the README explains the git hooks.
- Codex runs each hook once, and Cursor gains Grok 4.7. Smaller fixes reach Hermes, Pi, cli-jev and the deep loops.
The Skill Advisor
Claude, Codex, Cursor and Devin run the advisor through a hook shim, Pi runs it inside its own process and OpenCode runs it as a plugin. A comparison with pi-skill-orchestrator, a Pi extension that lets the model search for a skill on demand, recommended most of the hook changes below. Its larger ideas, such as a skill-search tool for Pi, wait on tests that have not run.
A Slow Advisor No Longer Empties the Turn
In Claude, Codex, Cursor and Devin a shim stops the hook after 2,500 ms, and the hook gave its advisor call the same 2,500 ms. A slow advisor hit the shim's limit first, so the turn got nothing, not even the fallback. The advisor now gets 2,200 ms, and the fallback arrives in time. Pi races the call against its own budget the same way, and a negative budget no longer throws away a live suggestion.
Less Work on Every Prompt
The check that keeps casual prompts away from the advisor had lost its only caller, so every prompt started an advisor call. The hook now runs it first, and /help or a short acknowledgement returns before any call starts. A replay of the labeled prompts confirmed that it turns away none that should route.
The hook also asked for a compiled route, a precomputed routing decision for a skill hub, and threw it away. It now tells the advisor to skip that work, which saved a median 65 ms per prompt on one measured build.
A Turn Without a Suggestion Says Why
A turn with no suggestion used to carry the bare fallback text, so the model could not tell an outage from a prompt that matched no skill. The fallback now opens with one line that names the case, and an outage line names the command to route by hand. Within a known session a repeat shrinks to that line, 26 bytes where the full text is 244, and the OpenCode plugin opens its fallback the same way.
Suggestions Lost to Old Paths, Short Names and Length
The advisor moved from .opencode to .skilled, and three parts of it kept the old habits. The hook's fast path looked under the old root and gave up on every prompt, Pi held back a changed suggestion as a repeat and the OpenCode plugin's cache never noticed a skill change. Each now finds the right root or compares the whole delivery.
The scorer also dropped every word of two characters or fewer, so "delegate this to pi" got no suggestion, and it now routes to cli-external-orchestration with a compiled route to cli-pi. Long prompts failed too, because the callers sent up to 64 KiB where the advisor accepts 10,000 characters. Callers now send at most 10,000 characters over standard input, which also keeps the prompt out of the process list.
A Healthy Advisor Stays Up
A launcher treated a live advisor as abandoned when its parent process was gone, so a second launcher deleted the lease and the healthy advisor shut down. A lease is now reclaimable only when its process is dead or its heartbeat is stale. SYSTEM_SKILL_ADVISOR_DB_DIR also moves every state file with it, so a test advisor can no longer mark the live one unavailable.
OpenCode Loads the Plugin and Stops Repeats
OpenCode refused the advisor plugin because the file exported helpers beside the factory it expects. The file now exports only its factory, and OpenCode sessions receive the suggestion and the status tool again. With deduplicateTransforms on, a second transform for the same message no longer delivers the block twice.
Diagnostics Name the Real Runtime
The hook now waits up to 100 ms for pending log writes, so a hook run as a separate process keeps its diagnostic record, and each record names its runtime. advisor_validate now accepts outcome events from Pi, Codex, Cursor and Devin as well.
Changelogs
What Every Entry Now Declares
Every entry opens with the five keys a spec document carries: title, description, trigger_phrases, importance_tier and contextType. Its first trigger phrases name the component and version, such as sk-git v1.0.0.0 and sk-git 1.0.0.0, or v4.0.0.2 release notes for a Skilled release. One or two topic phrases follow.
How to Look One Up
Run /speckit:search sk-git v1.0.0.0 --triggers, which matches your words against those phrases and ranks that entry first. Gate 1 runs the same lookup on every prompt. The index now also reads .skilled/changelog/skilled, so v4.0.0.0 release notes finds its release note.
The Validator Keeps It That Way
validate_document.py blocks an entry that lacks any of the five keys or a phrase that contains its version. The nested changelog generator gives every packet a phrase of its own, where all packets used to share one. It also escapes what it writes, so a quote in a spec title no longer damages a packet changelog.
Release Notes With a Home of Their Own
The system-spec-kit changelog had become the framework's release notes, and the skill's own history stopped at 3.9.0.0. The 45 release notes moved to .skilled/changelog/skilled/, and system-spec-kit writes entries about itself again. /create:changelog skilled writes a release entry, and only that line publishes a GitHub release.
One Format for Every History
Every older entry, the Skilled release notes included, was rewritten in the current format, so a skill's history reads the same from one version to the next. A fact check held each rewrite to what its original recorded.
The Goal System
One Limit and One Boundary
A top-level or phase-parent goal.md used to warn past 3,000 characters and fail past 4,000, and the lower number read as the limit. The warning is gone: up to 4,000 characters passes and past it fails.
The tools also disagreed on which folders the limit covers. They now agree that it applies to a top-level packet and to every phase parent, nested or not, and that only a phase child that is not itself a parent is exempt. Section 2 of the create-goal mode's references/budget-and-handoff.md states that boundary.
A Command That Writes Goals and a Checker
sk-doc gains the create-goal mode and its /create:goal command. It writes a top-level, phase-parent or phase-child goal from the packet's own documents, amends one and retrofits one onto a packet that has none with /create:goal <packet path> retrofit. It never sets or resends a session goal, which stays with the goal hooks.
check-goal.cjs is read-only and flags a phase the parent never bound, a leftover placeholder, a criterion count outside three to seven, a parent over budget and frontmatter that a --- line closed early. /create:goal hands back a parent's chat slice only after the checker passes.
A Sent Goal Carries Only Its Directive
Section 4 of budget-and-handoff.md is now the one rule for a goal sent in chat. A sent goal carries no frontmatter, comments, anchors, dividers, section numbers or author instructions, and it goes out only while its budget reads ok. The templates lost the text that used to travel with every sent goal, so an unfilled phase-parent goal fell from 2,823 characters in chat to 1,218.
Two Corrected Claims About Goals
create.sh --phase --with-goal writes child goals only, and --level phase-parent --with-goal writes the parent goal. The runtime injects a pointer to goal.md with the binding rule and the completion criteria on every turn, not the whole goal. system-spec-kit's phase docs and the create.sh help now send goal authoring to /create:goal.
Goal Commands That Fail Loudly
goal.cjs used to store `goal.cjs -...
v4.0.0.1 — The Foundation Hardened: Old Specs Forward and Quieter Sessions
v4.0.0.1 makes the v4 rebuild usable on the work you already have and quieter to work in. It is a hardening release, proven on this repository's own years of spec history rather than on a demo.
The headline is the migration path. Spec folders written under v3.x fail the strict validator almost without exception, and until now every workflow stopped at that failure. A new command repairs them in one pass with no language model involved. It fills in what a document lacks without rewriting a word you wrote and records what it cannot clear, so old findings stop looking like new mistakes.
The rest of the release closes the seams the first weeks of real use found. The question that binds a session to a spec folder now arrives once at your first save instead of on every turn, and it reaches every runtime. The tools that create, move and archive spec folders keep them consistent, saves stop touching files they do not own and the model rosters moved up a generation across every dispatch surface.
Documentation changed with them. Release notes are written in the narrative style this file follows, playbook prompts can no longer drift from their contracts and the always-loaded rules no longer contradict the prompting guides a model reads first.
Why This Release
A rebuild as large as v4.0.0.0 moves so much that its own history becomes the fastest way to find the seams. Everything here was measured against real packets, real sessions and real dispatches.
It also carries the one thing a person upgrading from v3.x needs most. Old spec folders should cost one command to adopt, not a repair project paid for in AI time. This release ships that command and the validator change that keeps legacy findings from blocking new work.
What's New at a Glance
- Old spec folders upgrade in one command. A tool with no language model fills what v3 documents lack and records the findings it cannot clear. Strict validation passes and old debt stops blocking new work.
- The spec-folder question asks once. The gate now waits for your first save instead of riding every turn, and it reaches every runtime.
- Model rosters moved up a generation. MiMo v2.6 and GPT-6 Luna and Sol replace their predecessors everywhere a dispatch can name them.
- MiMo runs through the gateway alone. The direct Xiaomi routes are gone and dispatches reach MiMo through LLM Gateway instead.
- Track roots list the packets they hold. A writer keeps every list true to disk and a push that would publish a lie stops at the hook.
- Archives keep packets with their family. Archiving a packet or a single phase stays inside its track or its parent and restores the same way.
- New packets land where the tools look. Scaffolding writes to the canonical specs root and numbers from the highest folder instead of restarting at 001.
- Saves write only what they own. A save records the fields it carries and leaves the files a whole track shares alone.
- The advisor brief is back on Pi. Every prompt carries its skill suggestion again and a recommendation that rotates mid-session reaches you.
- Dispatch rules say what they cover. The CLI packets now send only research and review through the shared runner and name the shape a single build dispatch takes.
- Release notes read like a story. The changelog workflow writes the narrative format this file follows and enforces the voice rules while it writes.
- Rules stop contradicting each other. The always-loaded guidance matches the current prompting guides and pauses before irreversible work only.
Spec Folders and Their Tools
Spec folders now behave like one system: created in the right place, numbered once, archived with their family and saved without collateral damage.
Old Specs Upgrade Without a Model
A person coming from v3.x holds spec folders that fail strict validation almost without exception, and every workflow stops at that failure. node .skilled/skills/system-spec-kit/runtime/cli/spec/upgrade-legacy.mjs fixes this in one pass and needs no language model. It works on packets in a top-level specs/ folder, where v4 reads them, and if yours still live in .opencode/specs it stops and prints the one move that fixes that. It is a dry run until you pass --apply. --roots <dir> narrows it to the folders you name and --include-archive brings archived packets in too.
For each active packet that fails, it fills in the frontmatter keys (the metadata block at the top of each document) a document lacks without rewriting any value already there. A document whose metadata block it cannot read is left as it is, and the run names it. It then runs the existing repair tools on the packet and writes every finding they cannot clear into the packet's upgrade-baseline.json.
The validator now reads that file. A finding listed there is a warning, so old findings stop blocking new work. A finding the file does not list is still an error, so a new mistake in an upgraded packet still fails. Findings about generated metadata are never listed for an active packet, because re-deriving the metadata clears them. Measured on snapshots of this repository's own v3 history, one run took two large trees from 2 of 170 and 0 of 995 active packets passing to 170 of 170 and 995 of 995. A second run changed nothing.
The same work tightens the backfill that regenerates a packet's derived metadata. It now exits with a failure when a folder cannot be re-derived, so the repair tool reports the failure instead of calling the packet repaired. The pre-commit hook therefore blocks a commit whose metadata cannot be re-derived. Set SPECKIT_SKIP_SPEC_REMINT=1 when you mean to bypass that check.
Track Roots List What They Hold
Fifteen of the eighteen track roots listed packets they did not hold or missed packets they did. Nothing wrote these lists and the one check compared counts and ran nowhere. A writer now sets each track root's list from disk and scaffolding adds a new packet to its track's list as it creates it. A push whose lists disagree with the folders beside them stops at the hook.
Archives Keep Packets at Home
Archiving used to send everything to one pile at the specs root. A track packet left its track, a restore dropped it at the root and same-numbered phases from different packets collided there. A packet now archives into the z_archive/ beside its own track and a phase into its parent's, and a restore returns each one where it came from. The track list refreshes after every move, so the push gate has nothing to block.
New Packets Land in the Right Place
New packets were briefly written to a tree nothing reads, and each new name started its numbering at 001. Scaffolding now writes to the canonical specs/ root and a packet made at the root takes the highest number already there. Making one no longer fetches or prunes remotes, so it never touches your network or your refs. It also works in a fresh checkout with no install, writing the same documents the full toolchain would and naming the missing build when a generated file cannot be written.
Saves Write What They Carry
A save used to ignore the continuity fields (the session bookkeeping a packet records) it was given and rewrite one file that every packet in the track shares, so concurrent sessions kept colliding and resume state went stale. The save now writes those fields into the packet that holds the work, and the track's pointer lives in the store that resume reads first. The reader also accepts hand-written continuity blocks as people actually write them, so a block you wrote by hand is read rather than rejected.
Level Upgrades Add Whole Sections
Upgrading a packet's level injected changed template lines along with the new sections, leaving a second title and broken tables behind. A level upgrade now adds whole sections only, and the upgraded packet passes strict validation straight away. The phase-map tool beside it now changes only the rows that disagree with their child and reports a completion mismatch instead of writing over it.
The Gate That Asks Once
The spec gate decides which folder a session's work belongs in. It grew quieter and more complete at once.
One Question at the First Save
The gate used to append its option menu to every turn that looked like a write, and one recent session carried 52 injections with most gate states still open. The question now arrives once per session at your first real file change, through whatever channel the runtime has: an interactive dialog on Pi, a tool-call notice elsewhere and a one-shot deferral where nothing better exists. The answer is remembered on disk, so a later write stays quiet. Answering also works for a folder you just named that does not exist yet.
Every Runtime Hears the Question
Moving the question left one runtime holding a gate that opened and never spoke, because the shared adapter had gone silent and the plugin's own mocks hid it. Hermes gets the same delivery the other runtimes have now. The pages that still described the old turn-time behavior were corrected against the code, including two claims that were wrong about when the question appears.
The Gate Covers Every Repository
The gate used to exempt anything under /tmp, which switched it off for a repository that itself lives there. A repository is now gated wherever it lives and scratch space outside it stays exempt. The CI workarounds that the old rule forced are gone, and the recorded probe scripts run under the shell standard the drift check expects.
Advisories Stop Crying Wolf
Two advisers warned about problems that were not there. The staging advice now stays quiet when a path is about to be expanded by the shell, and the completion sentinel no longer reads a cited document as a missing packet. The destructive checks still fire.
Models and Dispatch
Two m...
v4.0.0.0 — The Foundation Rebuilt: Spec Kit, Deep Loops and Parent Skills
v4 rebuilds the foundation of the framework. It gives an AI coding agent one path from a file change to a documented, validated result, one runtime for long-running research and review and one way to group related skills without hiding the work each mode does.
Spec Kit now owns the whole change record. Structured packets, continuity, lexical retrieval, workspace gates, derived metadata and completion checks live in one lifecycle. The specs moved to a physical top-level root, the retired memory database gave way to a committed trigger index and ripgrep and the completion gate learned to check scope and acceptance. A file change has a place to land, a record to follow it and a gate that can say whether it is finished.
Deep Loop now gives research, review, AI council and improvement one runtime. Each loop carries its state outside the chat, fans out to the executors you choose and records its passes in an append-only ledger with typed events, sealed artifacts and receipts. The runtime checks coverage and convergence before it permits a stop, so a finished-looking transcript is not proof that the work is finished.
Skills have a second shape now. A standalone skill still owns one job. A parent skill owns the route, reads the request and dispatches through mode-registry.json to a nested workflow or read-only surface. That gives code, documentation, design, MCP tooling and external CLI orchestration one identity each without flattening the behavior that belongs to their modes.
The rest of this release is the work that makes those three changes hold: new roots and compatibility paths, retrieval without the old database, stricter gates, safer dispatch, new executors and the runtime details that make a long session recoverable. The sections below explain the new shape and the breaking paths, defaults and implementation details that come with it.
Why This Release
Most of the framework's skills now present one identity instead of a scatter of separate skills. Seven families made the move: code, documentation, design, the deep loops, MCP tooling, external CLI orchestration and judgment transport. Prompt craft tried the parent shape during the cycle and returned to a standalone skill before release. sk-vision, sk-communication, sk-git, mcp-code-mode, the Spec Kit and the advisor stayed standalone.
A parent skill owns no workflow logic. It reads what you asked for and dispatches through a mode-registry.json to one of its nested modes, keyed by a workflowMode. A mode either does the work or supplies read-only evidence. The gain is practical:
- One place to maintain. One hub instead of near-duplicate homes.
- No slash-command bloat. A mode does not need its own command.
- Cleaner routing. Each domain presents one identity to the advisor.
- Easier to iterate. Change one mode without disturbing its neighbors.
sk-doc's create-skill-parent tooling stamps out each hub's router, modes, README and drift check the same verifiable way. The sections below explain what each family gained.
What's New at a Glance
- Spec Kit has one canonical core. Specs now live under
specs/, with level-gated core templates for specs, plans, tasks and implementation summaries. Old.opencode/specs/...paths still resolve through a compatibility symlink. - The memory database is retired. SQLite, embeddings and
memory_search/memory_saveare gone./speckit:searchuses a committed trigger index and ripgrep, returning a clean no-hit when the phrase is absent. - The Skill Advisor is CLI-only. Its MCP transport is gone.
node .skilled/bin/skill-advisor.cjsis the front door, and local scoring marks itself degraded when the daemon is unavailable. - Deep loops share one runtime. Research, review, AI council and improvement fan out across chosen executors. Typed ledger evidence and convergence checks control the stop.
- Cross-CLI dispatch has one parent.
cli-external-orchestrationroutes seven bridges. Cursor and Hermes join OpenCode, Claude Code, Codex, Devin and Pi as deep-loop executors. Each bridge checks its binary before dispatch. - Repo Rules are routed.
REPO RULES.mdloads scoped rules for evidence, blast radius, communication and hub routing before a first write. - Design has one hub.
sk-designroutes fundamentals, measuredDESIGN.mdextraction, charts and diagrams to the mode that owns the job. - Goals and hooks travel with the session. The goal plugin binds
goal.mdto OpenCode, Cursor, Pi and Devin. One switchable hook library carries shared guards into Hermes through one plugin. - Git stops before costly mistakes. Pushes, mass deletion and misnamed branches hit explicit tripwires. Parallel-session commits still reach the IDE checkout.
Spec Kit
The spec kit changed the most under the surface. The memory database that powered memory_search was retired outright, the specs folder moved, the runtime was renamed and nested, and the completion gate was made coherent. Three research rounds then cut the kit back to what a machine reads.
Most of the depth is in what left: an entire retrieval engine replaced by two small lexical tools.
Specs Move to the Top Level
Your spec paths have a new root. The specs folder moved from .opencode/specs/ to a physical top-level specs/ directory, so canonical spec paths are now specs/.... This is a breaking change, but the old location still resolves through a compatibility symlink, so anything still pointing at .opencode/specs keeps working while you catch up.
The Memory Database Retired
The Spec-Kit Memory engine is gone. The SQLite database, the embedder, the spec-memory MCP server, its daemon and the memory_search and memory_save tools were decommissioned end to end, first on a side branch and then landed on the release branch and main. The spec-memory daemon CLI under .opencode/bin/ and the system-spec-memory OpenCode plugin went down with it, since both were doors into the same engine.
What replaced them is deliberately small:
- A committed trigger index is generated from every document's
trigger_phrasesfrontmatter. - A lookup script reads it with no daemon.
- A set of ripgrep recipes handles free text.
All of it sits behind /speckit:search, and the kit's own context writer handles continuity. Retrieval is lexical now, and a miss is a clean no-hit rather than a degraded guess. The debt the decommission left, from dangling registrations to a package that still called itself an MCP server, was closed across six review passes.
The Runtime Renamed and Nested
The surviving package no longer carries an identity it lost. Its engine lives at .skilled/skills/system-spec-kit/runtime/cli/:
spec/validate.sh,spec/create.sh,spec/repair-derived.cjsandspec/recommend-level.shfor the packet lifecyclecontinuity/for savesspec-folder/for generated metadataretrieval/for the trigger index
The old scripts/ and mcp-server/ paths are gone, so anything you pinned to them needs repointing. A shared package underneath carries the Gate-3 classifier and the frontmatter parser that every other skill now imports rather than copies.
A Completion Gate That Tells the Truth
The gate that decides whether a packet may close now returns the same verdict whatever the environment, counts one fault once, and no longer asks a packet for what it cannot satisfy from inside itself. Forty rules are registered, and a warning is advice that does not fail a run.
It grew two checks of its own. Scope adherence refuses a packet whose changed files fall outside what its spec names, driven by SYSTEM_SCOPE_CHANGED_FILES and SYSTEM_SCOPE_BASE. Acceptance coverage runs on by default as an advisory.
Two documents joined the packet contract. acceptance-criteria.md decides closure at Levels 2 and 3, and an optional goal.md holds the durable directive you set as the session objective. A repair tool fixes the facts a packet records wrongly but a machine can recompute, such as its folder, level and generated fingerprints. It refuses to invent the facts only a person can write.
Smaller Templates, Same Output
The spec, plan, tasks and implementation-summary templates were consolidated into one shared core with level-gated addenda. The source dropped to 1,275 lines across four core templates without changing a single rendered byte, and the research template shrinks by level too.
A Level 1 spec now gets a short research doc instead of the full one, and a Level 1 research doc renders at 175 lines instead of 944. You see smaller, level-appropriate templates. What they produce is identical.
The Deep Loops, Unified and Extended
The deep loops finished collapsing into a single home and learned to run on any model you name, several at a time. The hub and its backend are one skill. The loops fan out across every CLI. Underneath, a new evidence-ledger runtime landed dark, proved parity and then, under an operator-ratified flip, became the authoritative record for every mode.
One Skill, Hub and Backend Together
The workflow hub and the runtime it ran on are one skill, system-deep-loop. Every downstream reference (commands, agents, READMEs, hooks, advisor routing) was repointed at the new home, and the modes you already use behave as before: research, review, ai-council, agent-improvement and the model benchmark. What changed is the address.
What left the tree along the way:
- The
deep-loop-workflowsanddeep-loop-runtimeskill identities no longer exist, so anything you script that still names one of them now points at a signpost that is not there. - The skill-benchmark lane was retired late in the cycle.
- The deep router agent was renamed
deeptodeep-loop, then retired with its routing folded into the/deep:*command bodies, so there is no router agent left to call.
An alignment mode was ...
v4.0.0.0 — Beta 1
First beta of the v4 development line, promoted from skilled/v4.0.0.0 to main as a fast-forward (2,711 commits ahead of the prior v3.6.0.0 release).
This is a pre-release for validation, not a stable cut.
Most recent work landed
- Injection-bloat reduction (
hooks/002) — candidate-004 route-only advisor delivery activated for Claude Code / Codex / Devin / OpenCode (~43 B on a proven same-epoch repeat vs ~806 B full, evidence-gated, fail-open; kill-switchSPECKIT_ROUTE_ONLY_ADVISOR_DISABLED=1). Plus Pi directive de-duplication (incl. the headless fallback), OpenCode directive single-source, a 006 compact-dispatch evaluation, and a measurement + rollback harness. Driven by a cli-pi/DeepSeek deep-research investigation. - The broader v4 line: spec-kit memory, skill-advisor, deep-loop, cli-external-orchestration, and the sk-* documentation/code/design/git hubs.
Beta caveats
- The advisor
mcp-servervitest suite has known environment-dependent failures in bare checkouts (launcher/daemon/scorer/vocabulary/parity/corpus) — run in a provisioned checkout for a clean gate. - Pi/Cursor route-only activation and the 006 compact dispatch directive are evaluated but not activated.
v3.6.0.0, Harder Core, Longer Reach
The center of this release is the part you rely on without thinking about it. Memory and search got smarter about telling you the truth, and the background services that hold them up got tougher under stress. A genuine match is cited again, search keeps working when its index is damaged, and the daemons (the long-running background services behind memory and search) now recover instead of duplicating themselves or corrupting their own data.
Around that core, two more things moved. The five separate deep-loop skills folded into one hub running on a shared backend, and a new set of terminal-driven design tools arrived. Both changes are real, but neither asks anything new of you.
None of this changes how you actually call the system. memory_search, memory_save, the /deep:* commands and the agent names all behave exactly as before. The work happened in the failure paths and the corners that never fit the happy case, not on the path you use every day.
What's New at a Glance
| Honest, citable answers | Search cites the matches it finds. A 0.89 match reads as a good match, cites its sources and reports real confidence. |
| Search survives a broken index | A corrupted slice of the semantic index is detected, quarantined and rebuilt, and the repair survives a restart. A memory-safe keyword engine covers the gap so search keeps answering. |
| Memory protects itself | A scrubber redacts secrets before anything is indexed, and a provenance guard keeps automated writers from overwriting fields a human edited. |
| Daemons that recover | The three background services reconnect to a live daemon instead of starting a second one, take a single-writer lock that keeps two of them from corrupting the same database, and each expose a scriptable command-line front door over its warm socket. |
| One hub for the deep loops | The five deep skills (context, research, review, ai-council, improvement) run as a single hub over a shared backend. Every /deep:* command and agent name is unchanged. |
| Design at the terminal | Three new skills move interface work into the terminal: the design-judgment skill, plus transports for the Open Design app and Figma. |
| And more | Kimi K2.7 joins as a small-model option, the design and MCP skills share one install-and-doctor setup, and the operating-discipline rules are now written into the framework docs. |
Search
Search picked up a handful of fixes this release, and most of them share a theme: a real match should be findable and the score on it should be honest.
Honest, Citable Answers
Search cites the matches it finds. A genuine 0.89 match reads as a good match, cites its results and reports a real confidence figure. A scoring mismatch had been holding good verdicts back, since the confidence step compared fusion scores (which cluster low near 0.03 because RRF, a method that blends two ranked result lists, compresses them) against thresholds tuned for cosine similarity at 0.7 and 0.4. A resolver now reads the right score for confidence while leaving the ranking order alone.
Memory-Safe Keyword Fallback
When the semantic path is unavailable, search falls back to keyword matching, and that fallback now stays light on memory. It packs its data tightly and weights title and trigger matches above body text, so the right entry ranks first. The payoff is concrete: a realistic warmup that once spiked to 687MB now holds at 137MB, well under budget, with the ranking unchanged.
Scoped Queries Return Their Full Set
A scoped search gives you everything that actually matches. The keyword engine gathers the full candidate set first, applies your filters and only then trims to your limit, so a narrow query comes back complete rather than short.
Scope Isolation
A scoped search stays inside its scope and cannot pull in results from other scopes. The broaden-and-rescue path (the lexical backfill and sibling injection that widens a thin result set) had re-queried the index without re-applying the spec-folder scope filter, so a scoped search could surface rows from outside its scope. Every injected row now re-applies the scope filter, backed by a regression test. This is separate from the full-set fix above: that one keeps a scoped query from coming back short within its scope, this one keeps it from leaking across scopes.
Safe Query Handling
A /memory:search query that contains shell-special characters now resolves exactly as typed, since the query argument is quoted, so a search for something with punctuation behaves predictably.
Cold and Archived Results
Older and archived entries stay within reach. Results from cold and deprecated tiers flow back into the keyword and trigger channels under a freshness-aware ranking from FSRS (a spaced-repetition scheduler that scores how relevant something still is), and the semantic lane fills in missing projections so cold entries surface when they match.
One Score Everywhere
The same match reads the same way on every surface. Rendered output reports a single score name on a plain 0-to-1 scale to two decimals, and the older confidence values and percentages are gone from what you see, so two callers reading the same result see one number.
Ranking Explanation Trace
When a result ranks where it does and you want to know why, you can now ask. An opt-in trace returns a breakdown of how the ranker scored a result, pulled from its own intermediate values, and an inline warning flags when one memory contradicts or supersedes another.
Memory Safety
A retrieval system you cannot trust to be honest, or to leave your edits alone, is worse than a slow one. Four protections close that gap, and the structural ones are on for everyone.
Secret Scrubber
Credentials should never reach the index, even by accident. A scrubber now sits at the very front of the pipeline, ahead of hashing, embedding and keyword indexing, and redacts API keys, tokens, bearer values and private-key blocks into typed redaction markers. It fails closed: if scrubbing cannot finish, the write is refused rather than stored in the clear.
Provenance Guard
When you edit a memory by hand, an automated process should not be able to quietly overwrite it. A new provenance column records who wrote each field, human or automated, and the write guard uses it to skip automated overwrites of human-authored content before the change lands. A companion tagger marks the automated writers, so the rule holds everywhere rather than at one checkpoint.
Safer Retention
The cleanup that ages out old entries checks an entry's tier and pinned state before deleting anything. A constitutional or critical entry that hits its expiry earns a logged refusal instead of being removed, so the cleanup stays within what it should touch.
Continuity Across Compaction
A long session resumes from fresh state instead of stale notes. An opt-in snapshot (the flag SPECKIT_AUTHORED_CONTINUITY_SNAPSHOT, off by default) refreshes the handover and continuity notes right before a long session is compacted, and startup now reports what it restored. It stays off until you turn it on.
Vector Self-Heal
The system detects, quarantines and rebuilds a damaged vector shard (a slice of the semantic search index), and the repair survives a restart. A persisted marker and a real completeness check decide between resuming the repair and clearing a stale flag, so a corrupted slice gets rebuilt rather than silently replaced with an empty one, even across a restart.
Resilient Daemons
A daemon is a background service that stays running and answers requests over a socket. The 026 release gave them a second way in. This release makes sure that when a connection drops, a daemon dies or a burst of traffic hits a limit, the system recovers instead of duplicating itself or corrupting its data. These protections are on for everyone, and the advisor parity takes effect on your next fresh session.
Advisor Launcher Parity
The skill-advisor launcher now recovers the way the memory and code services already do. A new session bridges the live daemon instead of spawning a rival, and a hung daemon is reaped and replaced under a lock. Until this release it could decide to reconnect on a dead socket and then act as if it had not, so a second session could start a second writer. That gap is closed.
Single-Writer Lock
A single-writer lock keeps two processes from writing the same database at once. The second writer loses the lock and exits cleanly, so the shared index stays intact.
Scriptable Command-Line Front Doors
All three daemons now expose their full tool set as scriptable command-line clients over the warm socket: 37 tools for spec-memory, 8 for code-index and 9 for skill-advisor. The normal connection stays primary and these are additive. They are read-only at prompt time, maintenance and mutation are blocked there, and advisor changes require an explicit trusted flag.
Higher Connection Cap
Busy fan-outs now have room: the limit on simultaneous connections to a daemon rose from eight to 64, so a burst of parallel work no longer runs into the ceiling.
Orphaned Helper Cleanup
Helper processes clean up after themselves. They exit when their input closes, when the connection drops or after a short watchdog timeout, so they no longer pile up once their parent is gone. Sixteen stranded processes were cleared in the work that landed the fix.
Config Alignment
Daemon re-election is now on by default in the launcher itself rather than only when a config turns it on, and the four runtime configs were cleaned and brought into line, including a fix for two of them that carried a JSON error that left them unparseable.
Advisor State Stays Out o...
v3.5.0.4 — The Deferred Lifecycle, Closed
v3.5.0.2 fixed the inert session-cleanup hook but said plainly that it was not the whole story: the owner/daemon lifecycle root causes were deferred. This release closes them. A converged two-model investigation (Opus 4.8 + gpt-5.5) found the still-active flap, and it was not a misclassification as first reported. The launcher's own shutdown handler explicitly kills the daemon child. When the owning session disposed, the supervisor scheduled a relaunch on a short backoff, the fresh daemon came up under a runtime that was already leaving, and it was killed again about a second later. Every session bridged to that shared daemon lost its transport.
Phase 017 stopped the dominant flap. The five hardening items v3.5.0.2 listed as deferred then shipped as 018-022, each flag-gated and default-safe where it changes process or daemon lifecycle, and each tested. Three independent models (a Claude subagent, gpt-5.5-fast via cli-opencode, and gpt-5.5 via cli-codex) cross-validated the result and caught two real bugs the first pass missed. Finally the operator-facing docs were brought into alignment with all of it.
What's New at a Glance
| Disposal flap guard | The relaunch timer re-checks at fire-time: if the launcher is shutting down or its owning runtime has gone away (ppid changed, or reparented to 1), it releases the lease and exits instead of respawning the daemon under a dying session. |
| Persistent launcher log | log() now also appends a bounded, best-effort durable line so a flap or disposal race is attributable from disk, not just from stderr the host may drop. |
| Reap hardening | A sibling reaps the lease owner and respawns only after N consecutive deep-probe failures, so a busy-but-alive owner mid-FTS-merge is no longer false-reaped into a duplicate daemon. |
| Code-index reconnect | mk-code-index now fronts its daemon with the same reconnecting session proxy as mk-spec-memory, so an owner change reattaches and replays read queries instead of a hard Connection closed. |
| Orphan-sweep activation | The Stop hook can now reap ownerless MCP daemons via the orphan-only sweeper when no session pid is available, flag-gated and default-off, never guessing the session pid. |
| Daemon re-election, on by default | The shared daemon now outlives its owning session: a disposing owner releases it for a live secondary to adopt instead of killing it. On by default in the runtime configs (SPECKIT_DAEMON_REELECTION), set 0 to revert. |
| Release test and loose ends | A hermetic integration test proves flag-on releases the detached daemon and flag-off kills it. The orphan-sweep LaunchAgent template now exists (dry-run by default) and cli sub-session session scoping is documented. |
| Live validation, and a fix it caught | A two-session durability test runs two real launchers in an isolated root to prove adoption end to end. It caught a fresh-session double-writer, where a cold start after disposal spawned a second daemon on the same database, now fixed by reaping the released daemon before respawn. |
| The audit closed the rest | An eleven-lineage multi-model review serialized the stale-reclaim reap under the respawn lock, then replaced it with true adoption: a fresh session bridges a live released daemon instead of reaping it, and reaps only a dead one. The release path also SIGKILL-escalates a wedged sidecar, and a tracked local settings file that leaked personal permissions was untracked. |
The Flap Was the Launcher, Not a Misclassification
What went wrong
The earlier investigation framed the relaunch as a clean-exit misclassification. Verifying against the code corrected that: the launcher already guards its own shutdown flag, and shutdownLauncherForSignal explicitly sends SIGTERM to the context-server child. The real gap was timing. The child-exit supervisor's 250 ms relaunch fired before the launcher's own shutdown signal landed, so a fresh daemon was spawned under a disposing runtime and immediately killed again.
The fix
The relaunch is now gated at the moment its backoff fires. The launcher captures its parent pid at startup. When the timer fires it re-checks reality, and if it is shutting down or has been orphaned it releases the lease and exits cleanly rather than respawning. Crash-recovery and RSS-recycle are untouched: both run with the owner alive and a matching ppid, so the gate is a no-op for them. The change is additive and was extracted into a pure, unit-tested predicate.
The Deferred Five
Observability first (018)
Before anything else, the launcher learned to leave a durable trace. log() keeps its stderr write and additionally appends a timestamped, pid-stamped line to a size-bounded file that rotates to a single previous generation at the cap. It is best-effort by design: a logging failure can never affect the launcher. Default-on, disablable with SPECKIT_LAUNCHER_LOG=0, path- and size-overridable.
Stop reaping the merely-slow (019)
The lease-holder check used to declare an owner dead on a single probe miss and respawn, which for a daemon momentarily blocked mid-merge meant a duplicate. It now requires N consecutive deep-probe failures (SPECKIT_LEASE_PROBE_RETRIES, default 1), with any single live probe short-circuiting to a bridge. The default budget stays inside the probe grace ceiling, and a genuinely dead socket still fails fast.
Give code-index a reconnect (020)
mk-code-index had the worst failure mode: a raw bridge with no reconnect, so an owner death surfaced as a hard Connection closed. It now uses the same reconnecting proxy as mk-spec-memory. The proxy's replay decision was generalized into a factory so each server passes its own replayable tool set. Code-graph read tools replay, while code_graph_scan, code_graph_apply and code_graph_verify (which persists a baseline) are never replayed.
Let the orphans be swept (021)
The Stop hook's session-cleanup no-ops without a session pid, and it deliberately refuses to guess one. When enabled, its no-session-pid branch now delegates to the orphan-only sweeper, which reaps only ownerless processes and so can never touch a live session. It ships default-off with a dry-run ramp.
A foundation for surviving disposal (022)
The complete fix, keeping the shared daemon alive across its owner's exit, landed as a flag-gated, default-off foundation. When enabled, the owner spawns the daemon detached and, on shutdown, releases it for a live secondary to adopt instead of killing it. Default-off is byte-identical to prior behavior. Secondary ownership adoption and the released daemon's terminal death are runtime-validation-gated. The flag stays off until that pass.
Three Models, Two Real Bugs
The implementation was cross-validated by three independent reviewers running the launcher suite and adversarially reading the diffs. They agreed the suite was green, and the harder lenses earned their keep. A gpt-5.5 review of the 022 diff caught a blocking bug: the process exit handler would have wiped the daemon lease the release path deliberately preserves, leaving a released daemon alive but unfindable. A later three-way audit found that code_graph_verify was wrongly in the code-index replayable set (it mutates when persisting a baseline) and that the launcher log skipped rotation for a custom path without a .log suffix. All three were fixed and re-verified.
The Docs Caught Up
A gpt-5.5 audit found the five features documented nowhere operator-facing. This release adds eight ENV_REFERENCE.md flag rows, five feature-catalog entries, five manual-testing-playbook scenarios and lifecycle rows in the relevant READMEs and the skill doc, with the playbook's file-count self-check and the catalog count reconciled and the repo-wide link check clean.
The Hardening Continued
The post-audit cross-check left two follow-ups, and both shipped next. The code-index launcher used to load env, force the maintainer index flags and write stderr at require time, because the entrypoint guard sat below those statements. Importing it now runs no side effects. Its env work moved into a bootstrap function called only when the file is the process entrypoint, and a new test asserts a require leaves the environment untouched. The comment-hygiene checker also grew the patterns it was missing. It now catches RC, single-number DR, hyphen phase and council-seat labels, and it scans inline trailing comments rather than only full-line ones. A blanket F-number pattern stayed out by design because almost all of its matches are function keys and figures.
With the checker stricter, the perishable labels it had let through were scrubbed. The daemon-reliability files cleared first, then a repo-wide sweep cleared the live skill, bin and plugin code, rewriting each comment to keep the reason and drop the rotting identifier. Archived specs, scratch directories and the pattern-defining tools were left alone, and git pre-commit wiring waits until that sweep settles so it does not block other sessions.
Re-election, Now On By Default
The foundation is now the default. The committed runtime configs set SPECKIT_DAEMON_REELECTION=1 for mk-spec-memory across all three runtimes, so a disposing session releases the shared daemon for another live session to adopt instead of killing it, and concurrent sessions keep their MCP transport. The launcher's code default stays off, so the configs are the single on-switch and 0 is a one-character revert.
Turning it on for everyone is safe because the downside is bounded. If a live secondary is connected when the owner disposes, the released daemon survives and that session keeps its transport. If no secondary is connected, the next fresh session reaps the released daemon before spawning a replacement, so th...
v3.5.0.3 — One Voice for Every Skill README
The skill READMEs had drifted apart. Each one was a tabular reference card with functional headers, a buried statistics block and no human entry point, and they ranged from 228 to 1084 lines, so a reader moving between skills re-learned the layout every time. New skills inherited the old shape because the sk-doc template still encoded it. This release rewrites the skill READMEs into one narrative voice, the same warm, problem-first voice the repo root README and these changelogs use, and updates the template so new skills start there.
The rewrites were grounded, not invented. Each README was rebuilt from the skill's real files: a two-iteration deep-context gather read the SKILL.md, the code and the references, then DeepSeek v4 Pro and MiMo v2.5 Pro each drafted the README and the orchestrator merged the result and verified it against source. That gather caught real drift along the way, from a wrong MCP integration recipe to a stale embedder claim, and every fix landed in the rewrite.
What's New at a Glance
| One Narrative Voice | The 22 skill READMEs now open with a one-line pitch, an at-a-glance table and a problem-first overview, then keep their reference detail below. |
| Template Updated | The sk-doc skill-README template now encodes the narrative skeleton plus the writing and validation rules, so new skills are authored in the voice from the start. |
| system-spec-kit Kept Its Depth | The 1084-line reference manual was restyled in place rather than compressed. The top was reframed and the house voice swept through while every reference block stayed. |
| Grounded in Source | Each README was gathered and verified against the skill's real files with a two-model deep-context pass, so the documented facts match the code. |
The Problem Was Consistency
What the READMEs looked like
The repo root README and these changelogs read in one voice: a problem-first opener, an at-a-glance table near the top, outcome-named section headers, sparse tables and a verification close. The skill READMEs did not. They led with an Overview or a feature inventory, hid their key facts in a Key Statistics block and named sections for their function rather than the reader's goal. Length drifted from 228 lines to 1084. Moving between two skills meant re-learning where everything lived.
The narrative standard
The standard now is the same for every skill. The README opens with a frontmatter block, an H1 and a one-line pitch. Section 1 is an At a Glance table a reader scans in five seconds. Section 2 is an Overview that states the problem before the solution. Quick Start, How It Works, Integration, Troubleshooting, FAQ and Related Documents follow as the skill needs them. Prose carries the explanation and tables appear only for genuine lookups. The voice rules hold throughout: no em dashes, no double-hyphen separators, no semicolons, no Oxford-comma lists and no filler words.
How Each README Was Rebuilt
Gather, draft, merge
Every skill ran the same recipe. A deep-context loop gathered the facts across two read-only iterations, building a verified map of the skill's tools, modes, key files and boundaries. Two models then drafted the README from that map against the locked template and the golden example. The orchestrator merged the stronger draft, grafted the best parts of the other, applied the voice rules, confirmed every cited path resolved and validated the structure before publishing. Nothing shipped on a model's word alone.
Depth was preserved where it mattered
The system-spec-kit README is the largest in the repo and carries dense reference content: the 37 memory tools, the five-channel search pipeline, the index schema history and the environment table. Regenerating that body would have risked dropping or paraphrasing facts, so the restyle reframed the top and swept the house voice through the existing text in place, keeping every reference block. The result reads in the narrative voice and stays a complete reference manual, moving from 1084 to 1067 lines.
Drift Caught Along the Way
A voice rewrite is also a fact audit. Because each README was checked against source, the gather surfaced stale claims that had outlived the code and corrected them in the same pass. An MCP integration recipe pointed at the wrong tool namespace and config key. A code skill listed a surface that no longer existed. A scorer README described a single shipped embedder when the shared registry now holds several. Version footers, tool counts and script counts had fallen behind their source. Each correction shipped with the README it belonged to, so the new voice and the current facts arrived together.
Verification Highlights
| Check | Result |
|---|---|
| Structure | validate_document.py --type readme reports 0 issues on every rewritten README |
| Voice | HVR prose scan clean across the batch: no em dashes, double-hyphen separators, prose semicolons or Oxford-comma lists |
| Accuracy | Tool counts, lane weights, documentation levels and cited paths verified against source. Every linked path resolves |
| Packet | validate.sh --strict PASS on each phase folder and on the phase parent |
Upgrade
No API changes. This release is documentation only. No SKILL.md, code or behavior changed, so nothing needs a rebuild or a restart.
New skills inherit the voice. Author a new skill README from the updated sk-doc template and it lands in the narrative shape by default.
The skills index is the closing step. The .opencode/skills/README.md catalog is the final piece of the standardization and ships as the last phase of the packet.
Full Details
The packet: .opencode/specs/skilled-agent-orchestration/135-skill-readme-standardization/
v3.5.0.2 — Sessions Keep Their Hands to Themselves
This patch closes a long-open case: why MCP transports kept dying mid-session even after the front-proxy work. The answer was not in the daemon at all. The session-cleanup hook, when it could not determine which session it belonged to, guessed — and under a shared terminal that guess resolved to an ancestor common to every open session, so one session ending would sweep away its siblings' MCP launchers by name. Dead launcher, dirty daemon exit, stale socket, refused connection: the whole familiar failure, from one bad fallback.
The probe work that followed found the damage was not hypothetical. The live memory index was already structurally corrupted from the earlier dirty shutdowns. It was salvaged the same evening at 9,888 of 9,890 rows.
What's New at a Glance
| Scoped Session Cleanup | The Stop-hook cleanup requires explicit session identity and re-proves each candidate's ancestry immediately before the kill. No identity, no action. |
| Post-Crash Integrity Gate | A boot after a dirty shutdown runs a whole-database quick_check. Corruption means a needs-rebuild sentinel and a refusal to serve — never a silently malformed index. |
| Index Salvaged | The live context-index.sqlite had real b-tree damage and no checkpoint. sqlite3 .recover brought back 9,888 of 9,890 rows; the original is preserved beside it. |
One Bad Fallback
What went wrong
session-cleanup.sh is supposed to reap only the MCP helpers belonging to the session that just ended. Its identity came from CLAUDE_SESSION_PID — or, when that was missing, the hook's parent pid. That fallback was the bug. With several sessions open under one terminal, the parent pid resolves to an ancestor they all share, and "walk my descendants" quietly becomes "walk everyone's descendants." The name-glob kill then did exactly what it was told.
The fix is deliberate paranoia
Missing identity now means the cleanup does nothing at all. And ancestry is no longer trusted from a snapshot: immediately before each kill, the script walks the candidate's live parent chain and only proceeds if it reaches this session's pid — a sibling session's process can never satisfy that. Every kill and every skip is logged with matched_by= and ancestor_ok= fields, so if a transport ever vanishes again the log answers who and why.
Proven with drills
Three drills locked the behavior in: no identity is a no-op, a foreign session's launcher survives, and the session's own helpers still die on schedule.
Refusing to Serve a Broken Database
The gate
The daemon already marked dirty shutdowns with an .unclean-shutdown marker; what it did with that knowledge was limited to a detect-only check of the FTS shadow index. Now a marker present at open also triggers a whole-database PRAGMA quick_check. On failure the store writes the checkpoint .needs-rebuild sentinel — the same crash-safe machinery checkpoints already use — and refuses to start, because failing loudly beats serving corrupted rows to every connected session. Clean shutdowns skip the probe entirely, so the common case pays nothing.
The salvage
The gate proved its worth immediately: pointed at the live index, it reported genuine structural corruption — invalid page numbers, doubly-referenced pages — the residue of the dirty shutdowns the cleanup bug had been causing. With no checkpoint to restore, sqlite3 .recover rebuilt a clean candidate that passed quick_check, kept 9,888 of 9,890 memory rows, and parked 368 orphaned rows in lost_and_found for triage. The swap was reversible by design: the corrupted original remains on disk beside the recovered file.
Verified, Not Re-Implemented
Two hardening items the forensics initially flagged — recording the daemon child pid in the launcher lease, and reclaiming a stale bootstrap lockdir — turned out to have shipped already in earlier daemon-reliability phases. They were verified in source and recorded as no-ops. The remaining root causes keep their existing owners: the bridge liveness probe, the provider dispose, and the exit watchdog are planned phases that this release deliberately did not touch.
Verification Highlights
| Check | Result |
|---|---|
| Cleanup scoping | 3 drills: no-identity no-op, foreign-session isolation, own-session kill with ancestry proof |
| Integrity gate | Corrupted scratch DB: FATAL + sentinel + refusal. Clean DB: normal boot. tsc and dist build clean |
| Live index | Post-salvage quick_check: ok; 9,888/9,890 rows; standalone re-index of saved docs succeeded against the recovered file |
| Honest gaps | vitest's runner would not start in this sandbox (environmental); the suite rerun is owed to the next dev session. An embedding-reconcile pass over the recovered index is the recorded follow-up |
Upgrade
No API changes. The cleanup fix is active on the next session stop; the integrity gate activates on the next daemon start (dist is rebuilt — recycle if a daemon is already running).
A refused boot is the gate working. If a future start stops with a needs-rebuild sentinel, the checkpoint rebuild path takes it from there.
/doctor memory remains the fastest health read after any daemon lifecycle event.
Full Details
The fix packet: .opencode/specs/system-spec-kit/026-graph-and-context-optimization/007-mcp-daemon-reliability/016-cross-session-kill-scoping/
v3.5.0.1 — Right Hooks, Honest Reverts
A small patch release about integration surfaces telling the truth. Codex sessions had been running Claude's hook scripts for every event; they now run their own adapters, and an event Codex never fires was unregistered rather than rewired. Two documentation lies were corrected along the way.
The release also contains a fix that isn't here: a repair to the broken OpenCode code-graph plugin bridge was applied, smoke-verified, and then deliberately reverted after an independent review showed the working bridge was more dangerous than the broken one.
What's New at a Glance
| Codex Hooks Corrected | Codex SessionStart and UserPromptSubmit now run the Codex-native adapters; the unsupported PreCompact registration was removed. |
| Doc Corrections | The Codex MCP config's code-graph database note had its default and legacy paths backwards; the Gemini hook catalog confused the shim for the implementation. Both fixed. |
| Honest Non-Fix | A code-graph plugin bridge repair was attempted, smoke-verified, and deliberately reverted — the runnable bridge was a second direct writer on the memory database. |
Codex Runs Its Own Hooks
The rewiring
The live .codex/hooks.json had drifted: it registered Claude's hook scripts for every event, so Codex sessions were primed through the wrong adapters. SessionStart and UserPromptSubmit now point at the Codex-native hooks, which read the Codex stdin envelope and answer in kind — both smoke-verified. The PreCompact registration was removed outright rather than rewired: the hook contract lists Codex compaction as unsupported, so the entry was dead weight pointing at a Claude-envelope script.
Two paper cuts
The Codex MCP config note claimed the code-graph database default was the legacy shared path and the skill-local path was legacy — it is exactly the other way around, and the note now matches the launcher source. And the skill-advisor's Gemini hook catalog pointed at the spec-kit shim as the implementation; it now names the active implementation and the shim separately.
The Bridge Fix That Was Correctly Un-Shipped
Working, racy, reverted
The OpenCode code-graph plugin bridge has been inert since a skill extraction moved three of its imports — it is why sessions report "Code Graph: unavailable." Re-pointing the imports made it run again and emit the right transport payload. A fresh-model review then flagged what the smoke test could not: the runnable bridge initializes the memory database directly in its own process, outside the daemon's single-writer lease — the same corruption class the daemon architecture exists to prevent.
The fix was reverted. Broken-and-inert is safer than working-and-racy; the bridge stays down until an IPC-backed transport replaces the direct imports properly.
The sweep that wasn't
A planned sweep of orphaned launcher processes ended the same honest way: parent-process classification showed all nine running launchers belonged to live sessions, so nothing was killed. The true orphan class — launchers that outlive an exited owner — is separately scoped daemon-reliability work.
Verification Highlights
| Check | Result |
|---|---|
| Codex hook rewiring | Both adapters smoke-tested with sample envelopes; hooks.json parses; PreCompact removed |
| Bridge revert | Working tree confirmed back to the inert state; the fix packet documents the dual-writer reasoning |
| Fresh-model review | Independent gpt-5.5 review of the uncommitted tree; its P0 and both P1s remediated before this release |
| Secret scan | Reviewer's focused scan found no credentials in the committed set |
Upgrade
No code changes ship in this release. The fixes are configuration and documentation.
The Codex hook correction takes effect on the next Codex session start. No action needed beyond starting a new session.
The code-graph OpenCode plugin remains unavailable by deliberate choice until an IPC-backed bridge transport lands.
Full Details
The fix packet: .opencode/specs/system-spec-kit/026-graph-and-context-optimization/008-runtime-defect-fixes/
v3.5.0.0 — Memory That Maintains Itself
The previous release was about honesty and hardening. This one is about the runtime getting out of its own way. Embeddings run locally first, on your own machine through Ollama or a Python-free local server, and only reach the cloud when you ask. The MCP servers are now three separate processes, and the memory one runs as a single shared daemon behind a front-proxy that hides a recycle so a restart stops surfacing as an error. The deep-loop engine behind deep-research and deep-review learned to fan out across several executors at once. And the prompt toolkit became forkable, with a small-model hub that drives open models through cli-opencode.
Underneath, memory still gets smarter on its own after every save, the index keeps itself reconciled, and the long-running 026 program documented every one of its roughly 634 phases into one canonical changelog tree before closing itself out. A cleanup pass also removed the vendored research repos that never belonged in the shipping tree, about 2.1 million lines of code that was never ours to ship.
What's New at a Glance
| Local-First Embeddings | Embeddings run on your machine through Ollama or a Python-free local server. Cloud is reached only by explicit choice. |
| Separate MCP Servers | Memory, code-graph and skill-advisor are three separate MCP servers, each with its own launcher. |
| One Shared Daemon | The memory daemon is a single shared background process. A front-proxy hides a recycle so a restart is invisible to clients. |
| Smarter by Default | Saves return immediately while enrichment and causal relation inference run in the background. |
| Checkpoints | Real database files now, with a crash-safe restore. |
| Deep-Loop Fan-Out | deep-research and deep-review can run several executor lineages from one loop, with clean recovery from a flaky one. |
| Deep Improvement Skill | A skill-benchmark lane plus a reusable, config-driven model-benchmark framework. |
| Forkable Prompt Toolkit | sk-prompt lifts out cleanly. sk-prompt-small-model drives open models through cli-opencode. |
Local-First Embeddings
The embedding stack runs locally first and only reaches the cloud when you ask it to. Resolution is your explicit EMBEDDINGS_PROVIDER when set, then a persisted Ollama embedder, then a Python-free local fallback server, and cloud providers like OpenAI or Voyage are never auto-selected. Running on a single small nomic model on your own machine, the common case never leaves the box.
This release closed the last place that rule was not honored. The primary resolver used to pick a cloud provider when Ollama was down and an API key happened to be set, the exact opposite of the intended priority. It now follows the local-first order everywhere, and the legacy cloud-preference tests were rewritten to assert it. The local foundation underneath, the Python-free local server and the consolidation onto a single nomic model, matured in the prior release. This release made the resolver match the design.
Separate MCP Servers, One Shared Daemon
A large part of this release was spent in the launcher and daemon lifecycle, because that surface was behind most of the prior release's friction and a hard-to-pin disconnect.
Three separate servers
Memory, code-graph and skill-advisor run as three separate MCP servers, each behind its own launcher (mk-spec-memory, mk-code-index, mk-skill-advisor) and sharing a common library. They start, recycle and fail independently of each other.
Many sessions, one daemon
Multiple sessions now share a single mk-spec-memory daemon instead of each spawning its own and contending on the database. A lease-holder launcher owns the daemon child, and other launchers bridge their own client to it through a session proxy. The owner stores its real IPC socket path in the lease, so a secondary launcher pointed at a different socket directory still finds the live daemon instead of failing with no bridge socket.
A front-proxy that hides a recycle
A front-proxy sits between clients and the daemon. When the daemon recycles, a connected client no longer sees a hard error. The launcher reconnects, re-handshakes on a protocol-version drift and keeps each connected client independently transparent. It is also hardened against an externally killed or slow-booting daemon. A boot-time full-text-search auto-heal plus a launcher clean-close barrier fixed the daemon-lifecycle root cause behind the recurring corruption-looking failures, and the launcher-lease integration suite, which had been failing every run because the test fixture copied the launcher but not its library tree, was repaired and un-skipped.
The Deep-Loop Runtime Learns to Fan Out
The engine behind deep-research and deep-review gained an opt-in fan-out layer. One loop can now drive several executor lineages, each an isolated sub-packet with its own state directory, behind a concurrency-capped pool with a status ledger. A consumer-specific merge folds the lineages back together: research dedups and attributes findings, review rolls up severity and fails closed if any lineage reports a P0. A salvage path recovers a lineage's work from captured output if its write fails. The upshot is that a deep-research or deep-review run can spread its work across executors and recover from a flaky one, instead of grinding through a single lineage at a time. The single-executor sequential path stays the default and is byte-identical to before.
Honest scope note: in this release the fan-out runs under the cap but still issues its lineages serially. The genuine-concurrency remediation was specced here and lands in the next release. Two parsers behind the loops also got real fixes: the review transcript parser no longer concatenates a model's prompt-echo into invalid JSON and zeroes out findings, and the convergence snapshot writer refuses to coalesce every snapshot to one id and silently overwrite history.
A Sharper deep-improvement Skill
The deep-improvement skill, renamed this cycle from deep-agent-improvement, gained two evaluation capabilities. A skill-benchmark lane measures how a real agent discovers, routes to and uses a skill in situ, scoring routing accuracy, unprompted discovery, efficiency and usefulness into a ranked, remediable report. And a reusable, config-driven model-benchmark framework replaced the one-off rigs: a single profile runs a framework bake-off or a model-versus-model comparison with no mode branches, correctness is a gate so a saturated score cannot crown a winner, and the verdict is WINNER, TIE or INCONCLUSIVE against a noise floor.
A Forkable Prompt Toolkit
The prompt toolkit was re-architected into three layers: framework craft in sk-prompt, per-model craft in sk-prompt-small-model and executor mechanics in the cli-* skills. The model registry and benchmarks moved into sk-prompt-small-model so the generic sk-prompt skill can be lifted out of this repo without dragging the small-model registry with it.
That small-model hub is what lets open models pay off. It holds the per-model prompt-craft profiles, the model registry and the benchmark run-data, so open models like DeepSeek, MiniMax and Xiaomi MiMo can be driven effectively through cli-opencode rather than prompted blind. MiniMax now defaults to its Token Plan provider and Xiaomi MiMo-V2.5-Pro was added as a selectable executor.
A run of supporting skill work landed alongside it: a new mcp-click-up skill for task work, surface-aware routing and mined animation principles in sk-code, a numbered wt/{NNNN}-{name} worktree convention for sk-git, and feature-catalog standards applied retroactively across roughly 370 catalog files.
Memory That Gets Smarter on Its Own
Enrichment runs in the background
Post-insert enrichment, reconsolidation and quality auto-fix are now on by default, and enrichment runs async. A save returns immediately with enrichmentStatus: deferred while a bounded background scheduler enriches graph and entity data after the commit. Reconsolidation stays opt-in. A SPECKIT_POST_INSERT_ENRICHMENT_SYNC escape hatch restores the old synchronous path when you need it.
Causal relation inference
Spec-document chains and lineage links are now promoted into typed created_by='auto' causal edges. The backfill is bounded, dry-run by default and reversible. Two opt-in collectors add a similarity supports signal and a structural contradicts signal. A conflict guard refuses to invalidate an existing valid edge, which closes a case where a committed contradicts run over a reciprocal lineage pair could silently undo the valid caused edge it had just created. The production backfill raised relation coverage from 39.91% to 43.59% and caused edges from 3 to 103, with zero conflicting edges skipped.
The index keeps itself reconciled
memory_index_scan graduated from a manual trigger to a surface that maintains itself. Redundant concurrent scans collapse into one, orphans get swept, a relocated spec folder is re-pointed instead of orphaned, and the index reports its state through memory_health. An active-row uniqueness guard and multi-tenant scope isolation close the door on duplicate rows and cross-scope bleed.
Smaller Durability and Isolation Wins
Durable checkpoints
Checkpoints are now real database files rather than in-process snapshots. Create writes a complete file with VACUUM INTO, restore swaps a finished file in behind a crash-safe journal, and a .needs-rebuild sentinel forces a clean rebuild if a restore cannot finish. The schema moved to v30.
Worktree-per-session isolation
Each session can run in its own worktree, reaped on close, which prevents the multi-session contention behind much of the prior release's daemo...