brew update && brew upgrade agentops · cd ~/.local/share/agentops && git pull --ff-only && ao skills link · checksums · verify provenance
Highlights
AgentOps 3.6 is the operations-layer alignment release, and most of the alignment was deletion. The product is now stated one way on every surface that reads it: AgentOps is the operations layer for agentic engineering, the topology is a federated integration graph, the interoperability contract is the semantic work-and-proof protocol, and the standard path is one RPI traversal. The knowledge-flywheel product surface is gone, the seven-move operating-loop workflow is a tombstone, and 25 skills plus every generated projection were realigned against the operating contract.
Two enforcement layers ride on top. Anti-ceremony guardrails fail closed when a process artifact cannot name its consumer, its decision, the observed defect, and its retirement condition, and rpi runs that STOP/CONTINUE guard once before Plan. The eval program gives the skill harness measurement instead of self-report: an eval architecture, a probe harness with a fail-closed v3 evidence contract, routing probes, a tier-2 flawed-plan corpus with hidden holdout scoring, and estate ablation sweeps with predictions registered before the runs. That same evidence contract is why measured probe coverage now reads an honest 0/12 rather than a flattering number.
Upgrade Notes
- Remove any
ao flywheelinvocation from scripts and CI. There is no replacement command; AgentOps no longer computes knowledge-compounding state. Learning stays available as an optional off-path consumer of durable verdicts through thelearnskill.flywheel:config keys can be left in place — they parse and are ignored. - If anything reads
.agents/handoff/*.jsonwritten byao session handoff, repoint it to.agents/ao/handoff/. Existing files stay put as read-only evidence andao session rehydratealready reads both roots, so no migration step is required. - The
ao init.gitignore block is idempotent on its begin marker, so a repository initialized before 3.6 keeps its old block and re-runningao initwill not update it. To adopt the current lines, edit the block by hand: replace.agents/ao/index/,.agents/ao/sessions/, and.agents/ao/provenance/with.agents/scratch/and.agents/projections/..agents/ao/intents/and.agents/ao/verdicts/remain deliberately un-ignored; whether loop evidence is committed is your repository's policy. ao gc preparenow writes Codex trust entries into$CODEX_HOME/config.tomlfor the Gas City directories that exist when it runs. Pass--codex-binifcodexis not on PATH; the same flag exists onao gc check.- Callers of
planthat used manifest mode should keep the multi-behavior specification in their own document and handplanone behavior at a time. - The behavioral probe ledger moved from the generated
skills/SKILL-TIERS.mdto the hand-maintainedevals/skill-probes/LEDGER.md.SKILL_PROBE_TIERS_FILEstill works as an alias; preferSKILL_PROBE_LEDGER_FILE. - The Go toolchain directive is 1.26.6. Local builds on an older toolchain will fetch or fail per your
GOTOOLCHAINsetting. - Otherwise no manual action is required. Existing installs upgrade in place with
brew upgrade agentopsor by re-running the install one-liner for your runtime.
Breaking Changes
This is a minor release, but it removes retired surfaces rather than freezing them. Each removal is a deliberate retirement with a migration pointer, not a regression.
ao flywheel(status,compare) is removed. The command exits 1 as an unknown command and prints a migration pointer. There is no replacement.planmanifest mode, added in 3.5.0, is removed along with its "plan manifest" trigger.- The seven-move
operating-loopworkflow is a tombstone; invoking it fails with a deterministic migration message. ao session handoffwrites new JSON to.agents/ao/handoff/instead of.agents/handoff/. Reads are unaffected —ao session rehydratesearches both roots and nothing moves or deletes legacy files.ao initno longer creates.agents/ao/sessions/,.agents/ao/index/,.agents/ao/provenance/, or.agents/handoff/. It creates only.agents/ao/intents/sha256and.agents/ao/verdicts/sha256.codebase-reconwrites new packs to.agents/scratch/codebase-recon/<run-id>/instead of.agents/recon/<run-id>/. Earlier packs stay where they are and remain discoverable; do not move them, because a delta manifest cites the prior pack's exact path.
At a Glance
| Product Area | Added | Changed | Refactored | Fixed | Deprecated/Removed |
|---|---|---|---|---|---|
| Install, Upgrade, and Distribution | 0 | 2 | 0 | 0 | 0 |
| CLI and Operator Commands | 1 | 3 | 0 | 1 | 1 |
| Daemon, Scheduling, and Factory | 2 | 0 | 0 | 0 | 0 |
| Skills and Workflows | 2 | 5 | 0 | 1 | 2 |
| Codex and Runtime Integrations | 0 | 2 | 0 | 1 | 0 |
| Eval, Validation, and Release Gates | 5 | 4 | 0 | 2 | 1 |
| Docs and Onboarding | 4 | 3 | 0 | 0 | 1 |
| Security, Privacy, and Supply Chain | 0 | 4 | 0 | 0 | 0 |
| Contributor/Internal Refactors | 2 | 1 | 0 | 0 | 2 |
Product Areas
Install, Upgrade, and Distribution
- Changed: The Homebrew formula description no longer advertises a knowledge-flywheel CLI. It reads "Optional AgentOps CLI for deterministic repository checks and evidence linking," which is what the binary actually does.
- Changed: The release workflow builds on Go 1.26.6, matching the toolchain directive in
cli/go.mod.
CLI and Operator Commands
- Added:
ao session prune-agentsapplies.agentsretention policies from Go. It is a read-only dry run by default;--executemutates and--quietprints only the summary. The global--dry-runseam always wins over--execute, andscripts/prune-agents.shbecomes a compatibility wrapper. - Changed:
ao initscaffolds only the requested-proof stores that have declared consumers —.agents/ao/intents/sha256and.agents/ao/verdicts/sha256. The former session-transcript, search-index, provenance, and handoff directories had no declared consumer and are no longer minted. - Changed:
ao session handoffwrites JSON to.agents/ao/handoff/.ao session rehydratesearches both the canonical and legacy roots and selects the newest lexical handoff id, with the canonical copy winning a filename tie. - Changed: The managed
.gitignoreblock written byao initcovers.agents/scratch/,.agents/projections/, and__pycache__/instead of the retired knowledge-store paths. - Fixed:
ao doctorGit-root detection no longer gets redirected by an empty or otherwise invalid ancestor.gitdirectory, which previously sent doctor artifacts to the wrong place. - Removed: The
ao flywheelcommand family (status,compare) and the flywheel-only metric helpers, app package, and build-profile bit. Existingflywheel:config sections still parse and are ignored.
Daemon, Scheduling, and Factory
- Added:
ao gc preparepre-seeds Codex trust for materialized Gas City session homes, writing both decisions Codex persists independently in$CODEX_HOME/config.toml— workspace trust under[projects."<dir>"]and one Codex-owned content hash per hook under[hooks.state."<hook-key>"]. Trusting a parent directory does not trust a session home, and hook identities come from Codex's ownhooks/listapp-server method rather than being reimplemented. Without this, the first Codex process in a fresh home could sit on the interactive trust dialog with a live pane that can never take dispatched work. - Added:
ao gc prepareandao gc checktake--codex-binto name the Codex CLI used to resolve hook trust identities; the default iscodexon PATH.
Skills and Workflows
- Added: The
anti-ceremonyskill, a clean-room artifact-free STOP/CONTINUE guard over the creation gate.rpitakes it as a hard dependency and invokes it exactly once before Plan. On STOP it dispatches no core phase and reportsNOT_PLANNEDwith the guard's reason; on CONTINUE the full traversal runs unchanged. - Added:
reverse-engineerships an output validator and a self-test, so its artifacts are checked by an executable rather than by assertion. - Changed: Skill contracts match the operating contract.
rpipreserves a durable caller-owned intent source by reference and digest and snapshots bytes only when no durable source exists.automation-shape-routingis advisory-only.bootstrapcreates durable verdict storage and project docs only when explicitly requested.implementandswarmreturn factual evidence instead of implicitly declaring semantic verdicts. - Changed:
codebase-reconbinds evidence to the exact Git commit its manifest declares. Citations must be safe repository-relative regular files at that commit, delta manifests must match Git's own changed-path diff rather than a self-reported boolean, and dirty source outside.agents/is refused. - Changed:
agent-mailrequires one mailbox owner and access mode per storage root, so an HTTP/MCP daemon and the directamCLI never contend for the same database. A busy activity lock or bounded read timeout is a degraded adapter result, not permission to restart the service. - Changed:
status,using-flywheel, andusing-gckeep the state classes separate. Factory-complete, checks-green, and AgentOps PASS are three different facts, never blended into one health claim. - Changed:
handoffdeclares its output contract as the caller-selected path or.agents/ao/handoff/, and states the compatibility rule in the owning skill: legacy artifacts stay as read-only evidence and no separate migration artifact is required. - Fixed: Skill audit grades are evidence-honest. Static readiness is scored and labeled separately from measured safety and effectiveness, so a static score can no longer be read as a behavioral one.
- Removed:
planmanifest mode and its "plan manifest" trigger.planshapes one active behavior in the existing intent source. - Removed: The seven-move operating-loop workflow is a tombstone that fails with a deterministic migration message instead of silently running retired doctrine.
Codex and Runtime Integrations
- Changed: Portable skill validation rejects duplicate YAML keys, enforces symlink containment, and applies the current optional-field rules, with negative and generator regression coverage behind it.
- Changed: The Codex plugin manifest states the operations-layer identity instead of the narrower fresh-context-validation framing.
- Fixed: Generated Codex descriptions are compacted without dropping canonical trigger text, so a compacted projection cannot lose the phrases that activate the skill. All owned Codex projections and hashes were regenerated against that rule.
Eval, Validation, and Release Gates
- Added: A 12-decision eval architecture for the skill harness, plus the probe harness that runs control-versus-treatment behavioral probes and records replayable scorecards.
- Added: Routing probes measure P(skill loaded | applicable task), the multiplier every skill-efficacy number depends on. Committed scenarios are templates with dispatch-time instantiation so an agent cannot read the fixture off disk and contaminate the run.
- Added: A tier-2 task-embedded corpus of flawed-plan execution fixtures with hidden holdout tests injected only at scoring, across four live flaw classes. The scorer's selftest against planted references caught a real scorer bug before any spend.
- Added: Estate ablation sweeps with predictions registered before results, applying delete-everything-and-measure-what-returns to this repository's own estate.
- Added: Between-releases release-path smoke.
.github/workflows/release-path-smoke.ymlruns a full GoReleaser snapshot nightly and on every change to a release-plumbing input, closing the commitment recorded indocs/audits/release-readiness-v3.3.0.md. The release path previously ran only on a realv*tag push, which is how a retired before-hook stayed broken for 20 days. CI and the negative witness share one script. - Changed: The probe-coverage gate reads a hand-maintained ledger at
evals/skill-probes/LEDGER.mdinstead of the generatedskills/SKILL-TIERS.md, which a regeneration had silently wiped. A row counts only if it names exactly one v3 scorecard whose fixture manifest, bound inputs, transcript hashes, canonical skill source, runtime identity, reps, treatment mode, and recomputed discriminator result all agree. - Changed: Cathedral-cut conformance scans linked skill references for retired identity claims, with Markdown fence, raw-HTML, and percent-encoding handling so the scan cannot be evaded by formatting.
- Changed:
scripts/extract-release-notes.shreflows hard-wrapped curated notes into logical Markdown lines before publishing, leaving headings, tables, code, HTML, block quotes, and explicit hard breaks intact. - Changed: The bats lane provisions Python with jsonschema and PyYAML once, instead of reinstalling PyYAML after
setup-pythonshadowed the runner interpreter. - Fixed: The eval command-surface fixture failed when it was actually executed, asserting
#{3,4}where the surface produces#{3,5}. The fixture now matches the real surface. - Fixed: A version-parity regression test now asserts the checked-in
versionfallback agrees with every version-bearing release surface, includingimages/claude/verify.sh. The existing manifest-consistency check compared only the two Claude manifests to each other and could not see that drift. - Removed: The vacuous retrieval-quality canary and the nightly job that ran it, the flywheel-compounding gate scripts and their bats suite, and the dormant
check-pillar-coverage.sh,check-thesis-stability.sh, andgenerate-index.shscripts.
Docs and Onboarding
- Added:
docs/contracts/ubiquitous-language.mdgains the product-architecture vocabulary and four forbidden conflations, including that check success is not a semantic PASS and that runtime completion is not validation. - Added:
docs/reference/skill-system-evolution.mdtraces the skill system across releases, names seven evolution regimes, diagrams the present authority graph, and records falsifiable predictions. - Added: Two research reports land as durable references — the skill-eval SOTA standards survey the eval architecture is built on, and a formal-verification-for-agent-code survey.
- Added: The operations-layer alignment plan is committed as the reference for what changed and why.
- Changed: The core architecture page is
docs/architecture/rpi-traversal.md, with a compatibility redirect at the old operating-loop path. - Changed:
AGENTS.mdis compressed against current-main evidence while preserving the honest-work and anti-ceremony doctrine, federated source authority, the factory and concurrency boundaries, and the Triggered Sources links. The estate-ablation aggregate counts that motivated the trim are labeled legacy-unverified and non-promotable. - Changed: ADR-0016 is amended for federated source authority and corrects an overclaim: the three-directory
.agents/layout is a target, not a currently enforced closed set. - Removed: The dated skill-quality audit report, which had no runtime consumer and bound a superseded audit snapshot. All functional skill-quality changes stayed.
Security, Privacy, and Supply Chain
- Changed: Handoff publishing is descriptor-anchored. It opens
.agents/ao/handoffone component at a time, refuses any component that is a symlink or not a real directory, re-verifies that each opened descriptor still identifies the component it inspected, and publishes by hard link rather than rename, so a parent-directory swap mid-write cannot redirect the write and an existing handoff id is never clobbered. - Changed: Prune path confinement runs traversal inside an
os.Rootwith a final current-path identity check immediately before each descriptor-rooted delete. The contract states its own boundary honestly: intermediate-directory traversal is bound, and it does not claim a final basename is immutable against an adversarial replacement. - Changed: Portable Codex skill validation enforces symlink containment and rejects duplicate YAML keys, so a projected skill package cannot reach outside its own tree or carry a key whose second definition silently wins.
- Changed: Pinned GitHub Actions digests were refreshed, including
actions/attest-build-provenanceon the release provenance step anddorny/paths-filterin the validate workflow.
Contributor/Internal Refactors
- Added: The canonical bd project identity is tracked at
.beads/identity.tomland bound to the live project id, so project identity travels with the repository instead of being reconstructed per checkout. - Added: A
.gitattributesfile lands so evidence fixtures and generated projections get deterministic treatment across checkouts. - Changed: Dependencies — Go toolchain directive 1.26.5 to 1.26.6,
golang.org/x/text0.40.0 to 0.41.0,github.com/santhosh-tekuri/jsonschema/v66.0.2 to 6.0.3, andgithub.com/BurntSushi/toml1.6.0 added for Codexconfig.tomlhandling. - Removed: Knowledge-flywheel internals were deleted with their tests:
cli/internal/flywheelapp,cli/cmd/ao/flywheel_composition.go,scripts/lib/flywheel-compile.sh,scripts/snapshot-flywheel-compounding.sh,scripts/log-telemetry.sh, andtests/skills/test-artifact-consistency.sh. - Removed: The consumer-free
dream:config block and its exclusive helpers. Existing config files still parse and the key is ignored. Remaining scratch-tier writers now write under.agents/scratch/, andbin/ralphstill resumes legacy.agents/ralph/checkpoints so the documented backwards-compatibility contract holds without a migration.
Known Issues
- Measured behavioral probe coverage is an honest 0/12 under the v3 evidence contract. The wave-1 classifications from earlier in this cycle are retained as
LEGACY-UNVERIFIEDrather than counted, because the probe harness did not isolate the skill corpus between control and treatment arms. Skill-efficacy claims in this release are directional, not proven. - The estate-ablation aggregate counts are labeled legacy-unverified and non-promotable: useful as hypotheses, not proof of causal effects, generalized token cost, or executor behavior.
tests/docs/validate-goal-count.shexits 1 with no diagnostic because its pipe-table parser predates the current prose-plus-listGOALS.md. It is wired intotests/run-all.sh, which is not in CI or the release path, so it gates nothing.scripts/check-product-freshness.shis advisory and wired into no gate.
Full changelog
AgentOps 3.6 is the operations-layer alignment release, and most of the alignment was deletion. The product is now stated one way on every surface that reads it: AgentOps is the operations layer for agentic engineering, the topology is a federated integration graph, the interoperability contract is the semantic work-and-proof protocol, and the standard path is one RPI traversal (Plan, Implement, fresh Validate, report and stop). The knowledge-flywheel product surface is gone — command family, app package, build profile, gates, config block — the seven-move operating-loop workflow is a tombstone, .agents/ writers are narrowed to declared destinations, and 25 skills plus every generated projection were realigned against the operating contract.
Two enforcement layers ride on top. Anti-ceremony guardrails fail closed when a process artifact cannot name its consumer, its decision, the observed defect, and its retirement condition, and rpi runs that STOP/CONTINUE guard once before Plan. The eval program gives the skill harness measurement instead of self-report: a 12-decision eval architecture, a probe harness with a fail-closed v3 evidence contract, routing probes, a tier-2 flawed-plan corpus with hidden holdout scoring, and estate ablation sweeps whose predictions were registered before the runs. That same evidence contract is why measured probe coverage now reads an honest 0/12 — the earlier wave-1 classifications are retained as LEGACY-UNVERIFIED rather than counted.
Added
- The
anti-ceremonyskill: a clean-room, artifact-free STOP/CONTINUE guard over the creation gate — name the consumer, the decision it gates, the observed defect, and the retirement condition, or do not create it.rpitakes it as a hard dependency and invokes it exactly once before Plan. On STOP it dispatches no core phase and reportsNOT_PLANNEDwith the guard's reason; on CONTINUE the full Plan, Implement, fresh Validate path runs unchanged. AgentOps now fails closed when process artifacts are manufactured, when GREEN comes from weakening the oracle, and when repeated control artifacts replace implementation evidence. ao session prune-agentsapplies.agentsretention policies from Go. It is a read-only dry run by default;--executemutates and--quietprints only the summary. The global--dry-runseam always wins over--execute.scripts/prune-agents.shbecomes a compatibility wrapper.ao gc preparepre-seeds Codex trust for materialized Gas City session homes. Codex persists two independent decisions in$CODEX_HOME/config.toml— workspace trust under[projects."<dir>"]and one Codex-owned content hash per hook under[hooks.state."<hook-key>"]— and trusting a parent directory does not trust a session home. Hook identities and hashes come from Codex's ownhooks/listapp-server method rather than being reimplemented. Without this, the first Codex process in a fresh home could sit on the interactive trust dialog with a live pane that can never take dispatched work.ao gc prepareandao gc checktake--codex-binto name the Codex CLI; the default iscodexon PATH.- The eval program for the skill harness: a 12-decision eval architecture (
docs/architecture/eval-architecture.md) and the probe harness that runs control-versus-treatment behavioral probes and records replayable scorecards (bash scripts/probe-skill.sh --probe <id> --replay). - Routing probes measure P(skill loaded | applicable task), the multiplier every skill-efficacy number depends on. Committed scenarios are templates with dispatch-time instantiation so an agent cannot read the fixture off disk and contaminate the run. Task-noun skills route and judgment-posture skills do not; the Codex runtime routed 4/4 against 2/6 for Claude-in-repo, because
.claude/rulesauto-load already delivers the content without routing. - A tier-2 task-embedded corpus of flawed-plan execution fixtures with hidden holdout tests injected only at scoring, across four live flaw classes (quiet edge, vacuous green, burned compatibility bridge, opaque sentinel errors). The scorer emits
visible_pass,hidden_pass,claimed_done,false_pass, andflagged_gap, and its selftest against planted references caught a real scorer bug before any spend. - Estate ablation sweeps that apply delete-everything-and-measure-what-returns to this repository's own estate, with predictions registered before results.
- A dual-case anti-ceremony behavior probe that separates a justified process artifact from an unjustified one, with its discriminator under test.
- Between-releases release-path smoke:
.github/workflows/release-path-smoke.ymlruns a full GoReleaser snapshot (goreleaser release --snapshot --clean --skip=publish) nightly and on every change to a release-plumbing input (.goreleaser.yml,.github/workflows/release.yml,cli/go.mod,cli/Formula/**). This closes the process commitment recorded indocs/audits/release-readiness-v3.3.0.md: the release path was previously executed only by a realv*tag push, which is how a retired.goreleaser.ymlbefore-hook stayed broken for 20 days. The smoke and its negative witness share one script (tests/scripts/lib/release-snapshot-smoke.sh, proven able to fail bytests/scripts/release-path-smoke.bats). reverse-engineerships an output validator and a self-test, so its artifacts are checked by an executable rather than by assertion.docs/contracts/ubiquitous-language.mdgains the product-architecture vocabulary and four forbidden conflations, including that check success is not a semantic PASS and that runtime completion is not validation.docs/reference/skill-system-evolution.mdtraces the skill system across releases, names seven evolution regimes, diagrams the present authority graph, and records falsifiable predictions.- The canonical bd project identity is tracked at
.beads/identity.toml, so project identity travels with the repository instead of being reconstructed per checkout.
Changed
- The core architecture page is
docs/architecture/rpi-traversal.md, with a compatibility redirect left at the old operating-loop path. - Skill contracts match the operating contract.
rpipreserves a durable caller-owned intent source by reference and digest and snapshots bytes only when no durable source exists.automation-shape-routingis advisory-only and returns the chosen owner without copying or starting the delegated workflow.bootstrapcreates durable verdict storage and project docs only when explicitly requested.implementandswarmreturn factual evidence instead of implicitly revising or declaring semantic verdicts. ao initscaffolds only the requested-proof stores that have declared consumers:.agents/ao/intents/sha256and.agents/ao/verdicts/sha256. The former session-transcript, search-index, provenance, and handoff directories had no declared consumer and are no longer minted. The managed.gitignoreblock now covers.agents/scratch/,.agents/projections/, and__pycache__/instead of the retired knowledge-store paths.ao session handoffwrites JSON to.agents/ao/handoff/instead of.agents/handoff/.ao session rehydratesearches both roots and selects the newest lexical handoff id, with the canonical copy winning a filename tie. No command moves or deletes the legacy files.codebase-reconbinds evidence to the exact Git commit its manifest declares. Citations must be safe repository-relative regular files at that commit, a supplied line number must exist in the committed blob, delta manifests must match Git's own prior-to-current changed-path diff rather than a self-reported boolean, and dirty tracked, staged, or untracked source outside.agents/is refused. New packs default to.agents/scratch/codebase-recon/<run-id>/; earlier packs stay where they are and remain discoverable.agent-mailrequires one mailbox owner and access mode per storage root, so an HTTP/MCP daemon and the directamCLI never contend for the same database. A busy activity lock or bounded read timeout is a degraded adapter result, not permission to restart the service or switch roots.status,using-flywheel, andusing-gckeep the state classes separate. Factory-complete, checks-green, and AgentOps PASS are reported as three different facts, never blended into one health claim.- The probe-coverage gate reads a hand-maintained ledger at
evals/skill-probes/LEDGER.mdinstead of the generatedskills/SKILL-TIERS.md, which a regeneration had silently wiped. A row counts only if it names exactly one v3 scorecard whose fixture manifest, bound inputs, prompt events, transcript hashes, canonical skill source, non-overrideable runtime identity, reps, treatment mode, and recomputed discriminator result all agree. - Cathedral-cut conformance scans linked skill references for retired identity claims, with Markdown fence, raw-HTML, and percent-encoding handling so the scan cannot be evaded by formatting.
scripts/extract-release-notes.shreflows hard-wrapped curated notes into logical Markdown lines before publishing, leaving headings, tables, code, HTML, block quotes, and explicit hard breaks intact. GitHub's renderer preserves soft line breaks, so copying the hard-wrapped source byte-for-byte produced a narrow ragged column in the published release.- Handoff publishing is descriptor-anchored: it opens
.agents/ao/handoffone component at a time, refuses any component that is a symlink or not a real directory, re-verifies that each opened descriptor still identifies the component it inspected, and publishes by hard link rather than rename, so a parent-directory swap mid-write cannot redirect the write and an existing handoff id is never clobbered. - Prune path confinement runs traversal inside an
os.Rootwith a final current-path identity check immediately before each descriptor-rooted delete. The contract states its own boundary honestly: intermediate-directory traversal is bound, and it does not claim a final basename is immutable against an adversarial replacement. - Portable Codex skill validation rejects duplicate YAML keys, enforces symlink containment, and applies the current optional-field rules.
AGENTS.mdis compressed against current-main evidence (about 10.5KB to 8.3KB) while preserving the honest-work and anti-ceremony doctrine, federated source authority, the factory and concurrency boundaries, and the Triggered Sources links thatvalidate-agents-splitrequires. The estate-ablation aggregate counts that motivated the trim are labeled legacy-unverified and non-promotable.- ADR-0016 is amended for federated source authority and corrects an overclaim: the three-directory
.agents/layout is a target, not a currently enforced closed set. - The Homebrew formula and the Codex plugin manifest state the operations-layer identity instead of advertising a knowledge-flywheel CLI.
- Dependencies: Go toolchain directive 1.26.5 to 1.26.6,
golang.org/x/text0.40.0 to 0.41.0,github.com/santhosh-tekuri/jsonschema/v66.0.2 to 6.0.3,github.com/BurntSushi/toml1.6.0 added for Codexconfig.tomlhandling, and the pinned GitHub Actions digests refreshed.
Fixed
- Generated Codex descriptions are compacted without dropping canonical trigger text, so a compacted projection cannot lose the phrases that activate the skill. All owned Codex projections and hashes were regenerated against that rule (portable conformance 52/52).
- Skill audit grades are evidence-honest: static readiness is scored and labeled separately from measured safety and effectiveness, so a static score can no longer be read as a behavioral one.
ao doctorGit-root detection no longer gets redirected by an empty or otherwise invalid ancestor.gitdirectory, which previously sent doctor artifacts to the wrong place.- The eval command-surface fixture failed when it was actually executed: it asserted
#{3,4}where the surface produces#{3,5}. The fixture now matches the real surface. - The bats lane provisions Python with jsonschema and PyYAML once, instead of reinstalling PyYAML in a later job after
setup-pythonshadowed the runner interpreter.
Removed
- The
ao flywheelcommand family (status,compare) and the whole live knowledge-flywheel product surface:cli/internal/flywheelapp, the flywheel-only metric helpers and types, theflywheelbuild-profile bit, theflywheel-compoundinggate scripts, and theflywheel:config block (existing config files still parse; the key is ignored). AgentOps no longer computes or reports knowledge-compounding state (COMPOUNDING,DECAYING, escape velocity). Invokingao flywheelnow exits 1 as an unknown command and prints a migration pointer (seedocs/MIGRATION.md). Learning remains an optional off-path consumer of durable verdicts via thelearnskill. The compatibility baseline records the family as intentionallyretiredrather than freezing the old product claim. planmanifest mode, added in 3.5.0, along with its "plan manifest" trigger.planshapes one active behavior in the existing intent source; a multi-behavior specification stays a document the caller owns.- The seven-move operating-loop workflow is a tombstone that fails with a deterministic migration message instead of silently running retired doctrine.
- The consumer-free
dream:config block (the retired overnight subsystem's settings; existing config files still parse and the key is ignored) and its exclusive helpers, the vacuous retrieval-quality canary and the nightly job that ran it (every Go test it named was already deleted; the retrieval-comparison contract is now explicitly dormant and the blocking manifest-paths gate remains the live retrieval guard), the dormantcheck-pillar-coverage.shGOALS.yaml script,check-thesis-stability.sh, andgenerate-index.sh. Remaining scratch-tier writers (mine-all-sessions.sh,team-runner.sh,bin/ralph) now write under.agents/scratch/.bin/ralphstill resumes legacy.agents/ralph/checkpoints, so the documented backwards-compatibility contract holds without a migration. - The dated skill-quality audit report, which had no runtime consumer and bound a superseded audit snapshot; all functional skill-quality changes stayed.
Full Changelog: v3.5.0...v3.6.0