Skip to content

v1.2.0 — Agent Reliability Hardening

Choose a tag to compare

@lasswellt lasswellt released this 17 Apr 01:32
· 274 commits to main since this release

Added

Three new shared protocols under skills/_shared/:

  • subagent-types.md — authoritative guidance on choosing between Explore, general-purpose, Plan, and the six blitz plugin agents. Documents the foot-gun where the SDK picks read-only Explore for write-required work. Covers model inheritance rules (including the [1m] propagation that broke v1.1.3).
  • agent-workload-sizing.md — Light/Medium/Heavy weight class table with hard caps on file reads, web searches, tool calls, output lines, and wall-clock budget. Mandatory HEARTBEAT + PARTIAL protocol for Heavy class. Banned patterns (unbounded files, unbounded diff, write-at-end). Fail-fast rationale grounded in the fact that all tokens are billed regardless of outcome.
  • waves.md — dependency-DAG execution protocol extracted from sprint-dev Phase 1.4 into a shared spec. Gates adoption on "do you actually have a DAG?" so flat-pool orchestrators don't cargo-cult waves.

Fixed

Silent-failure sites where orchestrators consumed missing/empty agent output without validation:

  • sprint-plan Phase 2.4 — now validates all research-agent output files exist and are non-empty; aborts if 2+ of 4 agents failed; retries one failure with narrower scope.
  • sprint-review Phase 2.6 — same validation pattern for reviewer agents. Security-domain miss is now an explicit abort condition.
  • roadmap Phases 5 & 7 — domain-spec and epic-gen agents now validated before proceeding to synthesis.
  • fix-issue Phase 1.4 — root-cause analysis now refuses to proceed on missing/empty research output; downgrades confidence and asks user.

Subagent-type declarations added to 4 at-risk skills (research, sprint-plan, sprint-review, codebase-audit). Every agent spawn site now explicitly specifies subagent_type: general-purpose for write-required work, preventing the SDK from routing to read-only Explore.

Documentation

Three research docs under docs/_research/ driving this release:

  • 2026-04-16_plugin-agent-strategy.md — architectural audit of agent usage across 31 skills.
  • 2026-04-16_subagent-type-selection.md — Explore-vs-general-purpose foot-gun investigation.
  • 2026-04-16_agent-reliability.md — timeout/context-exhaustion taxonomy and blitz workload audit.

Upgrade Notes

No breaking changes. Existing skill invocations continue to work. The new shared docs are additive; skills that don't link them continue to function but miss the workload-sizing guidance. Sprint 2 (v1.3.0 target) will roll out the HEARTBEAT/PARTIAL protocol across more skills and refactor codebase-map, integration-check, and quality-metrics to parallel workers.