Concrete upgrade proposals: a mechanical story-closure gate, evidence receipts, config echo, a domain toolchain, and more #135
jiangshoujie
started this conversation in
Ideas
Replies: 1 comment
|
leave it with me mate, some absolute gold here. Will let you know about how we will go about some of these |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hi Donchitos —
I've read through CLAUDE.md,
context-management.md, the three largest skills (story-done / dev-story / gate-check), all 12 hooks, and the scripts in.claude/scripts/. Short version: this is the most thoroughly engineered Claude Code template I've come across — the measuredminimaldefault, the two-regionactive.mdschema, the locked verdict vocabulary ("A parse check is not a run"), and the "scripts emit observations, never verdicts" principle (the header comment ofproject-coherence.shis a model in itself) are all things I haven't seen elsewhere.Below are nine concrete improvement proposals from working through the system. Each is self-contained — problem, design (written against the existing directory layout and conventions), acceptance criteria, and rough effort. If any of them appeal to you, I'm happy to send PRs.
Proposal 1 — Story closure must be executed by a script (closing the direct-Status-edit bypass)
Problem. All of
/story-done's checks are executed by the model itself (grep/glob). The model can skip the checks — or skip/story-doneentirely and just edit a story'sStatus:field to Done. Nothing intercepts that today (story-status.shis a manual helper, not a guard).Design.
.claude/scripts/close-story.sh <story-path>becomes the only writer ofStatus: Done. Flow:yaml-helper.sh resolve_configfor the story's system →qa.level/testing.strictEVIDENCE_MISSING/RECEIPT_STALE/BUILD_MISMATCH/SIGNOFF_INCOMPLETE). That output doubles as the BLOCKED basis of the completion report..claude/hooks/guard-story-status.sh, wired insettings.jsonunder PreToolUse (Write|Edit): deny when the target matchesproduction/epics/**/*.md, the content containsStatus: Done, and the write is not coming from the close-story flow — with a "Run /story-done" hint.Acceptance. A direct Status edit is blocked by the hook; incomplete evidence makes
close-story.shexit 1 with itemized reason codes.Effort. S (2–3 days; the path matching can copy
validate-assets.sh).Proposal 2 — Evidence receipts (from "exists" to "provable")
Problem. Today's evidence check globs
production/qa/evidence/for a png + sign-off doc. Reusing another story's screenshot or copying a placeholder image is undetectable.Run result: OBSERVEDonly requires that the model Read some image — there is no binding of "this screenshot ← this build". Test evidence explicitly reports only "test file present".Design. Two wrapper scripts; receipts land in the existing evidence directory:
.claude/scripts/run-tests.sh <test-path>: resolves the engine test command fromproject.yaml, runs it, writestest-receipt.json:{ "kind": "test", "story": "TR-combat-013", "command": "godot --headless --script tests/unit/combat/test_damage.gd", "exit_code": 0, "passed": 12, "failed": 0, "timestamp": "…", "build_id": "git:8f3c2e1", "output_sha256": "…" }.claude/scripts/run-and-capture.sh: launches the game per engine and captures a screenshot (dev-story already referencesScreenshotOnArg.cs— wire that existing convention into the script), writingscreenshot-receipt.jsonwithimage_sha256+build_id+story.close-story.sh(Proposal 1) reads receipts instead of globbing:image_sha256must match the file on disk, andbuild_idmust match the current HEAD — otherwiseRECEIPT_STALE. Reusing an old screenshot stops working.Acceptance. Reusing an old screenshot fails closure with
RECEIPT_STALE; never-run tests yield no receipt →EVIDENCE_MISSING.Effort. M (3–5 days; per-OS screenshot capture is the bulk — Godot can start with an autoload-based capture).
Proposal 3 — resolve_config echo + value validation
Problem. Two traps your own comments document: (a) reading
project.yamlalone silently ignoresproject.local.yamloverrides; (b)testing.strictvalues likeyes/1/maybeare silently treated as unset. In both cases gates are silently waived while the model and the user believe the config is in effect.Design.
yaml-helper.sh resolve_config --explain: prints, per key, "effective value ← source file" alongside the resolution, and the skill preamble injects it. One line of cost to make "which gates are actually in effect" visible to both the model and the human./settingsvalidates on write: a boolean knob receiving anything outside its legal values → refuse the write and list the legal values (moving story-done's "treated as unset" tolerance to a write-time rejection).Acceptance.
/settings testing.strict.logic=1is refused; the skill preamble showstesting.strict.logic: true ← project.local.yaml.Effort. S (1–2 days).
Proposal 4 — Independent review receipts
Problem. The Phase 5 Lead Programmer review under
review_mode: solois the same model grading its own work.log-agent.shrecords that an agent ran, but not what the review concluded.Design.
review-receipton completion (extend the existingreview-receipts.sh):{reviewer_agent, model, checklist: [{item, pass, note}], verdict}. Checklist items come from the machine-checkable rules inrules/(binary verification, not vibes).Reviewed by: <agent>@<model> — receipt: <path>; undersoloit explicitly readsReviewed by: self — no independent review.Acceptance. Every completion report traces to a reviewer identity and per-item results.
Effort. S–M (2–4 days).
Proposal 5 — Compile the conditional matrices ("run cards")
Problem.
dev-story's body is 44KB (~11k tokens) containing a five-layer conditional matrix (rigor × qa.level × story_granularity × testing.strict × system_overrides) expressed entirely in prose and tables. The weaker the model, the more branches get dropped — most visible on non-Claude models. The!yaml-helper resolve_configinjection at the top of dev-story is already the right pattern; it just doesn't cover the matrix itself.Design. Backward-compatible and incremental:
matrix.yaml(knobs + branches); branch-specific prose moves into.claude/skills/<name>/branches/<knob-value>.mdfragments;SKILL.mdkeeps only the branch-agnostic skeleton.yaml-helper.sh compile <skill>: resolves config → assembles the run card for the one effective path (target ≤8KB) → injected via!in place of the matrix prose. The model only ever sees a single path; it never reads a table again.Acceptance. The three skills' bodies are ≤8KB; running a
standardstory shows nofull-branch STOP text anywhere in the resolved card.Effort. L (1–2 weeks, incremental per skill).
Proposal 6 — Upgrade /skill-test to behavioral regression (CI for prompts)
Problem. /skill-test today is a static linter + rubric; it cannot verify that gate behavior actually happens. Proposals 1, 2 and 5 each need a regression net.
Design. Seed fixture projects under
CCGS Skill Testing Framework/fixtures/, one per behavior:stale-screenshot/: complete evidence but an outdated build_id → assert close-story reportsRECEIPT_STALEdirect-status-edit/: simulated bypass → assert the hook denieslocal-override-ignored/: local.yaml override → assert--explainshows the sourcePlus a GitHub Actions workflow (all hooks are bash — they run on CI as-is).
Acceptance.
/skill-test behaviorruns all fixtures in one command and reports a pass rate.Effort. M (3–5 days).
Proposal 7 — Model declarations → role mapping (decouple from Anthropic ids)
Problem. Agent frontmatter hardcodes
model: opus / sonnet— ids bound to Anthropic product names. Sessions running via Bedrock / Vertex / compatibility gateways with third-party or custom models read meaningless values. And there is no config hook for tiered scheduling ("directors on the flagship, specialists on the workhorse").Design.
project.yamlgains:Frontmatter model fields become
inherit(= use the session model). Thereviewerkey doubles as the mount point for Proposal 4's "reviewer on a different model".Acceptance. Switch model backend or adjust the three-tier mapping without touching any agent file.
Effort. S (1–2 days).
Proposal 8 — Content supply lines (four independently shippable pieces)
{source_model, prompt_hash, license, approved_by}receipts into the asset audit — provenance for AI-generated assets is becoming a hard requirement, and it is isomorphic to Proposal 2.project.yaml's commands section + a GitHub Actions template (Proposal 6's fixtures reuse it directly).notify.shgains osascript/notify-send branches; UPGRADING gets a migration script (migrate-v1-config.shis a good start).Proposal 9 — Domain toolchain: world tech / save systems / dialogue runtime / animation
Problem. After auditing all 74 skills and the relevant agents: the template encodes process knowledge (workflow / review / QA / release) comprehensively, but several technical systems every indie game needs have no domain scaffolding —
world-builderowns lore (factions / cultures / history);level-designerowns layout docs (encounter layouts / pacing) — neither touches the technical side. Across all 74 skills there is no procgen coverage (noise terrain / WFC / dungeon generation / tilemap autotiling) and no seamless-streaming coverage (chunk streaming, LOD, UE5 World Partition, Unity Addressables scene grouping, Godot chunking). A user aiming for a seamless world has no domain guidance after /design-system.Why it fits. The team-* / domain-guide pattern is proven (/team-combat etc.); these four are the same pattern extended to the technical-content side. All four are engine-agnostic (per-engine specifics as subsections, matching the existing engine-specialist approach), their outputs flow through the existing GDD → ADR → story pipeline, and they are verified by the existing /story-done gates — no new mechanisms introduced.
Design. Four new skills, written in the style of design-system / create-architecture:
/world-tech(or split/procgen+/world-streaming): procgen decision tree (deterministic vs runtime, seed management, reproducibility), streaming budget checklist (memory / loading / culling), per-engine subsections/save-system: save schema design, version-migration strategy, cloud-save notes; matching ADR template/dialogue-system: format comparison, localization key conventions (bridging /localize), VO pipeline/animation-pipeline: state machine / blend tree / IK decisions, capture hooks (bridging Proposal 2's screenshot receipts)Acceptance. Each skill's outputs enter the existing GDD/ADR/story pipeline and are covered by the existing gates (no new mechanisms).
Effort. M each (2–4 days apiece, shippable individually); suggest /world-tech and /save-system first.
Priorities (from the maintainer's perspective)
Closing
One principle across all of this: new capabilities must come with their own receipts and evidence — never add another step that relies on the model's honesty. Happy to send PRs for any of these; Proposals 1+2 would be my starting point.
All reactions