Skip to content

test(prose): discovery's epic arm, to the same durability boundary#571

Open
leeovery wants to merge 1 commit into
docs/conventions-reference-directed-routingfrom
test/prose-discovery-epic-arm
Open

test(prose): discovery's epic arm, to the same durability boundary#571
leeovery wants to merge 1 commit into
docs/conventions-reference-directed-routingfrom
test/prose-discovery-epic-arm

Conversation

@leeovery

@leeovery leeovery commented Jul 26, 2026

Copy link
Copy Markdown
Owner

Summary

The same empty project and the same walk as the feature case, committed as an epic instead. The two diverge at exactly one point — confirm-trigger's routing — and everything else follows from it.

A feature concludes there. An epic continues into Step 7, which runs the discovery gateway, and that call is the proof of which arm was taken: it sits in calls_include and in the ordering chain behind manifest exists and workunit create, so a walk can't reach it by another route.

Three claims belong to the epic alone and can't pass by accident:

  • the session log records the map as empty at a first session, not the "not applicable" a single-topic type carries
  • the engine sets active_session — state a feature never acquires
  • the shaping conversation is carried forward rather than restarted, so no second session log is authored

calls_exclude covers topic start, discussion-map and the entry skills: nothing may begin while topics are still unnamed.

Provenance

This case's first run is what surfaced the walker wander at discovery Step 5, and through that the asserter's inability to say whose fault a step was. Both are fixed beneath this in the stack (#569, #570). The case is unchanged since that run, so it wants a run against those fixes before it's trusted — expected outcome is a clean pass with the wander gone, but that's a prediction, not a result.

Test plan

  • Corpus valid at 10 cases; snapshot rebuilds byte-identical
  • Re-run pending a plugin reload, since both fixes are agent/prose changes

🤖 Generated with Claude Code

Stack

  1. docs(design): prose-tests programme design log #544
  2. feat(prose-tests): the framework — cases, worlds, runner, skill #545
  3. test(prose): feature happy-path corpus — five worlds, seven cases #546
  4. test(prose): bugfix corpus — the investigation-centric surfaces #548
  5. test: retry recursive teardown removals — kill a class of phantom failures #549
  6. fix(entry-skills): close the handoff fences — six files render their arms wrong #550
  7. docs: a contributing page for working on the system #551
  8. fix(entry-skills): every handoff arm says to invoke the skill #552
  9. fix(implementation): environment setup belongs to the setup reference alone #553
  10. fix(prose-tests): the asserter is told which substitutions were armed #554
  11. feat(prose-tests): the mid-flow substitution, and a world only prose can describe #555
  12. test(prose): claims assert consequences, not what was displayed #556
  13. feat(prose-tests): record everything the agents do, results included #557
  14. fix(discussion-entry): the handoff reports the source it actually had #558
  15. fix(prose-tests): the stop hook records, and names the model that walked #559
  16. fix(prose-tests): command output was never actually recorded #560
  17. feat(prose-tests): judge the walk as told, not the summary returned #561
  18. feat(prose-tests): decide in code what an agent should not be deciding #562
  19. test(prose): a case starts where a session starts #563
  20. feat(prose-tests): walk on Sonnet, judge on Opus, escalate a failure #564
  21. test(prose): give the eight read-only cases something that can fail #565
  22. test(prose): only walks that can be observed, and checks that survive the trip #566
  23. fix(prose-tests): the verdict names only the model the record names #567
  24. test(prose): discovery, walked to the point where work first exists #568
  25. fix(prose-tests): the asserter judges which of prose or walker was at fault #569
  26. docs(conventions): a step whose reference routes every exit still signposts #570
  27. test(prose): discovery's epic arm, to the same durability boundary #571 👈 current

The same empty project and the same walk as the feature case, committed
as an epic instead. The two diverge at one point — confirm-trigger's
routing — and everything downstream follows from it.

A feature concludes there. An epic continues into Step 7, which runs the
discovery gateway, and that call is what proves which arm was taken: it
sits in calls_include and in the ordering chain behind manifest exists
and workunit create, so a walk cannot reach it by another route.

Three claims are the epic's alone and cannot pass by accident. The
session log records the map as empty at a first session rather than the
not-applicable a single-topic type carries. The engine sets
active_session, which a feature never acquires. And the shaping
conversation is carried forward rather than restarted, so no second log
is authored.

calls_exclude covers topic start, the discussion map and the entry
skills: nothing may begin while the topics are still unnamed.

Its first run is what surfaced the walker wander at Step 5, and through
that the asserter's inability to say whose fault a step was. Both are
fixed beneath this in the stack; the case is unchanged since, so it wants
a run against them before it is trusted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This was referenced Jul 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant