Skip to content

test: retry recursive teardown removals — kill a class of phantom failures#549

Open
leeovery wants to merge 1 commit into
prose-tests/bugfix-corpusfrom
fix/test-teardown-rmsync-retries
Open

test: retry recursive teardown removals — kill a class of phantom failures#549
leeovery wants to merge 1 commit into
prose-tests/bugfix-corpusfrom
fix/test-teardown-rmsync-retries

Conversation

@leeovery

@leeovery leeovery commented Jul 25, 2026

Copy link
Copy Markdown
Owner

Summary

  • A full-gate run went red on ENOTEMPTY inside an afterEach rmSync, in a suite unrelated to the change being made, while four background agents were churning temp directories. The same suite passed 3/3 in isolation.
  • Cause: the engine spawns git subprocesses; git drops lock/temp files a moment after the command returns, so a recursive remove that has just emptied a directory then fails to remove the directory itself. Pure timing — it only shows up under parallel load.
  • Every recursive teardown removal in tests/scripts now passes maxRetries: 5, retryDelay: 100, which is what the prose-test world builder already used for exactly this reason. 64 call sites across 25 suites, mechanical, no behavioural change.

Test plan

  • npm test green: 1699 tests, 0 fail.
  • The originally-failing suite re-run 3× in isolation before the change: clean each time, confirming a race rather than a logic fault.

🤖 Generated with Claude Code

Stack

  1. docs(design): prose-tests programme design log #544
  2. feat(prose-tests): the framework — cases, worlds, runner, skill #545
  3. test(prose): feature happy-path corpus — five worlds, seven cases #546
  4. test(prose): bugfix corpus — the investigation-centric surfaces #548
  5. test: retry recursive teardown removals — kill a class of phantom failures #549 👈 current
  6. fix(entry-skills): close the handoff fences — six files render their arms wrong #550
  7. docs: a contributing page for working on the system #551
  8. fix(entry-skills): every handoff arm says to invoke the skill #552
  9. fix(implementation): environment setup belongs to the setup reference alone #553
  10. fix(prose-tests): the asserter is told which substitutions were armed #554
  11. feat(prose-tests): the mid-flow substitution, and a world only prose can describe #555
  12. test(prose): claims assert consequences, not what was displayed #556
  13. feat(prose-tests): record everything the agents do, results included #557
  14. fix(discussion-entry): the handoff reports the source it actually had #558
  15. fix(prose-tests): the stop hook records, and names the model that walked #559
  16. fix(prose-tests): command output was never actually recorded #560
  17. feat(prose-tests): judge the walk as told, not the summary returned #561
  18. feat(prose-tests): decide in code what an agent should not be deciding #562
  19. test(prose): a case starts where a session starts #563
  20. feat(prose-tests): walk on Sonnet, judge on Opus, escalate a failure #564
  21. test(prose): give the eight read-only cases something that can fail #565
  22. test(prose): only walks that can be observed, and checks that survive the trip #566
  23. fix(prose-tests): the verdict names only the model the record names #567
  24. test(prose): discovery, walked to the point where work first exists #568
  25. fix(prose-tests): the asserter judges which of prose or walker was at fault #569
  26. docs(conventions): a step whose reference routes every exit still signposts #570
  27. test(prose): discovery's epic arm, to the same durability boundary #571

…lures

A full-gate run went red on ENOTEMPTY inside an afterEach rmSync, in a
suite unrelated to what was being changed, while four background agents
churned temp directories. It passed 3/3 in isolation: the engine spawns
git subprocesses, git drops lock and temp files a moment after the
command returns, and a recursive remove that has just emptied a
directory then fails to remove it.

Every recursive teardown removal in tests/scripts now passes
maxRetries/retryDelay, as the prose-test world builder already did. 64
call sites across 25 suites; no behavioural change, and a false red on
the gate costs more than the churn.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This was referenced Jul 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant