You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Two silent-failure fixes for long-lived runtimes (readiness gate; composition write-back), with tests
#3509
Two silent-failure fixes we have been carrying downstream, offered back with tests. Issues are off and this repository does not accept pull requests, so posting here — branches are linked below and can be cherry-picked directly.
Both were found in a long-lived deployment (an SDK runtime kept alive across many requests rather than one run per process). Both fail silently: nothing throws, nothing logs, and the symptom appears far from the cause.
1. The SDK readiness gate covers only initialize
packages/sdk/server/src/index.ts waits for the loader when the method is initialize. That covers the handshake, but the plugin tree can move again afterwards — a session-scoped composition swap remounts rows while the runtime keeps serving.
A request answered in that window composes its agent from a half-registered tool registry: the model receives whichever tools happen to be up at that instant. No error is raised anywhere. The visible symptom is a model that appears to have "forgotten" most of its tools for one run, which is very hard to trace back.
Fix: wait on every method except shutdown (a runtime asked to stop must not be held open by whatever is still loading).
The readiness test keeps its Loader fixture and tightens the assertion: while the delayed entry is applying, no frame is answered at all — not just initialize. That also removes a scheduler-timing assumption, since the fixture already proves the delayed entry is mid-apply before the requests are written.
2. A failed load can write an included composition file back as []
This one destroys user configuration on disk. A single bad row in an included subtree could leave the composition file containing [], so the next start loaded an empty tree.
Two independent gaps combine to produce it — and each one alone is enough to keep the file safe, which is why the regression test only reproduces with both removed:
Include publishes before its commit point.data is assigned exactly when an apply commits, so while it is undefined the tree is not a representation of the file at all. A transactional apply failure leaves the root group rolled back to its previous state, which for a first apply is EMPTY. A row disposed during that rollback persists itself and flushes that empty tree over the real file.
Loader's case-6 check looks at one entry. A nested group's rows are disposed by the GROUP entry's fiber and never pass through their own Entry._dispose, so rolling back a failed group update looks exactly like every child calling ctx.fiber.dispose() — that is, like a runtime decision to leave, which the self-modification feature is supposed to persist.
Fix: Include returns early from write() while data is undefined; Loader's case-6 walks the ancestor entry chain.
Adds packages/boot/app-boot/tests/config-writeback.spec.ts (5 cases): a failed first apply leaves the included file byte-identical, both for a nested-group sibling and for a top-level row; a rolled-back update never persists a live row as disabled; and runtime self-disposal still writes back disabled: true, at top level and inside a nested group.
Both branches are cut from dsh-v0.1.0-rc.8 and touch nothing else. Happy to reshape either one to fit your preferred structure if that helps — and if there is a better channel for this kind of report, please point me at it.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Two silent-failure fixes we have been carrying downstream, offered back with tests. Issues are off and this repository does not accept pull requests, so posting here — branches are linked below and can be cherry-picked directly.
Both were found in a long-lived deployment (an SDK runtime kept alive across many requests rather than one run per process). Both fail silently: nothing throws, nothing logs, and the symptom appears far from the cause.
1. The SDK readiness gate covers only
initializepackages/sdk/server/src/index.tswaits for the loader when the method isinitialize. That covers the handshake, but the plugin tree can move again afterwards — a session-scoped composition swap remounts rows while the runtime keeps serving.A request answered in that window composes its agent from a half-registered tool registry: the model receives whichever tools happen to be up at that instant. No error is raised anywhere. The visible symptom is a model that appears to have "forgotten" most of its tools for one run, which is very hard to trace back.
Fix: wait on every method except
shutdown(a runtime asked to stop must not be held open by whatever is still loading).The readiness test keeps its Loader fixture and tightens the assertion: while the delayed entry is applying, no frame is answered at all — not just
initialize. That also removes a scheduler-timing assumption, since the fixture already proves the delayed entry is mid-apply before the requests are written.Branch: https://github.com/jrsdhr/deepseek-harness/tree/fix/sdk-server-loader-readiness-every-request
2. A failed load can write an included composition file back as
[]This one destroys user configuration on disk. A single bad row in an included subtree could leave the composition file containing
[], so the next start loaded an empty tree.Two independent gaps combine to produce it — and each one alone is enough to keep the file safe, which is why the regression test only reproduces with both removed:
datais assigned exactly when an apply commits, so while it isundefinedthe tree is not a representation of the file at all. A transactional apply failure leaves the root group rolled back to its previous state, which for a first apply is EMPTY. A row disposed during that rollback persists itself and flushes that empty tree over the real file.Entry._dispose, so rolling back a failed group update looks exactly like every child callingctx.fiber.dispose()— that is, like a runtime decision to leave, which the self-modification feature is supposed to persist.Fix: Include returns early from
write()whiledatais undefined; Loader's case-6 walks the ancestor entry chain.Adds
packages/boot/app-boot/tests/config-writeback.spec.ts(5 cases): a failed first apply leaves the included file byte-identical, both for a nested-group sibling and for a top-level row; a rolled-back update never persists a live row asdisabled; and runtime self-disposal still writes backdisabled: true, at top level and inside a nested group.Branch: https://github.com/jrsdhr/deepseek-harness/tree/fix/vendor-writeback-only-committed-trees
Both branches are cut from
dsh-v0.1.0-rc.8and touch nothing else. Happy to reshape either one to fit your preferred structure if that helps — and if there is a better channel for this kind of report, please point me at it.All reactions