Settle pending SQLite WASM requests before replacing workers - #6994
Conversation
🦋 Changeset detectedLatest commit: 08f0ec3 The changes in this PR will be included in the next version bump. This PR includes changesets to release 30 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
There was a problem hiding this comment.
Important
This PR's title and description promise an implementation fix, but the diff only adds the reproduction test. packages/sql/sqlite-wasm/src/SqliteClient.ts is unchanged, so worker replacement still leaves callbacks in the shared pending map suspended and vulnerable to id collisions. CI fails as expected. The reproduction test also needs a correction before it can serve as the regression spec.
Reviewed changes
- Added a focused reproduction test in
packages/sql/sqlite-wasm/test/Client.test.tsfor pending-query settlement on worker error. - The implementation fix to settle or remove callbacks in
packages/sql/sqlite-wasm/src/SqliteClient.tsis not present in this diff.
⚠️ Implementation fix missing
The PR body and title describe settling pending requests before replacing workers, but SqliteClient.ts was not modified. onError still only calls ScopedRef.set(connectionRef, makeConnection), leaving the shared pending map untouched. The new connection restarts currentId at 0, so replacement request ids can overwrite abandoned callbacks. The fix needs to fail or remove every callback owned by the replaced connection before the replacement proceeds.
Technical details
# Settle pending callbacks on worker error
## Affected sites
- `packages/sql/sqlite-wasm/src/SqliteClient.ts:348` — `onError` replaces the worker without touching `pending`
- `packages/sql/sqlite-wasm/src/SqliteClient.ts:309` — `pending` is shared across all connections created by `makeConnection`
- `packages/sql/sqlite-wasm/src/SqliteClient.ts:312` — each new connection resets `currentId` to `0`
## Required outcome
- When a worker errors, every callback currently in `pending` must be settled (e.g., with `Exit.fail(SqlError)`) or removed before the replacement connection starts.
- After the fix, a new connection's id `0` must not invoke a stale callback from the previous connection.
## Suggested approach
- Inside `onError`, before calling `ScopedRef.set`, copy the current entries from `pending`, clear the map, and invoke each resume with `Exit.fail(...)` describing the worker failure.
- Run the focused test from the PR body to confirm the reproduction now passes.⚠️ Test assertion does not match expected behavior
Effect.timeoutOption preserves failures from the source effect. When the implementation correctly fails the in-flight query with SqlError, Fiber.join(fiber).pipe(Effect.timeoutOption(...)) will fail with SqlError, so Option.isSome(result) is never reached. The test should capture the fiber's Exit and assert both that it settled and that it failed.
@v0 or keep the SHA fresh with Dependabot | Fix all ➔ | Fix 👍s ➔ | View workflow run | Using Kimi K2 (free via Pullfrog for OSS) | 𝕏
# Conflicts: # packages/sql/sqlite-wasm/test/Client.test.ts
There was a problem hiding this comment.
ℹ️ No critical issues — one correctness observation inline.
Reviewed changes
- Moved
pendingfrom the outermakescope intomakeConnectionso each worker/connection owns its own callback map. - Updated the worker
errorhandler to fail every pending request withSqlErrorbefore replacing the connection. - Corrected the reproduction test to assert both that the in-flight query settles and that it fails.
- Added a patch changeset for
@effect/sql-sqlite-wasm.
@v0 or keep the SHA fresh with Dependabot | Fix all ➔ | Fix 👍s ➔ | View workflow run | Using Kimi K2 (free via Pullfrog for OSS) | 𝕏
| function onError(cause: Event) { | ||
| const exit = Exit.fail( | ||
| new SqlError({ reason: classifyError(cause, "SQLite WASM worker failed", "worker") }) | ||
| ) | ||
| const requests = Array.from(pending.values()) | ||
| pending.clear() | ||
| for (const resume of requests) { | ||
| resume(exit) |
There was a problem hiding this comment.
The snapshot-then-clear ordering leaves a small race: if a message response completes a callback between Array.from(pending.values()) and pending.clear(), onError will still call that resume again. Effect.callback resumes are not idempotent, so a double-resume could defect.
Consider iterating entries and deleting each from the map before calling resume, or clearing the map first and then resuming the captured callbacks. Either removes the window.
| function onError(cause: Event) { | |
| const exit = Exit.fail( | |
| new SqlError({ reason: classifyError(cause, "SQLite WASM worker failed", "worker") }) | |
| ) | |
| const requests = Array.from(pending.values()) | |
| pending.clear() | |
| for (const resume of requests) { | |
| resume(exit) | |
| function onError(cause: Event) { | |
| const exit = Exit.fail( | |
| new SqlError({ reason: classifyError(cause, "SQLite WASM worker failed", "worker") }) | |
| ) | |
| const requests = Array.from(pending.values()) | |
| pending.clear() | |
| for (const resume of requests) { | |
| resume(exit) | |
| } | |
| Effect.runFork(ScopedRef.set(connectionRef, makeConnection)) | |
| } |
Bundle Size AnalysisGenerated from PR build output; treat the content below as untrusted.
|

Summary
Replacing a failed worker leaves existing requests suspended forever and allows replacement request IDs to overwrite abandoned callbacks.
Important
This PR starts with focused failing reproduction tests. Add the implementation fix to this same branch; CI is expected to fail until that fix is included.
Worker replacement abandons in-flight requests
Module:
sqlite-wasm/SqliteClientAudit ID:
sql-adapters-sw-1Severity / confidence: high / high
What happens
Replacing a failed worker leaves existing requests suspended forever and allows replacement request IDs to overwrite abandoned callbacks.
Why it happens
Replacement neither settles nor removes callbacks in the shared pending map, while each new connection restarts its request ID counter at zero.
Expected behavior
A worker-backed client restart must settle all operations owned by the replaced scoped connection.
Relevant implementation
These links and excerpts are pinned to audit base
c9b56ab507f224426ee8388dc450da447ec4715f.packages/sql/sqlite-wasm/src/SqliteClient.ts:309packages/sql/sqlite-wasm/src/SqliteClient.ts:348-353packages/sql/sqlite-wasm/src/SqliteClient.ts:369-373View problematic code at
packages/sql/sqlite-wasm/src/SqliteClient.ts:309View exact lines on GitHub
View problematic code at
packages/sql/sqlite-wasm/src/SqliteClient.ts:348-353View exact lines on GitHub
View problematic code at
packages/sql/sqlite-wasm/src/SqliteClient.ts:369-373View exact lines on GitHub
Reproduction
pnpm test --run packages/sql/sqlite-wasm/test/Client.test.tsObserved failure: The in-flight query remained pending after worker replacement.
Implementation handoff
The initial reproduction tests on this branch are the regression specification for the implementation fix that should follow in this PR.
pnpm test --run packages/sql/sqlite-wasm/test/Client.test.tsAudit provenance
c9b56ab507f224426ee8388dc450da447ec4715f8f9499f562729f5f7b08d8bcc4db86b4aeff8a21sql-adapters-sw-1Closes EFF-431