Repository navigation
Release v0.9.0
Release v0.9.0
Minor Changes
- #151 Thanks @Hazzng! - fix(lock): abort the exec on definitive lease loss before commit (F2-L1)
The distributed exec lock wrappers ran the critical section to completion and only
then checked the loss flag, so a writer whose lease lapsed mid-script still
committed its script-tx and bumped the version before throwingELOCKLOST— a
write that durably happened surfaced as an error, causing retrying agents to
double-apply.
The lock now wires its DEFINITIVE-loss signal (lease expiry / ownership taken —
not transient renew blips) into anAbortControllerthat is plumbed through to
bash.exec. On a definitive loss the in-flight exec is aborted, its script-tx
rolls back BEFORE any commit (noINCR), and the client receives a clean,
retryableELOCKLOST(now mapped to 503). Because just-bash treats an aborted
run as a resolved result rather than a rejection, the runtime explicitly rolls
back and re-raisesLockLostErrorwhen the lock-lost signal fired, instead of
committing the partial script. A plain timeout abort still commits (unchanged,
audit L7) — only the dedicated lock-lost signal triggers rollback.
This is Layer 1 of the F2 fix; the complete epoch/version fence is tracked
separately (#131). - #148 Thanks @Hazzng! - fix(lock): circuit-break Redis acquire to stop the 300s outage fuse (F5)
The distributed lock acquire loops conflated "lock busy" (contention) with "Redis
unreachable" (a thrown connection error): both retried untilacquireTimeoutMs
(default 300 s), so a Redis outage hung every exec/file op for ~5 minutes on an
otherwise-healthy Postgres.- New process-wide Redis circuit breaker (
src/redis/circuit-breaker.ts) wired
into the lock ACQUIRE paths only (distributed-rw-lock.tsshared/exclusive +
waitReadersDrained, legacydistributed-lock.ts). After K (=5) consecutive
connection-class failures it opens and acquire fast-fails 503 immediately; a
successful eval/PING closes it. Renew/release paths are untouched (they keep
tolerating transient errors to avoid dropping leases / leaking keys). - Separate short per-call error budget (
errorBudgetMs, default 4 s) that
advances only on thrown errors, so genuine contention still uses the full
acquireTimeoutMswindow. commandTimeout: 2000on the ioredis client so commands reject promptly during
an outage instead of queueing on the offline queue./readyznow PINGs Redis and returns 503 when Redis is configured but
unreachable.- Fixed the stale
ownership.tsdocstring (the readOnly path DOES take a shared
distributed lock).
- New process-wide Redis circuit breaker (
- #149 Thanks @Hazzng! - fix(blob): commit the CAS blob upsert in its own short, self-committing
transaction (own connection, no advisory lock) BEFORE the inode/dirent composite,
and run the composite without itsblob_insertCTE. This removes hot-blob
contention (F6): previously theON CONFLICT (sha256) DO UPDATE SET last_referenced_at = now()tuple lock on a deduplicated hot blob (empty file,
.gitkeep, common lockfiles) was held for the whole script, serializing
unrelated sandboxes within one tenant DB and risking pool-exhaustion → 503. The
touch stays unconditional so the GC grace window protects the freshly-committed
blob until its inode commits; the blob-gc REPEATABLE READ + 40001 re-adoption
handshake is preserved. Applies towriteFile,appendFile, andbulkIngest. - #150 Thanks @Hazzng! - Surface a retryable
ESESSIONCLOSING(HTTP 503) instead of a generic 500 when a request
loses the reaper-vs-straggler race (F9c). A request that captured a session reference just
before the idle/overBudget reaper marked itclosingcould run the pre-lock
ensureFreshCacheprobe against a Postgres pool being disconnected, producing an unmapped
error (e.g.PostgresDialect: not connected, which carries nocode) that defaulted to a 500. Both pre-lock probe sites (withSessionEntry,withSessionReadEntry) now re-check the
session state on probe failure and convert it into a clean, retryableESESSIONCLOSING
(already mapped to 503) so clients retry instead of seeing a non-retryable 500. No
concurrency-model change.
Image: ghcr.io/hazzng/sql-fs:v0.9.0