Repository navigation
Release v1.0.1
Patch Changes
- #161 Thanks @Hazzng! - Fence stale sandbox writers with durable epochs: script scopes pin
sandboxes.versionunder the advisory lock and every composite mutation conditionally advances it, so a writer whose lease lapsed before its first write now fails withESTALEinstead of silently overwriting a live writer's changes. Sandbox deletion persists a tombstone epoch so ID reuse cannot reset the fence. - #208 Thanks @Hazzng! - Stop holding a Postgres transaction open across a user's bash script (#166): metadata mutations are buffered in memory and flushed in one short advisory-locked transaction at scope end, while file bytes still commit eagerly via
commitBlob. Idle-in-transaction age dropped from 2.98s to 0s for a 3s script, and from 8.60s to 0.50s across 40 concurrent writers on a real Neon pooler. Constraint errors now surface at flush instead of at the failing command, a lost cross-replica race fails the loser withESTALE, and hitting the buffer cap fails the whole script closed withESCRIPTBUFFER(413). Toggle viaSCRIPT_TX_BUFFERED(default true). - #201 Thanks @Hazzng! - Stop
POST /writeFilesand thefilesmap on sandbox creation from accepting a wider batch than a single write allows:MAX_BULK_WRITE_BYTESnow defaults toMAX_FILE_WRITE_BYTESinstead of its own 128 MiB default, both routes share one set of limits, and/writeFilesgains a streamed body cap (callers currently sending 50-128 MiB batches now get413).GET /readyzalso exposes the event-loop lag histogram, including a newp999Msand anevent_loop_stallcritical log line aboveEVENT_LOOP_STALL_THRESHOLD_MS. - #202 Thanks @Hazzng! - Cap the file size a sandbox
execscript may read whole or produce with one write asMAX_EXEC_FILE_BYTES(default 8 MiB,EFBIGto 413): just-bash's text utilities rebuild strings synchronously on the main thread, and past roughly 2s a stall starts timing out other tenants' in-flight Redis commands. Applies only insidebash.exec; the file is still retrievable whole viaGET .../files/{path}. Holds per file, not per script — a pipeline likecat a b | wc -ccan still exceed it; the structural fix is movingbash.execoff the main thread (#198). - #194 Thanks @Hazzng! - Stop a Postgres connection dying mid-transaction from crashing the replica:
postgres.jsthrows a fatal, uncaughtTypeErrorfrom a baresetImmediatewhen a reaped backend's buffered write flushes to a nulled socket.PG_DRIVER_FAULT_GUARD(default true) recognizes only that exact stack frame, logsdriver_socket_fault, and fails the stuck DB awaits withEDRIVERFAULTto 503 after a grace window (5s) instead of taking every other in-flight request down with it. A condemned script scope can no longer commit partial work, and boot migrations get the same guard. - #195 Thanks @Hazzng! - Repair a structurally invalid Redis version key (
WRONGTYPE, non-integerSET) in place instead of returning503 ECOHERENCEon every write to that sandbox forever. The key resets to the current epoch in milliseconds — never 1, which could equal a warm replica'slastSeenVersionand mask staleness — except the F7DESTROYEDtombstone, which is left alone. - #196 Thanks @Hazzng! - Require
maxmemory-policy allkeys-lru(orallkeys-lfu) on the Redis backing the blob cache and warn at boot when it isn't set. Redis's defaultnoevictionmakes a full instance a permanent outage — writes refused forever, since a 24h blob-cache TTL doesn't age out fast enough to recover (measured 97.6% 5xx, no recovery). The boot check (redis_eviction_policy_unsafe, critical) never fails startup and is skipped when the data client carries no data-plane state. - #204 Thanks @Hazzng! - Put every filesystem mutation behind the epoch fence and make all of them advance
sandboxes.version, not just the four composite writes. Fourteen other call sites (bulkIngest,mkdir -p,rm -r,cp,cp -r,link,symlink,chmod,utimes, non-composite fallbacks) previously left the counter untouched, so a live writer using only those left a stale peer's pin still matching.cp,chmod,ln,touch,mkdir -p,rm -rand ingest can now fail with409 ESTALE, which they never could before. - #179 Thanks @Hazzng! - Stop leaking raw driver/SQLSTATE error codes (
ECONNRESET, bare codes like53300) to clients: the allowlist that already redacted error messages now also redacts thecodefield, shared by the global error handler and the SSE error frame. Connection-class SQLSTATEs (08xxx,53300,53400,57P03) now map to a retryable503 EUNAVAILABLEinstead of500. - #209 Thanks @Hazzng! - Add a
retryableboolean to every error body (and the SSEerrorframe) so clients can tell an applied-but-unacknowledged write from one that never landed.ELOCKLOST_APPLIED(503, not retryable) now covers a lease lost after commit,ECOHERENCE_UNAPPLIEDcovers a rolled-back turn, and a read-only request no longer inherits a previous turn's stranded version-publish failure. The guarantee is per-transaction: multi-step routes like sandbox creation can commit an earlier step before a later one fails retryably. - #183 Thanks @Hazzng! - Lower the default exec-lock acquire timeout from 300s to 75s (lease plus ~15s reap margin, so the 503 reaches the client before typical ingress timeouts sever the connection), and refuse to boot when it's below
REDIS_EXEC_LOCK_LEASE_MSorREDIS_RWLOCK_READER_LEASE_MS, since a shorter window turns crashed-holder recovery into a permanent 503. - #182 Thanks @Hazzng! - Abort a running script when the client of
POST /v1/sandboxes/:id/exec-syncdisconnects, instead of holding the sandbox's exclusive exec lock for the rest of its timeout — the route never wired upc.req.raw.signal, unlike its SSE and batch siblings. Work already committed before the abort stays committed. - #178 Thanks @Hazzng! - Make the startup migration runner safe under transaction-mode connection pooling: the whole run is now one transaction opened with
pg_advisory_xact_lockas its first statement, since a session-scoped lock taken outside a transaction doesn't hold across a pooler reassigning connections per transaction (a second booter previously acquired the lock in 267-335ms instead of waiting — mutual exclusion silently wasn't holding). Migrations are now atomic as a side effect.DATABASE_DIRECT_URLis no longer needed by the server, only bydrizzle-kit, and its deployment secret is removed. - #183 Thanks @Hazzng! - Make
GET /v1/sandboxes/:id/files/*andGET /v1/sandboxes/:id/treetake the shared session lock instead of the exclusive write lock, so concurrent reads of one sandbox run in parallel instead of serializing like writes (a burst of 32 concurrentGET /treerequests dropped from ~150ms to 5ms). A GET still waits behind an in-flight writer, unchanged. - #207 Thanks @Hazzng! - Split the Redis connection by role (
control: locks, version counter, session state;data: blob cache, path snapshot) so multi-MiB blob writes can no longer head-of-line block latency-critical lock/version commands, scope the circuit breaker per role, cap in-flight blob-cache backfills (drop rather than queue over cap), and log every breaker transition. On a 6s Redis pause, 5xx fell from 90.0% to 54.3% on one shared instance, and to 0% for data-plane-only outages onceREDIS_DATA_URLpoints at a separate Redis. - #190 Thanks @Hazzng! - Give four test apps' hand-rolled
onErrorproduction's actual error contract instead of a copy of the pre-#174 leaky one, plus a source scan that fails if the hand-rolled fallback reappears. Test-only; no shipped behavior changes. - #210 Thanks @Hazzng! - Make the vitest excludes path-independent (
**/comparison/**,**/.claude/**) so a git worktree checked out inside the repo — where Claude Code's agent isolation puts them — no longer contributes a duplicatesrc/suite and a failingcomparison/copy topnpm test:unit. Tooling only. - #205 Thanks @Hazzng! - Make the integration suites actually run: 29 tests across seven files that silently skipped everywhere because nothing said
REDIS_URLwas required now execute, 5sql-fstests that failed against a real database because the epoch fence (#161) madesetSandboxContextWithLockrequire a real sandbox row are fixed, andbeforeAllmigration races between files were replaced with a schema check plus serial file execution.
Image: ghcr.io/hazzng/sql-fs:v1.0.1