Skip to content

Release v1.0.1

Latest

Choose a tag to compare

@github-actions github-actions released this 20 Sep 08:06
b814ca8

Release v1.0.1

Patch Changes

  • #161 Thanks @Hazzng! - Fence stale sandbox writers with durable epochs: script scopes pin sandboxes.version under the advisory lock and every composite mutation conditionally advances it, so a writer whose lease lapsed before its first write now fails with ESTALE instead of silently overwriting a live writer's changes. Sandbox deletion persists a tombstone epoch so ID reuse cannot reset the fence.
  • #208 Thanks @Hazzng! - Stop holding a Postgres transaction open across a user's bash script (#166): metadata mutations are buffered in memory and flushed in one short advisory-locked transaction at scope end, while file bytes still commit eagerly via commitBlob. Idle-in-transaction age dropped from 2.98s to 0s for a 3s script, and from 8.60s to 0.50s across 40 concurrent writers on a real Neon pooler. Constraint errors now surface at flush instead of at the failing command, a lost cross-replica race fails the loser with ESTALE, and hitting the buffer cap fails the whole script closed with ESCRIPTBUFFER (413). Toggle via SCRIPT_TX_BUFFERED (default true).
  • #201 Thanks @Hazzng! - Stop POST /writeFiles and the files map on sandbox creation from accepting a wider batch than a single write allows: MAX_BULK_WRITE_BYTES now defaults to MAX_FILE_WRITE_BYTES instead of its own 128 MiB default, both routes share one set of limits, and /writeFiles gains a streamed body cap (callers currently sending 50-128 MiB batches now get 413). GET /readyz also exposes the event-loop lag histogram, including a new p999Ms and an event_loop_stall critical log line above EVENT_LOOP_STALL_THRESHOLD_MS.
  • #202 Thanks @Hazzng! - Cap the file size a sandbox exec script may read whole or produce with one write as MAX_EXEC_FILE_BYTES (default 8 MiB, EFBIG to 413): just-bash's text utilities rebuild strings synchronously on the main thread, and past roughly 2s a stall starts timing out other tenants' in-flight Redis commands. Applies only inside bash.exec; the file is still retrievable whole via GET .../files/{path}. Holds per file, not per script — a pipeline like cat a b | wc -c can still exceed it; the structural fix is moving bash.exec off the main thread (#198).
  • #194 Thanks @Hazzng! - Stop a Postgres connection dying mid-transaction from crashing the replica: postgres.js throws a fatal, uncaught TypeError from a bare setImmediate when a reaped backend's buffered write flushes to a nulled socket. PG_DRIVER_FAULT_GUARD (default true) recognizes only that exact stack frame, logs driver_socket_fault, and fails the stuck DB awaits with EDRIVERFAULT to 503 after a grace window (5s) instead of taking every other in-flight request down with it. A condemned script scope can no longer commit partial work, and boot migrations get the same guard.
  • #195 Thanks @Hazzng! - Repair a structurally invalid Redis version key (WRONGTYPE, non-integer SET) in place instead of returning 503 ECOHERENCE on every write to that sandbox forever. The key resets to the current epoch in milliseconds — never 1, which could equal a warm replica's lastSeenVersion and mask staleness — except the F7 DESTROYED tombstone, which is left alone.
  • #196 Thanks @Hazzng! - Require maxmemory-policy allkeys-lru (or allkeys-lfu) on the Redis backing the blob cache and warn at boot when it isn't set. Redis's default noeviction makes a full instance a permanent outage — writes refused forever, since a 24h blob-cache TTL doesn't age out fast enough to recover (measured 97.6% 5xx, no recovery). The boot check (redis_eviction_policy_unsafe, critical) never fails startup and is skipped when the data client carries no data-plane state.
  • #204 Thanks @Hazzng! - Put every filesystem mutation behind the epoch fence and make all of them advance sandboxes.version, not just the four composite writes. Fourteen other call sites (bulkIngest, mkdir -p, rm -r, cp, cp -r, link, symlink, chmod, utimes, non-composite fallbacks) previously left the counter untouched, so a live writer using only those left a stale peer's pin still matching. cp, chmod, ln, touch, mkdir -p, rm -r and ingest can now fail with 409 ESTALE, which they never could before.
  • #179 Thanks @Hazzng! - Stop leaking raw driver/SQLSTATE error codes (ECONNRESET, bare codes like 53300) to clients: the allowlist that already redacted error messages now also redacts the code field, shared by the global error handler and the SSE error frame. Connection-class SQLSTATEs (08xxx, 53300, 53400, 57P03) now map to a retryable 503 EUNAVAILABLE instead of 500.
  • #209 Thanks @Hazzng! - Add a retryable boolean to every error body (and the SSE error frame) so clients can tell an applied-but-unacknowledged write from one that never landed. ELOCKLOST_APPLIED (503, not retryable) now covers a lease lost after commit, ECOHERENCE_UNAPPLIED covers a rolled-back turn, and a read-only request no longer inherits a previous turn's stranded version-publish failure. The guarantee is per-transaction: multi-step routes like sandbox creation can commit an earlier step before a later one fails retryably.
  • #183 Thanks @Hazzng! - Lower the default exec-lock acquire timeout from 300s to 75s (lease plus ~15s reap margin, so the 503 reaches the client before typical ingress timeouts sever the connection), and refuse to boot when it's below REDIS_EXEC_LOCK_LEASE_MS or REDIS_RWLOCK_READER_LEASE_MS, since a shorter window turns crashed-holder recovery into a permanent 503.
  • #182 Thanks @Hazzng! - Abort a running script when the client of POST /v1/sandboxes/:id/exec-sync disconnects, instead of holding the sandbox's exclusive exec lock for the rest of its timeout — the route never wired up c.req.raw.signal, unlike its SSE and batch siblings. Work already committed before the abort stays committed.
  • #178 Thanks @Hazzng! - Make the startup migration runner safe under transaction-mode connection pooling: the whole run is now one transaction opened with pg_advisory_xact_lock as its first statement, since a session-scoped lock taken outside a transaction doesn't hold across a pooler reassigning connections per transaction (a second booter previously acquired the lock in 267-335ms instead of waiting — mutual exclusion silently wasn't holding). Migrations are now atomic as a side effect. DATABASE_DIRECT_URL is no longer needed by the server, only by drizzle-kit, and its deployment secret is removed.
  • #183 Thanks @Hazzng! - Make GET /v1/sandboxes/:id/files/* and GET /v1/sandboxes/:id/tree take the shared session lock instead of the exclusive write lock, so concurrent reads of one sandbox run in parallel instead of serializing like writes (a burst of 32 concurrent GET /tree requests dropped from ~150ms to 5ms). A GET still waits behind an in-flight writer, unchanged.
  • #207 Thanks @Hazzng! - Split the Redis connection by role (control: locks, version counter, session state; data: blob cache, path snapshot) so multi-MiB blob writes can no longer head-of-line block latency-critical lock/version commands, scope the circuit breaker per role, cap in-flight blob-cache backfills (drop rather than queue over cap), and log every breaker transition. On a 6s Redis pause, 5xx fell from 90.0% to 54.3% on one shared instance, and to 0% for data-plane-only outages once REDIS_DATA_URL points at a separate Redis.
  • #190 Thanks @Hazzng! - Give four test apps' hand-rolled onError production's actual error contract instead of a copy of the pre-#174 leaky one, plus a source scan that fails if the hand-rolled fallback reappears. Test-only; no shipped behavior changes.
  • #210 Thanks @Hazzng! - Make the vitest excludes path-independent (**/comparison/**, **/.claude/**) so a git worktree checked out inside the repo — where Claude Code's agent isolation puts them — no longer contributes a duplicate src/ suite and a failing comparison/ copy to pnpm test:unit. Tooling only.
  • #205 Thanks @Hazzng! - Make the integration suites actually run: 29 tests across seven files that silently skipped everywhere because nothing said REDIS_URL was required now execute, 5 sql-fs tests that failed against a real database because the epoch fence (#161) made setSandboxContextWithLock require a real sandbox row are fixed, and beforeAll migration races between files were replaced with a schema check plus serial file execution.

Image: ghcr.io/hazzng/sql-fs:v1.0.1