Skip to content

Releases: Hazzng/sql-fs

Release v1.0.1

Choose a tag to compare

@github-actions github-actions released this 20 Sep 08:06
b814ca8

Release v1.0.1

Patch Changes

  • #161 Thanks @Hazzng! - Fence stale sandbox writers with durable epochs: script scopes pin sandboxes.version under the advisory lock and every composite mutation conditionally advances it, so a writer whose lease lapsed before its first write now fails with ESTALE instead of silently overwriting a live writer's changes. Sandbox deletion persists a tombstone epoch so ID reuse cannot reset the fence.
  • #208 Thanks @Hazzng! - Stop holding a Postgres transaction open across a user's bash script (#166): metadata mutations are buffered in memory and flushed in one short advisory-locked transaction at scope end, while file bytes still commit eagerly via commitBlob. Idle-in-transaction age dropped from 2.98s to 0s for a 3s script, and from 8.60s to 0.50s across 40 concurrent writers on a real Neon pooler. Constraint errors now surface at flush instead of at the failing command, a lost cross-replica race fails the loser with ESTALE, and hitting the buffer cap fails the whole script closed with ESCRIPTBUFFER (413). Toggle via SCRIPT_TX_BUFFERED (default true).
  • #201 Thanks @Hazzng! - Stop POST /writeFiles and the files map on sandbox creation from accepting a wider batch than a single write allows: MAX_BULK_WRITE_BYTES now defaults to MAX_FILE_WRITE_BYTES instead of its own 128 MiB default, both routes share one set of limits, and /writeFiles gains a streamed body cap (callers currently sending 50-128 MiB batches now get 413). GET /readyz also exposes the event-loop lag histogram, including a new p999Ms and an event_loop_stall critical log line above EVENT_LOOP_STALL_THRESHOLD_MS.
  • #202 Thanks @Hazzng! - Cap the file size a sandbox exec script may read whole or produce with one write as MAX_EXEC_FILE_BYTES (default 8 MiB, EFBIG to 413): just-bash's text utilities rebuild strings synchronously on the main thread, and past roughly 2s a stall starts timing out other tenants' in-flight Redis commands. Applies only inside bash.exec; the file is still retrievable whole via GET .../files/{path}. Holds per file, not per script — a pipeline like cat a b | wc -c can still exceed it; the structural fix is moving bash.exec off the main thread (#198).
  • #194 Thanks @Hazzng! - Stop a Postgres connection dying mid-transaction from crashing the replica: postgres.js throws a fatal, uncaught TypeError from a bare setImmediate when a reaped backend's buffered write flushes to a nulled socket. PG_DRIVER_FAULT_GUARD (default true) recognizes only that exact stack frame, logs driver_socket_fault, and fails the stuck DB awaits with EDRIVERFAULT to 503 after a grace window (5s) instead of taking every other in-flight request down with it. A condemned script scope can no longer commit partial work, and boot migrations get the same guard.
  • #195 Thanks @Hazzng! - Repair a structurally invalid Redis version key (WRONGTYPE, non-integer SET) in place instead of returning 503 ECOHERENCE on every write to that sandbox forever. The key resets to the current epoch in milliseconds — never 1, which could equal a warm replica's lastSeenVersion and mask staleness — except the F7 DESTROYED tombstone, which is left alone.
  • #196 Thanks @Hazzng! - Require maxmemory-policy allkeys-lru (or allkeys-lfu) on the Redis backing the blob cache and warn at boot when it isn't set. Redis's default noeviction makes a full instance a permanent outage — writes refused forever, since a 24h blob-cache TTL doesn't age out fast enough to recover (measured 97.6% 5xx, no recovery). The boot check (redis_eviction_policy_unsafe, critical) never fails startup and is skipped when the data client carries no data-plane state.
  • #204 Thanks @Hazzng! - Put every filesystem mutation behind the epoch fence and make all of them advance sandboxes.version, not just the four composite writes. Fourteen other call sites (bulkIngest, mkdir -p, rm -r, cp, cp -r, link, symlink, chmod, utimes, non-composite fallbacks) previously left the counter untouched, so a live writer using only those left a stale peer's pin still matching. cp, chmod, ln, touch, mkdir -p, rm -r and ingest can now fail with 409 ESTALE, which they never could before.
  • #179 Thanks @Hazzng! - Stop leaking raw driver/SQLSTATE error codes (ECONNRESET, bare codes like 53300) to clients: the allowlist that already redacted error messages now also redacts the code field, shared by the global error handler and the SSE error frame. Connection-class SQLSTATEs (08xxx, 53300, 53400, 57P03) now map to a retryable 503 EUNAVAILABLE instead of 500.
  • #209 Thanks @Hazzng! - Add a retryable boolean to every error body (and the SSE error frame) so clients can tell an applied-but-unacknowledged write from one that never landed. ELOCKLOST_APPLIED (503, not retryable) now covers a lease lost after commit, ECOHERENCE_UNAPPLIED covers a rolled-back turn, and a read-only request no longer inherits a previous turn's stranded version-publish failure. The guarantee is per-transaction: multi-step routes like sandbox creation can commit an earlier step before a later one fails retryably.
  • #183 Thanks @Hazzng! - Lower the default exec-lock acquire timeout from 300s to 75s (lease plus ~15s reap margin, so the 503 reaches the client before typical ingress timeouts sever the connection), and refuse to boot when it's below REDIS_EXEC_LOCK_LEASE_MS or REDIS_RWLOCK_READER_LEASE_MS, since a shorter window turns crashed-holder recovery into a permanent 503.
  • #182 Thanks @Hazzng! - Abort a running script when the client of POST /v1/sandboxes/:id/exec-sync disconnects, instead of holding the sandbox's exclusive exec lock for the rest of its timeout — the route never wired up c.req.raw.signal, unlike its SSE and batch siblings. Work already committed before the abort stays committed.
  • #178 Thanks @Hazzng! - Make the startup migration runner safe under transaction-mode connection pooling: the whole run is now one transaction opened with pg_advisory_xact_lock as its first statement, since a session-scoped lock taken outside a transaction doesn't hold across a pooler reassigning connections per transaction (a second booter previously acquired the lock in 267-335ms instead of waiting — mutual exclusion silently wasn't holding). Migrations are now atomic as a side effect. DATABASE_DIRECT_URL is no longer needed by the server, only by drizzle-kit, and its deployment secret is removed.
  • #183 Thanks @Hazzng! - Make GET /v1/sandboxes/:id/files/* and GET /v1/sandboxes/:id/tree take the shared session lock instead of the exclusive write lock, so concurrent reads of one sandbox run in parallel instead of serializing like writes (a burst of 32 concurrent GET /tree requests dropped from ~150ms to 5ms). A GET still waits behind an in-flight writer, unchanged.
  • #207 Thanks @Hazzng! - Split the Redis connection by role (control: locks, version counter, session state; data: blob cache, path snapshot) so multi-MiB blob writes can no longer head-of-line block latency-critical lock/version commands, scope the circuit breaker per role, cap in-flight blob-cache backfills (drop rather than queue over cap), and log every breaker transition. On a 6s Redis pause, 5xx fell from 90.0% to 54.3% on one shared instance, and to 0% for data-plane-only outages once REDIS_DATA_URL points at a separate Redis.
  • #190 Thanks @Hazzng! - Give four test apps' hand-rolled onError production's actual error contract instead of a copy of the pre-#174 leaky one, plus a source scan that fails if the hand-rolled fallback reappears. Test-only; no shipped behavior changes.
  • #210 Thanks @Hazzng! - Make the vitest excludes path-independent (**/comparison/**, **/.claude/**) so a git worktree checked out inside the repo — where Claude Code's agent isolation puts them — no longer contributes a duplicate src/ suite and a failing comparison/ copy to pnpm test:unit. Tooling only.
  • #205 Thanks @Hazzng! - Make the integration suites actually run: 29 tests across seven files that silently skipped everywhere because nothing said REDIS_URL was required now execute, 5 sql-fs tests that failed against a real database because the epoch fence (#161) made setSandboxContextWithLock require a real sandbox row are fixed, and beforeAll migration races between files were replaced with a schema check plus serial file execution.

Image: ghcr.io/hazzng/sql-fs:v1.0.1

Release v1.0.0

Choose a tag to compare

@github-actions github-actions released this 18 Sep 13:37
093f1d5

Release v1.0.0

Major Changes

  • #158 Thanks @Hazzng! - Add a sandbox git command backed by just-git, export server GITHUB_TOKEN into sandbox GitHub-compatible Git/curl env, and let MCP-created sandboxes request network access for clone/fetch/push. A per-request env.GITHUB_TOKEN re-points git's HTTP credentials at that token, refusing plaintext http:// remotes, validating each redirect hop by hand, dropping credentials on cross-origin redirects, and rewriting a redirected POST to GET on 301/302/303 the way fetch does — so a push packfile is never replayed at a host it was merely forwarded to.

Minor Changes

  • #162 Thanks @Hazzng! - Add file access to MCP — file_read, file_write, file_edit — plus PATCH /v1/sandboxes/:id/files/*path for exact-string edits.
    file_edit requires oldString to match exactly once (replaceAll opts into multiple), rejecting an ambiguous match as EDIT_NOT_UNIQUE rather than guessing; it runs in one script-tx scope on Postgres so a concurrent reader never sees the file mid-edit, refuses non-UTF-8 content and unpaired surrogates, and preserves file mode and a leading BOM. file_read pages via offset/limit/byteOffset, bounding both the file it opens (MAX_MCP_READ_FILE_BYTES) and the reply it returns (MAX_MCP_READ_RESPONSE_BYTES). file_write creates parent directories and refuses to clobber a directory on every backend. HTTP and MCP share one implementation in src/api/lib/file-ops.ts.

Patch Changes

  • #162 Thanks @Hazzng! - Document container sizing for the write cap: a single write costs roughly 7x the file size in external memory above steady state, so a 50 MiB write needs on the order of 700 MB of headroom (a 512 MiB container OOM'd on one request; 768 MiB survived).
  • #162 Thanks @Hazzng! - Apply the single-file write limit to each entry of a bulk write, not just the combined total, so one oversized entry in POST /writeFiles can no longer land a blob the contentCache cannot hold.
  • #162 Thanks @Hazzng! - Roll back a PATCH edit or bulk write when the distributed exec lock is definitively lost mid-request, instead of committing it and reporting a retryable ELOCKLOST. Both routes now go through runInScriptTx, which checks for lock loss before the commit.
  • #162 Thanks @Hazzng! - Reject an edit whose oldString/newString carries an unpaired surrogate as EDIT_LONE_SURROGATE, and enforce the write limit on the encoded result rather than on a projected size that assumed one match encodes to the bytes it replaces.
  • #162 Thanks @Hazzng! - Clean up the destination of a git clone that fails partway (e.g. on a symlink, since allowSymlinks defaults to false), instead of leaving a half-checked-out tree whose complete index made every missing file look like a staged deletion — an agent then following up with git add -A && git commit && git push turned that into a real destructive commit.
  • #162 Thanks @Hazzng! - Refuse to replay a git request body across origins on a 307/308 redirect, since both preserve the method and body — for git, the packfile being pushed — and could otherwise forward a whole push to an attacker-chosen host.
  • #162 Thanks @Hazzng! - Budget the whole file_read reply, not just the content string, against MAX_MCP_READ_RESPONSE_BYTES, and normalize the path MCP tools echo back instead of only prefixing a slash.
  • #162 Thanks @Hazzng! - Size a file_read page against the reply as the transport actually serializes it (which re-escapes once more), and stop calling split("\n") on the whole file just to count lines — both scanning-based fixes remove a real-world 170 KiB overshoot and a multi-million-element allocation.
  • #162 Thanks @Hazzng! - Enforce the file-write limit on PUT /v1/sandboxes/:id/files/*path as the body streams, instead of trusting Content-Length and buffering the whole request first.
  • #162 Thanks @Hazzng! - Keep a leading UTF-8 BOM in what file_read and fs_export return, matching the ignoreBOM decoding editFile already used, so content and stat.size round-trip byte for byte.
  • #162 Thanks @Hazzng! - Return RESPONSE_BUDGET_TOO_SMALL instead of a reply that can never fit any content when MAX_MCP_READ_RESPONSE_BYTES is configured below the size of the response envelope, which previously left a client resuming a page in an infinite loop.
  • #162 Thanks @Hazzng! - Keep file_read paging identical to the split/join it replaced when a file ends in a newline, instead of returning a trailing newline the old implementation would have dropped.
  • #162 Thanks @Hazzng! - Ignore a relative PWD when recording a session's working directory instead of rooting it into a path that never existed.
  • #162 Thanks @Hazzng! - Build a replaceAll edit by assembling flushed chunks instead of split(oldString).join(newString), cutting peak RSS on a worst-case near-limit file from 1184 MB to 357 MB with no regression on ordinary single-match edits.
  • #162 Thanks @Hazzng! - Stop an abort that races script-tx opening from rejecting a promise with no listener, which was fatal under Node's default --unhandled-rejections=throw.
  • #162 Thanks @Hazzng! - Stop a late-arriving transaction open from adopting into a finished scope (each open now carries a generation checked before adopting), and refuse cache-served reads (stat, readFile, readdir, exists, getAllPaths) once a script-tx is lost.
  • #162 Thanks @Hazzng! - Fail every remaining operation in a script scope once its transaction's connection is lost, instead of letting a write silently self-commit outside the scope on a reconnected-but-transactionless connection — previously a 600-file bulk write could answer HTTP 500 with 599 of them durable.
  • #162 Thanks @Hazzng! - Write a whole file through one shared transactional path (writeFileAtPath) on both the MCP file_write tool and PUT /v1/sandboxes/:id/files/*, so parent directories and file content commit together and PUT matches MCP in refusing to clobber a directory (400 EISDIR).
  • #162 Thanks @Hazzng! - Default the single-file write limit to the contentCache cap (50 MiB) instead of 64 MiB — load testing found memory cost doubles just past the cache cap, so the old default pinned 256 MB per warm session for a single large read.

Image: ghcr.io/hazzng/sql-fs:v1.0.0

Release v0.10.0

Choose a tag to compare

@github-actions github-actions released this 20 Jun 09:18
541eb7a

Release v0.10.0

Minor Changes

  • #157 Thanks @Hazzng! - feat(observability): event-loop lag monitoring for the Redis leases (F8).
    The exec-lock writer lease, the RW-lock writer flag, and the RW-lock reader ZSET
    scores are all kept alive by setTimeout heartbeats that silently assume timers
    fire on schedule. A long event-loop stall (a V8 GC pause or a pathological
    synchronous bash stretch) can fire a renewal past the lease, voiding it — Lock 3
    keeps Postgres consistent, so this was always an observability gap, not a
    correctness bug, but nothing measured it.
    New src/api/event-loop-monitor.ts (purely observational, no behavior change):
    • A perf_hooks.monitorEventLoopDelay histogram started at boot, sampled every
      EVENT_LOOP_MONITOR_INTERVAL_MS (default 10s) and logged as
      event:"event_loop_lag" (p50Ms/p99Ms/maxMs/meanMs), then reset.
    • Per-heartbeat gap measurement wired into all three lease sites: each heartbeat
      reports actual-minus-expected fire time as event:"heartbeat_gap" at
      severity:"warn" (gap > renewMs) or "critical" (gap > leaseMs), tagged with
      the lock kind (exec/rw-writer/rw-reader) and key.
      Alert thresholds are documented in DEVELOPER.md ("Lock observability"). End-to-end
      smoke tests reproduce a >lease stall on each lease and assert the critical
      heartbeat_gap fires (with a no-stall control proving no false positives).
  • #156 Thanks @Hazzng! - Heal stranded cross-replica version publishes after a Redis INCR failure (F3): a background drainer and reap-time best-effort publish flush the bump even if no further client traffic arrives or the session is idle-evicted.
  • #155 Thanks @Hazzng! - fix(session): destroy now reaches warm sessions on other replicas (F7).
    Destroying a sandbox on one replica previously left warm sessions on other
    replicas serving ghost state: a written session would reload a deleted tree
    into an empty pathCache (surfacing as a non-zero exit + garbage stderr inside an
    HTTP 200 exec), and a never-written session would never reload at all because
    the deleted version key read as 0 and matched its lastSeenVersion === 0.
    Two layered fixes:
    • Primary (Redis-independent): SqlFs.reload() now detects a zero-row
      loadAllPaths — which for a live sandbox always returns at least its root dir
      — and throws a typed ESANDBOXGONE instead of installing an empty pathCache.
      The session manager catches it, tears the stale warm session down (drops it
      from the pool and disconnects the per-session Postgres pool), and surfaces a
      clean ENOENT → 404.
    • Secondary (tombstone): destroy now writes a distinct DESTROYED sentinel to
      the version key (with the version-key TTL) instead of deleting it.
      ensureFreshCache recognises the sentinel before the numeric parse and tears
      the session down — covering the never-written variant. Re-creating a
      tombstoned sandbox clears the sentinel and starts cleanly at version 0.
  • #153 Thanks @Hazzng! - fix(lock): add bounded jitter + tunable retry to the distributed acquire loops (F9d, #141)
    The distributed exec lock and RW lock polled Redis on a flat acquireRetryMs
    (default 50 ms) interval, leaving competing replicas phase-aligned so a
    cross-replica writer could be repeatedly passed over (bounded by
    acquireTimeoutMs, then 503). Every acquire/drain poll now sleeps a jittered
    retryMs/2 + random()*retryMs/2 (range [retryMs/2, retryMs]) to
    de-synchronize pollers. The retry interval is now configurable via
    REDIS_EXEC_LOCK_ACQUIRE_RETRY_MS (previously hardcoded — server.ts omitted
    it). Circuit-breaker / error-budget behavior is unchanged. The FIFO ZSET ticket
    queue is deferred as a follow-up.
  • #154 Thanks @Hazzng! - perf(cache): O(1) pathCache byte accounting to avoid full-map scans (F9e, #142)
    SqlFs now maintains an incremental #pathCacheBytes counter, adjusted on
    every pathCache set/delete and reset on reload()/ready(), and exposes
    getPathCacheBytes(). SessionManager's path-cache memory budget calls it
    instead of re-walking the entire pathCache (Σ path.length + 100) on every
    dirty exec. The value equals the previous full-walk exactly. Falls back to
    the full walk for backends that do not expose the counter.
    The #childrenByParent children index (part B of #142) is deferred to a
    follow-up; it is benchmark-gated and (A) delivers the higher-value, lower-risk
    win without touching readdir correctness.

Image: ghcr.io/hazzng/sql-fs:v0.10.0

Release v0.9.0

Choose a tag to compare

@github-actions github-actions released this 13 Jun 08:58
bf56de6

Release v0.9.0

Minor Changes

  • #151 Thanks @Hazzng! - fix(lock): abort the exec on definitive lease loss before commit (F2-L1)
    The distributed exec lock wrappers ran the critical section to completion and only
    then checked the loss flag, so a writer whose lease lapsed mid-script still
    committed its script-tx and bumped the version before throwing ELOCKLOST — a
    write that durably happened surfaced as an error, causing retrying agents to
    double-apply.
    The lock now wires its DEFINITIVE-loss signal (lease expiry / ownership taken —
    not transient renew blips) into an AbortController that is plumbed through to
    bash.exec. On a definitive loss the in-flight exec is aborted, its script-tx
    rolls back BEFORE any commit (no INCR), and the client receives a clean,
    retryable ELOCKLOST (now mapped to 503). Because just-bash treats an aborted
    run as a resolved result rather than a rejection, the runtime explicitly rolls
    back and re-raises LockLostError when the lock-lost signal fired, instead of
    committing the partial script. A plain timeout abort still commits (unchanged,
    audit L7) — only the dedicated lock-lost signal triggers rollback.
    This is Layer 1 of the F2 fix; the complete epoch/version fence is tracked
    separately (#131).
  • #148 Thanks @Hazzng! - fix(lock): circuit-break Redis acquire to stop the 300s outage fuse (F5)
    The distributed lock acquire loops conflated "lock busy" (contention) with "Redis
    unreachable" (a thrown connection error): both retried until acquireTimeoutMs
    (default 300 s), so a Redis outage hung every exec/file op for ~5 minutes on an
    otherwise-healthy Postgres.
    • New process-wide Redis circuit breaker (src/redis/circuit-breaker.ts) wired
      into the lock ACQUIRE paths only (distributed-rw-lock.ts shared/exclusive +
      waitReadersDrained, legacy distributed-lock.ts). After K (=5) consecutive
      connection-class failures it opens and acquire fast-fails 503 immediately; a
      successful eval/PING closes it. Renew/release paths are untouched (they keep
      tolerating transient errors to avoid dropping leases / leaking keys).
    • Separate short per-call error budget (errorBudgetMs, default 4 s) that
      advances only on thrown errors, so genuine contention still uses the full
      acquireTimeoutMs window.
    • commandTimeout: 2000 on the ioredis client so commands reject promptly during
      an outage instead of queueing on the offline queue.
    • /readyz now PINGs Redis and returns 503 when Redis is configured but
      unreachable.
    • Fixed the stale ownership.ts docstring (the readOnly path DOES take a shared
      distributed lock).
  • #149 Thanks @Hazzng! - fix(blob): commit the CAS blob upsert in its own short, self-committing
    transaction (own connection, no advisory lock) BEFORE the inode/dirent composite,
    and run the composite without its blob_insert CTE. This removes hot-blob
    contention (F6): previously the ON CONFLICT (sha256) DO UPDATE SET last_referenced_at = now() tuple lock on a deduplicated hot blob (empty file,
    .gitkeep, common lockfiles) was held for the whole script, serializing
    unrelated sandboxes within one tenant DB and risking pool-exhaustion → 503. The
    touch stays unconditional so the GC grace window protects the freshly-committed
    blob until its inode commits; the blob-gc REPEATABLE READ + 40001 re-adoption
    handshake is preserved. Applies to writeFile, appendFile, and bulkIngest.
  • #150 Thanks @Hazzng! - Surface a retryable ESESSIONCLOSING (HTTP 503) instead of a generic 500 when a request
    loses the reaper-vs-straggler race (F9c). A request that captured a session reference just
    before the idle/overBudget reaper marked it closing could run the pre-lock
    ensureFreshCache probe against a Postgres pool being disconnected, producing an unmapped
    error (e.g. PostgresDialect: not connected, which carries no code) that defaulted to a 500. Both pre-lock probe sites (withSessionEntry, withSessionReadEntry) now re-check the
    session state on probe failure and convert it into a clean, retryable ESESSIONCLOSING
    (already mapped to 503) so clients retry instead of seeing a non-retryable 500. No
    concurrency-model change.

Image: ghcr.io/hazzng/sql-fs:v0.9.0

Release v0.8.0

Choose a tag to compare

@github-actions github-actions released this 13 Jun 06:25
b06bb65

Release v0.8.0

Minor Changes

  • #143 Thanks @Hazzng! - Guard publishVersionIfDirty with a cache-poison flag (F1): when a correlated Postgres failure fails both the script-tx COMMIT and the recovery reload, the session no longer publishes a version/snapshot of uncommitted phantom state — it suppresses the INCR, forces a reload on next use, and surfaces ECOHERENCE.
  • #145 Thanks @Hazzng! - fix(lock): when REDIS_RWLOCK_ENABLED=false, readers now take the same legacy single-key lock as writers (closing the F4 reader/writer race during rolling deploys), and SqlFs.reload() is a no-op while a script scope is open so a concurrent reload can never clobber an open writer's in-memory cache.
  • #144 Thanks @Hazzng! - fix(cache): writeFile now evicts the displaced inode's contentCache entry on overwrite (including empty-file overwrite), preventing orphaned LRU weight (F9a, #138).
  • #146 Thanks @Hazzng! - Boot-assert that the session idle window (SESSION_IDLE_MS / MCP_SESSION_IDLE_MS) stays at or below half the Redis version-key TTL when Redis is enabled, failing fast on misconfiguration that would break cache coherence (audit F9b).

Image: ghcr.io/hazzng/sql-fs:v0.8.0

Release v0.7.0

Choose a tag to compare

@github-actions github-actions released this 09 Jun 00:59
01412ab

Release v0.7.0

Minor Changes

  • #120 Thanks @NeilMazumdar! - Add static-header (API-key) auth for the MCP endpoint so external clients that can only send fixed headers — e.g. LibreChat — can connect without minting a per-request JWT. Set MCP_API_KEY to accept a pre-shared Authorization: Bearer <key>; the sandbox owner (sub) is derived from a forwarded identity header (MCP_IDENTITY_HEADER, default x-librechat-user-id), giving each end-user an isolated sandbox. New env vars: MCP_API_KEY, MCP_IDENTITY_HEADER, MCP_DEFAULT_SUB, MCP_STATIC_TENANT. Static auth is additive and off unless MCP_API_KEY is set — JWT clients on /mcp and all /v1/* routes are unchanged.
    Startup hardening: MCP_IDENTITY_HEADER cannot be a reserved transport header (authorization, cookie, content-type, accept, mcp-session-id, mcp-protocol-version, last-event-id) — otherwise every request would derive owner from a shared value and collapse all users into one sandbox. The mcp_static_auth_enabled startup log records only whether a fallback owner is configured (hasDefaultSub), never the MCP_DEFAULT_SUB value.
  • #125 Thanks @Hazzng! - feat(gc): multi-tenant orphan-blob garbage collection via pnpm db:gc.
    Restores the pnpm db:gc CLI as a real, multi-tenant orphan-blob sweep for an external scheduler (cron / k8s CronJob). Orphan blobs (rows in blobs referenced by zero inodes) previously accumulated forever.
    • New migration 0006 adds blobs.last_referenced_at (instant, catalog-only — legacy rows stay NULL and are treated as ancient/collectible). Every blob reference (insert + dedup re-adoption) now bumps it via ON CONFLICT (sha256) DO UPDATE, which also touches the blob so the grace window tracks real usage.
    • gcOrphanBlobs rewritten to a null-safe NOT EXISTS anti-join with a grace window (minAgeMs), returning the deleted sha256s. It runs with no sandbox context (RLS escape) so the anti-join sees every inode; a blob referenced by another sandbox survives.
    • The sweep runs at REPEATABLE READ with bounded retries to close the dedup re-adoption race: under READ COMMITTED a concurrent writer that re-adopts an existing orphan blob could leave its committed inode without content (the GC's NOT EXISTS re-check keeps a stale snapshot of inodes). REPEATABLE READ turns that conflict into a serialization failure that is retried, so even --min-age-ms 0 is safe under concurrent writes.
    • Deleted blobs are purged from the tenant-scoped Redis blob cache (RedisBlobCache.mdel, fail-open).
    • New env BLOB_GC_MIN_AGE_MS (default 3h) sets the grace window; pnpm db:gc -- --min-age-ms 0 collects all orphans now, --tenant <id> restricts to one tenant.

Patch Changes

  • #122 Thanks @Hazzng! - Fix local Postgres dev setup for the integration test suite (#119).
    • Add docker-compose.local.yml (Postgres 16 + Redis 7). Its initdb script
      (scripts/initdb/00-create-app-role.sql) provisions a non-superuser sqlfs_app
      role that owns the sqlfs database — required because migration 0005 enables
      FORCE ROW LEVEL SECURITY, which a superuser silently bypasses (the RLS isolation
      tests fail under the default postgres superuser). The file was previously
      referenced by the README but .gitignored, so it could never be committed.
    • Document the non-superuser-owner requirement and the full local-DB workflow in
      CONTRIBUTING.md (new "Local database" section).
    • Correct the DATABASE_DIRECT_URL row in the README env table: it is optional and
      used only by drizzle-kit (pnpm db:generate); the server's boot-time migration
      runner uses DATABASE_URL.
    • Update .env.example defaults to match the compose stack.
    • Remove the broken, unused pnpm db:migrate script (drizzle-kit migrate with no
      journal). Migrations are applied automatically on server boot.
  • #125 Thanks @Hazzng! - fix(fs): delete inodes when their link count reaches zero (no nlink=0 tombstones).
    The Postgres rmComposite, writeFileComposite, and mvComposite paths decremented an inode's nlink and deleted it (when it hit 0) within a single CTE statement. Postgres applies only the UPDATE when a row is both updated and deleted in one statement, so the inode was left at nlink=0 instead of being removed — a tombstone that still referenced content_sha256. This pinned the blob (defeating the new orphan-blob GC, whose anti-join saw the tombstone) and leaked inode rows on every file delete, overwrite, and move-overwrite.
    Each path now splits the work into two mutually-exclusive branches against the statement snapshot — delete when nlink <= 1, decrement when nlink > 1 — so each inode row is touched exactly once. gcOrphanBlobs additionally ignores nlink = 0 inodes so blobs pinned by tombstones left behind by older builds become collectible. Hardlinked inodes are unaffected (still decremented, not deleted, while other links remain).

Image: ghcr.io/hazzng/sql-fs:v0.7.0

TypeScript SDK v0.3.1

Choose a tag to compare

@github-actions github-actions released this 08 Jun 03:54
d02fd75

TypeScript SDK v0.3.1

Added

  • ingestFiles(..., { allowOversized }) — rejects files larger than 8 MiB with
    ValidationError (code EFILE_TOO_LARGE_FOR_CPYTHON) before anything is sent.
    The python3 runtime (CPython WASM) reads sandbox files through an 8 MiB IPC
    bridge, so open() fails on larger files. Pass allowOversized: true to
    ingest anyway (the bytes stay usable from bash and js-exec; only python3 open() can't read them), or split the file into <8 MiB chunks.

Install via npm:

npm install sql-fs-sdk@0.3.1

Python SDK v0.3.1

Choose a tag to compare

@github-actions github-actions released this 08 Jun 03:54
d02fd75

Python SDK v0.3.1

Added

  • ingest_files(..., allow_oversized=False) — rejects files larger than 8 MiB
    with ValidationError(code="EFILE_TOO_LARGE_FOR_CPYTHON") before anything is
    sent. The python3 runtime (CPython WASM) reads sandbox files through an 8 MiB
    IPC bridge, so open() fails on larger files. Pass allow_oversized=True to
    ingest anyway (the bytes stay usable from bash and js-exec; only python3 open() can't read them), or split the file into <8 MiB chunks.

Install via pip:

pip install sql-fs-sdk==0.3.1

TypeScript SDK v0.3.0

Choose a tag to compare

@github-actions github-actions released this 05 Jun 12:43
5dccd5d

TypeScript SDK v0.3.0

Added

  • Initial public release of the TypeScript SDK
  • Client for authentication and sandbox CRUD
  • Sandbox helpers for sync, batch, and streaming exec
  • File read/write, bulk write, mkdir, tree, delete, and base64 ingest operations
  • Client-side maxFileSize guard
  • Idempotency-aware retries and typed API errors
  • Full TypeScript types and ESM-only distribution

Install via npm:

npm install sql-fs-sdk@0.3.0

Release v0.6.3

Choose a tag to compare

@github-actions github-actions released this 04 Jun 15:17
56f833c

Release v0.6.3

Patch Changes

  • #113 Thanks @Hazzng! - Fix POST /v1/sandboxes/:id/ingest-files returning 500 (Internal Server Error) for files larger than ~750 KB. isValidBase64 ran a structural regex whose (?:[A-Za-z0-9+/]{4})* quantifier overflowed V8's call stack (RangeError: Maximum call stack size exceeded) on base64 strings beyond ~1 MB — failing during request validation, before any database work. The regex is now skipped for strings over 1 MB, relying solely on the canonical round-trip check (Buffer.from(s, "base64").toString("base64") === s), which is native and never overflows. Ingesting multi-MB files now succeeds.

Image: ghcr.io/hazzng/sql-fs:v0.6.3