Repository navigation
Releases: Hazzng/sql-fs
Releases · Hazzng/sql-fs
Release list
Release v1.0.1
Release v1.0.1
Patch Changes
- #161 Thanks @Hazzng! - Fence stale sandbox writers with durable epochs: script scopes pin
sandboxes.versionunder the advisory lock and every composite mutation conditionally advances it, so a writer whose lease lapsed before its first write now fails withESTALEinstead of silently overwriting a live writer's changes. Sandbox deletion persists a tombstone epoch so ID reuse cannot reset the fence. - #208 Thanks @Hazzng! - Stop holding a Postgres transaction open across a user's bash script (#166): metadata mutations are buffered in memory and flushed in one short advisory-locked transaction at scope end, while file bytes still commit eagerly via
commitBlob. Idle-in-transaction age dropped from 2.98s to 0s for a 3s script, and from 8.60s to 0.50s across 40 concurrent writers on a real Neon pooler. Constraint errors now surface at flush instead of at the failing command, a lost cross-replica race fails the loser withESTALE, and hitting the buffer cap fails the whole script closed withESCRIPTBUFFER(413). Toggle viaSCRIPT_TX_BUFFERED(default true). - #201 Thanks @Hazzng! - Stop
POST /writeFilesand thefilesmap on sandbox creation from accepting a wider batch than a single write allows:MAX_BULK_WRITE_BYTESnow defaults toMAX_FILE_WRITE_BYTESinstead of its own 128 MiB default, both routes share one set of limits, and/writeFilesgains a streamed body cap (callers currently sending 50-128 MiB batches now get413).GET /readyzalso exposes the event-loop lag histogram, including a newp999Msand anevent_loop_stallcritical log line aboveEVENT_LOOP_STALL_THRESHOLD_MS. - #202 Thanks @Hazzng! - Cap the file size a sandbox
execscript may read whole or produce with one write asMAX_EXEC_FILE_BYTES(default 8 MiB,EFBIGto 413): just-bash's text utilities rebuild strings synchronously on the main thread, and past roughly 2s a stall starts timing out other tenants' in-flight Redis commands. Applies only insidebash.exec; the file is still retrievable whole viaGET .../files/{path}. Holds per file, not per script — a pipeline likecat a b | wc -ccan still exceed it; the structural fix is movingbash.execoff the main thread (#198). - #194 Thanks @Hazzng! - Stop a Postgres connection dying mid-transaction from crashing the replica:
postgres.jsthrows a fatal, uncaughtTypeErrorfrom a baresetImmediatewhen a reaped backend's buffered write flushes to a nulled socket.PG_DRIVER_FAULT_GUARD(default true) recognizes only that exact stack frame, logsdriver_socket_fault, and fails the stuck DB awaits withEDRIVERFAULTto 503 after a grace window (5s) instead of taking every other in-flight request down with it. A condemned script scope can no longer commit partial work, and boot migrations get the same guard. - #195 Thanks @Hazzng! - Repair a structurally invalid Redis version key (
WRONGTYPE, non-integerSET) in place instead of returning503 ECOHERENCEon every write to that sandbox forever. The key resets to the current epoch in milliseconds — never 1, which could equal a warm replica'slastSeenVersionand mask staleness — except the F7DESTROYEDtombstone, which is left alone. - #196 Thanks @Hazzng! - Require
maxmemory-policy allkeys-lru(orallkeys-lfu) on the Redis backing the blob cache and warn at boot when it isn't set. Redis's defaultnoevictionmakes a full instance a permanent outage — writes refused forever, since a 24h blob-cache TTL doesn't age out fast enough to recover (measured 97.6% 5xx, no recovery). The boot check (redis_eviction_policy_unsafe, critical) never fails startup and is skipped when the data client carries no data-plane state. - #204 Thanks @Hazzng! - Put every filesystem mutation behind the epoch fence and make all of them advance
sandboxes.version, not just the four composite writes. Fourteen other call sites (bulkIngest,mkdir -p,rm -r,cp,cp -r,link,symlink,chmod,utimes, non-composite fallbacks) previously left the counter untouched, so a live writer using only those left a stale peer's pin still matching.cp,chmod,ln,touch,mkdir -p,rm -rand ingest can now fail with409 ESTALE, which they never could before. - #179 Thanks @Hazzng! - Stop leaking raw driver/SQLSTATE error codes (
ECONNRESET, bare codes like53300) to clients: the allowlist that already redacted error messages now also redacts thecodefield, shared by the global error handler and the SSE error frame. Connection-class SQLSTATEs (08xxx,53300,53400,57P03) now map to a retryable503 EUNAVAILABLEinstead of500. - #209 Thanks @Hazzng! - Add a
retryableboolean to every error body (and the SSEerrorframe) so clients can tell an applied-but-unacknowledged write from one that never landed.ELOCKLOST_APPLIED(503, not retryable) now covers a lease lost after commit,ECOHERENCE_UNAPPLIEDcovers a rolled-back turn, and a read-only request no longer inherits a previous turn's stranded version-publish failure. The guarantee is per-transaction: multi-step routes like sandbox creation can commit an earlier step before a later one fails retryably. - #183 Thanks @Hazzng! - Lower the default exec-lock acquire timeout from 300s to 75s (lease plus ~15s reap margin, so the 503 reaches the client before typical ingress timeouts sever the connection), and refuse to boot when it's below
REDIS_EXEC_LOCK_LEASE_MSorREDIS_RWLOCK_READER_LEASE_MS, since a shorter window turns crashed-holder recovery into a permanent 503. - #182 Thanks @Hazzng! - Abort a running script when the client of
POST /v1/sandboxes/:id/exec-syncdisconnects, instead of holding the sandbox's exclusive exec lock for the rest of its timeout — the route never wired upc.req.raw.signal, unlike its SSE and batch siblings. Work already committed before the abort stays committed. - #178 Thanks @Hazzng! - Make the startup migration runner safe under transaction-mode connection pooling: the whole run is now one transaction opened with
pg_advisory_xact_lockas its first statement, since a session-scoped lock taken outside a transaction doesn't hold across a pooler reassigning connections per transaction (a second booter previously acquired the lock in 267-335ms instead of waiting — mutual exclusion silently wasn't holding). Migrations are now atomic as a side effect.DATABASE_DIRECT_URLis no longer needed by the server, only bydrizzle-kit, and its deployment secret is removed. - #183 Thanks @Hazzng! - Make
GET /v1/sandboxes/:id/files/*andGET /v1/sandboxes/:id/treetake the shared session lock instead of the exclusive write lock, so concurrent reads of one sandbox run in parallel instead of serializing like writes (a burst of 32 concurrentGET /treerequests dropped from ~150ms to 5ms). A GET still waits behind an in-flight writer, unchanged. - #207 Thanks @Hazzng! - Split the Redis connection by role (
control: locks, version counter, session state;data: blob cache, path snapshot) so multi-MiB blob writes can no longer head-of-line block latency-critical lock/version commands, scope the circuit breaker per role, cap in-flight blob-cache backfills (drop rather than queue over cap), and log every breaker transition. On a 6s Redis pause, 5xx fell from 90.0% to 54.3% on one shared instance, and to 0% for data-plane-only outages onceREDIS_DATA_URLpoints at a separate Redis. - #190 Thanks @Hazzng! - Give four test apps' hand-rolled
onErrorproduction's actual error contract instead of a copy of the pre-#174 leaky one, plus a source scan that fails if the hand-rolled fallback reappears. Test-only; no shipped behavior changes. - #210 Thanks @Hazzng! - Make the vitest excludes path-independent (
**/comparison/**,**/.claude/**) so a git worktree checked out inside the repo — where Claude Code's agent isolation puts them — no longer contributes a duplicatesrc/suite and a failingcomparison/copy topnpm test:unit. Tooling only. - #205 Thanks @Hazzng! - Make the integration suites actually run: 29 tests across seven files that silently skipped everywhere because nothing said
REDIS_URLwas required now execute, 5sql-fstests that failed against a real database because the epoch fence (#161) madesetSandboxContextWithLockrequire a real sandbox row are fixed, andbeforeAllmigration races between files were replaced with a schema check plus serial file execution.
Image: ghcr.io/hazzng/sql-fs:v1.0.1
Release v1.0.0
Release v1.0.0
Major Changes
- #158 Thanks @Hazzng! - Add a sandbox
gitcommand backed by just-git, export serverGITHUB_TOKENinto sandbox GitHub-compatible Git/curl env, and let MCP-created sandboxes request network access for clone/fetch/push. A per-requestenv.GITHUB_TOKENre-points git's HTTP credentials at that token, refusing plaintexthttp://remotes, validating each redirect hop by hand, dropping credentials on cross-origin redirects, and rewriting a redirected POST to GET on 301/302/303 the wayfetchdoes — so a push packfile is never replayed at a host it was merely forwarded to.
Minor Changes
- #162 Thanks @Hazzng! - Add file access to MCP —
file_read,file_write,file_edit— plusPATCH /v1/sandboxes/:id/files/*pathfor exact-string edits.
file_editrequiresoldStringto match exactly once (replaceAllopts into multiple), rejecting an ambiguous match asEDIT_NOT_UNIQUErather than guessing; it runs in one script-tx scope on Postgres so a concurrent reader never sees the file mid-edit, refuses non-UTF-8 content and unpaired surrogates, and preserves file mode and a leading BOM.file_readpages viaoffset/limit/byteOffset, bounding both the file it opens (MAX_MCP_READ_FILE_BYTES) and the reply it returns (MAX_MCP_READ_RESPONSE_BYTES).file_writecreates parent directories and refuses to clobber a directory on every backend. HTTP and MCP share one implementation insrc/api/lib/file-ops.ts.
Patch Changes
- #162 Thanks @Hazzng! - Document container sizing for the write cap: a single write costs roughly 7x the file size in
externalmemory above steady state, so a 50 MiB write needs on the order of 700 MB of headroom (a 512 MiB container OOM'd on one request; 768 MiB survived). - #162 Thanks @Hazzng! - Apply the single-file write limit to each entry of a bulk write, not just the combined total, so one oversized entry in
POST /writeFilescan no longer land a blob the contentCache cannot hold. - #162 Thanks @Hazzng! - Roll back a
PATCHedit or bulk write when the distributed exec lock is definitively lost mid-request, instead of committing it and reporting a retryableELOCKLOST. Both routes now go throughrunInScriptTx, which checks for lock loss before the commit. - #162 Thanks @Hazzng! - Reject an edit whose
oldString/newStringcarries an unpaired surrogate asEDIT_LONE_SURROGATE, and enforce the write limit on the encoded result rather than on a projected size that assumed one match encodes to the bytes it replaces. - #162 Thanks @Hazzng! - Clean up the destination of a
git clonethat fails partway (e.g. on a symlink, sinceallowSymlinksdefaults to false), instead of leaving a half-checked-out tree whose complete index made every missing file look like a staged deletion — an agent then following up withgit add -A && git commit && git pushturned that into a real destructive commit. - #162 Thanks @Hazzng! - Refuse to replay a git request body across origins on a 307/308 redirect, since both preserve the method and body — for git, the packfile being pushed — and could otherwise forward a whole push to an attacker-chosen host.
- #162 Thanks @Hazzng! - Budget the whole
file_readreply, not just the content string, againstMAX_MCP_READ_RESPONSE_BYTES, and normalize the path MCP tools echo back instead of only prefixing a slash. - #162 Thanks @Hazzng! - Size a
file_readpage against the reply as the transport actually serializes it (which re-escapes once more), and stop callingsplit("\n")on the whole file just to count lines — both scanning-based fixes remove a real-world 170 KiB overshoot and a multi-million-element allocation. - #162 Thanks @Hazzng! - Enforce the file-write limit on
PUT /v1/sandboxes/:id/files/*pathas the body streams, instead of trustingContent-Lengthand buffering the whole request first. - #162 Thanks @Hazzng! - Keep a leading UTF-8 BOM in what
file_readandfs_exportreturn, matching theignoreBOMdecodingeditFilealready used, so content andstat.sizeround-trip byte for byte. - #162 Thanks @Hazzng! - Return
RESPONSE_BUDGET_TOO_SMALLinstead of a reply that can never fit any content whenMAX_MCP_READ_RESPONSE_BYTESis configured below the size of the response envelope, which previously left a client resuming a page in an infinite loop. - #162 Thanks @Hazzng! - Keep
file_readpaging identical to thesplit/joinit replaced when a file ends in a newline, instead of returning a trailing newline the old implementation would have dropped. - #162 Thanks @Hazzng! - Ignore a relative
PWDwhen recording a session's working directory instead of rooting it into a path that never existed. - #162 Thanks @Hazzng! - Build a
replaceAlledit by assembling flushed chunks instead ofsplit(oldString).join(newString), cutting peak RSS on a worst-case near-limit file from 1184 MB to 357 MB with no regression on ordinary single-match edits. - #162 Thanks @Hazzng! - Stop an abort that races script-tx opening from rejecting a promise with no listener, which was fatal under Node's default
--unhandled-rejections=throw. - #162 Thanks @Hazzng! - Stop a late-arriving transaction open from adopting into a finished scope (each open now carries a generation checked before adopting), and refuse cache-served reads (
stat,readFile,readdir,exists,getAllPaths) once a script-tx is lost. - #162 Thanks @Hazzng! - Fail every remaining operation in a script scope once its transaction's connection is lost, instead of letting a write silently self-commit outside the scope on a reconnected-but-transactionless connection — previously a 600-file bulk write could answer HTTP 500 with 599 of them durable.
- #162 Thanks @Hazzng! - Write a whole file through one shared transactional path (
writeFileAtPath) on both the MCPfile_writetool andPUT /v1/sandboxes/:id/files/*, so parent directories and file content commit together andPUTmatches MCP in refusing to clobber a directory (400 EISDIR). - #162 Thanks @Hazzng! - Default the single-file write limit to the contentCache cap (50 MiB) instead of 64 MiB — load testing found memory cost doubles just past the cache cap, so the old default pinned 256 MB per warm session for a single large read.
Image: ghcr.io/hazzng/sql-fs:v1.0.0
Release v0.10.0
Release v0.10.0
Minor Changes
- #157 Thanks @Hazzng! - feat(observability): event-loop lag monitoring for the Redis leases (F8).
The exec-lock writer lease, the RW-lock writer flag, and the RW-lock reader ZSET
scores are all kept alive bysetTimeoutheartbeats that silently assume timers
fire on schedule. A long event-loop stall (a V8 GC pause or a pathological
synchronous bash stretch) can fire a renewal past the lease, voiding it — Lock 3
keeps Postgres consistent, so this was always an observability gap, not a
correctness bug, but nothing measured it.
Newsrc/api/event-loop-monitor.ts(purely observational, no behavior change):- A
perf_hooks.monitorEventLoopDelayhistogram started at boot, sampled every
EVENT_LOOP_MONITOR_INTERVAL_MS(default 10s) and logged as
event:"event_loop_lag"(p50Ms/p99Ms/maxMs/meanMs), then reset. - Per-heartbeat gap measurement wired into all three lease sites: each heartbeat
reports actual-minus-expected fire time asevent:"heartbeat_gap"at
severity:"warn"(gap > renewMs) or"critical"(gap > leaseMs), tagged with
the lock kind (exec/rw-writer/rw-reader) and key.
Alert thresholds are documented in DEVELOPER.md ("Lock observability"). End-to-end
smoke tests reproduce a >lease stall on each lease and assert the critical
heartbeat_gap fires (with a no-stall control proving no false positives).
- A
- #156 Thanks @Hazzng! - Heal stranded cross-replica version publishes after a Redis INCR failure (F3): a background drainer and reap-time best-effort publish flush the bump even if no further client traffic arrives or the session is idle-evicted.
- #155 Thanks @Hazzng! - fix(session): destroy now reaches warm sessions on other replicas (F7).
Destroying a sandbox on one replica previously left warm sessions on other
replicas serving ghost state: a written session would reload a deleted tree
into an empty pathCache (surfacing as a non-zero exit + garbage stderr inside an
HTTP 200 exec), and a never-written session would never reload at all because
the deleted version key read as 0 and matched itslastSeenVersion === 0.
Two layered fixes:- Primary (Redis-independent):
SqlFs.reload()now detects a zero-row
loadAllPaths— which for a live sandbox always returns at least its root dir
— and throws a typedESANDBOXGONEinstead of installing an empty pathCache.
The session manager catches it, tears the stale warm session down (drops it
from the pool and disconnects the per-session Postgres pool), and surfaces a
cleanENOENT→ 404. - Secondary (tombstone):
destroynow writes a distinctDESTROYEDsentinel to
the version key (with the version-key TTL) instead of deleting it.
ensureFreshCacherecognises the sentinel before the numeric parse and tears
the session down — covering the never-written variant. Re-creating a
tombstoned sandbox clears the sentinel and starts cleanly at version 0.
- Primary (Redis-independent):
- #153 Thanks @Hazzng! - fix(lock): add bounded jitter + tunable retry to the distributed acquire loops (F9d, #141)
The distributed exec lock and RW lock polled Redis on a flatacquireRetryMs
(default 50 ms) interval, leaving competing replicas phase-aligned so a
cross-replica writer could be repeatedly passed over (bounded by
acquireTimeoutMs, then 503). Every acquire/drain poll now sleeps a jittered
retryMs/2 + random()*retryMs/2(range[retryMs/2, retryMs]) to
de-synchronize pollers. The retry interval is now configurable via
REDIS_EXEC_LOCK_ACQUIRE_RETRY_MS(previously hardcoded —server.tsomitted
it). Circuit-breaker / error-budget behavior is unchanged. The FIFO ZSET ticket
queue is deferred as a follow-up. - #154 Thanks @Hazzng! - perf(cache): O(1) pathCache byte accounting to avoid full-map scans (F9e, #142)
SqlFs now maintains an incremental#pathCacheBytescounter, adjusted on
every pathCache set/delete and reset onreload()/ready(), and exposes
getPathCacheBytes(). SessionManager's path-cache memory budget calls it
instead of re-walking the entire pathCache (Σ path.length + 100) on every
dirty exec. The value equals the previous full-walk exactly. Falls back to
the full walk for backends that do not expose the counter.
The#childrenByParentchildren index (part B of #142) is deferred to a
follow-up; it is benchmark-gated and (A) delivers the higher-value, lower-risk
win without touching readdir correctness.
Image: ghcr.io/hazzng/sql-fs:v0.10.0
Release v0.9.0
Release v0.9.0
Minor Changes
- #151 Thanks @Hazzng! - fix(lock): abort the exec on definitive lease loss before commit (F2-L1)
The distributed exec lock wrappers ran the critical section to completion and only
then checked the loss flag, so a writer whose lease lapsed mid-script still
committed its script-tx and bumped the version before throwingELOCKLOST— a
write that durably happened surfaced as an error, causing retrying agents to
double-apply.
The lock now wires its DEFINITIVE-loss signal (lease expiry / ownership taken —
not transient renew blips) into anAbortControllerthat is plumbed through to
bash.exec. On a definitive loss the in-flight exec is aborted, its script-tx
rolls back BEFORE any commit (noINCR), and the client receives a clean,
retryableELOCKLOST(now mapped to 503). Because just-bash treats an aborted
run as a resolved result rather than a rejection, the runtime explicitly rolls
back and re-raisesLockLostErrorwhen the lock-lost signal fired, instead of
committing the partial script. A plain timeout abort still commits (unchanged,
audit L7) — only the dedicated lock-lost signal triggers rollback.
This is Layer 1 of the F2 fix; the complete epoch/version fence is tracked
separately (#131). - #148 Thanks @Hazzng! - fix(lock): circuit-break Redis acquire to stop the 300s outage fuse (F5)
The distributed lock acquire loops conflated "lock busy" (contention) with "Redis
unreachable" (a thrown connection error): both retried untilacquireTimeoutMs
(default 300 s), so a Redis outage hung every exec/file op for ~5 minutes on an
otherwise-healthy Postgres.- New process-wide Redis circuit breaker (
src/redis/circuit-breaker.ts) wired
into the lock ACQUIRE paths only (distributed-rw-lock.tsshared/exclusive +
waitReadersDrained, legacydistributed-lock.ts). After K (=5) consecutive
connection-class failures it opens and acquire fast-fails 503 immediately; a
successful eval/PING closes it. Renew/release paths are untouched (they keep
tolerating transient errors to avoid dropping leases / leaking keys). - Separate short per-call error budget (
errorBudgetMs, default 4 s) that
advances only on thrown errors, so genuine contention still uses the full
acquireTimeoutMswindow. commandTimeout: 2000on the ioredis client so commands reject promptly during
an outage instead of queueing on the offline queue./readyznow PINGs Redis and returns 503 when Redis is configured but
unreachable.- Fixed the stale
ownership.tsdocstring (the readOnly path DOES take a shared
distributed lock).
- New process-wide Redis circuit breaker (
- #149 Thanks @Hazzng! - fix(blob): commit the CAS blob upsert in its own short, self-committing
transaction (own connection, no advisory lock) BEFORE the inode/dirent composite,
and run the composite without itsblob_insertCTE. This removes hot-blob
contention (F6): previously theON CONFLICT (sha256) DO UPDATE SET last_referenced_at = now()tuple lock on a deduplicated hot blob (empty file,
.gitkeep, common lockfiles) was held for the whole script, serializing
unrelated sandboxes within one tenant DB and risking pool-exhaustion → 503. The
touch stays unconditional so the GC grace window protects the freshly-committed
blob until its inode commits; the blob-gc REPEATABLE READ + 40001 re-adoption
handshake is preserved. Applies towriteFile,appendFile, andbulkIngest. - #150 Thanks @Hazzng! - Surface a retryable
ESESSIONCLOSING(HTTP 503) instead of a generic 500 when a request
loses the reaper-vs-straggler race (F9c). A request that captured a session reference just
before the idle/overBudget reaper marked itclosingcould run the pre-lock
ensureFreshCacheprobe against a Postgres pool being disconnected, producing an unmapped
error (e.g.PostgresDialect: not connected, which carries nocode) that defaulted to a 500. Both pre-lock probe sites (withSessionEntry,withSessionReadEntry) now re-check the
session state on probe failure and convert it into a clean, retryableESESSIONCLOSING
(already mapped to 503) so clients retry instead of seeing a non-retryable 500. No
concurrency-model change.
Image: ghcr.io/hazzng/sql-fs:v0.9.0
Release v0.8.0
Release v0.8.0
Minor Changes
- #143 Thanks @Hazzng! - Guard
publishVersionIfDirtywith a cache-poison flag (F1): when a correlated Postgres failure fails both the script-tx COMMIT and the recovery reload, the session no longer publishes a version/snapshot of uncommitted phantom state — it suppresses the INCR, forces a reload on next use, and surfaces ECOHERENCE. - #145 Thanks @Hazzng! - fix(lock): when
REDIS_RWLOCK_ENABLED=false, readers now take the same legacy single-key lock as writers (closing the F4 reader/writer race during rolling deploys), andSqlFs.reload()is a no-op while a script scope is open so a concurrent reload can never clobber an open writer's in-memory cache. - #144 Thanks @Hazzng! - fix(cache): writeFile now evicts the displaced inode's contentCache entry on overwrite (including empty-file overwrite), preventing orphaned LRU weight (F9a, #138).
- #146 Thanks @Hazzng! - Boot-assert that the session idle window (
SESSION_IDLE_MS/MCP_SESSION_IDLE_MS) stays at or below half the Redis version-key TTL when Redis is enabled, failing fast on misconfiguration that would break cache coherence (audit F9b).
Image: ghcr.io/hazzng/sql-fs:v0.8.0
Release v0.7.0
Release v0.7.0
Minor Changes
- #120 Thanks @NeilMazumdar! - Add static-header (API-key) auth for the MCP endpoint so external clients that can only send fixed headers — e.g. LibreChat — can connect without minting a per-request JWT. Set
MCP_API_KEYto accept a pre-sharedAuthorization: Bearer <key>; the sandbox owner (sub) is derived from a forwarded identity header (MCP_IDENTITY_HEADER, defaultx-librechat-user-id), giving each end-user an isolated sandbox. New env vars:MCP_API_KEY,MCP_IDENTITY_HEADER,MCP_DEFAULT_SUB,MCP_STATIC_TENANT. Static auth is additive and off unlessMCP_API_KEYis set — JWT clients on/mcpand all/v1/*routes are unchanged.
Startup hardening:MCP_IDENTITY_HEADERcannot be a reserved transport header (authorization,cookie,content-type,accept,mcp-session-id,mcp-protocol-version,last-event-id) — otherwise every request would deriveownerfrom a shared value and collapse all users into one sandbox. Themcp_static_auth_enabledstartup log records only whether a fallback owner is configured (hasDefaultSub), never theMCP_DEFAULT_SUBvalue. - #125 Thanks @Hazzng! - feat(gc): multi-tenant orphan-blob garbage collection via
pnpm db:gc.
Restores thepnpm db:gcCLI as a real, multi-tenant orphan-blob sweep for an external scheduler (cron / k8s CronJob). Orphan blobs (rows inblobsreferenced by zeroinodes) previously accumulated forever.- New migration
0006addsblobs.last_referenced_at(instant, catalog-only — legacy rows stay NULL and are treated as ancient/collectible). Every blob reference (insert + dedup re-adoption) now bumps it viaON CONFLICT (sha256) DO UPDATE, which also touches the blob so the grace window tracks real usage. gcOrphanBlobsrewritten to a null-safeNOT EXISTSanti-join with a grace window (minAgeMs), returning the deleted sha256s. It runs with no sandbox context (RLS escape) so the anti-join sees every inode; a blob referenced by another sandbox survives.- The sweep runs at REPEATABLE READ with bounded retries to close the dedup re-adoption race: under READ COMMITTED a concurrent writer that re-adopts an existing orphan blob could leave its committed inode without content (the GC's
NOT EXISTSre-check keeps a stale snapshot ofinodes). REPEATABLE READ turns that conflict into a serialization failure that is retried, so even--min-age-ms 0is safe under concurrent writes. - Deleted blobs are purged from the tenant-scoped Redis blob cache (
RedisBlobCache.mdel, fail-open). - New env
BLOB_GC_MIN_AGE_MS(default 3h) sets the grace window;pnpm db:gc -- --min-age-ms 0collects all orphans now,--tenant <id>restricts to one tenant.
- New migration
Patch Changes
- #122 Thanks @Hazzng! - Fix local Postgres dev setup for the integration test suite (#119).
- Add
docker-compose.local.yml(Postgres 16 + Redis 7). Itsinitdbscript
(scripts/initdb/00-create-app-role.sql) provisions a non-superusersqlfs_app
role that owns thesqlfsdatabase — required because migration0005enables
FORCE ROW LEVEL SECURITY, which a superuser silently bypasses (the RLS isolation
tests fail under the defaultpostgressuperuser). The file was previously
referenced by the README but.gitignored, so it could never be committed. - Document the non-superuser-owner requirement and the full local-DB workflow in
CONTRIBUTING.md(new "Local database" section). - Correct the
DATABASE_DIRECT_URLrow in the README env table: it is optional and
used only by drizzle-kit (pnpm db:generate); the server's boot-time migration
runner usesDATABASE_URL. - Update
.env.exampledefaults to match the compose stack. - Remove the broken, unused
pnpm db:migratescript (drizzle-kitmigratewith no
journal). Migrations are applied automatically on server boot.
- Add
- #125 Thanks @Hazzng! - fix(fs): delete inodes when their link count reaches zero (no
nlink=0tombstones).
The PostgresrmComposite,writeFileComposite, andmvCompositepaths decremented an inode'snlinkand deleted it (when it hit 0) within a single CTE statement. Postgres applies only the UPDATE when a row is both updated and deleted in one statement, so the inode was left atnlink=0instead of being removed — a tombstone that still referencedcontent_sha256. This pinned the blob (defeating the new orphan-blob GC, whose anti-join saw the tombstone) and leaked inode rows on every file delete, overwrite, and move-overwrite.
Each path now splits the work into two mutually-exclusive branches against the statement snapshot — delete whennlink <= 1, decrement whennlink > 1— so each inode row is touched exactly once.gcOrphanBlobsadditionally ignoresnlink = 0inodes so blobs pinned by tombstones left behind by older builds become collectible. Hardlinked inodes are unaffected (still decremented, not deleted, while other links remain).
Image: ghcr.io/hazzng/sql-fs:v0.7.0
TypeScript SDK v0.3.1
TypeScript SDK v0.3.1
Added
ingestFiles(..., { allowOversized })— rejects files larger than 8 MiB with
ValidationError(codeEFILE_TOO_LARGE_FOR_CPYTHON) before anything is sent.
Thepython3runtime (CPython WASM) reads sandbox files through an 8 MiB IPC
bridge, soopen()fails on larger files. PassallowOversized: trueto
ingest anyway (the bytes stay usable from bash andjs-exec; onlypython3 open()can't read them), or split the file into <8 MiB chunks.
Install via npm:
npm install sql-fs-sdk@0.3.1
Python SDK v0.3.1
Python SDK v0.3.1
Added
ingest_files(..., allow_oversized=False)— rejects files larger than 8 MiB
withValidationError(code="EFILE_TOO_LARGE_FOR_CPYTHON")before anything is
sent. Thepython3runtime (CPython WASM) reads sandbox files through an 8 MiB
IPC bridge, soopen()fails on larger files. Passallow_oversized=Trueto
ingest anyway (the bytes stay usable from bash andjs-exec; onlypython3 open()can't read them), or split the file into <8 MiB chunks.
Install via pip:
pip install sql-fs-sdk==0.3.1
TypeScript SDK v0.3.0
TypeScript SDK v0.3.0
Added
- Initial public release of the TypeScript SDK
Clientfor authentication and sandbox CRUDSandboxhelpers for sync, batch, and streaming exec- File read/write, bulk write, mkdir, tree, delete, and base64 ingest operations
- Client-side
maxFileSizeguard - Idempotency-aware retries and typed API errors
- Full TypeScript types and ESM-only distribution
Install via npm:
npm install sql-fs-sdk@0.3.0
Release v0.6.3
Release v0.6.3
Patch Changes
- #113 Thanks @Hazzng! - Fix
POST /v1/sandboxes/:id/ingest-filesreturning 500 (Internal Server Error) for files larger than ~750 KB.isValidBase64ran a structural regex whose(?:[A-Za-z0-9+/]{4})*quantifier overflowed V8's call stack (RangeError: Maximum call stack size exceeded) on base64 strings beyond ~1 MB — failing during request validation, before any database work. The regex is now skipped for strings over 1 MB, relying solely on the canonical round-trip check (Buffer.from(s, "base64").toString("base64") === s), which is native and never overflows. Ingesting multi-MB files now succeeds.
Image: ghcr.io/hazzng/sql-fs:v0.6.3