Skip to content

Releases: ChristianKohlberg/backlot

backlot 0.10.0

Choose a tag to compare

@ChristianKohlberg ChristianKohlberg released this 30 Jul 14:36
8fe5440

One feature, and it is about the accounts your stack seeds: auth.logins now takes a
list, so a stack can advertise one login per role instead of the single account
consumers then drove everything as. CI green on macOS + Linux, and additionally
verified by hand against a real 0.10.0 daemon: a two-login manifest reports the primary
in logins and the full roster in allLogins (through up and a later ctx alike),
and a manifest using the old single-object form reports the same object in both.

A stack may advertise several logins, each with a purpose

auth.logins admitted exactly one {user, password}, and the schema is
additionalProperties: false at both levels — so a stack seeding admin, auditor,
partner and read-only accounts had no way to say so: not as a list, not as an extra
key, not as a sibling. The other logins existed only in that repo's seed
documentation, which the agents reading ctx never see. The practical result is the
part worth fixing: everything got driven as the admin login, which is the one account
that can never expose a scoping bug.

It now takes a single login (unchanged) or a list, and a login gains two optional
fields — role, naming the {{role}} your auth.token hook takes, and
description, saying what the login is for:

auth:
  logins:
    - { user: qa-admin,    password: Demo!1234, role: admin, description: "all rights, all branches" }
    - { user: qa-auditor,  password: Demo!1234, role: auditor, description: "read-only across branches" }
    - { user: qa-partner,  password: Demo!1234, role: partner, description: "own branch only — proves scoping" }
  token: scripts/mint-token --role {{role}} --json

ctx reports both, and the redundancy is the feature:

  • logins — the primary login, always a single object, always manifest entry 0.
  • allLogins — every declared login, in manifest order.

So nothing existing breaks. A consumer reading ctx.logins.user keeps reading the
same thing when a stack grows a list, and a single-login stack reports the same object
in both places — code that wants to enumerate never has to ask which form the manifest
used.

description is the point rather than decoration: a bare roster of usernames moves the
question instead of answering it, because a consumer still cannot tell which login
proves a scoping bug and which merely passes. An empty list is rejected — it claims
logins are declared while declaring none; omitting the key remains how a stack says it
has none.

Still not an agent feature. backlot does not create these logins, verify them, mint
sessions for them, or know what a "role" means. Your seed makes them, the manifest
declares what exists, ctx reports it. Same passthrough logins always was, widened.
See decision 0026.

Also in this release

  • The shipped skill now says to read the roster. The Claude Code plugin's skill
    (bumped to 0.6.0) tells an agent to pick from allLogins by description instead of
    defaulting to the admin — the behaviour the feature exists to change.
  • The documented example did not survive a YAML parser. In a flow mapping,
    description: all rights, all branches ends the value at the comma and makes
    all branches a second key, which additionalProperties: false then rejects. So
    copy-pasting the manifest reference shipped in 0.9.1 produced
    the backlot manifest is invalid from the docs themselves. Both descriptions are now
    quoted, and both blocks are verified by parsing them and validating against
    schema/backlot.schema.json.

Upgrading

npm i -g backlot@0.10.0
backlot update             # restart the running daemon onto the installed build

Adopting the list form is version-coupled to the daemon, and this is the first time
skew reaches the manifest: a stack that uses a list fails validation on a pre-0.10.0
backlot with the backlot manifest is invalid, which reads as a broken manifest rather
than an old install. Upgrade the machine first, then edit the manifest. The single-login
form validates on every version, indefinitely.

Known open

  • #56 — a datastore's preset is fixed at first provision; --preset is silently
    ignored on every later bind, --reset-data included.
  • #44 — a 160k-file bind on macOS is ~an order of magnitude slower than on Linux. Needs
    a phase-split measurement; tracked, not a regression.

backlot 0.9.1

Choose a tag to compare

@ChristianKohlberg ChristianKohlberg released this 28 Jul 22:19
cc2339d

A one-fix patch, and it is the one 0.9.0 needed: backlot update could not update a daemon older than itself.

The fix

0.9.0 shipped backlot update to replace a running daemon with the installed build. It called update-plan unconditionally, first — and no pre-0.9.0 daemon has that verb. So every install upgrading from 0.8.0 or earlier got

$ npm i -g backlot@0.9.0 && backlot update
backlot: [env-error] daemon does not know verb 'update-plan' (rpc)

from the verb whose entire purpose is replacing an old daemon. It worked for nobody, and the way out — backlot daemon stop, then any verb — is exactly the manual step update exists to remove.

Both new RPCs now fall back to verbs every version has:

  • update-plan unknown → the daemon is pre-0.9.0 by definition, and status exists in every version, so the plan is rebuilt from it: the holders who must rebind are still named, and the pid to wait on still comes from the daemon itself. It reports busy: [] with a note that an old daemon cannot be asked what is in flight — the in-flight refusal genuinely cannot apply, and pretending otherwise would be a lie in the output.
  • daemon-restart unknown → shutdown, which every version has, and which is what daemon stop has always done.

backlot update --check against such a daemon now reports rather than failing.

Why 0.9.0's tests did not catch it

Skew is provoked with BACKLOT_FAKE_VERSION, which runs the current build while merely claiming an old version — so the stand-in always knew the new verbs. That is the same defect as the ping stand-ins corrected while the skew gate was being built ("a stand-in must model the thing it stands in for"), one file over.

The regression test now uses a stand-in that answers the way a real pre-0.9.0 daemon does: ping without a version, the old status shape, the daemon's verbatim "does not know verb" error for anything newer, and shutdown by closing the socket so the new daemon takes it — reporting a pid that really exits, because update waits for the old process as well as the socket. Verified additionally against a genuinely installed backlot@0.8.0: update takes its daemon from pre-0.9.0 to this build, after which ordinary verbs work.

Upgrading

npm i -g backlot@0.9.1
backlot update

If you are on 0.9.0 with an older daemon still running, backlot daemon stop is the one-command equivalent — the next verb autospawns the installed build.

Known open

  • #44 — a 160k-file bind on macOS is ~an order of magnitude slower than on Linux. Needs a phase-split measurement; tracked, not a regression.

backlot 0.9.0

Choose a tag to compare

@ChristianKohlberg ChristianKohlberg released this 28 Jul 21:31
b0e2b2b

Upgrading to this release takes two steps, and the second one is new: backlot update. CI green on macOS + Linux; the pool fixes were additionally verified by driving the reported scenarios by hand.

Highlights

  • backlot update — make the running daemon be the installed build. The CLI spawns the daemon from its own dist/, so installing a new backlot never replaced a daemon already in memory: it went on serving old code for the rest of its life, and the socket carried no version, so nothing on either side could notice. There was no backlot --version either. Both are fixed, and version skew is now refused (infra-error, exit 3) rather than warned about — an old daemon does not reject a flag it has never heard of, it ignores it, so up --data-only against a 0.8.0 daemon booted the whole application into what the caller believed was a database-only lease and reported success. The MCP adapter enforces the same gate independently, since that is the surface agents actually drive. A restart keeps every lease: services stop, each holder's next verb rebinds. Only an in-flight run or a downgrade is refused. backlot never installs itself — it prints the command for your install (decision 0024).

  • A cold environment no longer holds a machine-wide pool slot forever (#46). Idle reclamation quiesces heat, not the environment, and the row is what counts toward poolMaxTotal — so on a 12-core host with six cold, unleased environments across seven worktrees, a seventh stack was refused pool at capacity (6/6) indefinitely, while nothing was running. The ceiling measured history rather than load. When the machine-wide cap binds, backlot now evicts the least-recently-used cold environment and takes its slot. Never a leased or busy one; the victim pays one cold provision on its next bind, and every eviction is logged as pool-evict. (Simply not counting cold rows — the obvious fix — under-protects: both caps gate creation only, reuse is never capacity-checked, so the row count is exactly what bounds worst-case concurrent load.)

  • A capacity refusal names the cap that actually bound (#47). pool at capacity (6/6) was not even a count — it was the per-stack cap printed twice — and the remedy it gave (BACKLOT_POOL_MAX) cannot clear a machine-wide block; following it changed nothing, as it could not. Both refusals now quote real counts for every ceiling, say which one refused, and name the knob that moves it. They also fail immediately where waiting is provably useless: a machine-wide block never clears by waiting, because the count is of environments and a release leaves the row behind.

  • A data-only lease is priced like a catalog, not like a stack (#48). poolMax/poolMaxTotal come from min(cores/2, memGB/4) because they bound running services; a data-only environment starts none. It is now counted against its own disk-shaped ceiling, BACKLOT_POOL_MAX_DATA_ONLY (default max(4, 2 × the heuristic)), and against neither application cap — so a test lane on every integration run stops competing with the interactive leases people use to look at the app, which is the contention --data-only existed to remove. Changing an environment's shape is now a metered capacity event; switching your own lease both ways still works (decision 0025).

Fixes

  • A quiesce in flight is transient, not a permanent capacity block. The idle quiesce runs under the environment lock, which marks the environment busy, and decision 0021 keeps its mid-quiesce state as plain hot — so while the sweeper reclaimed heat from the only candidate, a new stack got an immediate refusal whose whole claim is that waiting cannot help. Caught by CI on macOS, where teardown is slower.
  • The journal stamps PRAGMA user_version, and a daemon refuses to open a state root written by a newer build rather than reading a default where the newer build stored meaning. The precedent: the sha256 env-id migration stranded rows that then held ports and counted against POOL_MAX_TOTAL forever, because nothing on disk said what wrote them.
  • A daemon that fails to start names the reason instead of only pointing at daemon.log.
  • A capacity refusal reports why each environment could not be given up — leased and by whom, busy, an unexpected state, or genuinely too recent — instead of labelling all four "too recent to evict".
  • src/core/version.ts is the single source of version truth; the MCP adapter's second reader is gone (a hand-maintained one had already drifted, shipping 0.4.0 as 0.5.0).

Upgrading

npm i -g backlot@0.9.0   # or however it is installed — backlot never installs itself
backlot update           # restart the daemon onto the build you just installed

Order matters: install first, then update, or the restart just brings the old build back. Skip the second step and every verb except update, doctor and daemon stop refuses with infra-error naming both versions, rather than quietly serving you 0.8.0 behaviour. backlot update --check reports the versions and who would have to rebind, and changes nothing.

Known open

  • #44 — a 160k-file bind on macOS is ~an order of magnitude slower than on Linux (94.8–142.3s vs ~8.5s). Needs a phase-split measurement; tracked, not a regression.

backlot 0.8.0

Choose a tag to compare

@ChristianKohlberg ChristianKohlberg released this 28 Jul 16:56
d3056e3

Three reports from one long agent-fleet session turned out to be a single chain; this release breaks it, and unbundles the environment downward for test lanes. CI green on macOS + Linux.

Highlights

  • up --data-only — lease a database, not an application (#39) — binds an ordinary pooled environment's datastores alone: sync and upkeep run, the namespace is created or restored at its preset, no service starts and no build runs. For a test lane that needs a seeded database per run: read .datastores.<name>.url from ctx, reset-data between runs. Two lanes get two namespaces and neither sees the other's writes — same pool, same lease, same retention, same teardown, no second pool and no new lease kind. The consumer that prompted it was paying container-start + full-backup-restore per test collection outside backlot: 7m34s of actual assertions became ~2h of wall clock with three agents active. Such an environment is published warm, since warm already means "nothing running, everything else intact" (decision 0023).
  • Agent-safe leases (#41) — BACKLOT_HOLDER_PID=$$ backlot up cannot work from an agent harness: each command runs in a fresh shell, so $$ names a shell that has already exited, the lease is born holderAlive: false, and the sweeper's dead-holder rule frees it while the agent is still working. The next binder took the environment and the first caller was left looking at a different, unseeded store through the same URL — read by ~30 agents over 14 hours as a stale seed template, the wrong subsystem entirely. A bind naming an already-dead pid is now refused (exit 64 from the CLI, work-error over the RPC). --ttl is documented as the form for agents and scripts, --holder-pid as the form for callers that outlive the command.
  • pool recycle <env-id> recycles that environment (#40) — the id was parsed and then dropped, so naming one environment recycled the whole pool, including five siblings' live leases. A named target now recycles exactly that one; an unknown id is a usage error rather than a pool-wide teardown; a leased target refuses and names its holder instead of being silently skipped. --force is the accurate name for the override that takes a live lease (--all still works), and a bare recycle reports what it left behind and why.
  • Escaped service processes are reaped eagerly, on every teardown path (#34) — stopAll() is not a stop: a service that called setsid() or spawned a detached grandchild escapes the -pgid signal and keeps holding its port. Deferring the reap to the next bind was the bug — a quiesced env can sit cold for hours, and a stopping daemon has no next anything. bindAndStart, teardownClaimed, the quiesce path, shutdown(), and crash recover() are all now bound by it.

Fixes

  • daemon stop no longer SIGKILLs an in-flight run — a check is spawned detached and carries its env tag, so the eager shutdown reap matched the very process the caller was polling for a verdict from. busy is now respected there as it already was in claimForTeardown, both sweeper branches, and pool gc.
  • The cwd-based teardown sweep no longer reaps a process someone is sitting in front of — killGroupVerified signals the matched pid's whole process group, so a developer who ran cd <env-tree> to look around would have lost their shell and every job in it. scanByCwd now skips any process with a controlling terminal; backlot's own services never have one, so the leak it exists to catch is untouched.
  • teardownClaimed no longer prefers the journal's recorded pid over the live supervisor's — a restart landing in that window would have reaped the stale pid and left the live process behind.
  • All three --data-only contradictions (a named service, --watch, a manifest with no datastore) are refused in the engine, so every RPC client gets them, not just the CLI.
  • status describes a provisioning environment as "being created" rather than "free and quiesced".

Docs & tooling

  • Official Claude Code plugin and marketplace manifest (#38) — /plugin marketplace add ChristianKohlberg/backlot && /plugin install backlot. Skill only; backlot is CLI-only.
  • The event-loop test's ceiling was calibrated for tmpfs, not for CI: measured 94.8s/111.3s on green macOS runs against a 120s timeout, versus ~8.5s on tmpfs-backed Linux. Raised to 300s, with both figures recorded in the test so the next person tempted to trim the tree does not have to re-derive them.
  • The data-only tests now count service processes portably — they asserted liveness via the Linux-only tag scan, which made every "no services started" assertion pass vacuously on macOS.

Known open

  • #44 — a 160k-file bind on macOS is ~an order of magnitude slower than on Linux (94.8–142.3s vs ~8.5s). Needs a phase-split measurement; tracked, not a regression in this release.

backlot 0.7.0

Choose a tag to compare

@ChristianKohlberg ChristianKohlberg released this 22 Jul 18:36
a1cd2bf

Verified by driving a real session before release; CI green on macOS + Linux.

Highlights

  • Selective service startup — backlot up <service...> brings up only the named slice and genuinely stops what it excludes; a released subset hands the next holder the whole app, not the leftover slice, and each slice builds per service rather than gating on the whole-source stamp (#35)
  • Sleep pardon that actually fires on Apple Silicon — darwin sweeps now read the kernel's own record (kern.sleeptime/kern.waketime) and pardon the real sleep gap; the previous wall-vs-mono detector was inert on Apple Silicon (confirmed by a lid-close test). Every pardon logs a pardon event naming the gap and detector (#28)
  • sync no longer full-rebinds on a one-line edit — routes through the 2s watch projection when every service declares hot_reload: true; any undeclared service keeps the full rebind (projecting under a non-watching process would serve stale code) (#28)
  • bind --ref preserves the lease clock — re-pointing a ref keeps the existing TTL; new bind --ttl overrides it explicitly (#36)

Fixes

  • Reap escaped service processes before rebinding env ports — closes the rebind-time leak of detached grandchildren (#26)

Docs

  • README fronts humans over release notes; objections page; landscape/findings refresh; npm/CI/release badges (#27, #30–#33)

Known open

  • #34 — deleted-cwd orphan reap has a remaining edge case (escaped grandchildren on lazy rebind); tracked, not closed by this release.

backlot 0.6.0

Choose a tag to compare

@ChristianKohlberg ChristianKohlberg released this 19 Jul 21:37
b3f35d7

The manifest is now backlot.yml. stack.yaml keeps loading (when both exist, backlot.yml wins), so existing consumers upgrade without a flag day. The schema ships as schema/backlot.schema.json.

Upgrade note: stack identity is now a full-path hash — sibling worktrees like agent-1/myapp and agent-2/myapp no longer silently share one pool. Environments created by ≤0.5.0 are keyed under old ids; the sweeper reaps them automatically, or run backlot pool recycle --all once after upgrading.

Highlights

  • Big binds no longer stall the daemon — sync runs on a worker thread; a 24k-file bind used to block every concurrent verb for its full duration
  • --watch two-stage reload — saves project files into the environment without bouncing services (dev-server watchers do the reload); saves touching upkeep triggers (lockfiles, migrations) deliberately take the full bind
  • Every repo-declared command is bounded — exec/token/build/probes/upkeep/retention can no longer wedge an environment past --force; BACKLOT_CMD_TIMEOUT_S overrides class defaults
  • Fairness and safety — per-stack capacity queues, holder rebinds bypass the queue, lapsed leases can't jump it, and a quiesce is no longer published as a teardown (decision 0021: a crash mid-quiesce keeps your lease and caches)
  • MCP detached runs — backlot_run_detach / backlot_job / backlot_job_ls
  • mssql template_restore proven live (BACKUP/RESTORE bake+restore, docker-gated test)
  • Progress while queued — a verb waiting on a busy environment heartbeats instead of going silent
  • Journal crash-atomicity (transactions + dangling-lease pruning), sun_path guard, idempotent pull, clean-slate sweeps that respect declared caches, macOS CI fixture fixes, a fully deflaked suite, coverage tooling (npm run coverage), and a nightly soak harness

Verified by use: a full driven session against the founding .NET+Angular+MSSQL monorepo before release.