Skip to content

v0.8.0 — Multi-replica support & dependency remediation

Choose a tag to compare

@anshu8858 anshu8858 released this 21 Sep 03:02
· 15 commits to main since this release
2998e93

Closes the three items left open by 0.7.0, and fixes two problems found while doing them.

🔁 Multi-replica support

Background jobs now take a database-backed lease before running, so the scheduler, the metrics alert sweep and the nightly backup execute on exactly one replica.

Until now the only thing standing between you and double-probing the entire fleet — and duplicate notifications, and two SQLite backups racing each other — was instances: 1 in the PM2 config. docker compose up --scale api=2 broke it immediately.

A holder killed mid-job recovers automatically once its lease expires (JOB_LOCK_TTL_MS, default 120s), with no operator cleanup. The lease is renewed on a heartbeat so a long sweep against a large fleet cannot lose it mid-run.

Verified under a 25-way stampede: exactly one winner in the insert race, exactly one in the expired-lease race, zero winners against a live lease, and 25 impostors all failed to release it. A lock that is only usually exclusive would be worse than none — the failure would be intermittent and nearly undiagnosable.

🔑 Host-key administration

GET /api/v1/ssh-host-keys (editor+) and DELETE /api/v1/ssh-host-keys/:id (admin, audited with the discarded fingerprint).

0.7.0 documented a migration to SSH_HOST_POLICY=tofu that was not practical: reviewing what had been pinned meant sqlite3 inside the container, and every legitimate rebuild meant deleting a row by hand. The listing also flags sharedWithOtherEndpoints — normal for a cluster built from one image, worth investigating otherwise.

🔐 Legacy credentials upgrade themselves

A secret still sealed in the pre-0.7 v1 envelope is re-encrypted to v3 when its row is written for another reason. Never on read, never over a value the request is itself supplying, and never for a vault-wrapped v2 blob. A failed upgrade leaves the stored credential untouched.

📦 Dependencies

24 of 27 Dependabot alerts cleared: hono 4.12.25 → 4.13.8 (six advisories), better-auth 1.6.18 → 1.6.22 (high), nodemailer 9.0.1 → 9.1.1 (high), @hono/node-server → 1.19.17, and transitively nanoid → 3.3.19, postcss → 8.5.28, browserslist → 4.29.0.

🐛 Two bugs this surfaced

2FA verification would have failed at runtime. better-auth 1.6.21 added an account lockout that is enabled by default (10 attempts, 15 minutes) and writes failedVerificationCount, lockedUntil and verified on every verification attempt. The TwoFactor model had none of those columns. The test suite does not cover the 2FA verify path, so a green suite did not clear this — it was found by reading the installed package.

Container start would have aborted on upgrade. The 0.7.0 adoption step reconciles a pre-existing database with db push, which creates every table including ones belonging to later migrations, but then recorded only the baseline as applied. migrate deploy would then fail trying to create a table that already existed, exit non-zero, and the container's start chain would never reach the application. Adoption now records the full migration history.

Also fixed

  • A service's auth token could never be updated — authToken was not destructured in updateService, so it reached Prisma as an unknown field and the update threw, despite the schema accepting it.
  • Rate limiting collapsed to one shared bucket behind two or more proxy hops. better-auth refuses to guess which hop is the client and returns no IP for a multi-value X-Forwarded-For; every caller then shared a single key. trustedProxies now defaults to loopback plus the RFC1918 ranges, overridable with TRUSTED_PROXY_CIDRS.
  • nginx: the 24-hour read timeout is scoped to the SSH WebSocket path rather than every API request; Connection: upgrade is only sent when the client asks (it was breaking keep-alive on every ordinary request); client_max_body_size raised to 25 MB so XLSX imports are not rejected at nginx's 1 MB default.
  • compose: the healthcheck targets /health/ready instead of /health/live, which returned a literal ok and could never fail; added start_period so first-boot migrations do not exhaust the retries; the web container has a healthcheck; both rotate logs; Postgres binds to loopback rather than every interface.
  • The test suite was tripping better-auth's own sign-in limiter — 15 logins in one shared process against a cap of 10 per minute — so suites intermittently failed to collect with 429, hitting a different victim each run.

Upgrading

git pull
docker compose up -d --build

Migrations are applied automatically, including adoption of databases created by pre-0.7 images. Take a backup first, as always.

To actually run more than one API replica, scale after upgrading — the lease makes it safe:

docker compose up -d --scale api=2

Verification

pnpm build, pnpm -r typecheck and pnpm test all pass. 258 tests, up from 208 in 0.7.0 and 35 before this work began.

Both migration paths were re-verified against a copy of a real 125-server database: fresh install applies cleanly, an existing database adopts with all rows intact and exits zero, and a second run is a no-op.

Still open

  • deepmerge-ts (high) is pinned exactly by @prisma/config@6.19.3; reaching the fixed version requires a Prisma major upgrade.
  • vitest 3 → 4 (medium, devDependency) removes poolOptions, which this project relies on to keep every spec in one process against a shared SQLite test database. Deferred rather than destabilise the suite at release time.

Full changelog: https://github.com/deziss/rackmap/blob/v0.8.0/CHANGELOG.md