Skip to content

v1.0.0 — Operations console: automation, alerting, patching, drift & time-boxed access

Latest

Choose a tag to compare

@anshu8858 anshu8858 released this 26 Sep 03:27
· 5 commits to main since this release

RackMap 1.0 turns the inventory into an operations console: cron editing and monitoring, runbooks across the fleet,
pluggable alerting, systemd control, patch management, drift detection and time-boxed access — on PostgreSQL 18,
with the OS-user management bugs fixed and the security gaps found along the way closed.

Upgrading from 0.8.x? Read MIGRATION.md first — see Breaking changes below.

Breaking changes

  • PostgreSQL 18 is the only database. SQLite support is gone. Copy an existing SQLite database across with
    pnpm --filter @inv/api db:migrate:postgres (see MIGRATION.md).
  • Docker Compose runs PostgreSQL and requires POSTGRES_PASSWORD (URL-safe). Set POSTGRES_HOST_PORT if 5432 is
    already taken on the host.
  • The API image runs as the non-root node user. Volumes written by older root-run images must be re-owned
    (chown -R node:node /data) or SSH keys under /data become unreadable.
  • Editing and deleting OS users now needs the remote_os_users license feature, like creating them.
  • Checkout and license activation are admin-only; completing a checkout requires BILLING_MODE=simulated, and
    unverified license keys are refused in production.
  • POST /servers/:id/test-alert needs alertChannel:manage (admin) and only sends a test event.
  • Rebuild the web image: nginx changes fix an empty Servers page on 0.8.0 images and add long timeouts for host
    actions.

Added

  • Cron job editor on each server page (editor+). Lists user crontabs from the spool directory,
    /etc/crontab, /etc/cron.d/*, and systemd timers (read-only). Schedule builder with a plain-English
    description and the next run times in the host's time zone, raw editor, diff preview, and "run now".
    Saves are compare-and-set on the file's hash (409 if it changed on the host) and back the previous
    version up to /var/backups/rackmap-cron on the host. Root, system files, and users that are
    root-equivalent on the host (sudo/wheel/docker/… membership or a sudoers rule) need admin.
  • Cron heartbeat monitoring. "Monitor this job" wraps a crontab entry so the host checks in at
    PUBLIC_BASE_URL/api/v1/ping/<token> with the job's exit code; missed, late, and failing runs raise
    alerts and recover automatically. Heartbeats can also be created by hand for any job that can call a URL.
  • Runbooks. Saved, parameterised scripts run against servers chosen by tag, environment, location, or
    name, with target preview, dry run, per-host output, cancel/rerun, and schedules. Parameters reach the
    script as exported environment variables, never string-substituted. Authoring is admin-only; root runs
    requested by editors wait for an admin, and nobody can approve their own run.
  • Alert channels (Settings → Alerts): Slack, Microsoft Teams, Discord, PagerDuty (incidents open and
    resolve), Telegram, email, and HMAC-signed webhooks, each subscribed to chosen events. Delivery goes
    through a database outbox with retries, backoff, Retry-After handling, and a delivery log. Outbound
    requests are pinned to the resolved IP, never follow redirects, and refuse private, link-local, and
    metadata addresses unless allowed. NOTIFY_WEBHOOK_URL / NOTIFY_TELEGRAM_* keep working and appear as
    read-only channels (the webhook body is unchanged).
  • SSL certificates are scanned daily (SSL_SCAN_CRON) and alert at 30, 14, 7, and 1 days before expiry.
  • systemd Services tab — list units, view details and the journal, and start/stop/restart/reload/enable/disable
    them. Protected units (SSH, networking, D-Bus, Docker, systemd-*, targets, mounts) need an admin; aliases are
    resolved on the host so a protected unit cannot be acted on under another name.
  • Patch management — nightly fleet scan of pending and security updates, reboot-required hosts, and kernels for
    apt/dnf/yum/zypper (PATCH_SCAN_CRON), with a fleet page, per-server card, and admin-only apply.
  • Drift detection — nightly configuration snapshots (DRIFT_SCAN_CRON) compared with an accepted baseline, with
    severity-ranked events, acknowledgement, and drift_detected alerts.
  • Time-boxed access grants — temporary OS accounts and SSH keys revoked automatically at expiry, with host-side
    expiry as a backstop and alerts when a revoke keeps failing.
  • Prometheus service discovery (GET /api/v1/prometheus/sd), new exporter series (heartbeats, runbooks, alert
    deliveries, patches, drift, access grants), and an example scrape config plus Grafana dashboard in contrib/.
  • Scheduled PostgreSQL backups with pg_dump (BACKUP_CRON, BACKUP_KEEP); /health/ready reports the
    last backup.
  • Status probe history stores far less. A row is written only when a server's status changes, or once
    per STATUS_SAMPLE_INTERVAL_MS (default 15 min) otherwise — about 96 rows per server per day instead of
    about 1,440. The table is capped at STATUS_MAX_ROWS (default 10,000) on top of the retention days, and
    admins can see its size and clean it (older than N days, or keep the newest N rows) under
    Settings → Maintenance. On upgrade the first prune trims existing history to the cap.
  • A per-account sign-in limit (AUTH_LOGIN_ACCOUNT_RATE_LIMIT_MAX, default 10/min) alongside the per-IP one.
  • CI runs the test suite against PostgreSQL 18 and fails when migrations drift from the schema.

Fixed

  • Creating, editing, or deleting an OS user hung forever. The request never returned and leaked the SSH
    connection. The commands also ignored the server's SSH password, so they only worked with passwordless
    sudo.
  • Deleting an OS user always failed with 400. The query-string booleans were rejected.
  • When sudo rejects the stored server password (RackMap usually logs in with a key, so a stale password only shows up
    on root actions), the OS-user dialogs ask for the sudo password once and retry; it is sent as X-Sudo-Password
    for that request only and never stored.
  • The Servers page showed no servers on 0.8.0 images (nginx 301-redirected the list to a URL the API does not route),
    and the web container always reported unhealthy.
  • The committed PostgreSQL baseline migration was truncated. It was missing two tables and every foreign
    key and index, and was not valid SQL, so fresh installs could not migrate. If migrate deploy already
    failed on it, see MIGRATION.md.
  • docker compose up could not start: the defaults still pointed at SQLite, and the PostgreSQL volume was
    mounted where the PostgreSQL 18 image refuses to start.
  • Nightly backups silently did nothing on PostgreSQL.
  • ensure-baseline.mjs queried SQLite system tables and silently skipped on PostgreSQL.
  • SSL-expiry emails were never sent (lib/mail.ts only logged). Telegram alerts failed for hostnames
    containing _.
  • The test suite no longer reuses rows across runs, and refuses to run against a database whose name does
    not end in _test.

Security

  • The server's SSH password is no longer placed in the remote command line (it was visible in ps on the
    host). Remote scripts are uploaded over stdin into a private temp file, and sudo is probed before any
    password is sent.
  • Editors can no longer grant sudo or privileged group membership (sudo, wheel, docker, …), change or delete
    root-equivalent accounts, or set a password containing a newline (which injected extra chpasswd lines).
    Password changes are no longer written to the audit log.
  • Creating or editing an OS user without server:sudo is also refused, on the host, when a requested supplementary
    group is root-equivalent there — a group with its own sudoers rule (%deploy ALL=…) or gid 0 — even if it is not on
    the static privileged list.
  • Checkout and license activation are admin-only. Checkout completes only with BILLING_MODE=simulated,
    and in production any-key license activation is refused without a license server.
  • POST /servers/:id/test-alert requires alertChannel:manage and sends a test event only.

Dependencies

  • Resolved the open Dependabot alerts: vitest 4.1 (test runner config migrated), deepmerge-ts 8 via a pnpm override
    (Prisma 6.19 pins 7.x), and a single zod 4.6 across packages; better-auth, nodemailer, hono, postcss, browserslist
    and nanoid were already on patched versions.
  • Radix UI, TanStack Router, @vitejs/plugin-react and jspdf-autotable 5 (PDF export moved to the functional API).
  • GitHub Actions: actions/checkout v7, docker/metadata-action v6, docker/setup-qemu-action v4,
    docker/setup-buildx-action v4, docker/build-push-action v7.
  • Docker images stay on Node.js 24 LTS; Dependabot now skips Node majors until the next even release reaches LTS.

Changed

  • PostgreSQL 18 is the only supported database. Compose runs it by default and requires POSTGRES_PASSWORD.
  • New server:cron, alertChannel, heartbeat, and runbook permissions; license features
    remote_cron and runbooks (Pro and Enterprise).

Full changelog: v0.8.0...v1.0.0