Skip to content

v0.2.0

Choose a tag to compare

@github-actions github-actions released this 12 May 23:15
· 48 commits to main since this release
702c989

Use in your Dockerfile

COPY --from=ghcr.io/0ploy/zpinit:latest /usr/local/bin/zpinit /usr/local/bin/zpctl /usr/local/bin/

Pin to this exact release with ghcr.io/0ploy/zpinit:0.2.0.

Try it standalone

docker run --rm -it ghcr.io/0ploy/zpinit:0.2.0 sh
# then: zpinit --version, zpinit --check-config /etc/zpinit, zpctl help

Manual binary download

Linux only. Pick the binary for your architecture below and verify against checksums.txt.


Features

  • replicas = N runs N supervised copies of a service. Set
    replicas = N on any service in services/*.toml to run N
    first-class supervised children of the same command. Each replica
    has its own PID and crash budget, and shows up as <name>/<index>
    in zpctl status. By default all replicas share the same log file
    (Linux O_APPEND is atomic for line-sized writes, so they don't
    tear); operators wanting per-replica files opt in via {index} in
    the path (/var/log/api-{index}.log becomes api-0.log,
    api-1.log, ...). Each replica sees ZPINIT_REPLICA_INDEX in env.
    zpctl verbs take a svc/N form to target one replica
    (zpctl tail consumer/2, zpctl restart api/0); the bare service
    name fans out to all replicas. Listener workloads share a port via
    reusePort: true in their listen() call (Node >= 22.12.0, Bun,
    Deno); see the README's new "Node.js clustering" section for the
    full table of trade-offs vs PM2 cluster mode.

  • zpinit --doctor pre-flight environment audit. New flag runs
    a read-only superset of --check-config: confirms the config dir
    layout, parses and validates every service, checks each
    command[0] resolves on PATH (or as an absolute path), warns when
    a declared node service has replicas > 1 and the configured node
    binary is below 22.12.0 (the reusePort floor), surfaces Bun/Deno
    versions when those runtimes are referenced, and reports whether a
    zpinit instance is already attached to the control socket. Doctor
    probes the configured node binary, not whatever resolves on PATH,
    so a config that points at /opt/node-v20/bin/node is checked
    against THAT binary even if PATH has a newer one. Exit code 0 on
    green (warnings allowed), 1 on any FAIL, 2 on WARN-only;
    --doctor-quiet suppresses OK rows.

Security

  • ZPINIT_REPLICA_INDEX is reserved. Setting this key in
    globals [env] or per-service [env] is now a config-load error.
    The supervisor injects per-replica values at spawn time; allowing
    user-supplied overrides via the env-merge chain would let one
    service shadow every replica's identity with a single static
    value, breaking sharding, log attribution, and any logic keyed on
    the index.

Bug Fixes

  • Entrypoint scripts and readiness probes no longer hang boot when
    a child wedges in uninterruptible kernel sleep.
    Previously
    entrypoint.runOne and the readiness prober both blocked on the
    child's reap channel after issuing SIGKILL on timeout or
    cancellation. A child pinned in D state (e.g. wedged NFS, broken
    FUSE, stuck device I/O) cannot be SIGKILL'd until its kernel
    syscall completes, so the wait was effectively unbounded and pinned
    PID 1 with no higher-level deadline to bail it out (the entrypoint
    phase runs before the centralized reaper, so per-script timeouts
    could not interrupt the post-kill wait). Both sites now cap the
    post-kill wait at 5s and log when the bound trips. Memory is
    unchanged: the reaper's per-PID channel is buffered, so the
    eventual reap completes correctly even after the wait is
    abandoned.

  • exit_code_from retarget on SIGHUP reload no longer risks
    shutting the supervisor down for the wrong service.
    The previous
    watcher cancel did not synchronize with the old goroutine's
    progress: if the old target reached terminal state at the same
    instant a reload retargeted exit_code_from to a different
    service, the old watcher could observe the terminal state and
    trigger early-shutdown for a service the new config no longer
    references. Watcher installations now carry a generation counter;
    the goroutine re-checks under the lock that it is still current
    before firing shutdown.

  • zpctl tail no longer emits a leading partial log line. The
    8KB tail window almost always starts mid-line; the first chunk
    rendered was the tail of a line whose head was beyond the window.
    When the window starts mid-file, the leading partial fragment is
    now trimmed at the first newline so operators only see whole
    lines.

  • Replica shutdown no longer scales linearly with replica count.
    Previously stopAll and the reload remove/restart paths walked
    every runner serially, waiting up to stop_timeout + reapGrace
    per replica. With replicas = 64 and the default 10s timeout, a
    single stuck service could burn ~16 minutes of shutdown budget.
    Teardown now groups by filename: between groups stays
    sequential (preserves filename-encoded dependency order), within
    a group all replicas are signaled and awaited in parallel. The
    budget calculation matches the new schedule, so ShutdownBudget
    reports one (stop_timeout + reapGrace) per logical service
    rather than per process.