v0.2.0
Use in your Dockerfile
COPY --from=ghcr.io/0ploy/zpinit:latest /usr/local/bin/zpinit /usr/local/bin/zpctl /usr/local/bin/Pin to this exact release with ghcr.io/0ploy/zpinit:0.2.0.
Try it standalone
docker run --rm -it ghcr.io/0ploy/zpinit:0.2.0 sh
# then: zpinit --version, zpinit --check-config /etc/zpinit, zpctl helpManual binary download
Linux only. Pick the binary for your architecture below and verify against checksums.txt.
Features
-
replicas = Nruns N supervised copies of a service. Set
replicas = Non any service inservices/*.tomlto run N
first-class supervised children of the same command. Each replica
has its own PID and crash budget, and shows up as<name>/<index>
inzpctl status. By default all replicas share the same log file
(LinuxO_APPENDis atomic for line-sized writes, so they don't
tear); operators wanting per-replica files opt in via{index}in
the path (/var/log/api-{index}.logbecomesapi-0.log,
api-1.log, ...). Each replica seesZPINIT_REPLICA_INDEXin env.
zpctl verbs take asvc/Nform to target one replica
(zpctl tail consumer/2,zpctl restart api/0); the bare service
name fans out to all replicas. Listener workloads share a port via
reusePort: truein theirlisten()call (Node >= 22.12.0, Bun,
Deno); see the README's new "Node.js clustering" section for the
full table of trade-offs vs PM2 cluster mode. -
zpinit --doctorpre-flight environment audit. New flag runs
a read-only superset of--check-config: confirms the config dir
layout, parses and validates every service, checks each
command[0]resolves on PATH (or as an absolute path), warns when
a declared node service hasreplicas > 1and the configured node
binary is below 22.12.0 (the reusePort floor), surfaces Bun/Deno
versions when those runtimes are referenced, and reports whether a
zpinit instance is already attached to the control socket. Doctor
probes the configured node binary, not whatever resolves on PATH,
so a config that points at/opt/node-v20/bin/nodeis checked
against THAT binary even if PATH has a newer one. Exit code 0 on
green (warnings allowed), 1 on any FAIL, 2 on WARN-only;
--doctor-quietsuppresses OK rows.
Security
ZPINIT_REPLICA_INDEXis reserved. Setting this key in
globals[env]or per-service[env]is now a config-load error.
The supervisor injects per-replica values at spawn time; allowing
user-supplied overrides via the env-merge chain would let one
service shadow every replica's identity with a single static
value, breaking sharding, log attribution, and any logic keyed on
the index.
Bug Fixes
-
Entrypoint scripts and readiness probes no longer hang boot when
a child wedges in uninterruptible kernel sleep. Previously
entrypoint.runOneand the readiness prober both blocked on the
child's reap channel after issuing SIGKILL on timeout or
cancellation. A child pinned inDstate (e.g. wedged NFS, broken
FUSE, stuck device I/O) cannot be SIGKILL'd until its kernel
syscall completes, so the wait was effectively unbounded and pinned
PID 1 with no higher-level deadline to bail it out (the entrypoint
phase runs before the centralized reaper, so per-script timeouts
could not interrupt the post-kill wait). Both sites now cap the
post-kill wait at 5s and log when the bound trips. Memory is
unchanged: the reaper's per-PID channel is buffered, so the
eventual reap completes correctly even after the wait is
abandoned. -
exit_code_fromretarget onSIGHUPreload no longer risks
shutting the supervisor down for the wrong service. The previous
watcher cancel did not synchronize with the old goroutine's
progress: if the old target reached terminal state at the same
instant a reload retargetedexit_code_fromto a different
service, the old watcher could observe the terminal state and
trigger early-shutdown for a service the new config no longer
references. Watcher installations now carry a generation counter;
the goroutine re-checks under the lock that it is still current
before firing shutdown. -
zpctl tailno longer emits a leading partial log line. The
8KB tail window almost always starts mid-line; the first chunk
rendered was the tail of a line whose head was beyond the window.
When the window starts mid-file, the leading partial fragment is
now trimmed at the first newline so operators only see whole
lines. -
Replica shutdown no longer scales linearly with replica count.
PreviouslystopAlland the reload remove/restart paths walked
every runner serially, waiting up tostop_timeout + reapGrace
per replica. Withreplicas = 64and the default 10s timeout, a
single stuck service could burn ~16 minutes of shutdown budget.
Teardown now groups by filename: between groups stays
sequential (preserves filename-encoded dependency order), within
a group all replicas are signaled and awaited in parallel. The
budget calculation matches the new schedule, soShutdownBudget
reports one(stop_timeout + reapGrace)per logical service
rather than per process.