Skip to content

Startup Lifecycle and Readiness

itcmsgr edited this page Jul 16, 2026 · 1 revision

Startup Lifecycle & Readiness

Type: Architecture (Operator Reference) Scope: Daemon startup phases, systemd readiness contract, startup-pending diagnostic, and the lifecycle status surface Since: v1.221.0 (see status note below) Authority: nftband daemon (cmd/nftband) Terminology: Glossary & Vocabulary


Status: merged for upcoming v1.221.0; not available in v1.220.10. The behavior on this page describes the nftband daemon after the startup-lifecycle change. It applies once v1.221.0 is installed. On v1.220.10 and earlier, the structured phase logs, the lifecycle status object, and the startup-pending diagnostic are not present.


Purpose

This page explains how the nftband daemon starts, how it decides it is ready, and how to read that from logs and status output. It is written for operators who need to answer one question quickly: is the daemon up, and if not, where did it stop?

Startup is modelled as an ordered sequence of lifecycle phases. Each phase transition is written to the journal, and the current phase plus the readiness signals are exposed through the status IPC. You do not need to attach a debugger or inspect the binary to see where startup is.


Lifecycle phases

The daemon moves through 14 phases from process start to shutdown. These are journal-visible lifecycle phases — not user-configurable modules. You cannot enable, disable, reorder, or skip them; they describe the daemon's own progress.

Phase Meaning
PROCESS_START Process entered; runtime context created
RUNTIME_PATHS_INIT Required directories and the PID file created
NFT_MANAGER_INIT nftables backend connection established
OPQUEUE_INIT Async operation queue initialised
CONFIG_LOAD Configuration hash computed for reload tracking
MODULES_INIT Protection/monitoring modules registered and initialised
IPC_SOCKET_INIT Unix control socket bound and its accept loop serving
HTTP_INIT Local HTTP API started
WORKERS_START Modules, stats collector, and watchdog started
READINESS_PRECONDITIONS Mandatory readiness prerequisites evaluated
SYSTEMD_NOTIFY_READY READY=1 sent to systemd (when running under systemd)
RUNNING Steady state; waiting for a shutdown signal
SHUTDOWN_BEGIN Graceful shutdown started
SHUTDOWN_COMPLETE Cleanup finished; process exiting

Each phase is logged as a single structured line:

event=startup_phase phase=MODULES_INIT state=begin  pid=1234 component=registry elapsed_ms=0  total_elapsed_ms=42 error=-
event=startup_phase phase=MODULES_INIT state=complete pid=1234 component=registry elapsed_ms=8 total_elapsed_ms=50 error=-

state is one of begin, complete, failed, or degraded. total_elapsed_ms is measured from process start, so you can see how long the daemon has spent reaching the current phase.


Readiness semantics

Readiness depends on how the daemon is run.

Systemd service (Type=notify):
  NOTIFY_SOCKET present
  READY=1 must be delivered successfully
  a delivery failure is FATAL — the daemon exits rather than run un-ready

Direct execution (from a terminal):
  NOTIFY_SOCKET absent
  the daemon may run without sending a systemd readiness notification

Under systemd, the daemon does not stay alive while failing to report readiness. If READY=1 cannot be delivered, startup fails and systemd sees a start failure — it will not sit in activating indefinitely.

Three points to keep in mind when reading the signals:

  • ready_sent=true means systemd accepted the readiness notification. This is the authoritative "the service is up" signal under systemd.
  • ready_attempted=true on its own does not mean the service is ready. It only means the daemon reached the notification step; the result is carried by ready_sent.
  • A bound IPC socket may exist before the daemon is fully ready. ipc_bound=true means the control socket exists; it is not a readiness signal on its own. The daemon distinguishes ipc_bound (socket exists) from ipc_accepting (the accept loop is serving), and full readiness additionally requires the modules and HTTP API to have started.

Mandatory prerequisites must all hold before READY=1 is sent: runtime paths, module initialisation, IPC socket bound and accepting, HTTP API, and required modules started.


Startup-pending diagnostic

If the daemon has not reported readiness by a threshold, it emits one diagnostic line naming where it is stuck, so the blocked phase is captured in the journal before systemd's start timeout is reached:

event=startup_pending last_completed_phase=CONFIG_LOAD current_phase=MODULES_INIT \
  phase_elapsed_ms=3002 total_elapsed_ms=3002 ipc_bound=false ipc_accepting=false ready_sent=false
  • Default threshold: 60 seconds.
  • It is intended to fire before the default 90-second systemd startup timeout, so the blocked phase is recorded while the unit is still activating.
  • The event is emitted at most once, and is cancelled automatically once the daemon becomes ready, fails, or begins shutdown.
  • NFTBAN_STARTUP_PENDING_SEC is primarily a diagnostic / test override for the threshold. Invalid, zero, negative, or malformed values fall back safely to the 60-second default — the diagnostic cannot be accidentally disabled. Production units do not set it.

Status IPC — the lifecycle object

nftban status (and the daemon status IPC it calls) exposes an additive lifecycle object. Existing status fields are unchanged and remain compatible; the lifecycle object is added alongside them.

Principal fields:

Field Meaning
phase Current lifecycle phase
last_completed_phase Last phase that completed
state begin / complete / failed / degraded
ready Daemon has reached full readiness
ready_attempted The readiness notification step was reached
ready_sent systemd accepted READY=1 (authoritative under systemd)
notify_expected NOTIFY_SOCKET was present (running under systemd)
ipc_bound Control socket exists
ipc_accepting Control socket accept loop is serving
nft_ready nftables backend usable
opqueue_ready Async operation queue usable
required_modules_started Required modules started
degraded_components List of degradable components currently unavailable

Degraded components

Some components are degradable: if they are unavailable, the daemon may still reach readiness rather than refuse to start. These are surfaced, not hidden:

  • nft
  • opqueue
  • watchdog

When a degradable component is unavailable it appears in degraded_components, and the relevant phase is logged with state=degraded.

Degraded readiness does not mean full health. A daemon can be ready=true while running degraded. When triaging, inspect degraded_components — an empty list means no degradable component was reported unavailable at startup.


Child process environment isolation

Child processes launched by nftband do not inherit the daemon's systemd notification variables (NOTIFY_SOCKET, WATCHDOG_USEC, WATCHDOG_PID). Only the main daemon process may send readiness, watchdog, stopping, or status notifications to systemd. This prevents helper processes from emitting notifications on the daemon's behalf, which systemd would otherwise reject and log.


Operator troubleshooting

1. Ask systemd what state the unit is in:

systemctl show nftband.service \
  -p ActiveState \
  -p SubState \
  -p StatusText \
  -p MainPID

StatusText reflects the live phase during startup (NFTBan startup: <PHASE>) and NFTBan ready once readiness is accepted.

2. Read the lifecycle phases from the journal:

journalctl -u nftban.service \
  --since "-10 minutes" \
  --no-pager | grep -E 'event=startup_phase|event=startup_pending'

3. Ask the daemon directly:

nftban status

Interpreting the output

Startup stalled during a phase:

current_phase=MODULES_INIT
last_completed_phase=CONFIG_LOAD
ready_sent=false

The daemon completed configuration load and is blocked or delayed during module initialisation. ipc_bound=false here confirms it has not yet reached the socket phase. Pair this with the unit state — activating with this snapshot means startup is still in progress or stuck at MODULES_INIT.

Healthy, ready daemon:

phase=RUNNING
ready=true
ready_sent=true
ipc_accepting=true

The daemon completed the mandatory readiness prerequisites and systemd accepted READY=1. If degraded_components is non-empty in this state, the daemon is ready but running degraded — investigate the listed components.


See also

Clone this wiki locally