Releases: rahanahu/wgft
Release list
v1.2.1
Notes
v1.2.1 is a patch release of the v1.2 series with two fixes to the agent's userspace mode, the default mode.
- Security fix: a client could end a userspace-mode agent process by connecting to a published TCP port and resetting the connection at once. When the reset arrived while the connection waited to be accepted, the agent's relay read a missing remote address and panicked. The agent now refuses such a connection and keeps serving the rule; it logs these refusals at most once a minute per port. The refused connection does not reach the Admission Policy and takes no flow slot. Windows and macOS agents always run in userspace mode and are affected. Agents in kernel mode do not relay TCP themselves and are not affected. Advisory: GHSA-w3wj-h525-39q7.
- Fix: the agent's keepalive ping through the tunnel no longer leaves an ICMP endpoint behind on every ping. Before, each ping kept one endpoint and one ICMP identifier until the agent rebuilt the tunnel or restarted. Once the identifiers ran out, every ping failed without a log line, and if the VPS side lost the WireGuard session, the tunnel waited for the key lifetime to come back instead of the next ping.
- Upgrade the agent: both fixes are in the agent. Upgrade the agent, the binary or the
wgft-agentimage, to get them. The server's behaviour does not change in practice. - Upgrade and revert: a binary replacement with no schema change. Reverting to v1.2.0 works with the same database and agent credentials.
Changelog
v1.2.0
Notes
-
New: kernel mode for the Linux home agent. With
WGFT_MODE=kernel, the agent creates a kernel WireGuard interface,wgft0, andtable inet wgft_agent, and the kernel forwards to the LAN targets with DNAT. The agent process relays nothing, so forwarding continues while the agent is stopped or restarting. The default stays userspace mode.- The provided
agent.servicestays unprivileged. The new drop-indeploy/agent.kernel.confaddsCAP_NET_ADMINand nothing else. - Kernel mode forwards only to IPv4 targets and refuses loopback targets. The agent sets
net.ipv4.ip_forwardto 1 when it is 0, so the home host routes packets for any traffic. wgft agent teardownis new. It removeswgft0, the table, the conntrack entries of the flows the agent forwarded and the kernel-mode records inagent.json, so the agent can go back to userspace mode. It refuses while the agent runs.wgft agent doctorchecks the interface, the table and IP forwarding of a kernel-mode agent.- See README, "Kernel mode on the home agent", and docs/setup.md, "Run the agent in kernel mode".
- The provided
-
New: agent disable and enable.
wgft agent disable <name>stops forwarding for all of an agent's rules and keeps the registration, the keys and the rules.wgft agent enable <name>undoes it. The Web UI offers both.wgft status,wgft server doctorandwgft agent doctorshow a disabled agent. See README, "Disabling an agent", and docs/cli.md. -
New in the Web UI:
- The dashboard's rule list shows a diagnosis mark next to each rule's state. It names the node where the rule's traffic stops and opens that rule's diagnostics page. The error counts in the header and the group headings count these marks.
- Each agent has a detail page. Deleting an agent moved there, into a danger zone that asks for the agent's name. It can delete the agent's rules together with the agent.
- The dashboard shows a band for rules left by a deleted agent, with a button that deletes them.
-
New: last UDP reply seen. For a UDP rule,
wgft server doctoradds a line undertargetsaying when the server last saw a reply from the target, and the Web UI diagnostics page shows the same. It never changes a status: a UDP target still reads NOT TESTED. -
New: rule IDs as
rule lsshows them. Rule commands also accept the shortened ID fromrule lswith its trailing…or.... Commands that doctor and status suggest now use the full ID. -
Changes to machine-readable values. Read these if a script reads
--jsonor exit codes.- Healthy UDP rule,
server doctor --json: therule.targetcheck'schecks[].statuschanges fromoktonot_tested, with the newreasonudp_listener_onlyandnextconfirm the service from a real client.rules[].status, the top-levelstatusand the exit code 0 do not change. TCP rules do not change. The agent's report for a UDP rule shows only that its listener is open, which does not meet the documented meaning of OK, "observed to succeed". Under design.md 7a.11 this is a bug fix: the value contradicted a meaning that design.md defined before this change. (#203) - An agent's own listener bind failure,
server doctor --json:rule.targetnow readslistener_bind_failedinstead oftarget_error, and its next step points at the agent's own earlier listener or connection instead of the target service. This is seen with userspace-mode agents older than v1.1.2, and with current userspace-mode agents when a rule is re-added or re-enabled at the same port shortly after a session in which the target closed the connection first while the client's side was still open. design.md had not defined either code's meaning before this change, so it is an explicit classification change in v1.2.0, not a 7a.11 correction. (#245) - A kernel-mode agent's rule whose target name stopped resolving: the agent keeps forwarding to the address from the last successful resolution and says so in the rule's reason.
server doctor --jsonnow readsrule.target_resolveasunknowninstead offailed; thereasonstaystarget_resolve_failed.rule.targetis judged from the rest of the reason: with nothing after it, the rule has nostopped_atand the exit code is 0 unless something else fails; a probe error after it still stops the rule attarget. The human output showsDEGRADEDon that line. design.md now definesDEGRADEDby meaning: a check that is not judged a stop but is confirmed to have degraded from the normal state.DEGRADEDnever appears in--json. Inagent doctor --json,dataplane.tablereadsunknownwithresolve_failedwith exit 0 when such rules are its only rule errors. design.md rewrote its case list in the same change, so this is an explicit classification change in v1.2.0, not a 7a.11 correction. Only kernel-mode agents, new in this release, send this report, so no existing deployment sees a value change.server doctorfrom v1.1.x still says traffic stops at "target resolve" for such a rule. (#260) - Situations that now get their own reason code: design.md 7a.11 treats these as additions to open sets; existing values keep their meaning.
- The agent host's
ip_forwardis 0:server doctorreadsrule.targetasagent_ip_forward_off. Only kernel-mode agents, new in this release, send this report. (#274) - The VPS's
ip_forwardis 0 while a kernel-mode server runs and at least one published rule is forwarded by the kernel:server doctornow failsserver.dataplaneand the public port of each such rule withip_forward_off, and exits 1.wgft statusmarks Server degraded in the same case and exits 1. Before, both reported no failure. Proxy-mode rules do not depend on the value, so a server with only proxy-mode rules gets a note, not a failure. The server does not set the value back to 1 while it runs. (#274) agent doctor, the agent's pinned server certificate does not match:stream.connectionreads FAILEDserver_cert_mismatch, where it read UNKNOWNreconnecting. The top-levelstatusand the exit code do not change. (#274)agent doctor, the data directory cannot be read: the reason is the newdata_dir_unreadable. It used to becredentials_unreadable, which design.md already defined asagent.jsonexisting and not being readable; that code is now given only in that case. The top-levelstatusand the exit code do not change. (#274)
- The agent host's
- Added values: design.md 7a.11 treats check ids and reason codes as open sets: a consumer must read an unknown value as unknown, not as a failure.
server doctor --json: the check idagent.enabled; the reasonsagent_disabled,target_loopback_unsupported,agent_ip_forward_offandip_forward_off; the optionallast_reply_at,reply_sinceandreply_not_observedon a UDP rule'srule.target.agent doctor --json: the check idsdataplane.interface,dataplane.tableandhost.forwarding, which a userspace-mode agent also reports, asnot_testedwith the reasonuserspace_mode; the reasonsagent_disabled,needs_cap_net_admin, the kernel-mode codes that design.md section 10.2c lists,server_cert_mismatchanddata_dir_unreadable.status --json:agents.disabledandrules.agent_disabled, always present.- Admin API:
POST /api/v1/agents/{name}/disableand/enable;disabledanddisabled_atin the agent list;udp_repliesand, on a kernel-mode server,ip_forwardin the rules response;ip_forwardinGET /api/v1/server;agent_disabledinGET /api/v1/agents/{name}/state, where a disabled agent's rules also readenabled: false.
- Agent-supplied text: the server now caps the text an agent sends in its heartbeat and replaces unprintable characters before storing it: 64 bytes for a state, 512 for a reason, 128 for a rule ID or an endpoint. This reaches
agent ls --jsonandrule ls --json. design.md 7a.11 treats this as processing that does not change the values' meaning. (#270)
- Healthy UDP rule,
-
Upgrade the server: v1.2.0 moves the server database to schema version 9 on its first start. v1.1.x and earlier refuse to open it. On a staging machine, v1.1.3 exited with code 3 and left the database unchanged. Reverting is not promised: back up the data directory before upgrading. Moving back to v1.1.1 through v1.1.3 needs a backup taken before the upgrade to v1.2.0, and moving back to v1.1.0 or earlier needs one taken before the upgrade to v1.1.1 or later.
-
Upgrade the agent: an agent without
WGFT_MODEstays in userspace mode. An agent that already receivesWGFT_MODE=kernel, for example from a settings file it shares with the server, ignored it up to v1.1.x. v1.2.0 starts it in kernel mode, and under the provided unprivileged unit it stops at startup with exit code 3 and namesCAP_NET_ADMIN. It stops before it uses the join string or records anything inagent.json. Give the agent noWGFT_MODE, orWGFT_MODE=userspace, and restart the unit. -
Move a kernel-mode agent back to v1.1.x: v1.1.x has no kernel mode for the agent. With the v1.2.0 binary, stop the agent, run
sudo wgft agent teardown, and removeWGFT_MODE=kerneland the drop-in. Then install v1.1.x and start the agent. This order was verified on a staging machine. Without the teardown, v1.1.3 started on a staging machine, but it dropped the kernel-mode records fromagent.jsonand leftwgft0andtable inet wgft_agentin the kernel. What they do to forwarding has not been verified. If v1.1.x already started without the teardown,wgft agent teardownof v1.2.0 still removeswgft0and the table, but no longer finds the conntrack entries or theip_forwardrecord. See docs/setup.md, "Run the agent in kernel mode". -
Behavior changes:
- Interface names: a server whose
WGFT_WG_INTERFACEcontains a byte the kernel refuses in an interface name, or one that makes the name differ from the one requested, now exits with...
- Interface names: a server whose
v1.1.3
Notes
- Fix: on a kernel-mode server, an Admission Policy rule's
source_denyorsource_allowlist with 1638 or more separate entries, counting adjacent or overlapping ones as one, could load into nftables wrongly, with no error. The server took the table as it was loaded, reporting the rule as active.- What was seen: lab testing found the set loaded empty or partial on kernel 6.1. A real VPS running v1.1.2 on kernel 6.12, with 1700 addresses in one list, loaded the set partial and also holding a wrong, very wide range: 62 elements ending in
100.64.0.123-255.255.255.255;server doctorsaid OK. In the lab, a 5000-entry list made the kernel refuse the whole update; that failure was reported, and the previous table stayed in place. - Effect for a deny list: listed sources could get through, or unrelated sources could be dropped.
- Effect for an allow list: sources outside the list could get through, or listed sources could be dropped.
- Affected versions: every release up to and including v1.1.2.
- What was seen: lab testing found the set loaded empty or partial on kernel 6.1. A real VPS running v1.1.2 on kernel 6.12, with 1700 addresses in one list, loaded the set partial and also holding a wrong, very wide range: 62 elements ending in
- Now: the server sends set elements in chunks and reads each source list's set back. It compares what was loaded with what was sent, and treats a mismatch as a failed publication instead of accepting it.
- Verified in the lab on kernel 6.1 with deny and allow lists of up to 20000 entries. Verified on a real kernel-mode VPS on kernel 6.12 with the v1.1.3 candidate: deny and allow lists of 1700 and 5000 entries loaded exactly, and real traffic was filtered correctly in both directions.
- After a failed publication, reapplies triggered by nftables notifications are spaced out from 1 s up to 30 s; this spacing was verified by unit tests only. Admin changes and newly found drift are applied at once, and the normal 30-second retry is unchanged.
- How to check whether you were affected (kernel mode, before upgrading):
wgft rule lsshows each rule's DENY and ALLOW counts; a list with fewer than 1638 CIDRs was not affected. For a larger one, find its set insudo wgft server nft(orsudo nft list table inet wgft): the row carrying the commentwgft:<rule id>:denyor:allowuses@deny_<n>or@allow_<n>. Thensudo nft list set inet wgft deny_<n>(orallow_<n>): an empty set, missing addresses, or a range your list does not contain, such as one ending in255.255.255.255, means the rule was affected. After the upgrade the server loads the full list, so check first. - Upgrade the server: upgrade the server binary. Userspace mode, including the server image, does not use nftables for these lists and is not affected. The agent does not change.
- Known effect: on the kernel 6.12 VPS running the v1.1.3 candidate, publishing several 5000-entry lists back to back briefly overflowed the server's nftables change watcher. It logs
watching the data plane for changes failed: nftables notifications: netlink receive: recvmsg: no buffer space available; subscribing again, and checking every few minutes meanwhile, resubscribes, and drift repair kept working afterward. It is harmless. - Upgrade and revert: a binary replacement with no schema change. Reverting to v1.1.2 works with the same database, but brings the bug back.
Changelog
v1.1.2
Notes
- Fix: when a TCP rule's target changes while a session is open, the rule now listens again at once. Before, the agent held the port and the rule stayed down for 60 s or more. On a kernel-mode VPS whose input firewall drops by default, disabling a rule with an open session and re-enabling it at once could keep it down for more than ten minutes (about 13.5 minutes was measured on a real VPS). It now comes back at once.
- Upgrade the agent: the fix is in the agent. Upgrade the agent, the binary or the
wgft-agentimage, to get it. The server's behaviour does not change in practice. - Sessions cut on purpose now end with a TCP reset between the agent and the VPS. What the client sees depends on the server mode: with a userspace-mode server the client should still see a normal close; with a kernel-mode server a target change reaches the client as a reset, and a disable or delete reaches it as nothing, so the client notices nothing until it next sends. On a real kernel-mode VPS whose input firewall drops by default, these were observed with v1.1.2: a target change reached the client as a reset, and a rule disabled and re-enabled, or deleted and added again, or an agent restart, reached it as nothing until it next sent, then as a reset. The userspace-mode case follows from the code and has not been observed directly.
- Known limit: a session that ended on its own with the target closing first, such as HTTP, is expected to hold the rule's port for about 60 s. If the rule is disabled and re-enabled or retargeted during that time, its listener opens on the first 30-second retry after the hold ends.
- Upgrade and revert: a binary replacement with no schema change. Reverting to v1.1.1 works with the same database.
Changelog
v1.1.1
Changelog
- def9fd0: chore(release): v1.1.1 (#202) (@rahanahu)
- 92f0f35: fix(cli): move agent doctor's value-only checks to Observed values (#201) (@rahanahu)
- 47defa6: fix(server): keep a dismissed ip-mismatch pair acknowledged (#199) (@rahanahu)
- 38b7a9c: fix(server): report a short rule-set generation lag as pending (#198) (@rahanahu)
- e5182ac: fix(ui): give rate limit fields accessible names and error links (#200) (@rahanahu)
v1.1.0
Changelog
- 9468d92: chore(release): v1.1.0 (#197) (@rahanahu)
- 892442a: ci: compute the Windows and macOS test packages (#183) (@rahanahu)
- e8f362b: docs(design): align agent doctor with the implementation (#173) (@rahanahu)
- dc202e4: docs(design): apply four design revisions to design.md (#167) (@rahanahu)
- 7005ae5: docs(design): design the agent-side doctor command (#163) (@rahanahu)
- ecfdd00: docs(server): fix section reference to docs/design.md (#171) (@rahanahu)
- 9b7f46f: docs(ui): retake the dashboard screenshots for the diagnostics link (#179) (@rahanahu)
- 4e7e43d: docs: describe diagnostics in the README and retake the screenshots (#192) (@rahanahu)
- 877b6e9: docs: lift the v1.0 internal structure freeze (#164) (@rahanahu)
- b535985: feat(agent): add --json to agent doctor and settle its reason codes (#189) (@rahanahu)
- 46d5244: feat(agent): add wgft agent doctor with the static checks (#175) (@rahanahu)
- 00df6e6: feat(agent): answer doctor on the agent control socket (#176) (@rahanahu)
- 9826beb: feat(agent): hold the control stream observation in the runtime (#165) (@rahanahu)
- 914fd08: feat(agent): read the running agent's state in agent doctor (#180) (@rahanahu)
- 8bf5f49: feat(relay): add the relay state agent doctor needs (#174) (@rahanahu)
- a93c748: feat(ui): add the diagnostics page to the Web UI (#177) (@rahanahu)
- c32913f: feat(ui): draw the diagnosis as a path of nodes (#190) (@rahanahu)
- 05b832d: feat(ui): save rule settings with one button and group the detail page (#187) (@rahanahu)
- f646e53: fix(agent): name the agent's user when doctor runs as root (#193) (@rahanahu)
- 2fbf08b: fix(agent): report host.privileges as unknown when doctor runs as root (#185) (@rahanahu)
- bc9d6a1: fix(cli): accept --help=true in the non-Linux server stub (#188) (@rahanahu)
- 2388997: fix(cli): exit 3 when the server group runs on a non-Linux build (#182) (@rahanahu)
- 0920562: fix(cli): report an unreadable config file in agent doctor (#178) (@rahanahu)
- b1daec3: fix(flock): inspect the lock state without creating the lock file (#172) (@rahanahu)
- af75255: fix(server): classify the agent's probe-timeout text as target_timeout (#191) (@rahanahu)
- 250f799: fix(server): stop claiming a specific rule has not arrived on generation lag (#194) (@rahanahu)
- cf348cc: fix(server): stop suggesting --probe for UDP rules in doctor (#186) (@rahanahu)
- bf04c6a: fix(ui): route PROXY and Gen through i18n and annotate CLI-only hints (#181) (@rahanahu)
- 7f317b2: fix: remove round parentheses from English tool output (#166) (@rahanahu)
- 1bec645: test(agent): stop scanning a fixed UDP port range in rebuild tests (#170) (@rahanahu)
- 7634b2a: test(server): start idle timer measurement before the request (#184) (@rahanahu)
v1.0.0
Changelog
- f859f3f: chore(release): v1.0.0 (#162) (@rahanahu)
- 32b0267: docs(design): record the doctor and status field test (#159) (@rahanahu)
- 22e629f: docs: point to screenshot script instead of repeating its instructions (#155) (@rahanahu)
- fd17f8f: fix(server): hold the startup when the first rule apply fails (#160) (@rahanahu)
- d495cc4: fix(server): stop claiming forwarding direction the hold cannot tell (#161) (@rahanahu)
- be982eb: fix(ui): validate reserved ports in the read-import confirmation (#158) (@rahanahu)
- 5937c48: test(cli): anchor status RunE fixture timestamps to wall clock (#157) (@rahanahu)
v0.7.0
Changelog
- 7effd3f: build: tag the Linux-only packages so Windows can build the tree (#110) (@rahanahu)
- c31913d: chore(release): v0.7.0 (#154) (@rahanahu)
- 7bcf77f: ci: skip the Go checks when a change touches only documents (#147) (@rahanahu)
- 3148b69: docs(agent): record remote-VPS ICMP measurement for Windows bind fix (#115) (@rahanahu)
- acc5595: docs(design): correct what happens when an apply keeps failing (#139) (@rahanahu)
- fca402e: docs(design): record that the macOS clock stops during sleep (#137) (@rahanahu)
- 0debdf5: docs(design): record the tunnel rebuild seen in the lab and on a real link (#140) (@rahanahu)
- 937915c: docs(design): say that a roaming agent also raises ip-flapping (#138) (@rahanahu)
- 14bcf58: docs(design): scope the v1.0 contract to Linux (#152) (@rahanahu)
- d4d88b8: docs(design): state that TCP targets are rechecked periodically (#123) (@rahanahu)
- f393e9d: docs(readme): draw the WireGuard tunnel as a two-way edge (#146) (@rahanahu)
- 6871d15: docs(setup): describe large rule sets in kernel mode (#134) (@rahanahu)
- 07a52e7: docs(setup): note public ports inside the ephemeral range in userspace mode (#129) (@rahanahu)
- 073bdaa: docs(setup): note the unsigned-binary warning on Windows (#111) (@rahanahu)
- 87106ad: docs(setup): verify the Windows download against its sha256 (#106) (@rahanahu)
- caa3861: docs(testing): note WSL2 mirrored networking caveats for D1 (#116) (@rahanahu)
- 157c297: docs(testing): update inprocess forwarding test status (#113) (@rahanahu)
- f911723: docs: replace the contract wording with plainer Japanese (#153) (@rahanahu)
- fe77cea: docs: replace word-by-word renderings with natural Japanese (#126) (@rahanahu)
- 1308492: docs: say 通し instead of a literal end-to-end rendering (#114) (@rahanahu)
- 1877293: feat(cli): add --dry-run to rule add and rule set (#150) (@rahanahu)
- 26db00e: feat(cli): add server doctor to find where forwarding stops (#148) (@rahanahu)
- afdf8bb: feat(cli): add status to summarize deployment health (#149) (@rahanahu)
- e2e284b: feat(cli): show the protocol version in version and agent ls (#151) (@rahanahu)
- 8cb6117: fix(agent): explain a control socket path that is too long (#120) (@rahanahu)
- 80332a5: fix(agent): keep the Windows tunnel receiving after a connection reset (#107) (@rahanahu)
- d3de564: fix(agent): notice a dead control stream in seconds, not minutes (#142) (@rahanahu)
- a41b35a: fix(agent): probe rule targets without holding the relay lock (#145) (@rahanahu)
- c1f5b56: fix(agent): rebuild the tunnel when its receive path dies (#122) (@rahanahu)
- ebee6fb: fix(agent): retry a tunnel build that failed while applying a state (#130) (@rahanahu)
- 67018b4: fix(cli): point the Windows double-click message at PowerShell (#112) (@rahanahu)
- ffe9cb5: fix(nft): size the netlink buffers to the table replacement (#132) (@rahanahu)
- c0248ca: fix(server): advance the generation when a rule moves to another agent (#144) (@rahanahu)
- 64f9e13: fix(server): cut proxy sessions when the effective target changes (#143) (@rahanahu)
- 3e52e43: fix(ui): refresh the rules panel and the health block, and keep the toggle working (#141) (@rahanahu)
- 7f5c4df: refactor: use the same natural Japanese terms in comments (#127) (@rahanahu)
- 55d0c0a: test(agent): forward TCP and UDP through both sides in one process (#102) (@rahanahu)
- 38d90f9: test(lab): add the scale check (C3) (#131) (@rahanahu)
- cebe78b: test(lab): check that a stopped agent is shown as a last report (#121) (@rahanahu)
- eb108c1: test(lab): reserve the lab's fixed ports from the ephemeral range (#128) (@rahanahu)
- 7708de4: test(lab): use v0.6.0 as the previous release in the upgrade and skew checks (#125) (@rahanahu)
- 85fe4d4: test(lab): wait for the settled heartbeat around the agent's stop (#133) (@rahanahu)
v0.6.0
This release changes which startup failures make systemd stop and which make it retry, gives rule ls the agent's side of a rule's status, and fixes several places where a failed read was shown as "nothing". No change to the wire protocol or the data format. The server and the agents can be upgraded in either order.
Changes that affect running deployments
- Exit codes at startup follow one rule now: can a retry fix it? The default is exit 1, which the shipped units retry. Exit 3, which they do not retry, is reserved for a cause that is the configured value itself, or that cannot go away without an operator. Compared with v0.5.1:
- Now exit 1 (retried), was exit 3: an interface of the configured name that is not wgft's; the WireGuard port held by another WireGuard interface or by another process; an address range that overlaps another interface; the conntrack timeout values still unreadable after the table was applied. The other owner may release the resource, and with exit 3 the server stayed down after the conflict was gone. The checks still stop before anything is written, and the message still says what to change; expect a log line every two seconds while the conflict lasts.
- Now exit 3 (not retried), was exit 1: an empty
WGFT_DATA_DIR; aWGFT_WG_INTERFACEthe kernel rejects;WGFT_MTUoutside 576 to 9216;WGFT_WG_PORT=0; aWGFT_WG_ENDPOINTorWGFT_AGENT_API_HOSTthat is not host:port; a database written by a newer wgft; and an agent's registration answered 400, 401 or 409 (a name the server rejects, a join string that is spent or rejected, a name already taken). - Now refused, was silently ignored: a
WGFT_ADMIN_TAILSCALEthat is not a boolean.tureused to mean "off".
- Refusals name their cause. Every exit-3 message reads
refusing to start [<category> <subject>]: <reason>, with the category one ofconfig,prerequisite,conflict,mode-gate. A script that matches the oldrefusing to start: ...text needs the bracket. - Exit 3 stops the restart loop under the shipped systemd units only. launchd restarts an agent on exit 3 as on any other exit (checked on macOS 27), and
restart: unless-stoppedin the compose files ignores exit codes. - Pages and API routes that showed invented content now answer 500. When the server's database could not be read, the dashboard, the rule page, the add-rule form and the two import routes answered 200 with generation 0, zero drop counts, "no agents registered" or "rule not found". The import guard could even match its own invented generation and let a stale import through; it now stops.
New
wgft rule lsshows each rule's status on its agent. A rule refused by the agent'sWGFT_AGENT_ALLOW_TARGETS, or whose listener or target has an error on the agent, used to show its reason only inwgft agent lsand the Web UI.rule lshas a newAGENT_STATEcolumn, andGET /api/v1/rulesandrule ls --jsonhaveagent_rule_states, one entry per rule withagentandconnectedalways present andstate,reasonandatonce the agent has reported on that rule.connected: falsemarks a disconnected agent's last report as history. This closes the known issue listed in the last two releases.- What v1.0 will keep compatible is written down, per surface, in the design document: the admin API's routes and fields,
--jsonas the only machine interface of the CLI (the tables, their columns and the help text are for people and may change), theWGFT_*names, the wire protocol's negotiation, the upgrade path, unit and artefact names. Fields are only added; values such asstateandapply_stateare open sets that a consumer must read as "unknown" when it does not know them.
Fixes
agent ls --jsonno longer prints"last_handshake":"0001-01-01T00:00:00Z"insidetunnelfor a tunnel that never shook hands: the key is absent, and when present it is RFC3339 like the timestamps beside it. The wire format between agent and server is unchanged.rule_states[id].active_generationis absent, not0, for a rule that was never published.wgft server teardownwarns when no interface name was recorded, instead of silently assumingwgft0; what it deletes is unchanged.- An unparseable agent address in the database fails the read instead of becoming the zero address in the WireGuard peer.
- The mode-change gate, the connection check and the CLI's socket hint test error types instead of matching error texts.
- Messages: the shutdown log names the real interface and mentions the table only in kernel mode; a failed apply in userspace mode no longer blames nftables; the agent calls its
agent.json"credentials" and the server calls its database "server database" in every message.
Known issues
- A rule's failure on the agent side is shown per rule, but a rule's server-side apply state and its flow-budget refusals are still separate fields (
rule_states,resource_refusals);rule ls --jsonhas all three. - In kernel mode, host-side conntrack sizing is the operator's responsibility:
wgft server checksays whennf_conntrack_maxis below 65536. - Under investigation, not confirmed on a real machine: a Windows agent may stop receiving on its tunnel after it sent handshakes to a server whose WireGuard port was closed (a userspace-mode server that was stopped, or a VPS while it reboots), and stay that way until the agent is restarted. If you encounter this and forwarding does not return after the server is back, restarting the agent is the current workaround.
Upgrading
- Back up the server's data directory before upgrading. Downgrading afterwards is not supported.
- Replace the binary (or container image) and restart the service. A rolling upgrade from v0.5.1, v0.5.0 or v0.4.0 is supported, in either order. Upgrade compatibility is covered by the lab tests; the v0.5.1 path is additionally exercised through the shipped systemd units and machine reboots in the VM test.
- Run
wgft server check, and confirm inwgft rule lsandwgft agent lsthat the rules areactive, thatAGENT_STATEshowsok, and that the agents showok. If your monitoring looks at the unit's state: a conflict with another WireGuard interface or port now shows as a restarting unit, not a failed one.
Changelog
- f63651e: chore(release): v0.6.0 (#105) (@rahanahu)
- b396ffe: ci: run platform, release and vulnerability jobs only when they apply (#100) (@rahanahu)
- ad88662: docs(design): state the v1.0 compatibility contract per surface (#94) (@rahanahu)
- 4abc465: docs: fix the internal structure until v1.0 (#92) (@rahanahu)
- d818b12: docs: record what the macOS agent check found on a real Mac (Claude Opus 5 (1M context) noreply@anthropic.com)
- 35e4c9a: fix(server): call the database a database and say only what a stop keeps (#99) (@rahanahu)
- ada2ca3: fix(server): report agent-side rule status and tidy two JSON fields (#101) (@rahanahu)
- 5b38397: fix(server): settle startup failures by whether a retry can fix them (#97) (@rahanahu)
- bc32196: fix(server): stop failed reads from looking like empty or zero (#96) (@rahanahu)
- b4e643f: refactor(server): drop unused parameters and types, correct three messages (#91) (@rahanahu)
- 4d36762: test(deploy): check the binary-only upgrade in the distributed VM test (#104) (@rahanahu)
- 9333ee9: test(lab): move the previous release in B7 and D4 to v0.5.0 (#95) (@rahanahu)
- 31ec1f7: test(lab): provision a Ubuntu or Fedora Lab Host and record the results (#103) (@rahanahu)
- 367af4d: test(lab): wait for convergence before wall-clock windows (#98) (@rahanahu)
v0.5.1
A patch release: fixes only, no new settings, no change to the wire protocol or the data format. The server and the agents can be upgraded in either order.
Fixes
- A momentary database error could park a healthy agent. When the server's database failed while it authenticated an agent's stream, the server answered
401, which the agent reads as "the permanent token was revoked": it stopped retrying and waited to be re-enrolled. The server now logs the cause and answers500, and the agent reconnects with its normal backoff. An invalid token still gets401. - A database error during registration no longer reads as key theft. The same kind of failure while the server checked an agent's WireGuard public key was reported to the agent as "public key belongs to another agent", with nothing in the server's log. It is now logged and reported as an internal error.
wgft rule lsno longer prints invented values. If reading the generation or the drop counters failed, the admin API returned200withgeneration 0and all-zeroDROPPEDcounts. It now returns500, andrule lsfails with the cause, which names the read that failed.- A disconnected agent is no longer shown as healthy. After an agent's stream dropped, the Web UI kept showing its last tunnel state as a green "OK" and compared two historical IP addresses, and
wgft agent lsprintedokunderTUNNEL. The Web UI now shows "Last reported (disconnected)" in a muted colour, hides the IP comparison, and always flags an old heartbeat.wgft agent lsputslast:in front ofTUNNELandRULESfor a disconnected agent, for examplelast:ok; a script that parses those columns should expect the prefix. The admin API's JSON is unchanged:connectedandlast_heartbeatwere already there. - Configuration mistakes no longer make systemd restart the server every two seconds. A malformed
WGFT_WG_ADDRESS,WGFT_AGENT_APIorWGFT_ADMIN(no port, a port that does not exist,unix://for the agent API, which is TCP only, orunix://with no path) is now rejected at startup with exit status 3, which the shipped unit does not restart. So is kernel mode started without root orCAP_NET_ADMIN, with a message naming both ways out: run with the capability as the shipped unit does, or setWGFT_MODE=userspace. A real bind failure for the admin or agent API, such as an address already being in use, still exits 1 and is retried, because that can fix itself. This closes the first known issue of v0.5.0. - The agent exits 3, not 1, when its join string is the problem: no credentials and no
WGFT_JOIN, a malformed join string, or one that was already used. The shippedagent.servicedoes not restart on 3. A network failure during registration still exits 1 and is retried. A usedWGFT_JOINleft in the compose file of an agent that is already registered is ignored, as before. On macOS, launchd has no counterpart toRestartPreventExitStatus; restart behaviour for this case has not been tested. - The server no longer fails to start when its database is briefly locked by another process: a restart overlapping the old process, a binary swap,
server teardown, or a CLI command at that moment. It waits up to five seconds. wgft server checkcould not open the database through a relative path. For example withWGFT_DATA_DIR=./data. Fixed;server runwith a relative data directory was not affected in v0.5.0.
Known issues
- A rule refused by the agent's
WGFT_AGENT_ALLOW_TARGETSshows its reason inwgft agent ls(RULES column) and in the Web UI, not inwgft rule ls, whoseREFUSEDcolumn counts only flow-budget refusals. - In kernel mode, host-side conntrack sizing is the operator's responsibility:
wgft server checktells you whennf_conntrack_maxis below 65536 and gives the command to raise it. - A database path on a Windows network share (
\\server\share\...) is refused with a clear error. The server runs on Linux only, so this affects development tools, not deployments.
Upgrading
- Back up the server's data directory before upgrading. Downgrading afterwards is not supported.
- Replace the binary (or container image) and restart the service. A rolling upgrade from v0.5.0 or v0.4.0 is supported, in either order.
- Run
wgft server check, and confirm inwgft rule lsandwgft agent lsthat the rules areactiveand the agents showok. If the unit had been restarting in a loop because of a configuration mistake, it now stops with status 3 instead:journalctl -u wgftsays which setting to fix.
Changelog
- 80f9435: chore(release): v0.5.1 (#89) (@rahanahu)
- 5c3264a: docs: make the Sandbox the lab's isolation unit in the test policy (#83) (@rahanahu)
- 118fe89: docs: state the no-downgrade caveat and the conntrack prerequisite (#80) (@rahanahu)
- 69d8d32: fix(server): exit 3 on config and privilege errors, wait out a held DB lock (#86) (@rahanahu)
- c93cb5a: fix(server): keep the cause of backend failures and mark stale agent state (#85) (@rahanahu)
- 81c3efe: refactor(server): drop test-only bookkeeping left by the migration (#84) (@rahanahu)
- 52cc6a5: test(deploy): add the distributed-artefacts VM test (B9) (#77) (@rahanahu)
- 68a7b2a: test(lab): check the in-place upgrade from the previous release (D4) (#79) (@rahanahu)
- 3cb04a0: test(lab): let every lab scenario run in its own sandbox (#81) (@rahanahu)
- bad7f5e: test(lab): run the lab suite as sandboxes inside one Lab Host VM (#82) (@rahanahu)