Skip to content

Releases: rahanahu/wgft

v1.2.1

Choose a tag to compare

@github-actions github-actions released this 29 Sep 15:57

Notes

v1.2.1 is a patch release of the v1.2 series with two fixes to the agent's userspace mode, the default mode.

  • Security fix: a client could end a userspace-mode agent process by connecting to a published TCP port and resetting the connection at once. When the reset arrived while the connection waited to be accepted, the agent's relay read a missing remote address and panicked. The agent now refuses such a connection and keeps serving the rule; it logs these refusals at most once a minute per port. The refused connection does not reach the Admission Policy and takes no flow slot. Windows and macOS agents always run in userspace mode and are affected. Agents in kernel mode do not relay TCP themselves and are not affected. Advisory: GHSA-w3wj-h525-39q7.
  • Fix: the agent's keepalive ping through the tunnel no longer leaves an ICMP endpoint behind on every ping. Before, each ping kept one endpoint and one ICMP identifier until the agent rebuilt the tunnel or restarted. Once the identifiers ran out, every ping failed without a log line, and if the VPS side lost the WireGuard session, the tunnel waited for the key lifetime to come back instead of the next ping.
  • Upgrade the agent: both fixes are in the agent. Upgrade the agent, the binary or the wgft-agent image, to get them. The server's behaviour does not change in practice.
  • Upgrade and revert: a binary replacement with no schema change. Reverting to v1.2.0 works with the same database and agent credentials.

Changelog

v1.2.0

Choose a tag to compare

@github-actions github-actions released this 26 Sep 01:13
d5f84d7

Notes

  • New: kernel mode for the Linux home agent. With WGFT_MODE=kernel, the agent creates a kernel WireGuard interface, wgft0, and table inet wgft_agent, and the kernel forwards to the LAN targets with DNAT. The agent process relays nothing, so forwarding continues while the agent is stopped or restarting. The default stays userspace mode.

    • The provided agent.service stays unprivileged. The new drop-in deploy/agent.kernel.conf adds CAP_NET_ADMIN and nothing else.
    • Kernel mode forwards only to IPv4 targets and refuses loopback targets. The agent sets net.ipv4.ip_forward to 1 when it is 0, so the home host routes packets for any traffic.
    • wgft agent teardown is new. It removes wgft0, the table, the conntrack entries of the flows the agent forwarded and the kernel-mode records in agent.json, so the agent can go back to userspace mode. It refuses while the agent runs.
    • wgft agent doctor checks the interface, the table and IP forwarding of a kernel-mode agent.
    • See README, "Kernel mode on the home agent", and docs/setup.md, "Run the agent in kernel mode".
  • New: agent disable and enable. wgft agent disable <name> stops forwarding for all of an agent's rules and keeps the registration, the keys and the rules. wgft agent enable <name> undoes it. The Web UI offers both. wgft status, wgft server doctor and wgft agent doctor show a disabled agent. See README, "Disabling an agent", and docs/cli.md.

  • New in the Web UI:

    • The dashboard's rule list shows a diagnosis mark next to each rule's state. It names the node where the rule's traffic stops and opens that rule's diagnostics page. The error counts in the header and the group headings count these marks.
    • Each agent has a detail page. Deleting an agent moved there, into a danger zone that asks for the agent's name. It can delete the agent's rules together with the agent.
    • The dashboard shows a band for rules left by a deleted agent, with a button that deletes them.
  • New: last UDP reply seen. For a UDP rule, wgft server doctor adds a line under target saying when the server last saw a reply from the target, and the Web UI diagnostics page shows the same. It never changes a status: a UDP target still reads NOT TESTED.

  • New: rule IDs as rule ls shows them. Rule commands also accept the shortened ID from rule ls with its trailing … or .... Commands that doctor and status suggest now use the full ID.

  • Changes to machine-readable values. Read these if a script reads --json or exit codes.

    • Healthy UDP rule, server doctor --json: the rule.target check's checks[].status changes from ok to not_tested, with the new reason udp_listener_only and next confirm the service from a real client. rules[].status, the top-level status and the exit code 0 do not change. TCP rules do not change. The agent's report for a UDP rule shows only that its listener is open, which does not meet the documented meaning of OK, "observed to succeed". Under design.md 7a.11 this is a bug fix: the value contradicted a meaning that design.md defined before this change. (#203)
    • An agent's own listener bind failure, server doctor --json: rule.target now reads listener_bind_failed instead of target_error, and its next step points at the agent's own earlier listener or connection instead of the target service. This is seen with userspace-mode agents older than v1.1.2, and with current userspace-mode agents when a rule is re-added or re-enabled at the same port shortly after a session in which the target closed the connection first while the client's side was still open. design.md had not defined either code's meaning before this change, so it is an explicit classification change in v1.2.0, not a 7a.11 correction. (#245)
    • A kernel-mode agent's rule whose target name stopped resolving: the agent keeps forwarding to the address from the last successful resolution and says so in the rule's reason. server doctor --json now reads rule.target_resolve as unknown instead of failed; the reason stays target_resolve_failed. rule.target is judged from the rest of the reason: with nothing after it, the rule has no stopped_at and the exit code is 0 unless something else fails; a probe error after it still stops the rule at target. The human output shows DEGRADED on that line. design.md now defines DEGRADED by meaning: a check that is not judged a stop but is confirmed to have degraded from the normal state. DEGRADED never appears in --json. In agent doctor --json, dataplane.table reads unknown with resolve_failed with exit 0 when such rules are its only rule errors. design.md rewrote its case list in the same change, so this is an explicit classification change in v1.2.0, not a 7a.11 correction. Only kernel-mode agents, new in this release, send this report, so no existing deployment sees a value change. server doctor from v1.1.x still says traffic stops at "target resolve" for such a rule. (#260)
    • Situations that now get their own reason code: design.md 7a.11 treats these as additions to open sets; existing values keep their meaning.
      • The agent host's ip_forward is 0: server doctor reads rule.target as agent_ip_forward_off. Only kernel-mode agents, new in this release, send this report. (#274)
      • The VPS's ip_forward is 0 while a kernel-mode server runs and at least one published rule is forwarded by the kernel: server doctor now fails server.dataplane and the public port of each such rule with ip_forward_off, and exits 1. wgft status marks Server degraded in the same case and exits 1. Before, both reported no failure. Proxy-mode rules do not depend on the value, so a server with only proxy-mode rules gets a note, not a failure. The server does not set the value back to 1 while it runs. (#274)
      • agent doctor, the agent's pinned server certificate does not match: stream.connection reads FAILED server_cert_mismatch, where it read UNKNOWN reconnecting. The top-level status and the exit code do not change. (#274)
      • agent doctor, the data directory cannot be read: the reason is the new data_dir_unreadable. It used to be credentials_unreadable, which design.md already defined as agent.json existing and not being readable; that code is now given only in that case. The top-level status and the exit code do not change. (#274)
    • Added values: design.md 7a.11 treats check ids and reason codes as open sets: a consumer must read an unknown value as unknown, not as a failure.
      • server doctor --json: the check id agent.enabled; the reasons agent_disabled, target_loopback_unsupported, agent_ip_forward_off and ip_forward_off; the optional last_reply_at, reply_since and reply_not_observed on a UDP rule's rule.target.
      • agent doctor --json: the check ids dataplane.interface, dataplane.table and host.forwarding, which a userspace-mode agent also reports, as not_tested with the reason userspace_mode; the reasons agent_disabled, needs_cap_net_admin, the kernel-mode codes that design.md section 10.2c lists, server_cert_mismatch and data_dir_unreadable.
      • status --json: agents.disabled and rules.agent_disabled, always present.
      • Admin API: POST /api/v1/agents/{name}/disable and /enable; disabled and disabled_at in the agent list; udp_replies and, on a kernel-mode server, ip_forward in the rules response; ip_forward in GET /api/v1/server; agent_disabled in GET /api/v1/agents/{name}/state, where a disabled agent's rules also read enabled: false.
    • Agent-supplied text: the server now caps the text an agent sends in its heartbeat and replaces unprintable characters before storing it: 64 bytes for a state, 512 for a reason, 128 for a rule ID or an endpoint. This reaches agent ls --json and rule ls --json. design.md 7a.11 treats this as processing that does not change the values' meaning. (#270)
  • Upgrade the server: v1.2.0 moves the server database to schema version 9 on its first start. v1.1.x and earlier refuse to open it. On a staging machine, v1.1.3 exited with code 3 and left the database unchanged. Reverting is not promised: back up the data directory before upgrading. Moving back to v1.1.1 through v1.1.3 needs a backup taken before the upgrade to v1.2.0, and moving back to v1.1.0 or earlier needs one taken before the upgrade to v1.1.1 or later.

  • Upgrade the agent: an agent without WGFT_MODE stays in userspace mode. An agent that already receives WGFT_MODE=kernel, for example from a settings file it shares with the server, ignored it up to v1.1.x. v1.2.0 starts it in kernel mode, and under the provided unprivileged unit it stops at startup with exit code 3 and names CAP_NET_ADMIN. It stops before it uses the join string or records anything in agent.json. Give the agent no WGFT_MODE, or WGFT_MODE=userspace, and restart the unit.

  • Move a kernel-mode agent back to v1.1.x: v1.1.x has no kernel mode for the agent. With the v1.2.0 binary, stop the agent, run sudo wgft agent teardown, and remove WGFT_MODE=kernel and the drop-in. Then install v1.1.x and start the agent. This order was verified on a staging machine. Without the teardown, v1.1.3 started on a staging machine, but it dropped the kernel-mode records from agent.json and left wgft0 and table inet wgft_agent in the kernel. What they do to forwarding has not been verified. If v1.1.x already started without the teardown, wgft agent teardown of v1.2.0 still removes wgft0 and the table, but no longer finds the conntrack entries or the ip_forward record. See docs/setup.md, "Run the agent in kernel mode".

  • Behavior changes:

    • Interface names: a server whose WGFT_WG_INTERFACE contains a byte the kernel refuses in an interface name, or one that makes the name differ from the one requested, now exits with...
Read more

v1.1.3

Choose a tag to compare

@github-actions github-actions released this 24 Sep 08:50
582e5bd

Notes

  • Fix: on a kernel-mode server, an Admission Policy rule's source_deny or source_allow list with 1638 or more separate entries, counting adjacent or overlapping ones as one, could load into nftables wrongly, with no error. The server took the table as it was loaded, reporting the rule as active.
    • What was seen: lab testing found the set loaded empty or partial on kernel 6.1. A real VPS running v1.1.2 on kernel 6.12, with 1700 addresses in one list, loaded the set partial and also holding a wrong, very wide range: 62 elements ending in 100.64.0.123-255.255.255.255; server doctor said OK. In the lab, a 5000-entry list made the kernel refuse the whole update; that failure was reported, and the previous table stayed in place.
    • Effect for a deny list: listed sources could get through, or unrelated sources could be dropped.
    • Effect for an allow list: sources outside the list could get through, or listed sources could be dropped.
    • Affected versions: every release up to and including v1.1.2.
  • Now: the server sends set elements in chunks and reads each source list's set back. It compares what was loaded with what was sent, and treats a mismatch as a failed publication instead of accepting it.
    • Verified in the lab on kernel 6.1 with deny and allow lists of up to 20000 entries. Verified on a real kernel-mode VPS on kernel 6.12 with the v1.1.3 candidate: deny and allow lists of 1700 and 5000 entries loaded exactly, and real traffic was filtered correctly in both directions.
    • After a failed publication, reapplies triggered by nftables notifications are spaced out from 1 s up to 30 s; this spacing was verified by unit tests only. Admin changes and newly found drift are applied at once, and the normal 30-second retry is unchanged.
  • How to check whether you were affected (kernel mode, before upgrading): wgft rule ls shows each rule's DENY and ALLOW counts; a list with fewer than 1638 CIDRs was not affected. For a larger one, find its set in sudo wgft server nft (or sudo nft list table inet wgft): the row carrying the comment wgft:<rule id>:deny or :allow uses @deny_<n> or @allow_<n>. Then sudo nft list set inet wgft deny_<n> (or allow_<n>): an empty set, missing addresses, or a range your list does not contain, such as one ending in 255.255.255.255, means the rule was affected. After the upgrade the server loads the full list, so check first.
  • Upgrade the server: upgrade the server binary. Userspace mode, including the server image, does not use nftables for these lists and is not affected. The agent does not change.
  • Known effect: on the kernel 6.12 VPS running the v1.1.3 candidate, publishing several 5000-entry lists back to back briefly overflowed the server's nftables change watcher. It logs watching the data plane for changes failed: nftables notifications: netlink receive: recvmsg: no buffer space available; subscribing again, and checking every few minutes meanwhile, resubscribes, and drift repair kept working afterward. It is harmless.
  • Upgrade and revert: a binary replacement with no schema change. Reverting to v1.1.2 works with the same database, but brings the bug back.

Changelog

v1.1.2

Choose a tag to compare

@github-actions github-actions released this 24 Sep 04:50
c391a6e

Notes

  • Fix: when a TCP rule's target changes while a session is open, the rule now listens again at once. Before, the agent held the port and the rule stayed down for 60 s or more. On a kernel-mode VPS whose input firewall drops by default, disabling a rule with an open session and re-enabling it at once could keep it down for more than ten minutes (about 13.5 minutes was measured on a real VPS). It now comes back at once.
  • Upgrade the agent: the fix is in the agent. Upgrade the agent, the binary or the wgft-agent image, to get it. The server's behaviour does not change in practice.
  • Sessions cut on purpose now end with a TCP reset between the agent and the VPS. What the client sees depends on the server mode: with a userspace-mode server the client should still see a normal close; with a kernel-mode server a target change reaches the client as a reset, and a disable or delete reaches it as nothing, so the client notices nothing until it next sends. On a real kernel-mode VPS whose input firewall drops by default, these were observed with v1.1.2: a target change reached the client as a reset, and a rule disabled and re-enabled, or deleted and added again, or an agent restart, reached it as nothing until it next sent, then as a reset. The userspace-mode case follows from the code and has not been observed directly.
  • Known limit: a session that ended on its own with the target closing first, such as HTTP, is expected to hold the rule's port for about 60 s. If the rule is disabled and re-enabled or retargeted during that time, its listener opens on the first 30-second retry after the hold ends.
  • Upgrade and revert: a binary replacement with no schema change. Reverting to v1.1.1 works with the same database.

Changelog

v1.1.1

Choose a tag to compare

@github-actions github-actions released this 23 Sep 23:45
def9fd0

Changelog

v1.1.0

Choose a tag to compare

@github-actions github-actions released this 23 Sep 22:05
9468d92

Changelog

v1.0.0

Choose a tag to compare

@github-actions github-actions released this 22 Sep 21:27
f859f3f

Changelog

v0.7.0

Choose a tag to compare

@github-actions github-actions released this 22 Sep 07:44
c31913d

Changelog

v0.6.0

Choose a tag to compare

@github-actions github-actions released this 20 Sep 23:03
f63651e

This release changes which startup failures make systemd stop and which make it retry, gives rule ls the agent's side of a rule's status, and fixes several places where a failed read was shown as "nothing". No change to the wire protocol or the data format. The server and the agents can be upgraded in either order.

Changes that affect running deployments

  • Exit codes at startup follow one rule now: can a retry fix it? The default is exit 1, which the shipped units retry. Exit 3, which they do not retry, is reserved for a cause that is the configured value itself, or that cannot go away without an operator. Compared with v0.5.1:
    • Now exit 1 (retried), was exit 3: an interface of the configured name that is not wgft's; the WireGuard port held by another WireGuard interface or by another process; an address range that overlaps another interface; the conntrack timeout values still unreadable after the table was applied. The other owner may release the resource, and with exit 3 the server stayed down after the conflict was gone. The checks still stop before anything is written, and the message still says what to change; expect a log line every two seconds while the conflict lasts.
    • Now exit 3 (not retried), was exit 1: an empty WGFT_DATA_DIR; a WGFT_WG_INTERFACE the kernel rejects; WGFT_MTU outside 576 to 9216; WGFT_WG_PORT=0; a WGFT_WG_ENDPOINT or WGFT_AGENT_API_HOST that is not host:port; a database written by a newer wgft; and an agent's registration answered 400, 401 or 409 (a name the server rejects, a join string that is spent or rejected, a name already taken).
    • Now refused, was silently ignored: a WGFT_ADMIN_TAILSCALE that is not a boolean. ture used to mean "off".
  • Refusals name their cause. Every exit-3 message reads refusing to start [<category> <subject>]: <reason>, with the category one of config, prerequisite, conflict, mode-gate. A script that matches the old refusing to start: ... text needs the bracket.
  • Exit 3 stops the restart loop under the shipped systemd units only. launchd restarts an agent on exit 3 as on any other exit (checked on macOS 27), and restart: unless-stopped in the compose files ignores exit codes.
  • Pages and API routes that showed invented content now answer 500. When the server's database could not be read, the dashboard, the rule page, the add-rule form and the two import routes answered 200 with generation 0, zero drop counts, "no agents registered" or "rule not found". The import guard could even match its own invented generation and let a stale import through; it now stops.

New

  • wgft rule ls shows each rule's status on its agent. A rule refused by the agent's WGFT_AGENT_ALLOW_TARGETS, or whose listener or target has an error on the agent, used to show its reason only in wgft agent ls and the Web UI. rule ls has a new AGENT_STATE column, and GET /api/v1/rules and rule ls --json have agent_rule_states, one entry per rule with agent and connected always present and state, reason and at once the agent has reported on that rule. connected: false marks a disconnected agent's last report as history. This closes the known issue listed in the last two releases.
  • What v1.0 will keep compatible is written down, per surface, in the design document: the admin API's routes and fields, --json as the only machine interface of the CLI (the tables, their columns and the help text are for people and may change), the WGFT_* names, the wire protocol's negotiation, the upgrade path, unit and artefact names. Fields are only added; values such as state and apply_state are open sets that a consumer must read as "unknown" when it does not know them.

Fixes

  • agent ls --json no longer prints "last_handshake":"0001-01-01T00:00:00Z" inside tunnel for a tunnel that never shook hands: the key is absent, and when present it is RFC3339 like the timestamps beside it. The wire format between agent and server is unchanged.
  • rule_states[id].active_generation is absent, not 0, for a rule that was never published.
  • wgft server teardown warns when no interface name was recorded, instead of silently assuming wgft0; what it deletes is unchanged.
  • An unparseable agent address in the database fails the read instead of becoming the zero address in the WireGuard peer.
  • The mode-change gate, the connection check and the CLI's socket hint test error types instead of matching error texts.
  • Messages: the shutdown log names the real interface and mentions the table only in kernel mode; a failed apply in userspace mode no longer blames nftables; the agent calls its agent.json "credentials" and the server calls its database "server database" in every message.

Known issues

  • A rule's failure on the agent side is shown per rule, but a rule's server-side apply state and its flow-budget refusals are still separate fields (rule_states, resource_refusals); rule ls --json has all three.
  • In kernel mode, host-side conntrack sizing is the operator's responsibility: wgft server check says when nf_conntrack_max is below 65536.
  • Under investigation, not confirmed on a real machine: a Windows agent may stop receiving on its tunnel after it sent handshakes to a server whose WireGuard port was closed (a userspace-mode server that was stopped, or a VPS while it reboots), and stay that way until the agent is restarted. If you encounter this and forwarding does not return after the server is back, restarting the agent is the current workaround.

Upgrading

  1. Back up the server's data directory before upgrading. Downgrading afterwards is not supported.
  2. Replace the binary (or container image) and restart the service. A rolling upgrade from v0.5.1, v0.5.0 or v0.4.0 is supported, in either order. Upgrade compatibility is covered by the lab tests; the v0.5.1 path is additionally exercised through the shipped systemd units and machine reboots in the VM test.
  3. Run wgft server check, and confirm in wgft rule ls and wgft agent ls that the rules are active, that AGENT_STATE shows ok, and that the agents show ok. If your monitoring looks at the unit's state: a conflict with another WireGuard interface or port now shows as a restarting unit, not a failed one.

Changelog

v0.5.1

Choose a tag to compare

@github-actions github-actions released this 20 Sep 17:31
80f9435

A patch release: fixes only, no new settings, no change to the wire protocol or the data format. The server and the agents can be upgraded in either order.

Fixes

  • A momentary database error could park a healthy agent. When the server's database failed while it authenticated an agent's stream, the server answered 401, which the agent reads as "the permanent token was revoked": it stopped retrying and waited to be re-enrolled. The server now logs the cause and answers 500, and the agent reconnects with its normal backoff. An invalid token still gets 401.
  • A database error during registration no longer reads as key theft. The same kind of failure while the server checked an agent's WireGuard public key was reported to the agent as "public key belongs to another agent", with nothing in the server's log. It is now logged and reported as an internal error.
  • wgft rule ls no longer prints invented values. If reading the generation or the drop counters failed, the admin API returned 200 with generation 0 and all-zero DROPPED counts. It now returns 500, and rule ls fails with the cause, which names the read that failed.
  • A disconnected agent is no longer shown as healthy. After an agent's stream dropped, the Web UI kept showing its last tunnel state as a green "OK" and compared two historical IP addresses, and wgft agent ls printed ok under TUNNEL. The Web UI now shows "Last reported (disconnected)" in a muted colour, hides the IP comparison, and always flags an old heartbeat. wgft agent ls puts last: in front of TUNNEL and RULES for a disconnected agent, for example last:ok; a script that parses those columns should expect the prefix. The admin API's JSON is unchanged: connected and last_heartbeat were already there.
  • Configuration mistakes no longer make systemd restart the server every two seconds. A malformed WGFT_WG_ADDRESS, WGFT_AGENT_API or WGFT_ADMIN (no port, a port that does not exist, unix:// for the agent API, which is TCP only, or unix:// with no path) is now rejected at startup with exit status 3, which the shipped unit does not restart. So is kernel mode started without root or CAP_NET_ADMIN, with a message naming both ways out: run with the capability as the shipped unit does, or set WGFT_MODE=userspace. A real bind failure for the admin or agent API, such as an address already being in use, still exits 1 and is retried, because that can fix itself. This closes the first known issue of v0.5.0.
  • The agent exits 3, not 1, when its join string is the problem: no credentials and no WGFT_JOIN, a malformed join string, or one that was already used. The shipped agent.service does not restart on 3. A network failure during registration still exits 1 and is retried. A used WGFT_JOIN left in the compose file of an agent that is already registered is ignored, as before. On macOS, launchd has no counterpart to RestartPreventExitStatus; restart behaviour for this case has not been tested.
  • The server no longer fails to start when its database is briefly locked by another process: a restart overlapping the old process, a binary swap, server teardown, or a CLI command at that moment. It waits up to five seconds.
  • wgft server check could not open the database through a relative path. For example with WGFT_DATA_DIR=./data. Fixed; server run with a relative data directory was not affected in v0.5.0.

Known issues

  • A rule refused by the agent's WGFT_AGENT_ALLOW_TARGETS shows its reason in wgft agent ls (RULES column) and in the Web UI, not in wgft rule ls, whose REFUSED column counts only flow-budget refusals.
  • In kernel mode, host-side conntrack sizing is the operator's responsibility: wgft server check tells you when nf_conntrack_max is below 65536 and gives the command to raise it.
  • A database path on a Windows network share (\\server\share\...) is refused with a clear error. The server runs on Linux only, so this affects development tools, not deployments.

Upgrading

  1. Back up the server's data directory before upgrading. Downgrading afterwards is not supported.
  2. Replace the binary (or container image) and restart the service. A rolling upgrade from v0.5.0 or v0.4.0 is supported, in either order.
  3. Run wgft server check, and confirm in wgft rule ls and wgft agent ls that the rules are active and the agents show ok. If the unit had been restarting in a loop because of a configuration mistake, it now stops with status 3 instead: journalctl -u wgft says which setting to fix.

Changelog