Skip to content

Deployment

Abdul Wasey edited this page Oct 1, 2026 · 11 revisions

Deployment

This walks through putting the engine under an existing FRR, checking that it is really carrying the traffic, and the settings worth changing afterwards.

1. Reserve the data-plane port

Do this first, before anything binds it.

# /etc/sysctl.d/50-xdp-bfd.conf   (the package ships this)
net.ipv4.ip_local_reserved_ports = 50700

bfdd connects to the engine on 50700. If the engine is not listening when bfdd retries, and the kernel has handed 50700 out as an ephemeral source port, bfdd's connect can self-connect: one socket ends up connected to itself, holding the port. SO_REUSEADDR does not recover from it and the engine can never bind again until something is killed. Reserving the port keeps the ephemeral allocator away from it.

Do not also reserve 65472-65535. The engine binds its own per-session source ports in that range, and ip_local_reserved_ports blocks explicit bind() as well as ephemeral allocation, so reserving it breaks the thing you are trying to protect.

2. Point bfdd at the engine

# /etc/frr/daemons
bfdd=yes
bfdd_options="  --daemon -A 127.0.0.1 --dplaneaddr ipv4c:127.0.0.1:50700"

The package drops this as examples/frr-daemons.snippet.

3. Start the engine before FRR

Order matters. bfdd tries its data-plane connection at startup, and if nothing answers it falls back to serving sessions in software; it will reconnect, but you will spend the gap wondering why nothing is offloaded.

# /etc/xdp-bfd/engine.conf
XDP_BFD_ARGS="--dplane 50700 --kernel-tx eth0"
sudo systemctl enable --now xdp-bfd
sudo systemctl restart frr

--kernel-tx names the interface to attach XDP to. Without it the engine still serves bfdd, but entirely in software, which forfeits the reason to run it. Give it the interface your BFD peers are reached over; others are attached as bfdd places sessions on them.

--stats-dump is not required, but without it SIGUSR1 has nowhere to write and Monitoring has nothing to read.

4. Check that the fast path is actually carrying it

The engine says what it attached to, and in which mode:

kernel-tx: dead-man gate at 3000000us
kernel-tx: XDP attached to eth0 (native mode, link)
dplane: listening on 127.0.0.1:50700 (bfdd: ipv4c:127.0.0.1:50700)
dplane: bfdd connected

native mode means the driver supports XDP directly. generic mode means it does not and the kernel is emulating, which still works and is still faster than userspace, but is not what the numbers elsewhere were measured on.

bfdd connected is the line that matters. Without it the engine is running and idle while bfdd serves everything itself.

Then confirm from the engine's own view, not just from show bfd peer, which looks identical either way:

sudo pkill -USR1 -x xdp-bfd
python3 -c 'import json;d=json.load(open("/tmp/bfd_tx_stats.json"));print(d["kernel_tx"], d["sessions_up"], "/", d["sessions_configured"])'
True 64 / 64

kernel_tx: true and a session count that matches bfdd's means the offload is live. If sessions_configured is short of what bfdd has, see Troubleshooting.

Tuning

Defaults are chosen to be safe rather than fast; most deployments change nothing.

--dp-hold <seconds> keeps sessions on the wire across a bfdd restart. The engine adopts them again by address pair, keeping their discriminators, and tears them down at the deadline if bfdd never comes back. Default 0, which drops and recreates. Set it to a few seconds if you restart FRR on a live box and do not want the peers to notice.

--pin <dir> keeps the program and its session state across an engine restart. Give it a directory on bpffs, such as /sys/fs/bpf/xdp-bfd, and --dp-hold as well. To restart or upgrade, start the new engine with the same arguments while the old one runs: it loads and verifies the object, has the old one hand over, and adopts its sessions until bfdd re-adds them. On a 1024-session testbed the old engine was gone in about 10 ms and no peer recorded a down event. A build whose maps differ discards the pinned state and starts fresh. SIGUSR2 makes the engine leave the same way and exit 75, for a supervisor that restarts on that status, though sessions with sub-second detection that depend on this host's own packets may then flap.

Under the packaged unit, put --dp-hold 30 --pin /run/xdp-bfd-pin in /etc/xdp-bfd/engine.conf; xdp-bfd-pin.service mounts that bpffs for the unit's user. Then systemctl reload xdp-bfd is the restart above: the new engine takes over and the peers see nothing. systemctl restart still stops first, sending AdminDown.

--deadman-us <usec> bounds how long the fast path keeps answering for an engine that has stopped making progress. If userspace wedges, the kernel would otherwise go on replying forever and the peer would never learn anything is wrong. Default 3 s; 0 disables it. Keep it above the kernel's real-time throttling period (sched_rt_period_us, 1 s by default): an engine starved by real-time load runs about once a period, and a shorter bound takes its sessions down while it is still alive.

--demand-poll-us <usec> bounds how long a demand-mode session may go without verifying the path. Default 1 s; 0 restores stock bfdd behaviour, which is to never verify. Never faster than the session's own detection budget.

Scheduling. The fast path answers the peer from softirq whatever userspace is doing, but detection, the notification to bfdd, the dead-man heartbeat, and every packet the fast path cannot clock off the peer's run in the engine's loop. Those are sessions with asymmetric timers, demand at one end, and polls.

The packaged unit runs that loop on a SCHED_DEADLINE reservation, 1 ms of CPU in every 10 ms. Deadline tasks run ahead of every SCHED_FIFO priority, so no real-time load starves the loop however high it runs. The engine takes the reservation itself (--sched-deadline 1000/10000, which systemd cannot set) and drops CAP_SYS_NICE once it has it. Where the kernel refuses one, because the unit is pinned to some of the CPUs or admission control is full, the engine logs why and stays at the unit's SCHED_FIFO 60, above threaded interrupt handlers (50).

A fixed priority only moves the problem. Below the load the engine never runs at all: RT throttling hands the spare slice of each period to normal tasks, not to lower real-time ones. At normal priority it runs about once a period. On bare metal, 1024 sessions under a SCHED_FIFO hog on every core:

engine hog flaps
SCHED_FIFO 10 (the unit until #26) 50 every session, once
normal priority 50 the asymmetric and one-sided demand sessions
SCHED_FIFO 60 (the unit until #28) 50 none
SCHED_FIFO 60 99 959 down events
SCHED_DEADLINE 1 ms in 10 ms 99 none

It uses about 2% of a core at 1024 sessions: 0.2 ms in 10 ms still held, 0.05 ms did not. XDP_BFD_SCHED= in /etc/xdp-bfd/engine.conf leaves the unit's SCHED_FIFO 60; a --sched-deadline in XDP_BFD_ARGS replaces the reservation. Started outside the unit, give it the option as root or with CAP_SYS_NICE.

Under a flood. The program answers BFD in the driver, but everything else the interface receives still goes up the stack, by default on the same CPU and in the same softirq as the ring BFD arrives on. A line-rate flood of other traffic plus load on the CPUs can back that ring up, and with it BFD. Measured on bare metal (Intel I210, 1022 sessions) over an hour of a 1.2M frames/s flood rotating through five kinds of traffic while stress-ng took turns at the CPUs, real-time priority and memory: 165,028 down events as installed, 4,148 with all three of these.

  • --spread-pass hands that other traffic to the stack on every CPU through a cpumap, by flow, so the RX CPU only runs the program. It changes the receive path for the whole interface, so it is off by default.
  • On igb, turn flow control off (ethtool -A <port> autoneg off rx off tx off). That is what makes the driver drop per queue: with it on, one queue that falls behind fills the NIC's shared buffer and every queue drops, BFD included.
  • Hash UDP on ports as well (ethtool -N <port> rx-flow-hash udp4 sdfn, and udp6). igb hashes UDP on addresses only, so a flood from one address lands on one queue however its ports vary.

What was left came from memory thrashed to the point of swapping, where one core cannot keep up with a single line-rate flow, and it fell on the 10 ms x3 sessions. Do not leave hardware RX timestamping set to every packet on an igb port (HWTSTAMP_FILTER_ALL, as some capture tools set it): each packet then costs a register read, and one queue handles a sixth of the rate.

Reply latency. The fast path answers in about 1.5 us; what the peer sees is set by the host's receive path. On the bare-metal testbed (Intel I210) the wire reply took 118 us at p50 and 250 us at p99, nearly all of it the NIC's PCIe link waking from ASPM L1. With L1 off it was 33 and 122 us, and with ethtool -C <port> rx-usecs 0 as well 31 and 52 us, for half a watt. Detection times are in milliseconds, so this changes nothing a peer acts on; turn it off only where microseconds count. lspci -vvv on the NIC shows LnkCtl: ASPM L1 Enabled. Where the firmware keeps ASPM to itself (dmesg: "FADT indicates ASPM is unsupported, using BIOS configuration"), the kernel's pcie_aspm options do nothing: disable it in the BIOS, or clear it with setpci at boot on the NIC and its root port. Holding the cores out of deep C-states trims only the far tail and cost 25 W there.

Two things that will bite you

Echo needs the neighbour to forward. An echo packet is self-addressed: it only comes back if the far end forwards for that family. With forwarding off, nothing returns and echo looks broken when it is configured correctly.

An FRR peer's IPv6 echo is the other way round: bfdd sends it to our address and expects it back, so the engine returns it for any session it holds, whatever the forwarding settings.

A keychain key needs its algorithm set explicitly. bfdd only selects a key whose algorithm is cleartext or hmac-sha-1, and a key defaults to neither. Configure key-string alone and the session runs unauthenticated while show bfd peer still reports authentication as configured:

key chain example
 key 1
  key-string secret
  cryptographic-algorithm hmac-sha-1   <- without this, nothing is signed

Also read Limitations before enabling authentication on an offloaded session: stock bfdd cannot carry the keys to the data plane yet, and this engine fails closed rather than sending them in the clear.

Clone this wiki locally