Repository navigation
Deployment
This walks through putting the engine under an existing FRR, checking that it is really carrying the traffic, and the settings worth changing afterwards.
Do this first, before anything binds it.
# /etc/sysctl.d/50-xdp-bfd.conf (the package ships this)
net.ipv4.ip_local_reserved_ports = 50700
bfdd connects to the engine on 50700. If the engine is not listening
when bfdd retries, and the kernel has handed 50700 out as an ephemeral
source port, bfdd's connect can self-connect: one socket ends up
connected to itself, holding the port. SO_REUSEADDR does not recover
from it and the engine can never bind again until something is killed.
Reserving the port keeps the ephemeral allocator away from it.
Do not also reserve 65472-65535. The engine binds its own per-session
source ports in that range, and ip_local_reserved_ports blocks explicit
bind() as well as ephemeral allocation, so reserving it breaks the
thing you are trying to protect.
# /etc/frr/daemons
bfdd=yes
bfdd_options=" --daemon -A 127.0.0.1 --dplaneaddr ipv4c:127.0.0.1:50700"
The package drops this as examples/frr-daemons.snippet.
Order matters. bfdd tries its data-plane connection at startup, and if nothing answers it falls back to serving sessions in software; it will reconnect, but you will spend the gap wondering why nothing is offloaded.
# /etc/xdp-bfd/engine.conf
XDP_BFD_ARGS="--dplane 50700 --kernel-tx eth0"
sudo systemctl enable --now xdp-bfd
sudo systemctl restart frr
--kernel-tx names the interface to attach XDP to. Without it the
engine still serves bfdd, but entirely in software, which forfeits the
reason to run it. Give it the interface your BFD peers are reached over;
others are attached as bfdd places sessions on them.
--stats-dump is not required, but without it SIGUSR1 has nowhere to
write and Monitoring has nothing to read.
The engine says what it attached to, and in which mode:
kernel-tx: dead-man gate at 3000000us
kernel-tx: XDP attached to eth0 (native mode, link)
dplane: listening on 127.0.0.1:50700 (bfdd: ipv4c:127.0.0.1:50700)
dplane: bfdd connected
native mode means the driver supports XDP directly. generic mode
means it does not and the kernel is emulating, which still works and is
still faster than userspace, but is not what the numbers elsewhere were
measured on.
bfdd connected is the line that matters. Without it the engine is
running and idle while bfdd serves everything itself.
Then confirm from the engine's own view, not just from show bfd peer,
which looks identical either way:
sudo pkill -USR1 -x xdp-bfd
python3 -c 'import json;d=json.load(open("/tmp/bfd_tx_stats.json"));print(d["kernel_tx"], d["sessions_up"], "/", d["sessions_configured"])'
True 64 / 64
kernel_tx: true and a session count that matches bfdd's means the
offload is live. If sessions_configured is short of what bfdd has, see
Troubleshooting.
Defaults are chosen to be safe rather than fast; most deployments change nothing.
--dp-hold <seconds> keeps sessions on the wire across a bfdd
restart. The engine adopts them again by address pair, keeping their
discriminators, and tears them down at the deadline if bfdd never comes
back. Default 0, which drops and recreates. Set it to a few seconds if
you restart FRR on a live box and do not want the peers to notice.
--pin <dir> keeps the program and its session state across an
engine restart. Give it a directory on bpffs, such as
/sys/fs/bpf/xdp-bfd, and --dp-hold as well. To restart or upgrade,
start the new engine with the same arguments while the old one runs: it
loads and verifies the object, has the old one hand over, and adopts its
sessions until bfdd re-adds them. On a 1024-session testbed the old
engine was gone in about 10 ms and no peer recorded a down event. A build
whose maps differ discards the pinned state and starts fresh. SIGUSR2
makes the engine leave the same way and exit 75, for a supervisor that
restarts on that status, though sessions with sub-second detection that
depend on this host's own packets may then flap.
Under the packaged unit, put --dp-hold 30 --pin /run/xdp-bfd-pin in
/etc/xdp-bfd/engine.conf; xdp-bfd-pin.service mounts that bpffs for
the unit's user. Then systemctl reload xdp-bfd is the restart above:
the new engine takes over and the peers see nothing. systemctl restart
still stops first, sending AdminDown.
--deadman-us <usec> bounds how long the fast path keeps answering
for an engine that has stopped making progress. If userspace wedges, the
kernel would otherwise go on replying forever and the peer would never
learn anything is wrong. Default 3 s; 0 disables it. Keep it above the
kernel's real-time throttling period (sched_rt_period_us, 1 s by
default): an engine starved by real-time load runs about once a period,
and a shorter bound takes its sessions down while it is still alive.
--demand-poll-us <usec> bounds how long a demand-mode session may
go without verifying the path. Default 1 s; 0 restores stock bfdd
behaviour, which is to never verify. Never faster than the session's own
detection budget.
Scheduling. The fast path answers the peer from softirq whatever userspace is doing, but detection, the notification to bfdd, the dead-man heartbeat, and every packet the fast path cannot clock off the peer's run in the engine's loop. Those are sessions with asymmetric timers, demand at one end, and polls.
The packaged unit runs that loop on a SCHED_DEADLINE reservation, 1 ms of
CPU in every 10 ms. Deadline tasks run ahead of every SCHED_FIFO
priority, so no real-time load starves the loop however high it runs. The
engine takes the reservation itself (--sched-deadline 1000/10000, which
systemd cannot set) and drops CAP_SYS_NICE once it has it. Where the
kernel refuses one, because the unit is pinned to some of the CPUs or
admission control is full, the engine logs why and stays at the unit's
SCHED_FIFO 60, above threaded interrupt handlers (50).
A fixed priority only moves the problem. Below the load the engine never
runs at all: RT throttling hands the spare slice of each period to normal
tasks, not to lower real-time ones. At normal priority it runs about once
a period. On bare metal, 1024 sessions under a SCHED_FIFO hog on every
core:
| engine | hog | flaps |
|---|---|---|
SCHED_FIFO 10 (the unit until #26) |
50 | every session, once |
| normal priority | 50 | the asymmetric and one-sided demand sessions |
SCHED_FIFO 60 (the unit until #28) |
50 | none |
SCHED_FIFO 60 |
99 | 959 down events |
SCHED_DEADLINE 1 ms in 10 ms |
99 | none |
It uses about 2% of a core at 1024 sessions: 0.2 ms in 10 ms still held,
0.05 ms did not. XDP_BFD_SCHED= in /etc/xdp-bfd/engine.conf leaves the
unit's SCHED_FIFO 60; a --sched-deadline in XDP_BFD_ARGS replaces the
reservation. Started outside the unit, give it the option as root or with
CAP_SYS_NICE.
Under a flood. The program answers BFD in the driver, but everything
else the interface receives still goes up the stack, by default on the same
CPU and in the same softirq as the ring BFD arrives on. A line-rate flood of
other traffic plus load on the CPUs can back that ring up, and with it BFD.
Measured on bare metal (Intel I210, 1022 sessions) over an hour of a
1.2M frames/s flood rotating through five kinds of traffic while
stress-ng took turns at the CPUs, real-time priority and memory: 165,028
down events as installed, 4,148 with all three of these.
-
--spread-passhands that other traffic to the stack on every CPU through a cpumap, by flow, so the RX CPU only runs the program. It changes the receive path for the whole interface, so it is off by default. - On igb, turn flow control off (
ethtool -A <port> autoneg off rx off tx off). That is what makes the driver drop per queue: with it on, one queue that falls behind fills the NIC's shared buffer and every queue drops, BFD included. - Hash UDP on ports as well (
ethtool -N <port> rx-flow-hash udp4 sdfn, andudp6). igb hashes UDP on addresses only, so a flood from one address lands on one queue however its ports vary.
What was left came from memory thrashed to the point of swapping, where one
core cannot keep up with a single line-rate flow, and it fell on the 10 ms x3
sessions. Do not leave hardware RX timestamping set to every packet on an igb
port (HWTSTAMP_FILTER_ALL, as some capture tools set it): each packet then
costs a register read, and one queue handles a sixth of the rate.
Reply latency. The fast path answers in about 1.5 us; what the peer
sees is set by the host's receive path. On the bare-metal testbed (Intel
I210) the wire reply took 118 us at p50 and 250 us at p99, nearly all of
it the NIC's PCIe link waking from ASPM L1. With L1 off it was 33 and
122 us, and with ethtool -C <port> rx-usecs 0 as well 31 and 52 us, for
half a watt. Detection times are in milliseconds, so this changes nothing
a peer acts on; turn it off only where microseconds count. lspci -vvv
on the NIC shows LnkCtl: ASPM L1 Enabled. Where the firmware keeps
ASPM to itself (dmesg: "FADT indicates ASPM is unsupported, using BIOS
configuration"), the kernel's pcie_aspm options do nothing: disable it
in the BIOS, or clear it with setpci at boot on the NIC and its root
port. Holding the cores out of deep C-states trims only the far tail and
cost 25 W there.
Echo needs the neighbour to forward. An echo packet is self-addressed: it only comes back if the far end forwards for that family. With forwarding off, nothing returns and echo looks broken when it is configured correctly.
An FRR peer's IPv6 echo is the other way round: bfdd sends it to our address and expects it back, so the engine returns it for any session it holds, whatever the forwarding settings.
A keychain key needs its algorithm set explicitly. bfdd only selects
a key whose algorithm is cleartext or hmac-sha-1, and a key defaults
to neither. Configure key-string alone and the session runs
unauthenticated while show bfd peer still reports authentication as
configured:
key chain example
key 1
key-string secret
cryptographic-algorithm hmac-sha-1 <- without this, nothing is signed
Also read Limitations before enabling authentication on an offloaded session: stock bfdd cannot carry the keys to the data plane yet, and this engine fails closed rather than sending them in the clear.
Running it
Reference
Development