Repository navigation
Security model
This document states the threat model xdp-bfd is built against, what the fast path does with a hostile packet, and what a packet flood costs. The numbers come from one testbed, so read the per-frame costs rather than the packets per second: a pps figure is a property of that machine's NIC and CPU, while nanoseconds per frame times your own line rate is the answer for yours. Where a defence has a trade, the trade is named rather than hidden.
Session counts below are given as a fraction of 64 because that is the size of the fabric they were measured on; what matters is whether any session was lost, not the denominator.
Adversary. Controls the directly attached neighbour, or any host on the same L2 segment, or any host that can route packets to the box. Can sniff every frame on the link, forge any frame with any source address, TTL and content, and send at line rate. Does not have the authentication key, and cannot execute code on the host.
Assets.
- Truthful session state. An Up session must mean the peer is reachable and a Down one that it is not; the adversary must not be able to make either claim false for an authenticated session.
- Availability of the engine. The host's other work, and BFD sessions to other peers, must survive what one neighbour can send.
- The key.
Out of scope. An adversary holding the key (BFD's own model gives up there), a compromised control plane (bfdd is trusted), and the host being taken down by attacks unrelated to BFD.
RFC 5880 without authentication is not resistant to an on-link forger: it can hold a dead session Up by forging the peer's packets, or take a live one Down by forging AdminDown, because every field is visible on the wire. That is the protocol, not this implementation. xdp-bfd's job is to make authenticated sessions actually immune and to make unauthenticated ones no weaker than the RFC allows.
Note that driving authenticated offloaded sessions needs a data plane protocol that carries keys; the capability is present here (RFC 5880 s6.7 keyed SHA1, computed as an HMAC to match bfdd), but stock FRR does not yet push keys over bffdp. See Limitations and FRR issue 23274.
Every rejection is counted (see the stats map) and, on a BFD port, the packet is dropped in the XDP program rather than passed to the socket. The BFD ports have no consumer on the host but the engine's own socket, so a packet the fast path will not honour has nowhere useful to go, and passing it only costs a syscall and, at a flood, evicts real datagrams from the shared socket queue.
| Reaches | Disposition | Counter |
|---|---|---|
| non-IP / non-UDP / non-BFD port | passed to the stack, untouched | seen |
| UDP options / v4 header options to a BFD port | dropped | ip-options |
| first IPv4 fragment to a BFD port | dropped (a control packet never fragments) | rejected |
| UDP behind a v6 extension header to a BFD port | dropped | v6-exthdr |
| TTL/hop-limit not 255 (GTSM), no multihop session | dropped | rejected |
| malformed header, or an envelope that lies about length, to a BFD port | dropped | malformed |
| A-bit / M-bit the session cannot honour | dropped |
auth-mismatch, unsupported-flags
|
| well-formed, no configured session (engine mode) | dropped | unknown-session |
| authenticated, replay window then digest | dropped on failure | auth-bad |
| authenticated, too many digest failures this interval | dropped before the digest | auth-ratelimited |
The promiscuous observer (xdp-bfd-observe) passes unconfigured packets to
userspace for debugging; it must never run on a production interface.
The headline metric is nanoseconds per frame on the drop path, not
packets per second: a pps figure is a property of one testbed's NIC, while
ns-per-frame times a NIC's line rate is the answer for any NIC. Measured on
the testbed (bpftool prog show run-time delta, kernel.bpf_stats_enabled),
64-session mesh live, single 5-tuple so RSS pins one CPU:
| flood | before hardening | after |
|---|---|---|
| malformed to 3784, ~590k pps | passed to socket, RcvbufErrors +67k, mesh 61/64 |
dropped in XDP, +0, 64/64, 42 ns/frame |
| valid, unknown pair, ~580k pps | passed to socket, +9k, mesh 63/64 | dropped in XDP, +0, 64/64, 80 ns/frame |
| valid at TTL 64 (GTSM), ~667k pps | already dropped in the driver, +0, 64/64 | unchanged (the reference) |
Before those two drops a forger could flap sessions it was not addressing, by
filling the shared socket queue so datagrams for sessions still coming up
were evicted. After, both flood arms match the GTSM reference: dropped in
the driver, nothing evicted, mesh holds 64/64, recovery within a second.
The drops are attributed, not inferred: under the unknown-pair flood the
unknown-session counter took every one of 3.66M frames while
RcvbufErrors stayed flat.
Floods that name a real session. Three more arms forge frames for an address pair the fast path actually serves, so they reach the demux, the authentication path and the echo reflector that an unknown-session drop would otherwise hide. Measured live against individual sessions of the 64-session mesh:
| flood (aimed at a real session) | result |
|---|---|
well-formed, wrong your_disc, ~513k frames |
every frame dropped by the demux rule, rejected 1:1, 195 ns/frame; the named session did not flap and the mesh held 64/64. A forger who knows the address pair but not the discriminator still cannot disturb the session (RFC 5880 s6.8.6). |
| A bit set, bad digest, at an authenticated session, ~554k frames | the per-session auth bucket let the expensive verify run 248 times and rate-limited the other 553,497 before the digest; the targeted session went down by design (its real packets share the spent bucket) while the other 63 were untouched, recovering to 64/64 in ~1s. A bad-auth flood costs a bounded amount of CPU and cannot spread past the single session it targets. |
| self-addressed echo from a known echo peer, ~253k frames | reflected one-for-one (XDP_TX), 204 ns/frame, mesh 64/64; UDP/3785 from any source that is not a peer of an echo-active session is counted declined and never reflected, which closes the reflection off as an amplification vector. |
The auth arm is the sharpest of the three: without the bucket every one of those 554k frames would have cost a sequence check and, in window, a full HMAC-SHA1 in softirq; with it the costly path ran 0.04% as often, and the blast radius is one session rather than the host.
Unrelated line-rate traffic. Spread-port traffic (not aimed at a BFD port) is left alone before it is even counted; ~1M pps of it dips the mesh to 56/64 by softirq saturation and it self-recovers within 10s, with the dead-man gate never tripping. On the virtualised testbed the ceiling is the guest's RX path (~1M pps), not the engine; the per-frame ns figures are what carry to a bare-metal host with more RX queues.
At 1024 sessions. The same arms, from the chaos VM against the mesh grown to 1024 sessions (60 at 50ms, 964 at 300ms), with the witness counting on the far side. Every arm ran 10s at the injector's full rate:
| flood | ns/frame | result |
|---|---|---|
| unrelated UDP, spread ports, ~530k pps | 86 | 1024 held on both sides |
| valid at TTL 64 (GTSM), ~670k pps | 88 | 1024 held |
| malformed to 3784, ~740k pps | 49 | 1024 held |
| valid, unknown pair, ~760k pps | 86 | 1024 held |
real pair, wrong your_disc, ~600k pps |
134 | 1024 held |
| bad-auth at an authenticated session, ~710k pps | 134 | 1024 held |
| self-addressed echo from a known peer, ~620k pps | 103 | the spoofed peer's session flapped, 1023 held |
| unknown pair naming a real discriminator, ~580k pps | 137 | the named session flapped, 1023 held |
| real pair and discriminator, the peer's timers random in every frame, ~610k pps | 242 | the named session flapped on the peer, 1023 held |
The last two arms found two ways to hurt sessions a forger does not name, both fixed by a budget in the program:
- A packet from an unknown pair that names one of our discriminators goes
to userspace, which demultiplexes on Your Discriminator (RFC 5880 s6.3).
Unbudgeted, 428k of them a second ran the engine at 98% and 221 sessions
flapped. Each CPU now passes at most 256 per 100ms (
moved-ratelimited). - The echo reflector returned a spoofed peer's frames one for one, the
transmit ring overflowed, and the fast path's own replies were dropped
with them: 20 to 40 sessions flapped on both sides. RFC 5880 s6.8.9 bars a
peer from sending echo faster than our Required Min Echo RX, so each peer
now gets four times that plus 16 per 100ms (
echo-ratelimited).
A third came from the last arm. The program answered every packet of a
session, so a forger with its pair and discriminator got one reply per
frame; 500k replies a second overflowed the transmit ring and up to 56
sessions flapped on both sides. With the fast path switched off by the
churning timers, the same frames went to the socket instead. RFC 5880
s6.8.7 bars the peer from sending faster than our Required Min RX, so the
program now takes a session's packets at most twice that fast, answered or
passed up (too-fast), with Poll and Final under a budget of their own.
In all three the named session still flaps: its real packets share the budget the forger spent. That is the same trade as the auth bucket, one session rather than the host.
An authenticated session checks the replay window before the digest, so a
forger must supply an in-window sequence (visible on the wire); each such
packet then costs a full HMAC-SHA1 in softirq, per packet, per CPU. The bucket caps
this: once a session sees more than eight digest failures inside one detect
interval it drops further A-bit packets before the digest until the interval
turns over, counted auth-ratelimited.
The trade, stated plainly. Under a sustained in-window bad-digest flood the bucket empties and the session's own real packets are then dropped too, so that one session can go Down. That is the honest outcome of an on-link attack on a single session, and it bounds the CPU either way. The bucket is per-session: the other sessions are untouched. A key rollover produces at most a handful of failures, well under the ceiling.
BFD's model concedes a neighbour that holds the key: it can take its own session Up or Down at will, and nothing below the key can stop that. The question this section answers is the one left over, whether such a neighbour can reach past its own session, to other peers' sessions, to the engine's memory, or to the engine's life. Reviewed against the code 2026-10-02.
Memory: it cannot grow it. Nothing a received packet does allocates. The
session table and every map are sized once at start. The only map write on
the fast path is the per-session state insert, and it is reachable only for
an address pair that is already configured; an unconfigured pair is dropped
(unknown-session) before it. Sessions are created only by the control
plane (bfdd over bffdp, or static config), never by a packet. So a neighbour,
key or no key, cannot make the engine hold state for any pair it was not
configured for.
The engine's life: the fast path is a verified program. A key-holder's valid packets run the same BPF-verified code as any other; there is no separate path to reach, and the per-packet work is bounded.
What it can do: force an HMAC per packet on its RX CPU. For an
authenticated session the fast path checks a failure budget, then verifies
the digest, then applies the arrival-rate gate (RFC 5880 s6.8.7) that caps
replies and socket delivery. The failure budget counts only bad digests,
to bound a forger without the key (eight HMACs per detect interval, then
dropped before the digest). A neighbour that holds the key sends good
digests: it never trips that budget, and because the arrival-rate gate sits
after the verify, it pays a full HMAC-SHA1 per packet before the gate
declines to answer it. Measured on the testbed (BPF_PROG_TEST_RUN,
program-time only), a valid verify costs about 2560 ns/frame against
74 ns on the drop path. One core saturates near 390k verifies a second,
inside a 1G line rate, so a key-holder flooding valid authenticated packets
at line rate can hold one RX CPU on HMAC. With four RX queues for a thousand
sessions, that degrades the sessions sharing that queue, not only the
neighbour's own.
The blast radius is still bounded: it is one RX CPU's worth of HMAC, the reply and socket floods are stopped by the gate as for any other flood, no other RX queue is touched, and it lasts only until a compromised neighbour is noticed. It cannot take the host or an unrelated queue's sessions down. But it is wider than one session, so it is a gap in asset 2 under a key-holder, and it is listed here rather than hidden.
Closed by #32 (1829351). The
verify is now budgeted before it runs: per session, one verify per half
the Required Min RX (twice the rate the peer may legally send), with a burst
of two so a Final that closely follows a regular packet still gets through.
The budget reads only arrival time, never the Poll or Final bits, which are
not yet authenticated. A failed verify gives its token back, so bad digests
stay the failure budget's and a forger without the key cannot spend the real
peer's verifies. Packets over the budget are dropped before the digest and
counted verify-limited. The same flood of good digests now averages
267 ns/frame, against 2560 before, so a key-holder's reach is back to its
own session.
If a flood saturates every CPU's softirq, userspace starves, the engine's heartbeat goes stale, and after three seconds the gate stops the fast path from answering, so peers take their sessions Down on their own timers. Without the gate the fast path would keep answering from softirq while the engine was frozen, holding sessions Up on a lie. A starved engine cannot tell bfdd anything, so Down is the honest answer, and it is what any userspace BFD daemon would produce under the same flood, only later. The gate is the trade between a late-but-true Down and a timely lie, resolved for the truth.
To keep the engine schedulable under softirq load the systemd unit runs it
SCHED_FIFO 60, above threaded interrupt handlers; pinning it away from the RX queues' CPUs with CPUAffinity
is a documented option for operators who have the cores.
The engine allocates nothing per packet or per session after start. The
session table, the maps, the 256 KB event ring and the 64 KB data-plane
queue are all fixed. Under host memory pressure the risk is the OOM killer
taking the engine, at which point the XDP link detaches and peers detect on
their own timers, which is correct; the unit sets OOMScoreAdjust low to
make it a late target.
Running it
Reference
Development