Skip to content

Security model

Abdul Wasey edited this page Oct 2, 2026 · 7 revisions

Security model

This document states the threat model xdp-bfd is built against, what the fast path does with a hostile packet, and what a packet flood costs. The numbers come from one testbed, so read the per-frame costs rather than the packets per second: a pps figure is a property of that machine's NIC and CPU, while nanoseconds per frame times your own line rate is the answer for yours. Where a defence has a trade, the trade is named rather than hidden.

Session counts below are given as a fraction of 64 because that is the size of the fabric they were measured on; what matters is whether any session was lost, not the denominator.

Threat model

Adversary. Controls the directly attached neighbour, or any host on the same L2 segment, or any host that can route packets to the box. Can sniff every frame on the link, forge any frame with any source address, TTL and content, and send at line rate. Does not have the authentication key, and cannot execute code on the host.

Assets.

  1. Truthful session state. An Up session must mean the peer is reachable and a Down one that it is not; the adversary must not be able to make either claim false for an authenticated session.
  2. Availability of the engine. The host's other work, and BFD sessions to other peers, must survive what one neighbour can send.
  3. The key.

Out of scope. An adversary holding the key (BFD's own model gives up there), a compromised control plane (bfdd is trusted), and the host being taken down by attacks unrelated to BFD.

Authentication is required under this model

RFC 5880 without authentication is not resistant to an on-link forger: it can hold a dead session Up by forging the peer's packets, or take a live one Down by forging AdminDown, because every field is visible on the wire. That is the protocol, not this implementation. xdp-bfd's job is to make authenticated sessions actually immune and to make unauthenticated ones no weaker than the RFC allows.

Note that driving authenticated offloaded sessions needs a data plane protocol that carries keys; the capability is present here (RFC 5880 s6.7 keyed SHA1, computed as an HMAC to match bfdd), but stock FRR does not yet push keys over bffdp. See Limitations and FRR issue 23274.

What the fast path does with a hostile packet

Every rejection is counted (see the stats map) and, on a BFD port, the packet is dropped in the XDP program rather than passed to the socket. The BFD ports have no consumer on the host but the engine's own socket, so a packet the fast path will not honour has nowhere useful to go, and passing it only costs a syscall and, at a flood, evicts real datagrams from the shared socket queue.

Reaches Disposition Counter
non-IP / non-UDP / non-BFD port passed to the stack, untouched seen
UDP options / v4 header options to a BFD port dropped ip-options
first IPv4 fragment to a BFD port dropped (a control packet never fragments) rejected
UDP behind a v6 extension header to a BFD port dropped v6-exthdr
TTL/hop-limit not 255 (GTSM), no multihop session dropped rejected
malformed header, or an envelope that lies about length, to a BFD port dropped malformed
A-bit / M-bit the session cannot honour dropped auth-mismatch, unsupported-flags
well-formed, no configured session (engine mode) dropped unknown-session
authenticated, replay window then digest dropped on failure auth-bad
authenticated, too many digest failures this interval dropped before the digest auth-ratelimited

The promiscuous observer (xdp-bfd-observe) passes unconfigured packets to userspace for debugging; it must never run on a production interface.

What a flood costs

The headline metric is nanoseconds per frame on the drop path, not packets per second: a pps figure is a property of one testbed's NIC, while ns-per-frame times a NIC's line rate is the answer for any NIC. Measured on the testbed (bpftool prog show run-time delta, kernel.bpf_stats_enabled), 64-session mesh live, single 5-tuple so RSS pins one CPU:

flood before hardening after
malformed to 3784, ~590k pps passed to socket, RcvbufErrors +67k, mesh 61/64 dropped in XDP, +0, 64/64, 42 ns/frame
valid, unknown pair, ~580k pps passed to socket, +9k, mesh 63/64 dropped in XDP, +0, 64/64, 80 ns/frame
valid at TTL 64 (GTSM), ~667k pps already dropped in the driver, +0, 64/64 unchanged (the reference)

Before those two drops a forger could flap sessions it was not addressing, by filling the shared socket queue so datagrams for sessions still coming up were evicted. After, both flood arms match the GTSM reference: dropped in the driver, nothing evicted, mesh holds 64/64, recovery within a second. The drops are attributed, not inferred: under the unknown-pair flood the unknown-session counter took every one of 3.66M frames while RcvbufErrors stayed flat.

Floods that name a real session. Three more arms forge frames for an address pair the fast path actually serves, so they reach the demux, the authentication path and the echo reflector that an unknown-session drop would otherwise hide. Measured live against individual sessions of the 64-session mesh:

flood (aimed at a real session) result
well-formed, wrong your_disc, ~513k frames every frame dropped by the demux rule, rejected 1:1, 195 ns/frame; the named session did not flap and the mesh held 64/64. A forger who knows the address pair but not the discriminator still cannot disturb the session (RFC 5880 s6.8.6).
A bit set, bad digest, at an authenticated session, ~554k frames the per-session auth bucket let the expensive verify run 248 times and rate-limited the other 553,497 before the digest; the targeted session went down by design (its real packets share the spent bucket) while the other 63 were untouched, recovering to 64/64 in ~1s. A bad-auth flood costs a bounded amount of CPU and cannot spread past the single session it targets.
self-addressed echo from a known echo peer, ~253k frames reflected one-for-one (XDP_TX), 204 ns/frame, mesh 64/64; UDP/3785 from any source that is not a peer of an echo-active session is counted declined and never reflected, which closes the reflection off as an amplification vector.

The auth arm is the sharpest of the three: without the bucket every one of those 554k frames would have cost a sequence check and, in window, a full HMAC-SHA1 in softirq; with it the costly path ran 0.04% as often, and the blast radius is one session rather than the host.

Unrelated line-rate traffic. Spread-port traffic (not aimed at a BFD port) is left alone before it is even counted; ~1M pps of it dips the mesh to 56/64 by softirq saturation and it self-recovers within 10s, with the dead-man gate never tripping. On the virtualised testbed the ceiling is the guest's RX path (~1M pps), not the engine; the per-frame ns figures are what carry to a bare-metal host with more RX queues.

At 1024 sessions. The same arms, from the chaos VM against the mesh grown to 1024 sessions (60 at 50ms, 964 at 300ms), with the witness counting on the far side. Every arm ran 10s at the injector's full rate:

flood ns/frame result
unrelated UDP, spread ports, ~530k pps 86 1024 held on both sides
valid at TTL 64 (GTSM), ~670k pps 88 1024 held
malformed to 3784, ~740k pps 49 1024 held
valid, unknown pair, ~760k pps 86 1024 held
real pair, wrong your_disc, ~600k pps 134 1024 held
bad-auth at an authenticated session, ~710k pps 134 1024 held
self-addressed echo from a known peer, ~620k pps 103 the spoofed peer's session flapped, 1023 held
unknown pair naming a real discriminator, ~580k pps 137 the named session flapped, 1023 held
real pair and discriminator, the peer's timers random in every frame, ~610k pps 242 the named session flapped on the peer, 1023 held

The last two arms found two ways to hurt sessions a forger does not name, both fixed by a budget in the program:

  • A packet from an unknown pair that names one of our discriminators goes to userspace, which demultiplexes on Your Discriminator (RFC 5880 s6.3). Unbudgeted, 428k of them a second ran the engine at 98% and 221 sessions flapped. Each CPU now passes at most 256 per 100ms (moved-ratelimited).
  • The echo reflector returned a spoofed peer's frames one for one, the transmit ring overflowed, and the fast path's own replies were dropped with them: 20 to 40 sessions flapped on both sides. RFC 5880 s6.8.9 bars a peer from sending echo faster than our Required Min Echo RX, so each peer now gets four times that plus 16 per 100ms (echo-ratelimited).

A third came from the last arm. The program answered every packet of a session, so a forger with its pair and discriminator got one reply per frame; 500k replies a second overflowed the transmit ring and up to 56 sessions flapped on both sides. With the fast path switched off by the churning timers, the same frames went to the socket instead. RFC 5880 s6.8.7 bars the peer from sending faster than our Required Min RX, so the program now takes a session's packets at most twice that fast, answered or passed up (too-fast), with Poll and Final under a budget of their own.

In all three the named session still flaps: its real packets share the budget the forger spent. That is the same trade as the auth bucket, one session rather than the host.

The forced-HMAC bound and its trade

An authenticated session checks the replay window before the digest, so a forger must supply an in-window sequence (visible on the wire); each such packet then costs a full HMAC-SHA1 in softirq, per packet, per CPU. The bucket caps this: once a session sees more than eight digest failures inside one detect interval it drops further A-bit packets before the digest until the interval turns over, counted auth-ratelimited.

The trade, stated plainly. Under a sustained in-window bad-digest flood the bucket empties and the session's own real packets are then dropped too, so that one session can go Down. That is the honest outcome of an on-link attack on a single session, and it bounds the CPU either way. The bucket is per-session: the other sessions are untouched. A key rollover produces at most a handful of failures, well under the ceiling.

A compromised neighbour that holds the key

BFD's model concedes a neighbour that holds the key: it can take its own session Up or Down at will, and nothing below the key can stop that. The question this section answers is the one left over, whether such a neighbour can reach past its own session, to other peers' sessions, to the engine's memory, or to the engine's life. Reviewed against the code 2026-10-02.

Memory: it cannot grow it. Nothing a received packet does allocates. The session table and every map are sized once at start. The only map write on the fast path is the per-session state insert, and it is reachable only for an address pair that is already configured; an unconfigured pair is dropped (unknown-session) before it. Sessions are created only by the control plane (bfdd over bffdp, or static config), never by a packet. So a neighbour, key or no key, cannot make the engine hold state for any pair it was not configured for.

The engine's life: the fast path is a verified program. A key-holder's valid packets run the same BPF-verified code as any other; there is no separate path to reach, and the per-packet work is bounded.

What it can do: force an HMAC per packet on its RX CPU. For an authenticated session the fast path checks a failure budget, then verifies the digest, then applies the arrival-rate gate (RFC 5880 s6.8.7) that caps replies and socket delivery. The failure budget counts only bad digests, to bound a forger without the key (eight HMACs per detect interval, then dropped before the digest). A neighbour that holds the key sends good digests: it never trips that budget, and because the arrival-rate gate sits after the verify, it pays a full HMAC-SHA1 per packet before the gate declines to answer it. Measured on the testbed (BPF_PROG_TEST_RUN, program-time only), a valid verify costs about 2560 ns/frame against 74 ns on the drop path. One core saturates near 390k verifies a second, inside a 1G line rate, so a key-holder flooding valid authenticated packets at line rate can hold one RX CPU on HMAC. With four RX queues for a thousand sessions, that degrades the sessions sharing that queue, not only the neighbour's own.

The blast radius is still bounded: it is one RX CPU's worth of HMAC, the reply and socket floods are stopped by the gate as for any other flood, no other RX queue is touched, and it lasts only until a compromised neighbour is noticed. It cannot take the host or an unrelated queue's sessions down. But it is wider than one session, so it is a gap in asset 2 under a key-holder, and it is listed here rather than hidden.

Closed by #32 (1829351). The verify is now budgeted before it runs: per session, one verify per half the Required Min RX (twice the rate the peer may legally send), with a burst of two so a Final that closely follows a regular packet still gets through. The budget reads only arrival time, never the Poll or Final bits, which are not yet authenticated. A failed verify gives its token back, so bad digests stay the failure budget's and a forger without the key cannot spend the real peer's verifies. Packets over the budget are dropped before the digest and counted verify-limited. The same flood of good digests now averages 267 ns/frame, against 2560 before, so a key-holder's reach is back to its own session.

The dead-man gate under a flood

If a flood saturates every CPU's softirq, userspace starves, the engine's heartbeat goes stale, and after three seconds the gate stops the fast path from answering, so peers take their sessions Down on their own timers. Without the gate the fast path would keep answering from softirq while the engine was frozen, holding sessions Up on a lie. A starved engine cannot tell bfdd anything, so Down is the honest answer, and it is what any userspace BFD daemon would produce under the same flood, only later. The gate is the trade between a late-but-true Down and a timely lie, resolved for the truth.

To keep the engine schedulable under softirq load the systemd unit runs it SCHED_FIFO 60, above threaded interrupt handlers; pinning it away from the RX queues' CPUs with CPUAffinity is a documented option for operators who have the cores.

Memory

The engine allocates nothing per packet or per session after start. The session table, the maps, the 256 KB event ring and the 64 KB data-plane queue are all fixed. Under host memory pressure the risk is the OOM killer taking the engine, at which point the XDP link detaches and peers detect on their own timers, which is correct; the unit sets OOMScoreAdjust low to make it a late target.

Clone this wiki locally