Skip to content

Limitations

Abdul Wasey edited this page Oct 1, 2026 · 8 revisions

Limitations

What this will not do, and why. Most of these are consequences of sitting underneath bfdd rather than replacing it.

Authentication needs FRR master

bfdd sends a session's keys to the data plane only since FRR #23331, merged on 2026-09-25 and not in any release yet. With master, an authenticated session is offloaded with every key in its chain and their send and accept lifetimes, so a rollover happens in the data plane without bfdd.

A released bfdd has nowhere to put a key and offloads an authenticated session anyway. This engine gets no keys for it and will not run it in the clear, so the session never comes up: loud and safe rather than quiet and wrong.

With a released FRR, keep authenticated sessions off the data plane.

FRR #23463, approved and not yet merged, closes that gap from bfdd's side: a data plane declares what it can run when it connects, and bfdd offloads an authenticated session only to one that declared authentication, keeping it in software otherwise. The engine declares it since #29; a bfdd without #23463 ignores the declaration.

Keyed SHA1 follows bfdd, not the RFC

RFC 5880 s6.7.4 places the key in the Auth Key/Hash field, takes a plain SHA1 over the whole packet, and writes the digest back into that field. bfdd zeroes the field and computes an HMAC instead. The two do not interoperate.

bfdd is the control plane on one side of every session here, so this follows bfdd. Against another FRR it is correct; against a conforming implementation it will not authenticate. Filed as FRR #23274.

No keyed MD5

Types 2 and 3 are not implemented, because no bfdd keychain algorithm maps onto them. Nothing that bfdd can configure can reach this path.

Demand mode verifies the path; stock bfdd does not

A system in demand mode holds its detection timer, which is what demand mode is for. Taken literally that means nothing can ever bring such a session down, because the timer that would notice is the one being held.

So a demanding session here polls periodically to confirm the path is still there. --demand-poll-us sets the bound, never faster than the session's own detection budget; 0 restores bfdd's behaviour exactly. This is a deliberate deviation from RFC 5880 s6.6 in the safe direction, and it is the only place this engine sends something bfdd would not.

1024 sessions by default, up to 8192

The table holds 1024 sessions unless --max-sessions says otherwise (64-8192); the program's maps and the source-port block are sized from it at start. bfdd will register more than fit, and the surplus stays in software or stays down; the engine logs session table full.

Measured against a software bfdd peer, 300 ms timers: 2048 sessions at 4% of a core and 4096 at 6%, with no flaps; at 8192 the engine used 12% but the peer, running every session in software on one core, fell behind and 38 sessions flapped. At that scale also check the neighbour table: it is shared by every network namespace, and past gc_thresh3 the host loses ARP for everything, its own management address included.

A released bfdd, through 10.7.1, registers every session in one burst when it connects and loses whatever does not fit its 8 KB output buffer: about 58 sessions reach the engine and the rest stay down, on every reconnect. #22645 fixes that on master. With master, 1024 single-hop sessions at 300 ms came up in 2 s. On the testbed VMs, with 60 sessions at 50 ms besides, the engine uses about 2% of one core, the program about 1.3 µs per packet, and a cut session is detected within its budget plus the 5 ms sweep.

Some sessions need the engine to be scheduled

The fast path answers each packet the peer sends, from softirq, so a session whose packets are all answers keeps going while userspace is starved. Some sessions send more than they receive: asymmetric timers, where we must send faster than the peer does, and demand mode at one end only, where the peer goes quiet. Their extra packets come from the engine's loop, and so do polls, detection and the notification to bfdd.

Under a real-time load on every core those need the engine above the load, which the packaged unit does (SCHED_FIFO 60, see Deployment). Run at normal priority, on bare metal with 1024 sessions, they were the only sessions that flapped.

Two of these must not face each other

Transmission from the fast path is clocked by reception: the program answers packets as they arrive. Two such engines pointed at each other have no clock between them and will transmit as fast as they can process frames.

Against bfdd this is bounded, because bfdd paces from its own timer, and that is the deployment this is built for. Standalone mode with --kernel-tx on both ends is a test configuration, not a supported one.

Clone this wiki locally