Repository navigation
Troubleshooting
Symptoms first. These are the ways this has actually gone wrong, not a list of things that might.
Check the far end has the peer configured at all. A session configured on one side only sits Down forever and looks like a data-plane fault:
# on this box
sudo vtysh -c 'show bfd peers brief' | awk '$4!="up"'
# on the peer
sudo vtysh -c 'show bfd peers brief' | grep <this box's address>
If the peer has no matching entry, nothing below will help.
dplane: session table full
The engine holds 1024 sessions. bfdd will happily register more, and
the ones that do not fit stay in software or stay down. Count what bfdd
has against sessions_configured in the snapshot; if bfdd has more than
1024, that is the whole story.
If the engine has about 58 and the rest are Down with no session table full, the bfdd is a release that lost its registration burst; see
Limitations.
bfdd re-registers sessions when the engine reconnects, but it does not re-send everything for every session, so an authenticated or echo session can come back without the state it needs. Restarting bfdd makes it re-register from scratch:
sudo systemctl restart frr
If you restart the engine on a live box, expect this, and prefer
--dp-hold so the sessions are adopted rather than rebuilt. With --pin
as well, restart by starting the new engine alongside the old one (see
Deployment): the sessions are adopted with their state and
the peers see nothing.
Three separate causes, in the order worth checking:
-
The key has no algorithm. See the end of
Deployment.
show bfd peerreports authentication as configured either way, so it is not a useful check. - A released bfdd cannot carry keys to a data plane. Only FRR master can, since #23331. With a release this engine refuses to run such a session rather than send it in the clear, so it never comes up. See Limitations.
- The far end is not FRR. Keyed SHA1 here follows bfdd, which computes an HMAC where RFC 5880 s6.7.4 specifies a plain SHA1 with the key embedded. A conforming implementation will not interoperate.
show bfd peer looks the same whether the kernel or userspace is
answering. Check the engine instead:
sudo pkill -USR1 -x xdp-bfd # bfd_tx if you built from source
python3 -m json.tool /tmp/bfd_tx_stats.json | head -12
-
"kernel_tx": false- no--kernel-tx, so it never attached. -
"sessions_configured": 0- bfdd never connected. Check--dplaneaddrmatches the engine's--dplane, and that the engine loggeddplane: bfdd connected. - A session whose
last_ktx_usstays 0 is being answered from userspace, not from the program.
It is written on SIGUSR1, not continuously, and only if the engine was
started with --stats-dump. A file with an old timestamp usually means
an engine that was restarted without the flag, leaving the previous
one's file behind, which will happily report every session up for days.
Check
the age before believing it:
stat -c %y /tmp/bfd_tx_stats.json
Something else attached. The most common something else is
xdp-bfd-observe, which attaches the same program in promiscuous mode
for debugging and replaces whatever was there:
ip -d link show dev eth0 | grep -o 'xdp.*id [0-9]*'
Never run the observer on an interface an engine is serving. It also reopens the flood path that the unknown-session drop exists to close.
The engine failed to bind 50700 because something else holds it, usually a self-connected bfdd socket from a previous attempt. See step 1 of Deployment; reserving the port prevents it recurring.
sudo ss -tnp | grep 50700
unknown-session rising steadily means something is sending well-formed
BFD for address pairs you have not configured. That is either a
misconfigured neighbour or someone probing; the packets are dropped in
the driver either way. Monitoring explains the rest.
Running it
Reference
Development