Repository navigation
node-aware reachability probes + infra-symptom fingerprinting (a real cross-node debugging saga Radar could have shortened) #1292
nadaverell
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
The motivating case
This came from a real thread in a DevOps community group (details anonymized, translated from Hebrew). An experienced operator spent two days debugging a fresh on-prem cluster:
Environment: RKE2 · Canal · vSphere (not Tanzu, manually provisioned VMs, no CPI) · RHEL · SELinux disabled · firewalld disabled · broken from the first deployment.
Symptoms:
nmapshowsfiltered;tcpdumpshows the packets never arrive at all — except ICMPThe triage journey the group went through (each step hours apart):
What it most likely actually is (fits every symptom, and the CNI swap reproducing it is the tell): the well-known vmxnet3 UDP tunnel offload issue on ESXi. Both Canal and default Cilium encapsulate with VXLAN (UDP 8472); the vmxnet3 NIC advertises
tx-udp_tnl-segmentationoffload, ESXi mishandles the encapsulated TCP/UDP packets on non-NSX VXLAN ports, and they're dropped at the sender's hypervisor — so the receiver's tcpdump sees nothing. ICMP doesn't ride the offload path, so ping survives. The classic fix, on every node's uplink interface:(The runner-up suspect in a managed environment like this is an NSX Distributed Firewall policy allowing ICMP but default-denying TCP/UDP.)
Why Radar (with #1037) wouldn't have saved them — yet
The reachability work in #1037 is deliberately honest about what it verifies, and here it would honestly report: config clean, live TCP flaky. True — but this operator established that in the first hour with nmap and tcpdump. The two facts that actually cracked the case were:
Today's in-cluster probes can't produce either: probe Jobs are scheduled wherever the scheduler puts them and dial the Service name (
internal/reachability/incluster.go), so on a cluster like this the result is nondeterministically green or red depending on where the probe pod lands and which endpoint kube-proxy picks. "Flaky" is less precise than what the operator already knew.Proposal
1. Node-aware probe matrix
Extend the in-cluster test with a topology mode:
MaxInClusterProbescap)When cross-node probes fail while same-node probes succeed, emit a first-class finding:
This fits #1037's honesty model precisely: the same-node/cross-node asymmetry is measured evidence, not inference, and it's exactly the fact that reframes the whole investigation from "my manifests" to "my infrastructure."
2. Symptom → known-cause shortlist
cross-node-only + TCP/UDP-only + ICMP-passes + config-cleanis a fingerprint with a short, famous list of causes. Once the matrix detects it, surface a curated shortlist as hypotheses (never verdicts — we can't see below the K8s API), each with the exact next command:ethtool -K <uplink> tx-udp_tnl-segmentation off tx-udp_tnl-csum-segmentation offtcpdump -ni <uplink> udp port 8472(flannel/Cilium) /4789(Calico VXLAN) on the receivernm-cloud-setuprke2-canal.confignore rules; disablenm-cloud-setupDetection can be sharpened by what we already have in the cache: node
providerID/nodeInfoidentifying vSphere VMs, the CNI identified from running DaemonSets — enough to rank "VMware offload bug" first for exactly this environment.Why it's worth it
This failure class (cross-node overlay breakage on on-prem/VM infrastructure) is one of the most common and most expensive K8s networking failures — and the thread above shows how it goes without tooling: a competent operator, two days, a group chat crowd-sourcing theories, and at the end they were converging on a wrong theory (promiscuous mode) while the likely fix was a one-line
ethtool. Radar could have produced the matrix in one probe run and put the ethtool command on screen in minute five.Rough implementation surface:
internal/reachability/(node pinning + per-endpoint targets),internal/trace/coverage.go(matrix projection + new finding), the Reachability tab matrix UI, and the MCPdiagnosepayload.Feedback welcome — especially from anyone running multi-node clusters on vSphere/Proxmox/bare metal who's hit this class of problem: what signal would have saved you the most time?
All reactions