Releases: EsDmitrii/kconmon-ng
Release list
v2.5.1
kconmon-ng v2.5.1
Changed
- Shorter alert texts. Every built-in alert's
descriptionis now one or
two sentences on what to check first, down from up to a thousand
characters, so a notification template that prints one per firing alert no
longer turns four pairs into a wall of text. The explanations moved to the
new Alert runbooks
page. Node-pair alerts no longer print zones, which read as(zone )on
clusters without zone labels, andPathMTUBlackHole's summary now reads
"only packets up to N bytes get through".
Added
runbook_urlon every built-in alert, pointing at its section of the
Alert runbooks page.namespacelabel on every built-in alert, set to the release namespace.
The aggregated rules had none, so notification templates showed
Namespace: unknown, and an Alertmanager inhibit rule with
equal: [namespace]treated them as matching every other alert without a
namespace.
Fixed
- The console's notice when
console.alerting.enabledis off read as if
alerting as a whole were off. It now says that only the rules built in the
console are not applied, and that the chart's built-in rules keep alerting
through Prometheus.
Upgrade notes
- Alertmanager routes and inhibit rules that match on
namespacenow see
kconmon-ng's alerts in the release namespace; check routes that send a
namespace to an application team. - Templates, silences or tests that match the old
summaryor
descriptiontext need the new wording. Expressions, thresholds and the
other labels are unchanged.
- Docs: https://esdmitrii.github.io/kconmon-ng/
- Release notes: https://esdmitrii.github.io/kconmon-ng/reference/release-notes/
- Helm chart: https://artifacthub.io/packages/helm/kconmon-ng/kconmon-ng
- Images:
ghcr.io/esdmitrii/kconmon-ng-agent:2.5.1,
ghcr.io/esdmitrii/kconmon-ng-controller:2.5.1,
ghcr.io/esdmitrii/kconmon-ng-console:2.5.1
v2.5.0
kconmon-ng v2.5.0
2.5.0 adds a path MTU probe for the failure small probes cannot see: a pair
where handshakes and pings cross while full-size packets vanish now turns red
with the size that still crosses, instead of staying green on every plane.
The rest of the release pays the debts a first outside user runs into:
node-level alerts, maintenance windows that hold the console's webhooks,
local users managed from the console, a config reload that survives the way
files are really replaced and applies what it reads, a console whose first
load is a fifth of what it was, and a set of security fixes.
Read the Upgrade notes before rolling out: the probe is on by default, two of
the four new rules are critical, and the chart's NetworkPolicy is split per
component.
Added
- Path MTU probe (
config.checkers.pmtu, on by default). Once a minute
every agent sends each peer's UDP echo port a 64-byte datagram and a
full-size one with Don't Fragment set. The full size is the MTU of the route
to the peer (the route's ownmtuwhen the CNI sets one, as Cilium does,
else the egress device's);sizeoverrides it. When the full size does not
come back, the agent bisects (at most 16 sizes,timeout500ms each) and
confirms both ends, so a lossy path does not pass for a black hole. A pair
readsok(the full size is echoed),reduced(the path answers ICMP
frag-needed: TCP adapts, UDP without its own path MTU discovery does not) or
blackhole(full-size datagrams vanish while the small one crosses, and
large transfers stall). A lost small datagram is a connectivity failure,
left to the UDP plane, and a pmtu failure does not trigger MTR. The agent
warns whenintervalis under 28 timeouts (14s) or 3m and more. New series:
kconmon_ng_pmtu_bytesandkconmon_ng_pmtu_probe_bytes(per pair: the
size that crossed, the size probed),kconmon_ng_pmtu_results_total
(successfor ok and reduced,failfor a black hole),
kconmon_ng_zone_pmtu_results_totalandkconmon_ng_agent_pmtu_probe_bytes.
Agents advertiseplane:pmtu. Walkthrough in
Catch an MTU black hole. PathMTUBlackHoleandZonePathMTUBlackHole
(prometheusRule.pathMtuBlackHole, warning): more than half of a pair's
probes failed over 10 minutes, or more thansustainedThreshold(0.1) over
30 minutes with at least two failures and one in the last 10, held for 5.
The second arm catches a black hole on one of several ECMP paths. The zone
rule fires only where no per-pair pmtu series exist
(agent.metrics.detail=zone-only). A network that carries less than its
routes say on purpose and clamps TCP MSS can setconfig.checkers.pmtu.size
or turn the rule off.NodeUnreachableandNodeIsolated(prometheusRule.nodeUnreachable,
.nodeIsolated, critical,for: 5m): most peers fail most of their TCP
probes to a node, or a node fails to reach most of its peers, with at least
minPeers(2) reporting. They catch a node that stays registered behind a
host firewall, a NetworkPolicy or a broken CNI datapath; the
inhibit rules
on the metrics page fold the per-pair alerts under them.- Path MTU on the dashboards. Overview opens with the key indicators in
two rows (agents, leader, pairs, pairs with failures, black-hole and
reduced-path pairs), then the worst-pair and MTR bars, with the charts and
tables below, pairs below their probe size among them. Node Detail: the path
MTU to and from each peer and black-hole probes by peer. Panels show the
smallest size of the last 10 minutes, so an ECMP-split black hole stays on
screen. - PMTU in the console and the CLI. The matrix gains a PMTU protocol: the
path MTU in bytes, green at full size, amber on a reduced path ("1400 of
1500") or while recovering, red while recent probes fail, dashed "Not run"
for a 2.4.x agent. The Overview, node and pair pages follow it. Run checks
acceptspmtubetween nodes, and
kubectl kconmon check <source> <destination> --type pmtuexits 2 on a
black hole. - Local users in the console. With
auth.mode=local, Settings > Users
adds, re-roles, resets, disables and deletes accounts under the new
users:managepermission (built-in admin only), which the last enabled
holder cannot lose. Every local user can change their own password. A
password change or reset ends the user's other sessions; a disable or delete
also revokes their API tokens. Routes under/api/v1/usersin the
Console API. - NetworkPolicy keys (Upgrade notes 4 to 6):
networkPolicy.dnsEgressreplaces the default DNS rule, for NodeLocal
DNSCache and other host-network resolvers;networkPolicy.ciliumKubeAPIEgress(auto) adds CiliumNetworkPolicies
for apiserver and node traffic, which no ipBlock matches on Cilium;networkPolicy.clusterCIDRscarves pod and Service CIDRs out of the
default0.0.0.0/0egress, which on Calico and Antrea matches pods;console.networkPolicy.prometheusTargetPortopens the pod port behind
a Prometheus Service that maps it (Thanos 9090 to 10902).
console.clientAddress.trustedProxyCIDRs: the proxies whose
X-Forwarded-Fornames the client for rate limits, the WebSocket cap and
the audit log, never for identity (Upgrade note 20).agent.tls.enabled(false): TLS verified against the system trust
pool with no other TLS field set, for an external agent whose gateway has a
publicly signed certificate.
Changed
- Maintenance windows hold the console's alert webhooks. An alert that
starts firing inside a matching window (fleet-wide, a node, the pair or its
target) is delivered only if it still fires when the window closes, and not
at all if it resolves inside. A restarted console keeps holding.
kconmon_ng_console_webhook_suppressed_total{event}counts what was held. - Config hot reload applies what it reads. The keys that go live and the
ones that wait for a restart are in Upgrade note 11 and
What reloads and what does not. - The node page covers both directions, so a node
NodeUnreachablenames
no longer reads Healthy there, and its peer breakdown switches between To
peers and From peers. Diagnostic runs list failed pairs first. - Console pages load on demand. The first page load drops from 935 kB to
198 kB gzipped, and charts load only the ECharts parts they draw with. - DNS probes against an explicit resolver ask the absolute name, so the
search list no longer multiplies the query or sends cluster names outside.
checkers.dns.timeoutdefaults to 2s (was 5s). - A departed peer's series go away ten minutes after it leaves the agent's
peer list. - The controller refuses what cannot run.
POST /api/v1/diagnostics
answers 400 for an external destination with a type other thantcp,
icmpormtror for aplaneother thanpod, 501 when the source
agent does not run the type, and 503leadership lostwhen the lease goes
mid-task.kubectl kconmon checkexits 1 on these, andkubectl kconmon
finds the leader itself. - So does the console. Runs, definitions and schedules toward a target or
an ad-hoc address take onlytcp,icmpandmtr(a continuous schedule
alsodnsandhttp), and an edit that would leave a stored schedule or
definition unable to run answers 422;enabled: falsealways saves. The
import follows the same rules, and the forms offer only what runs. - Console API input. Malformed input answers 400 or 422 instead of 502,
and alert rule names that become one Prometheus alert name are refused. - The console's agent-missing template renders the chart's
KconmonAgentsMissingexpression; existing rules change on the next sync. ZoneLossHighdescription. Below about ten node pairs between two zones
one broken link crosses the default 10% alone; raise
prometheusRule.zoneLossHigh.thresholdandzoneChecksFailing.threshold.- Continuous external checks fit the controller's 8 MiB limit: the console
leaves out whole definitions, newest first, and counts them in
kconmon_ng_console_external_specs_skipped_total{reason="over-budget"}. - Console UI. A silent MTR hop shows a dash, not 100% loss; refused forms
focus the refused field; a foreign rule import asks to confirm; phones get a
drawer Close button and tables that scroll inside their card. - Chart. The console gets
GOMEMLIMITfrom its memory limit. The GeoLite2
sidecar moves toghcr.io/maxmind/geoipupdate:v8.0.0. The schema refuses
what the binaries refuse at startup. The install notes flag an Ingress with
no trusted proxies and an external gateway neither on
externalTrafficPolicy: Localnor behindloadBalancerSourceRanges. - deb/rpm. The packaged agent config ships with the DNS checker off, and
the postinstall keeps anet.ipv4.ping_group_rangethe admin already set. - Release images.
:latest, the chart and the GitHub release follow a tag
only after e2e passes on its images, and only the newest stable release
moves:latest, the Latest badge and the krew index. - Build stack. Go 1.27.1, distroless
static-debian13, Vite 8. The unused
OpenTelemetry SDK is gone;observability.otel.*logs a warning.
Fixed
- Agents.
- Hot reload stopped for good after the first atomic replacement of the
file (an editor's save, a puppet file resource, a ConfigMap swap). logLevel: DEBUGandlogFormat: TEXTran at info in JSON.- An agent probed at a secondary IP, a multi-homed
advertiseAddressor a
VIP read...
- Hot reload stopped for good after the first atomic replacement of the
v2.5.0-rc.3
kconmon-ng v2.5.0
2.5.0 adds a path MTU probe for the failure small probes cannot see: a pair
where handshakes and pings cross while full-size packets vanish now turns red
with the size that still crosses, instead of staying green on every plane.
The rest of the release pays the debts a first outside user runs into:
node-level alerts, maintenance windows that hold the console's webhooks,
local users managed from the console, a config reload that survives the way
files are really replaced and applies what it reads, a console whose first
load is a fifth of what it was, and a set of security fixes.
Read the Upgrade notes before rolling out: the probe is on by default, two of
the four new rules are critical, and the chart's NetworkPolicy is split per
component.
Added
- Path MTU probe (
config.checkers.pmtu, on by default). Once a minute
every agent sends each peer's UDP echo port a 64-byte datagram and a
full-size one with Don't Fragment set. The full size is the MTU of the route
to the peer (the route's ownmtuwhen the CNI sets one, as Cilium does,
else the egress device's);sizeoverrides it. When the full size does not
come back, the agent bisects (at most 16 sizes,timeout500ms each) and
confirms both ends, so a lossy path does not pass for a black hole. A pair
readsok(the full size is echoed),reduced(the path answers ICMP
frag-needed: TCP adapts, UDP without its own path MTU discovery does not) or
blackhole(full-size datagrams vanish while the small one crosses, and
large transfers stall). A lost small datagram is a connectivity failure,
left to the UDP plane, and a pmtu failure does not trigger MTR. The agent
warns whenintervalis under 28 timeouts (14s) or 3m and more. New series:
kconmon_ng_pmtu_bytesandkconmon_ng_pmtu_probe_bytes(per pair: the
size that crossed, the size probed),kconmon_ng_pmtu_results_total
(successfor ok and reduced,failfor a black hole),
kconmon_ng_zone_pmtu_results_totalandkconmon_ng_agent_pmtu_probe_bytes.
Agents advertiseplane:pmtu. Walkthrough in
Catch an MTU black hole. PathMTUBlackHoleandZonePathMTUBlackHole
(prometheusRule.pathMtuBlackHole, warning): more than half of a pair's
probes failed over 10 minutes, or more thansustainedThreshold(0.1) over
30 minutes with at least two failures and one in the last 10, held for 5.
The second arm catches a black hole on one of several ECMP paths. The zone
rule fires only where no per-pair pmtu series exist
(agent.metrics.detail=zone-only). A network that carries less than its
routes say on purpose and clamps TCP MSS can setconfig.checkers.pmtu.size
or turn the rule off.NodeUnreachableandNodeIsolated(prometheusRule.nodeUnreachable,
.nodeIsolated, critical,for: 5m): most peers fail most of their TCP
probes to a node, or a node fails to reach most of its peers, with at least
minPeers(2) reporting. They catch a node that stays registered behind a
host firewall, a NetworkPolicy or a broken CNI datapath; the
inhibit rules
on the metrics page fold the per-pair alerts under them.- Path MTU on the dashboards. Overview opens with the key indicators in
two rows (agents, leader, pairs, pairs with failures, black-hole and
reduced-path pairs), then the worst-pair and MTR bars, with the charts and
tables below, pairs below their probe size among them. Node Detail: the path
MTU to and from each peer and black-hole probes by peer. Panels show the
smallest size of the last 10 minutes, so an ECMP-split black hole stays on
screen. - PMTU in the console and the CLI. The matrix gains a PMTU protocol: the
path MTU in bytes, green at full size, amber on a reduced path ("1400 of
1500") or while recovering, red while recent probes fail, dashed "Not run"
for a 2.4.x agent. The Overview, node and pair pages follow it. Run checks
acceptspmtubetween nodes, and
kubectl kconmon check <source> <destination> --type pmtuexits 2 on a
black hole. - Local users in the console. With
auth.mode=local, Settings > Users
adds, re-roles, resets, disables and deletes accounts under the new
users:managepermission (built-in admin only), which the last enabled
holder cannot lose. Every local user can change their own password. A
password change or reset ends the user's other sessions; a disable or delete
also revokes their API tokens. Routes under/api/v1/usersin the
Console API. - NetworkPolicy keys (Upgrade notes 4 to 6):
networkPolicy.dnsEgressreplaces the default DNS rule, for NodeLocal
DNSCache and other host-network resolvers;networkPolicy.ciliumKubeAPIEgress(auto) adds CiliumNetworkPolicies
for apiserver and node traffic, which no ipBlock matches on Cilium;networkPolicy.clusterCIDRscarves pod and Service CIDRs out of the
default0.0.0.0/0egress, which on Calico and Antrea matches pods;console.networkPolicy.prometheusTargetPortopens the pod port behind
a Prometheus Service that maps it (Thanos 9090 to 10902).
console.clientAddress.trustedProxyCIDRs: the proxies whose
X-Forwarded-Fornames the client for rate limits, the WebSocket cap and
the audit log, never for identity (Upgrade note 20).agent.tls.enabled(false): TLS verified against the system trust
pool with no other TLS field set, for an external agent whose gateway has a
publicly signed certificate.
Changed
- Maintenance windows hold the console's alert webhooks. An alert that
starts firing inside a matching window (fleet-wide, a node, the pair or its
target) is delivered only if it still fires when the window closes, and not
at all if it resolves inside. A restarted console keeps holding.
kconmon_ng_console_webhook_suppressed_total{event}counts what was held. - Config hot reload applies what it reads. The keys that go live and the
ones that wait for a restart are in Upgrade note 11 and
What reloads and what does not. - The node page covers both directions, so a node
NodeUnreachablenames
no longer reads Healthy there, and its peer breakdown switches between To
peers and From peers. Diagnostic runs list failed pairs first. - Console pages load on demand. The first page load drops from 935 kB to
198 kB gzipped, and charts load only the ECharts parts they draw with. - DNS probes against an explicit resolver ask the absolute name, so the
search list no longer multiplies the query or sends cluster names outside.
checkers.dns.timeoutdefaults to 2s (was 5s). - A departed peer's series go away ten minutes after it leaves the agent's
peer list. - The controller refuses what cannot run.
POST /api/v1/diagnostics
answers 400 for an external destination with a type other thantcp,
icmpormtror for aplaneother thanpod, 501 when the source
agent does not run the type, and 503leadership lostwhen the lease goes
mid-task.kubectl kconmon checkexits 1 on these, andkubectl kconmon
finds the leader itself. - So does the console. Runs, definitions and schedules toward a target or
an ad-hoc address take onlytcp,icmpandmtr(a continuous schedule
alsodnsandhttp), and an edit that would leave a stored schedule or
definition unable to run answers 422;enabled: falsealways saves. The
import follows the same rules, and the forms offer only what runs. - Console API input. Malformed input answers 400 or 422 instead of 502,
and alert rule names that become one Prometheus alert name are refused. - The console's agent-missing template renders the chart's
KconmonAgentsMissingexpression; existing rules change on the next sync. ZoneLossHighdescription. Below about ten node pairs between two zones
one broken link crosses the default 10% alone; raise
prometheusRule.zoneLossHigh.thresholdandzoneChecksFailing.threshold.- Continuous external checks fit the controller's 8 MiB limit: the console
leaves out whole definitions, newest first, and counts them in
kconmon_ng_console_external_specs_skipped_total{reason="over-budget"}. - Console UI. A silent MTR hop shows a dash, not 100% loss; refused forms
focus the refused field; a foreign rule import asks to confirm; phones get a
drawer Close button and tables that scroll inside their card. - Chart. The console gets
GOMEMLIMITfrom its memory limit. The GeoLite2
sidecar moves toghcr.io/maxmind/geoipupdate:v8.0.0. The schema refuses
what the binaries refuse at startup. The install notes flag an Ingress with
no trusted proxies and an external gateway neither on
externalTrafficPolicy: Localnor behindloadBalancerSourceRanges. - deb/rpm. The packaged agent config ships with the DNS checker off, and
the postinstall keeps anet.ipv4.ping_group_rangethe admin already set. - Release images.
:latest, the chart and the GitHub release follow a tag
only after e2e passes on its images, and only the newest stable release
moves:latest, the Latest badge and the krew index. - Build stack. Go 1.27.1, distroless
static-debian13, Vite 8. The unused
OpenTelemetry SDK is gone;observability.otel.*logs a warning.
Fixed
- Agents.
- Hot reload stopped for good after the first atomic replacement of the
file (an editor's save, a puppet file resource, a ConfigMap swap). logLevel: DEBUGandlogFormat: TEXTran at info in JSON.- An agent probed at a secondary IP, a multi-homed
advertiseAddressor a
VIP read...
- Hot reload stopped for good after the first atomic replacement of the
v2.4.0
kconmon-ng v2.4.0
External agents stop being second-class. A host outside the cluster now
tells its peers where it listens, gets scraped without a hand-written
target, and shows up as what it is in the console, the CLI and the Time
Machine. One rule comes with it, and it is the one to read before rolling
out: per-agent ports are honoured only by upgraded agents; keep one port
set until every agent, deb/rpm hosts included, runs 2.4.0. An older agent
reports no ports and dials every peer on its own configured values, so in a
fleet whose ports differ each old agent goes one-way red toward every peer
listening elsewhere, its on-demand diagnostics included. The full skew
matrix is under Upgrade notes at the end of this section.
Added
- Per-agent ports on the wire.
AgentMetagainshttp_port,
udp_portandmetrics_port. Every agent reports its three listener
ports at registration, and peers probe it on the ones it reported, for the
scheduled mesh and for on-demand tasks alike, so an external host no longer
has to mirror the cluster's port pair. Zero means "not reported" (an agent
older than 2.4.0): the prober then dials its own configured port, each port
falling back on its own. The controller refuses a port above 65535, peer
lists carry the two probe ports but notmetrics_port(nobody dials it),
and an agent never adopts ports from the controller's reply; zone stays the
only thing it takes from there. See
Ports. - Prometheus HTTP SD for external agents. A bare host has no Service for
a ServiceMonitor to select, so the controller, the one party that knows
the host registered and on which address, now publishes it:
GET /api/v1/prometheus/sd, served onhttpPortand onmetricsPort(the
port the chart's scrape NetworkPolicy already opens). The contract:- one target group per external agent,
<advertised address>:<metricsPort>,
sorted by node name and deduplicated by address; - a fixed label set,
node,zone,external="true"andagent_id. An
agent's own labels never reach Prometheus, so a host cannot inject
target labels; - with no external agent registered the body is the literal
[], and
the list is always served withCache-Control: no-store; - a standby answers
503 not the leader, never200 []: Prometheus reads
every 200 as the complete target set, so an empty one from a standby
would wipe every external target, while on a non-200 it keeps the list
it has; - an agent that reported no metrics port (older than 2.4.0) is published
on the controller's ownconfig.metricsPort, and the controller logs
metrics port assumed from controller configonce per agent, which is
the clue when such a host on another port sits atup == 0; controller.prometheusSD.enabled: falsecloses the route (404 on both
listeners). The key reaches the shared ConfigMap only when false, so an
older controller image never trips over it;- with
controller.replicaCount > 1the controller Service spreads
refreshes over all replicas and roughly half of them land on a standby:
prometheus_sd_http_failures_totalclimbs for the job while the targets
stay correct. Cosmetic, and written down so nobody chases it. Body and
semantics in the
HTTP API reference.
- one target group per external agent,
scrapeConfig.externalAgentsin the chart. Renders a Prometheus
OperatorScrapeConfig(needs thescrapeconfigs.monitoring.coreos.com
CRD) named<release>-agent-externalthat reads the SD route, with
labelsfor your Prometheus' selector (kube-prometheus-stack wants
release: <its release name>),jobName,refreshInterval(30s) and
interval. It applies the sameagent.metrics.detailvalve as the agent
ServiceMonitor, so an external host never returns per-pair detail the valve
drops for the pods, and the valve no longer insists on
serviceMonitor.enabledwhen this is on. The chart refuses the
ScrapeConfig withoutcontroller.externalGateway.enabled(nothing external
could register) or withcontroller.prometheusSD.enabled=false(every
refresh would 404), and the install notes remind you when the gateway is on
without it, or whenlabelsis empty. Plain-Prometheushttp_sd_configs
job and the reachability rules in
Scraping external agents.KconmonExternalAgentDownandkconmon_ng_controller_external_agents.
An optional warning (prometheusRule.externalAgentDown, off by default,
for: 5m) onup{job=~".*agent-external.*"} == 0: a host the controller
lists that Prometheus cannot scrape, usually the host firewall admitting the
Prometheus pod IP when the CNI NATs its egress to a node IP. The new
controller gauge counts registered agents that came through the gateway.agent.hostNetwork, for pod networks external hosts cannot route.
The DaemonSet moves into each node's network namespace: the agents
advertise the node IP (KCONMON_NG_POD_IPfromstatus.hostIP), declare
hostPorton all three ports, getdnsPolicy: ClusterFirstWithHostNet
unlessagent.dnsPolicysays otherwise, and label themselves
kconmon-ng.io/host-network=truefrom a 2.4.0 image. The chart stops
rendering theping_group_rangepod sysctl there, since the kubelet refuses
net.*sysctls in the host namespace. It changes what is measured, for
the whole DaemonSet: every in-cluster pair then probes node IP to node IP
over the underlay, and the CNI datapath (overlay, conntrack, NetworkPolicy
enforcement) is no longer on the probe path, so the breakage this tool
exists to catch can hide behind a green matrix. Turn it on only when the
goal is visibility between external agents and a cluster whose pod network
they cannot reach. Before you do: PSSprivilegedfor the namespace, TCP
8080, UDP 9090 and TCP 9091 (or yourconfig.*Portvalues) free on every
node,ping_group_rangeset by the node OS, and one agent per machine (a
host-network pod and a bare-host agent cannot share an IP). See
When the pod network does not route
and Host networking.networkPolicy.nodeCidrsandnetworkPolicy.externalPeerCidrs.
Host-network agents register from node IPs that no pod selector matches, so
withagent.hostNetworkandnetworkPolicy.enabledboth on, the chart
refuses to render the policy untilnodeCidrslists the node CIDRs; without
it every registration but the one from the controller's own node would drop
silently.externalPeerCidrscloses the old gap where an external agent
registered fine and every cell between it and the cluster stayed red: its
CIDRs join the agent-to-agent rules in both directions (UDPgrpcPort,
TCPhttpPort, the ports-less ICMP/MTR rule) and never the gateway rule.- External agents in the console. Everything keys off the
kconmon-ng.io/externalregistration label, which the console now passes
through from the controller's topology together with the agent's
capabilities (labelsandcapabilitiesonTopologyAgentin the Console
API).- Topology draws the host beside the cluster nodes in the lane of its
zone, with a neutral external badge (identity, never a health tier)
and "readiness unknown" for screen readers. - Node page swaps Pod IP for Advertised address, explains the Ready
dash, and lists the probe Planes the agent advertised. Agents now
advertiseplane:tcp,plane:udp,plane:icmp,plane:dns,
plane:httpandplane:mtr; an agent advertising none (older than
2.4.0) reads as "unknown", never as running nothing. - Overview badges the host in Worst pairs and adds "+N external agents"
beside Nodes ready without counting them in, since that tile is
Kubernetes readiness. - Matrix tells two silences apart from plain no-data. An external agent
Prometheus is not scraping keeps the no-data fill and aria text, but its
cells' tooltip and a note above the grid say why and link the scraping
docs, until the first measured cell appears. A protocol the source does
not run renders dashed like not probed, with its own legend row.
Precedence when a cell has no data: excluded by the plan, then
unsupported, then unscraped. See
Silence with a known cause
and External agents on the map.
- Topology draws the host beside the cluster nodes in the lane of its
- External agents in the Time Machine.
TopologyChangedevents carry the
agent's labels, so a replay badges a host the way the live view does, and a
reconstructed topology lists a bare host underagentsonly, never as a
presence-derived READY node. History recorded before the upgrade shows no
external badges: a 2.3.x controller wrote no labels, and such a host stays
an ordinary node in those instants. Historical responses never carry
capabilities, since no event records them.
Changed
KconmonAgentsMissingis no longer masked by external agents.
Registered agents include them and expected agents (schedulable nodes) never
did, so one external host hid one missing in-cluster agent. The expression
now subtractscontroller_external_agents, with anor registered * 0
stand-in so the rule keeps working against a controller image that predates
the gauge.kubectl kconmontables.agentsgains anEXTERNALcolu...
v2.3.1
kconmon-ng v2.3.1
Fixed
ZoneChecksFailingandZoneLossHighfailed every evaluation with "vector
cannot contain metrics with the same labelset" and raised
PrometheusRuleFailureson the cluster:rate()over a__name__regex
union drops the metric name and collapses the per-protocol families into
duplicate labelsets. The expressions now build the union with
label_replace(...) or label_replace(...), which keeps the branches
distinct and still tolerates a disabled checker's absent family.- CI now evaluation-tests every alert rule with
promtool test rulesagainst
synthetic series for all metric families, including a positive check that
each zone alert fires on staged bad data. Rendering and syntax checks never
execute the query engine, which is exactly where this defect lived.
v2.3.0
kconmon-ng v2.3.0
The sparse mesh changes WHAT "no data for a pair" means: under
topology.mode: sparsemost directed pairs are deliberately never probed.
Everything in this release that reads per-pair series learns to tell "not
planned" from "went dark" through one new metric,
kconmon_ng_probe_intended— and that metric comes from the AGENT: images
below appVersion 2.3.0 do not export it. The chart's rules degrade
honestly on an older fleet (see PairWentSilent below), but do not flip
topology.mode: sparseuntil controller AND agents run a 2.3.0 image —
the controller config key is emitted only when sparse precisely because an
older controller image rejects it and crashloops. The appVersion pin is
aligned when the app release ships.
Added
topology.*— the sparse probe mesh, by values.topology.mode: sparsetrims the full N×(N−1) probe matrix to a ring over sorted node
names (sparse.ringDegreesuccessors each, the connectivity guarantee)
plus HRW-chosen cross-zone chords (sparse.zoneChordsper directed zone
pair, which keep the zone metric family fully populated), so probed pairs
— and every per-pair series they export — scale ~linearly with node count
instead of quadratically.sparse.autoThresholdis the floor: fleets
smaller than it get the full mesh regardless of mode, because sparse only
pays for itself at scale. Default ismode: full, byte-identical
rendering to 2.2.0.kconmon_ng_probe_intended— the plan, scrapable. A gauge, value 1
for every directed pair the topology plan assigns
({source_node, destination_node}, exported by the source agent), preset
from the peer list at registration and pruned on every plan change —
stale pairs are deleted, not left at 1. In full-mesh mode it simply marks
every peer, so dashboards and rules can join on it without caring which
mode the fleet runs. It is the one honest way to distinguish "this pair
is not supposed to report" from "this pair went dark", which is why it
ships in the same release as sparse mode and not one later.investigateUrlon the two zone alerts.ZoneChecksFailingand
ZoneLossHighnow annotate a console deep link,
/investigate?kind=zone-pair&scope=<source>-><destination>, straight
into the Investigate page scoped to the firing zone pair. The link is
console-RELATIVE on purpose — the chart cannot know the console's
external URL (ingress is optional), so notification templates prepend
their own origin; the console normalises the typeable->into its
canonical pair arrow.
Changed
PairWentSilentjoins on the plan. The rule now fires only for pairs
present in the source agent'skconmon_ng_probe_intendedseries — the
hard rule of the sparse design, shipped in the same release: without the
join, every pair the plan trims would read as "went silent" for the hour
its results take to age out of the lookback window. The fallback is per
SOURCE, not global: asource_nodeexporting noprobe_intendedat all
keeps the old two-window behaviour, so a pre-2.3.0 agent image alerts
exactly as before, a mixed fleet mid-rollout gets each behaviour where it
applies — and an agent that dies outright takes itsprobe_intended
series with it, which lands its pairs in the same fallback and preserves
the alert's original purpose: catching an agent that stopped running or
stopped being scraped.
This release also carries everything prepared for the never-published 2.2.0
tag (its pipeline caught two release-tooling defects before anything went
out); those changes follow below, under their original heading kept for
upgrade notes.
Carried over from the unreleased 2.2.0
Everything in this release reads the new zone-level metric family
(kconmon_ng_zone_*), and that family comes from the AGENT, not the chart:
agents below appVersion 2.2.0 do not export it (this chart pins 2.2.0, so a
default install is fine — the warning is for fleets running an older agent
image behind a newer chart). Until the fleet runs an agent image that does, the two zone
alerts are silently inert (their expressions match no series), the Zone
Heatmap dashboard renders empty, andagent.metrics.detail: zone-only
would drop the per-pair series with nothing replacing them — Prometheus
goes dark on the mesh while the console keeps working. Upgrade the agent
image first, flip the valve second. The appVersion pin is aligned when the
app release ships.
Added
ZoneChecksFailingandZoneLossHigh. Two alerts on the zone plane,
with the same per-rule knobs as the rest
(prometheusRule.{zoneChecksFailing,zoneLossHigh}.{enabled,threshold,for,severity}).
ZoneChecksFailingis the failure ratio of all TCP, UDP and ICMP probes
between a zone pair, in one expression — the__name__union keeps a
disabled checker from blanking the ratio.ZoneLossHighcomputes loss as
(sent − received) / sentfrom the zone packet counters; averaging the
per-pair loss-ratio gauges into a zone would weight an idle pair the same
as a busy one, so the chart never does. Its default threshold is0.1,
lower than the per-pairUDPLossHighat0.5, because the zone aggregate
dilutes any single link by the pair count: sustained loss at that level
means the fabric, not one node. Both survive everyagent.metrics.detail
mode — that is the point of alerting on the zone family.agent.metrics.detail— the cardinality valve. A scrape-time knob
rendered asmetricRelabelingson the agent ServiceMonitor:
full(default, everything, ~70 series per directed pair),
counters-only(drops the four per-pair histograms, ~10/pair — every pair
alert keeps firing),zone-only(drops every series naming a
destination_node, ~0/pair; the zone family at ~74×Z² series and the
linear DNS/HTTP/external families remain). At 100 nodes that is ~0.7M →
~0.1M → practically N-independent, by configuration alone. Setting it
withoutserviceMonitor.enabledis refused at render time rather than
silently dropping nothing; plain-Prometheus equivalents are in
docs/metrics.md.controller.externalGateway— the external agent gateway, exposed by the
chart. The controller's second gRPC listener (same services, but TLS with
a bootstrap token, for agents OUTSIDE the cluster) gets a values block and
three templates.templates/controller/service-external.yamlis a
NodePort/LoadBalancer Service carrying the gateway port ALONE — the
plaintext in-cluster gRPC port authenticates by network position and never
appears on it, because a LoadBalancer in front of it would hand the whole
mesh to anything that can reach the address. The deployment mounts two
referenced Secrets read-only:tls.secretName(akubernetes.io/tls
serving pair;tls.clientCaKeynames the CA bundle key in the same Secret
and switches on client-cert identity pinning — empty is token-only mode,
where any token holder can impersonate any agent, and NOTES.txt says so at
install) andbootstrapToken.{secretName,key}. With
networkPolicy.enabled, ingress on the gateway port is opened from
networkPolicy.externalAgentCidrstoward the controller pods alone, and an
empty list is refused at render rather than shipping a gateway no packet
can reach; missing Secret names and a port colliding with
config.{httpPort,grpcPort,metricsPort}are refused the same way. Two
operational notes. Rotation: the gateway reads the certificate and token
ONCE at startup and the chart cannot checksum content it only references,
so rotating either Secret in place needs
kubectl rollout restart deploy/<release>-controller. Version skew: the
externalGatewayconfig key is emitted only when enabled, because a
controller image at appVersion 2.0.3 rejects the unknown key and
crashloops — upgrade the image before flipping the switch, same rule as
the zone family above.
Changed
- The Zone Heatmap dashboard reads the zone family. Every panel that
aggregated per-pair series into zones at query time now reads the
pre-aggregatedkconmon_ng_zone_*metrics, so the dashboard keeps working
in everyagent.metrics.detailmode and its queries stop scaling with the
pair count. Loss panels are packet-weighted from the sent/received counters
instead of averaging the per-pair ratio gauges. The one exception is the
"MTR traces triggered" panel: MTR has no zone-level family, its counter is
per-pair, and inzone-onlymode that panel reads zero — its description
now says so.
Performance and self-observability
- Peer probing fans out with a bounded pool (32 in flight per round): a dead
peer costs one timeout, not one timeout per peer in sequence, so probe
cadence holds through partitions. - Reactive MTR traces are bounded by a global semaphore (4 in flight) on top
of the existing per-pair cooldown; a mass partition trickles traces out
instead of forking one per broken pair. - The agent exports self-metrics under
kconmon_ng_agent_*: probe cycle
duration and overruns per checker, controller reconnects, peer-list age,
reactive-MTR in-flight and coalesced counters. All are fleet-size
independent. - The controller coalesces peer-list broadcasts (trailing edge, 200 ms): a
rollout's burst of registrations produces one broadcast, not one per
change; the peer message is built once per broadcast and carries only the
fields agents read.
v2.0.3
kconmon-ng v2.0.3
Fixed
- A config change restarts the pods that read it. The agent and the
controller share one ConfigMap and read it once at startup, and a mounted
ConfigMap changes under a running process without telling it — so
controller.events.enabled: trueapplied to a live release updated the object
and left the controller on the old file. It went on advertising no
capabilities, the Console's realtime ingester retried against a stream that was
configured but never started, and nothing anywhere reported an error: the Live
page was simply empty. Both workloads now carrychecksum/config, so a values
change rolls them; a change the ConfigMap does not carry still does not.
v2.0.2
kconmon-ng v2.0.2
Fixed
- The matrix no longer opens at half size. The grid measured the height its
own content had produced and fed that back into the fit, so a fresh render
saw the container's 256px minimum, decided the grid did not fit, shrank to
50%, and the smaller grid then held the box at 256px — a loop with no way out.
It measures the space available instead, and a seven-node fleet opens at 100%. - Zooming in gives the node names back. The shared prefix every node name
begins with is dropped to buy column width, which is right while the column is
narrower than the names and wrong the moment it is not: at 125% a label column
holdsadm-kuber-01with room over and still read…01. The elision is now
decided per axis at the current scale, and the note above the grid appears only
while an axis is actually eliding.
Added
- A favicon. The console had none, so every tab showed the browser's blank
square; it now wears the mark it wears in its own sidebar.
v2.0.1
kconmon-ng v2.0.1
Fixes a hole in 2.0.0:
auth.mode=oidcandauth.mode=headershipped with no
way to grant anybody a role. Both modes worked, and neither was usable.
Fixed
-
An OIDC or header install can grant roles at deploy time. Role bindings
live in the database and are created through an API that already requires
rbac:manage, so a fresh install had nobody able to make the first binding;
the only alternative wasauth.defaultRole, which is one role for every
authenticated subject. The way out that 2.0.0 left was to bring the console up
in local mode, log in, create a binding by hand and only then switch — a
workaround, published as if it were a procedure.console.auth.groupRolesmaps a group the identity provider asserts onto a
role this console grants, in the values file:console: auth: groupRoles: platform-oncall: admin everyone: viewer
Roles resolve as the union of that map and any binding made through the API, so
a grant by hand still adds to what the provider's groups carry. A group absent
from the map grants nothing. What the map grants cannot be revoked through the
API — that is what makes it declarative. -
A role store outage no longer costs an operator their access. The store's
half still fails closed, because an unreadable database is no evidence a
subject holds anything; a grant that came from the claim and the config was
never in doubt, and an outage is when the console is most needed.
v2.0.0
kconmon-ng v2.0.0
A chart that installs monitoring and nothing else, a console that survives more
than one replica, and one that tells the truth about time. The chart no longer
ships a database or a cache — point it at the ones you already run. The Time
Machine moved out of the top bar and into each page's own time controls, and
the charts pin their axis to the window you asked for rather than to the data
that happened to arrive. MTR gained a Runner, path history that reads as a
timeline, and external targets.
Breaking
- The chart no longer installs PostgreSQL or Valkey.
database.mode,
database.cnpg.*and the bundled subcharts are gone: set
database.existingSecretto a Secret holding apostgres://DSN and
redis.existingSecretto one holding aredis://DSN, and any managed
instance works — RDS, a StatefulSet, a CloudNativePG cluster you run yourself.
Every removed key fails the render with a message naming its replacement
(templates/_migrations.tpl), so no old value is silently honoured. console.database.*moved to the top-leveldatabase.*, and
console.*keys that described the bundled datastores went with it.
Added
- Chart 2.0.0, templates split per component — agent, controller, console,
shared and observability each own their directory, with a NetworkPolicy set
covering every component, fail-closed on external egress. Render-time guards
refuse a port collision, an OIDCredirectURLthe console would not start on,
and more than one console replica without a shared cache. - The console scales past one replica — sessions, the fixed-window rate-limit
counters and the realtime fan-out live in the Redis-compatible server, and the
controller elects a leader so exactly one replica drives the reconcilers. - MTR Runner and path history — start a trace from the Explorer itself,
with a settable cadence and duration; every distinct route the fleet has taken
is kept, diffed and drawn on a timeline of when it changed. - External targets — probe a destination that is not a fleet peer, gated by
config.checkers.external.allowedCidrsand the cluster's own egress policy.
The console refuses at create time a target no agent could ever reach. - Time in the Console's result table — every figure says when it was read.
Changed
- The Time Machine lives with the page's time filters, not in a strip across
the top of every route. It is offered only on the pages that resolve their
reads through?at=, and the engaged banner stays global because writes are
disabled console-wide. - Explore's axis is the window you picked — a 24h view draws 24 hours even
when Prometheus holds less, instead of quietly redrawing three. - MTR Explorer is sorted by name, both destinations and their sources, with
numbers read as numbers (m9beforem10). - OIDC identity is the
subclaim, namespaced asoidc:<sub>— the only
claim OIDC Core §5.7 allows as an identifier.auth.oidc.usernameClaimnow
decides the display name alone, so renaming a person no longer moves their
roles (Grafana's CVE-2023-3128 is what the old shape risked). Group membership
is re-read on every token refresh. Bindings made against a username stop
granting; the console names them at boot so they can be remapped. - The configuration bundle carries access control — custom roles and the
grant list, but only for a caller who holdsrbac:manage. Roles import;
bindings never do, because a grant names a person in the source console's own
identity namespace. - Only the chart under the cursor shows a tooltip. Its neighbours keep the
shared crosshair and mark their own samples with a dot, instead of each
covering its own curves with a box of numbers.
Fixed
- WebSocket topics are authorized per topic:
events:readno longer carries the
topology and matrix snapshots thattopology:readandmatrix:readgate, and
a permission taken away reaches a socket that is already open — the topics it
may no longer have are dropped, the rest of the connection is left alone. - The audit row describes the mutation that happened. A body could name one
thing for the handler and another for the audit log by spelling a key in a
different case, and a value carrying a NUL made the whole row unwritable — in
both directions the caller chose whether their own privileged action was
recorded. The extraction now matches keys the wayencoding/jsonmatches
struct fields, and is bounded before it is decoded, so a wide body on the
public login route can no longer take the replica past its memory limit. - A broken alert rule no longer freezes the whole bundle. Editing a deployed
rule into PromQL the apiserver rejects used to stop every other rule from
being applied, while the API answered 2xx and Prometheus kept evaluating the
stale set. The quarantine now keys on rendered content rather than rule ids,
offers each suspect to the cluster on its own, and removes the object only
when every rule was offered and every one refused. auth.mode=anonymousis not exempt from CSRF. Any page an operator's
browser visited could POST into a console kept off the internet; a
cross-origin write is now refused, while a script that sends noOriginis
unaffected.- The node-local HTTP checker verifies certificates. An expired certificate,
one issued for another hostname or an interceptor's CA all used to pass, so an
https check could not fail on the condition it was added to notice; opt out per
target withinsecureSkipVerify. - External metrics separate the checks on one target — the series carry
check_type, so an icmp and a tcp check on the same target no longer average
each other's failures away under theExternalChecksFailingrule. - A check no agent could run is refused when it is written, instead of being
dropped by every agent with nothing but a log line while the console listed it
as enabled. - The MTR destination listing is complete. It is paged behind a keyset
cursor rather than capped, so no pair is missing from the Explorer and no
per-destination total is short. - A subscriber that stops reading its peer-update stream is torn down rather than
holding a controller goroutine and its connection slot until TCP notices. - Shutdown finishes in-flight runs before tearing down the pipeline they publish
onto, so a rolling update no longer logs dropped frames that were delivered. - Every request body is capped, so one oversized POST can no longer take a
console replica past its memory limit. - The OIDC callback binds its
stateto the browser that started the flow. - A role-store failure now refuses rather than granting the default role.
- External TCP and UDP checks probe what was asked for instead of speaking the
agent's own protocol to something that is not an agent. - A user binding can no longer be resolved by a subject of another kind: role
resolution matches the caller's kind as well as their id. - Revoking a role binding is auditable — the audit row names the role and the
subject, read before the row is destroyed rather than after. - Path history says when it has reached the end instead of leaving a "Load
older" button that can never be pressed, and counts the routes it is showing
against the traces folded into them. - A probe tick on a diagnostics run leads with that probe — its sequence, its
clock, its latency or its error — so two ticks on an unchanged route are no
longer indistinguishable.