Skip to content

Releases: EsDmitrii/kconmon-ng

v2.5.1

Choose a tag to compare

@github-actions github-actions released this 27 Sep 10:19
Immutable release. Only release title and notes can be modified.
v2.5.1
f3994e8

kconmon-ng v2.5.1

Changed

  • Shorter alert texts. Every built-in alert's description is now one or
    two sentences on what to check first, down from up to a thousand
    characters, so a notification template that prints one per firing alert no
    longer turns four pairs into a wall of text. The explanations moved to the
    new Alert runbooks
    page. Node-pair alerts no longer print zones, which read as (zone ) on
    clusters without zone labels, and PathMTUBlackHole's summary now reads
    "only packets up to N bytes get through".

Added

  • runbook_url on every built-in alert, pointing at its section of the
    Alert runbooks page.
  • namespace label on every built-in alert, set to the release namespace.
    The aggregated rules had none, so notification templates showed
    Namespace: unknown, and an Alertmanager inhibit rule with
    equal: [namespace] treated them as matching every other alert without a
    namespace.

Fixed

  • The console's notice when console.alerting.enabled is off read as if
    alerting as a whole were off. It now says that only the rules built in the
    console are not applied, and that the chart's built-in rules keep alerting
    through Prometheus.

Upgrade notes

  1. Alertmanager routes and inhibit rules that match on namespace now see
    kconmon-ng's alerts in the release namespace; check routes that send a
    namespace to an application team.
  2. Templates, silences or tests that match the old summary or
    description text need the new wording. Expressions, thresholds and the
    other labels are unchanged.

v2.5.0

Choose a tag to compare

@github-actions github-actions released this 27 Sep 06:57
Immutable release. Only release title and notes can be modified.
v2.5.0
8ae3735

kconmon-ng v2.5.0

2.5.0 adds a path MTU probe for the failure small probes cannot see: a pair
where handshakes and pings cross while full-size packets vanish now turns red
with the size that still crosses, instead of staying green on every plane.
The rest of the release pays the debts a first outside user runs into:
node-level alerts, maintenance windows that hold the console's webhooks,
local users managed from the console, a config reload that survives the way
files are really replaced and applies what it reads, a console whose first
load is a fifth of what it was, and a set of security fixes.

Read the Upgrade notes before rolling out: the probe is on by default, two of
the four new rules are critical, and the chart's NetworkPolicy is split per
component.

Added

  • Path MTU probe (config.checkers.pmtu, on by default). Once a minute
    every agent sends each peer's UDP echo port a 64-byte datagram and a
    full-size one with Don't Fragment set. The full size is the MTU of the route
    to the peer (the route's own mtu when the CNI sets one, as Cilium does,
    else the egress device's); size overrides it. When the full size does not
    come back, the agent bisects (at most 16 sizes, timeout 500ms each) and
    confirms both ends, so a lossy path does not pass for a black hole. A pair
    reads ok (the full size is echoed), reduced (the path answers ICMP
    frag-needed: TCP adapts, UDP without its own path MTU discovery does not) or
    blackhole (full-size datagrams vanish while the small one crosses, and
    large transfers stall). A lost small datagram is a connectivity failure,
    left to the UDP plane, and a pmtu failure does not trigger MTR. The agent
    warns when interval is under 28 timeouts (14s) or 3m and more. New series:
    kconmon_ng_pmtu_bytes and kconmon_ng_pmtu_probe_bytes (per pair: the
    size that crossed, the size probed), kconmon_ng_pmtu_results_total
    (success for ok and reduced, fail for a black hole),
    kconmon_ng_zone_pmtu_results_total and kconmon_ng_agent_pmtu_probe_bytes.
    Agents advertise plane:pmtu. Walkthrough in
    Catch an MTU black hole.
  • PathMTUBlackHole and ZonePathMTUBlackHole
    (prometheusRule.pathMtuBlackHole, warning): more than half of a pair's
    probes failed over 10 minutes, or more than sustainedThreshold (0.1) over
    30 minutes with at least two failures and one in the last 10, held for 5.
    The second arm catches a black hole on one of several ECMP paths. The zone
    rule fires only where no per-pair pmtu series exist
    (agent.metrics.detail=zone-only). A network that carries less than its
    routes say on purpose and clamps TCP MSS can set config.checkers.pmtu.size
    or turn the rule off.
  • NodeUnreachable and NodeIsolated (prometheusRule.nodeUnreachable,
    .nodeIsolated, critical, for: 5m): most peers fail most of their TCP
    probes to a node, or a node fails to reach most of its peers, with at least
    minPeers (2) reporting. They catch a node that stays registered behind a
    host firewall, a NetworkPolicy or a broken CNI datapath; the
    inhibit rules
    on the metrics page fold the per-pair alerts under them.
  • Path MTU on the dashboards. Overview opens with the key indicators in
    two rows (agents, leader, pairs, pairs with failures, black-hole and
    reduced-path pairs), then the worst-pair and MTR bars, with the charts and
    tables below, pairs below their probe size among them. Node Detail: the path
    MTU to and from each peer and black-hole probes by peer. Panels show the
    smallest size of the last 10 minutes, so an ECMP-split black hole stays on
    screen.
  • PMTU in the console and the CLI. The matrix gains a PMTU protocol: the
    path MTU in bytes, green at full size, amber on a reduced path ("1400 of
    1500") or while recovering, red while recent probes fail, dashed "Not run"
    for a 2.4.x agent. The Overview, node and pair pages follow it. Run checks
    accepts pmtu between nodes, and
    kubectl kconmon check <source> <destination> --type pmtu exits 2 on a
    black hole.
  • Local users in the console. With auth.mode=local, Settings > Users
    adds, re-roles, resets, disables and deletes accounts under the new
    users:manage permission (built-in admin only), which the last enabled
    holder cannot lose. Every local user can change their own password. A
    password change or reset ends the user's other sessions; a disable or delete
    also revokes their API tokens. Routes under /api/v1/users in the
    Console API.
  • NetworkPolicy keys (Upgrade notes 4 to 6):
    • networkPolicy.dnsEgress replaces the default DNS rule, for NodeLocal
      DNSCache and other host-network resolvers;
    • networkPolicy.ciliumKubeAPIEgress (auto) adds CiliumNetworkPolicies
      for apiserver and node traffic, which no ipBlock matches on Cilium;
    • networkPolicy.clusterCIDRs carves pod and Service CIDRs out of the
      default 0.0.0.0/0 egress, which on Calico and Antrea matches pods;
    • console.networkPolicy.prometheusTargetPort opens the pod port behind
      a Prometheus Service that maps it (Thanos 9090 to 10902).
  • console.clientAddress.trustedProxyCIDRs: the proxies whose
    X-Forwarded-For names the client for rate limits, the WebSocket cap and
    the audit log, never for identity (Upgrade note 20).
  • agent.tls.enabled (false): TLS verified against the system trust
    pool with no other TLS field set, for an external agent whose gateway has a
    publicly signed certificate.

Changed

  • Maintenance windows hold the console's alert webhooks. An alert that
    starts firing inside a matching window (fleet-wide, a node, the pair or its
    target) is delivered only if it still fires when the window closes, and not
    at all if it resolves inside. A restarted console keeps holding.
    kconmon_ng_console_webhook_suppressed_total{event} counts what was held.
  • Config hot reload applies what it reads. The keys that go live and the
    ones that wait for a restart are in Upgrade note 11 and
    What reloads and what does not.
  • The node page covers both directions, so a node NodeUnreachable names
    no longer reads Healthy there, and its peer breakdown switches between To
    peers and From peers. Diagnostic runs list failed pairs first.
  • Console pages load on demand. The first page load drops from 935 kB to
    198 kB gzipped, and charts load only the ECharts parts they draw with.
  • DNS probes against an explicit resolver ask the absolute name, so the
    search list no longer multiplies the query or sends cluster names outside.
    checkers.dns.timeout defaults to 2s (was 5s).
  • A departed peer's series go away ten minutes after it leaves the agent's
    peer list.
  • The controller refuses what cannot run. POST /api/v1/diagnostics
    answers 400 for an external destination with a type other than tcp,
    icmp or mtr or for a plane other than pod, 501 when the source
    agent does not run the type, and 503 leadership lost when the lease goes
    mid-task. kubectl kconmon check exits 1 on these, and kubectl kconmon
    finds the leader itself.
  • So does the console. Runs, definitions and schedules toward a target or
    an ad-hoc address take only tcp, icmp and mtr (a continuous schedule
    also dns and http), and an edit that would leave a stored schedule or
    definition unable to run answers 422; enabled: false always saves. The
    import follows the same rules, and the forms offer only what runs.
  • Console API input. Malformed input answers 400 or 422 instead of 502,
    and alert rule names that become one Prometheus alert name are refused.
  • The console's agent-missing template renders the chart's
    KconmonAgentsMissing expression; existing rules change on the next sync.
  • ZoneLossHigh description. Below about ten node pairs between two zones
    one broken link crosses the default 10% alone; raise
    prometheusRule.zoneLossHigh.threshold and zoneChecksFailing.threshold.
  • Continuous external checks fit the controller's 8 MiB limit: the console
    leaves out whole definitions, newest first, and counts them in
    kconmon_ng_console_external_specs_skipped_total{reason="over-budget"}.
  • Console UI. A silent MTR hop shows a dash, not 100% loss; refused forms
    focus the refused field; a foreign rule import asks to confirm; phones get a
    drawer Close button and tables that scroll inside their card.
  • Chart. The console gets GOMEMLIMIT from its memory limit. The GeoLite2
    sidecar moves to ghcr.io/maxmind/geoipupdate:v8.0.0. The schema refuses
    what the binaries refuse at startup. The install notes flag an Ingress with
    no trusted proxies and an external gateway neither on
    externalTrafficPolicy: Local nor behind loadBalancerSourceRanges.
  • deb/rpm. The packaged agent config ships with the DNS checker off, and
    the postinstall keeps a net.ipv4.ping_group_range the admin already set.
  • Release images. :latest, the chart and the GitHub release follow a tag
    only after e2e passes on its images, and only the newest stable release
    moves :latest, the Latest badge and the krew index.
  • Build stack. Go 1.27.1, distroless static-debian13, Vite 8. The unused
    OpenTelemetry SDK is gone; observability.otel.* logs a warning.

Fixed

  • Agents.
    • Hot reload stopped for good after the first atomic replacement of the
      file (an editor's save, a puppet file resource, a ConfigMap swap).
    • logLevel: DEBUG and logFormat: TEXT ran at info in JSON.
    • An agent probed at a secondary IP, a multi-homed advertiseAddress or a
      VIP read...
Read more

v2.5.0-rc.3

v2.5.0-rc.3 Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 26 Sep 18:21
Immutable release. Only release title and notes can be modified.
v2.5.0-rc.3
8258c62

kconmon-ng v2.5.0

2.5.0 adds a path MTU probe for the failure small probes cannot see: a pair
where handshakes and pings cross while full-size packets vanish now turns red
with the size that still crosses, instead of staying green on every plane.
The rest of the release pays the debts a first outside user runs into:
node-level alerts, maintenance windows that hold the console's webhooks,
local users managed from the console, a config reload that survives the way
files are really replaced and applies what it reads, a console whose first
load is a fifth of what it was, and a set of security fixes.

Read the Upgrade notes before rolling out: the probe is on by default, two of
the four new rules are critical, and the chart's NetworkPolicy is split per
component.

Added

  • Path MTU probe (config.checkers.pmtu, on by default). Once a minute
    every agent sends each peer's UDP echo port a 64-byte datagram and a
    full-size one with Don't Fragment set. The full size is the MTU of the route
    to the peer (the route's own mtu when the CNI sets one, as Cilium does,
    else the egress device's); size overrides it. When the full size does not
    come back, the agent bisects (at most 16 sizes, timeout 500ms each) and
    confirms both ends, so a lossy path does not pass for a black hole. A pair
    reads ok (the full size is echoed), reduced (the path answers ICMP
    frag-needed: TCP adapts, UDP without its own path MTU discovery does not) or
    blackhole (full-size datagrams vanish while the small one crosses, and
    large transfers stall). A lost small datagram is a connectivity failure,
    left to the UDP plane, and a pmtu failure does not trigger MTR. The agent
    warns when interval is under 28 timeouts (14s) or 3m and more. New series:
    kconmon_ng_pmtu_bytes and kconmon_ng_pmtu_probe_bytes (per pair: the
    size that crossed, the size probed), kconmon_ng_pmtu_results_total
    (success for ok and reduced, fail for a black hole),
    kconmon_ng_zone_pmtu_results_total and kconmon_ng_agent_pmtu_probe_bytes.
    Agents advertise plane:pmtu. Walkthrough in
    Catch an MTU black hole.
  • PathMTUBlackHole and ZonePathMTUBlackHole
    (prometheusRule.pathMtuBlackHole, warning): more than half of a pair's
    probes failed over 10 minutes, or more than sustainedThreshold (0.1) over
    30 minutes with at least two failures and one in the last 10, held for 5.
    The second arm catches a black hole on one of several ECMP paths. The zone
    rule fires only where no per-pair pmtu series exist
    (agent.metrics.detail=zone-only). A network that carries less than its
    routes say on purpose and clamps TCP MSS can set config.checkers.pmtu.size
    or turn the rule off.
  • NodeUnreachable and NodeIsolated (prometheusRule.nodeUnreachable,
    .nodeIsolated, critical, for: 5m): most peers fail most of their TCP
    probes to a node, or a node fails to reach most of its peers, with at least
    minPeers (2) reporting. They catch a node that stays registered behind a
    host firewall, a NetworkPolicy or a broken CNI datapath; the
    inhibit rules
    on the metrics page fold the per-pair alerts under them.
  • Path MTU on the dashboards. Overview opens with the key indicators in
    two rows (agents, leader, pairs, pairs with failures, black-hole and
    reduced-path pairs), then the worst-pair and MTR bars, with the charts and
    tables below, pairs below their probe size among them. Node Detail: the path
    MTU to and from each peer and black-hole probes by peer. Panels show the
    smallest size of the last 10 minutes, so an ECMP-split black hole stays on
    screen.
  • PMTU in the console and the CLI. The matrix gains a PMTU protocol: the
    path MTU in bytes, green at full size, amber on a reduced path ("1400 of
    1500") or while recovering, red while recent probes fail, dashed "Not run"
    for a 2.4.x agent. The Overview, node and pair pages follow it. Run checks
    accepts pmtu between nodes, and
    kubectl kconmon check <source> <destination> --type pmtu exits 2 on a
    black hole.
  • Local users in the console. With auth.mode=local, Settings > Users
    adds, re-roles, resets, disables and deletes accounts under the new
    users:manage permission (built-in admin only), which the last enabled
    holder cannot lose. Every local user can change their own password. A
    password change or reset ends the user's other sessions; a disable or delete
    also revokes their API tokens. Routes under /api/v1/users in the
    Console API.
  • NetworkPolicy keys (Upgrade notes 4 to 6):
    • networkPolicy.dnsEgress replaces the default DNS rule, for NodeLocal
      DNSCache and other host-network resolvers;
    • networkPolicy.ciliumKubeAPIEgress (auto) adds CiliumNetworkPolicies
      for apiserver and node traffic, which no ipBlock matches on Cilium;
    • networkPolicy.clusterCIDRs carves pod and Service CIDRs out of the
      default 0.0.0.0/0 egress, which on Calico and Antrea matches pods;
    • console.networkPolicy.prometheusTargetPort opens the pod port behind
      a Prometheus Service that maps it (Thanos 9090 to 10902).
  • console.clientAddress.trustedProxyCIDRs: the proxies whose
    X-Forwarded-For names the client for rate limits, the WebSocket cap and
    the audit log, never for identity (Upgrade note 20).
  • agent.tls.enabled (false): TLS verified against the system trust
    pool with no other TLS field set, for an external agent whose gateway has a
    publicly signed certificate.

Changed

  • Maintenance windows hold the console's alert webhooks. An alert that
    starts firing inside a matching window (fleet-wide, a node, the pair or its
    target) is delivered only if it still fires when the window closes, and not
    at all if it resolves inside. A restarted console keeps holding.
    kconmon_ng_console_webhook_suppressed_total{event} counts what was held.
  • Config hot reload applies what it reads. The keys that go live and the
    ones that wait for a restart are in Upgrade note 11 and
    What reloads and what does not.
  • The node page covers both directions, so a node NodeUnreachable names
    no longer reads Healthy there, and its peer breakdown switches between To
    peers and From peers. Diagnostic runs list failed pairs first.
  • Console pages load on demand. The first page load drops from 935 kB to
    198 kB gzipped, and charts load only the ECharts parts they draw with.
  • DNS probes against an explicit resolver ask the absolute name, so the
    search list no longer multiplies the query or sends cluster names outside.
    checkers.dns.timeout defaults to 2s (was 5s).
  • A departed peer's series go away ten minutes after it leaves the agent's
    peer list.
  • The controller refuses what cannot run. POST /api/v1/diagnostics
    answers 400 for an external destination with a type other than tcp,
    icmp or mtr or for a plane other than pod, 501 when the source
    agent does not run the type, and 503 leadership lost when the lease goes
    mid-task. kubectl kconmon check exits 1 on these, and kubectl kconmon
    finds the leader itself.
  • So does the console. Runs, definitions and schedules toward a target or
    an ad-hoc address take only tcp, icmp and mtr (a continuous schedule
    also dns and http), and an edit that would leave a stored schedule or
    definition unable to run answers 422; enabled: false always saves. The
    import follows the same rules, and the forms offer only what runs.
  • Console API input. Malformed input answers 400 or 422 instead of 502,
    and alert rule names that become one Prometheus alert name are refused.
  • The console's agent-missing template renders the chart's
    KconmonAgentsMissing expression; existing rules change on the next sync.
  • ZoneLossHigh description. Below about ten node pairs between two zones
    one broken link crosses the default 10% alone; raise
    prometheusRule.zoneLossHigh.threshold and zoneChecksFailing.threshold.
  • Continuous external checks fit the controller's 8 MiB limit: the console
    leaves out whole definitions, newest first, and counts them in
    kconmon_ng_console_external_specs_skipped_total{reason="over-budget"}.
  • Console UI. A silent MTR hop shows a dash, not 100% loss; refused forms
    focus the refused field; a foreign rule import asks to confirm; phones get a
    drawer Close button and tables that scroll inside their card.
  • Chart. The console gets GOMEMLIMIT from its memory limit. The GeoLite2
    sidecar moves to ghcr.io/maxmind/geoipupdate:v8.0.0. The schema refuses
    what the binaries refuse at startup. The install notes flag an Ingress with
    no trusted proxies and an external gateway neither on
    externalTrafficPolicy: Local nor behind loadBalancerSourceRanges.
  • deb/rpm. The packaged agent config ships with the DNS checker off, and
    the postinstall keeps a net.ipv4.ping_group_range the admin already set.
  • Release images. :latest, the chart and the GitHub release follow a tag
    only after e2e passes on its images, and only the newest stable release
    moves :latest, the Latest badge and the krew index.
  • Build stack. Go 1.27.1, distroless static-debian13, Vite 8. The unused
    OpenTelemetry SDK is gone; observability.otel.* logs a warning.

Fixed

  • Agents.
    • Hot reload stopped for good after the first atomic replacement of the
      file (an editor's save, a puppet file resource, a ConfigMap swap).
    • logLevel: DEBUG and logFormat: TEXT ran at info in JSON.
    • An agent probed at a secondary IP, a multi-homed advertiseAddress or a
      VIP read...
Read more

v2.4.0

Choose a tag to compare

@github-actions github-actions released this 15 Sep 19:12
Immutable release. Only release title and notes can be modified.
v2.4.0
dbd9930

kconmon-ng v2.4.0

External agents stop being second-class. A host outside the cluster now
tells its peers where it listens, gets scraped without a hand-written
target, and shows up as what it is in the console, the CLI and the Time
Machine. One rule comes with it, and it is the one to read before rolling
out: per-agent ports are honoured only by upgraded agents; keep one port
set until every agent, deb/rpm hosts included, runs 2.4.0.
An older agent
reports no ports and dials every peer on its own configured values, so in a
fleet whose ports differ each old agent goes one-way red toward every peer
listening elsewhere, its on-demand diagnostics included. The full skew
matrix is under Upgrade notes at the end of this section.

Added

  • Per-agent ports on the wire. AgentMeta gains http_port,
    udp_port and metrics_port. Every agent reports its three listener
    ports at registration, and peers probe it on the ones it reported, for the
    scheduled mesh and for on-demand tasks alike, so an external host no longer
    has to mirror the cluster's port pair. Zero means "not reported" (an agent
    older than 2.4.0): the prober then dials its own configured port, each port
    falling back on its own. The controller refuses a port above 65535, peer
    lists carry the two probe ports but not metrics_port (nobody dials it),
    and an agent never adopts ports from the controller's reply; zone stays the
    only thing it takes from there. See
    Ports.
  • Prometheus HTTP SD for external agents. A bare host has no Service for
    a ServiceMonitor to select, so the controller, the one party that knows
    the host registered and on which address, now publishes it:
    GET /api/v1/prometheus/sd, served on httpPort and on metricsPort (the
    port the chart's scrape NetworkPolicy already opens). The contract:
    • one target group per external agent, <advertised address>:<metricsPort>,
      sorted by node name and deduplicated by address;
    • a fixed label set, node, zone, external="true" and agent_id. An
      agent's own labels never reach Prometheus, so a host cannot inject
      target labels;
    • with no external agent registered the body is the literal [], and
      the list is always served with Cache-Control: no-store;
    • a standby answers 503 not the leader, never 200 []: Prometheus reads
      every 200 as the complete target set, so an empty one from a standby
      would wipe every external target, while on a non-200 it keeps the list
      it has;
    • an agent that reported no metrics port (older than 2.4.0) is published
      on the controller's own config.metricsPort, and the controller logs
      metrics port assumed from controller config once per agent, which is
      the clue when such a host on another port sits at up == 0;
    • controller.prometheusSD.enabled: false closes the route (404 on both
      listeners). The key reaches the shared ConfigMap only when false, so an
      older controller image never trips over it;
    • with controller.replicaCount > 1 the controller Service spreads
      refreshes over all replicas and roughly half of them land on a standby:
      prometheus_sd_http_failures_total climbs for the job while the targets
      stay correct. Cosmetic, and written down so nobody chases it. Body and
      semantics in the
      HTTP API reference.
  • scrapeConfig.externalAgents in the chart. Renders a Prometheus
    Operator ScrapeConfig (needs the scrapeconfigs.monitoring.coreos.com
    CRD) named <release>-agent-external that reads the SD route, with
    labels for your Prometheus' selector (kube-prometheus-stack wants
    release: <its release name>), jobName, refreshInterval (30s) and
    interval. It applies the same agent.metrics.detail valve as the agent
    ServiceMonitor, so an external host never returns per-pair detail the valve
    drops for the pods, and the valve no longer insists on
    serviceMonitor.enabled when this is on. The chart refuses the
    ScrapeConfig without controller.externalGateway.enabled (nothing external
    could register) or with controller.prometheusSD.enabled=false (every
    refresh would 404), and the install notes remind you when the gateway is on
    without it, or when labels is empty. Plain-Prometheus http_sd_configs
    job and the reachability rules in
    Scraping external agents.
  • KconmonExternalAgentDown and kconmon_ng_controller_external_agents.
    An optional warning (prometheusRule.externalAgentDown, off by default,
    for: 5m) on up{job=~".*agent-external.*"} == 0: a host the controller
    lists that Prometheus cannot scrape, usually the host firewall admitting the
    Prometheus pod IP when the CNI NATs its egress to a node IP. The new
    controller gauge counts registered agents that came through the gateway.
  • agent.hostNetwork, for pod networks external hosts cannot route.
    The DaemonSet moves into each node's network namespace: the agents
    advertise the node IP (KCONMON_NG_POD_IP from status.hostIP), declare
    hostPort on all three ports, get dnsPolicy: ClusterFirstWithHostNet
    unless agent.dnsPolicy says otherwise, and label themselves
    kconmon-ng.io/host-network=true from a 2.4.0 image. The chart stops
    rendering the ping_group_range pod sysctl there, since the kubelet refuses
    net.* sysctls in the host namespace. It changes what is measured, for
    the whole DaemonSet
    : every in-cluster pair then probes node IP to node IP
    over the underlay, and the CNI datapath (overlay, conntrack, NetworkPolicy
    enforcement) is no longer on the probe path, so the breakage this tool
    exists to catch can hide behind a green matrix. Turn it on only when the
    goal is visibility between external agents and a cluster whose pod network
    they cannot reach. Before you do: PSS privileged for the namespace, TCP
    8080, UDP 9090 and TCP 9091 (or your config.*Port values) free on every
    node, ping_group_range set by the node OS, and one agent per machine (a
    host-network pod and a bare-host agent cannot share an IP). See
    When the pod network does not route
    and Host networking.
  • networkPolicy.nodeCidrs and networkPolicy.externalPeerCidrs.
    Host-network agents register from node IPs that no pod selector matches, so
    with agent.hostNetwork and networkPolicy.enabled both on, the chart
    refuses to render the policy until nodeCidrs lists the node CIDRs; without
    it every registration but the one from the controller's own node would drop
    silently. externalPeerCidrs closes the old gap where an external agent
    registered fine and every cell between it and the cluster stayed red: its
    CIDRs join the agent-to-agent rules in both directions (UDP grpcPort,
    TCP httpPort, the ports-less ICMP/MTR rule) and never the gateway rule.
  • External agents in the console. Everything keys off the
    kconmon-ng.io/external registration label, which the console now passes
    through from the controller's topology together with the agent's
    capabilities (labels and capabilities on TopologyAgent in the Console
    API).
    • Topology draws the host beside the cluster nodes in the lane of its
      zone, with a neutral external badge (identity, never a health tier)
      and "readiness unknown" for screen readers.
    • Node page swaps Pod IP for Advertised address, explains the Ready
      dash, and lists the probe Planes the agent advertised. Agents now
      advertise plane:tcp, plane:udp, plane:icmp, plane:dns,
      plane:http and plane:mtr; an agent advertising none (older than
      2.4.0) reads as "unknown", never as running nothing.
    • Overview badges the host in Worst pairs and adds "+N external agents"
      beside Nodes ready without counting them in, since that tile is
      Kubernetes readiness.
    • Matrix tells two silences apart from plain no-data. An external agent
      Prometheus is not scraping keeps the no-data fill and aria text, but its
      cells' tooltip and a note above the grid say why and link the scraping
      docs, until the first measured cell appears. A protocol the source does
      not run renders dashed like not probed, with its own legend row.
      Precedence when a cell has no data: excluded by the plan, then
      unsupported, then unscraped. See
      Silence with a known cause
      and External agents on the map.
  • External agents in the Time Machine. TopologyChanged events carry the
    agent's labels, so a replay badges a host the way the live view does, and a
    reconstructed topology lists a bare host under agents only, never as a
    presence-derived READY node. History recorded before the upgrade shows no
    external badges
    : a 2.3.x controller wrote no labels, and such a host stays
    an ordinary node in those instants. Historical responses never carry
    capabilities, since no event records them.

Changed

  • KconmonAgentsMissing is no longer masked by external agents.
    Registered agents include them and expected agents (schedulable nodes) never
    did, so one external host hid one missing in-cluster agent. The expression
    now subtracts controller_external_agents, with an or registered * 0
    stand-in so the rule keeps working against a controller image that predates
    the gauge.
  • kubectl kconmon tables. agents gains an EXTERNAL colu...
Read more

v2.3.1

Choose a tag to compare

@github-actions github-actions released this 31 Aug 17:52
Immutable release. Only release title and notes can be modified.
v2.3.1
4bdfb9b

kconmon-ng v2.3.1

Fixed

  • ZoneChecksFailing and ZoneLossHigh failed every evaluation with "vector
    cannot contain metrics with the same labelset" and raised
    PrometheusRuleFailures on the cluster: rate() over a __name__ regex
    union drops the metric name and collapses the per-protocol families into
    duplicate labelsets. The expressions now build the union with
    label_replace(...) or label_replace(...), which keeps the branches
    distinct and still tolerates a disabled checker's absent family.
  • CI now evaluation-tests every alert rule with promtool test rules against
    synthetic series for all metric families, including a positive check that
    each zone alert fires on staged bad data. Rendering and syntax checks never
    execute the query engine, which is exactly where this defect lived.

v2.3.0

Choose a tag to compare

@github-actions github-actions released this 31 Aug 16:13
Immutable release. Only release title and notes can be modified.
v2.3.0
a42b296

kconmon-ng v2.3.0

The sparse mesh changes WHAT "no data for a pair" means: under
topology.mode: sparse most directed pairs are deliberately never probed.
Everything in this release that reads per-pair series learns to tell "not
planned" from "went dark" through one new metric,
kconmon_ng_probe_intended — and that metric comes from the AGENT: images
below appVersion 2.3.0 do not export it. The chart's rules degrade
honestly on an older fleet (see PairWentSilent below), but do not flip
topology.mode: sparse until controller AND agents run a 2.3.0 image —
the controller config key is emitted only when sparse precisely because an
older controller image rejects it and crashloops. The appVersion pin is
aligned when the app release ships.

Added

  • topology.* — the sparse probe mesh, by values. topology.mode: sparse trims the full N×(N−1) probe matrix to a ring over sorted node
    names (sparse.ringDegree successors each, the connectivity guarantee)
    plus HRW-chosen cross-zone chords (sparse.zoneChords per directed zone
    pair, which keep the zone metric family fully populated), so probed pairs
    — and every per-pair series they export — scale ~linearly with node count
    instead of quadratically. sparse.autoThreshold is the floor: fleets
    smaller than it get the full mesh regardless of mode, because sparse only
    pays for itself at scale. Default is mode: full, byte-identical
    rendering to 2.2.0.
  • kconmon_ng_probe_intended — the plan, scrapable. A gauge, value 1
    for every directed pair the topology plan assigns
    ({source_node, destination_node}, exported by the source agent), preset
    from the peer list at registration and pruned on every plan change —
    stale pairs are deleted, not left at 1. In full-mesh mode it simply marks
    every peer, so dashboards and rules can join on it without caring which
    mode the fleet runs. It is the one honest way to distinguish "this pair
    is not supposed to report" from "this pair went dark", which is why it
    ships in the same release as sparse mode and not one later.
  • investigateUrl on the two zone alerts. ZoneChecksFailing and
    ZoneLossHigh now annotate a console deep link,
    /investigate?kind=zone-pair&scope=<source>-><destination>, straight
    into the Investigate page scoped to the firing zone pair. The link is
    console-RELATIVE on purpose — the chart cannot know the console's
    external URL (ingress is optional), so notification templates prepend
    their own origin; the console normalises the typeable -> into its
    canonical pair arrow.

Changed

  • PairWentSilent joins on the plan. The rule now fires only for pairs
    present in the source agent's kconmon_ng_probe_intended series — the
    hard rule of the sparse design, shipped in the same release: without the
    join, every pair the plan trims would read as "went silent" for the hour
    its results take to age out of the lookback window. The fallback is per
    SOURCE, not global: a source_node exporting no probe_intended at all
    keeps the old two-window behaviour, so a pre-2.3.0 agent image alerts
    exactly as before, a mixed fleet mid-rollout gets each behaviour where it
    applies — and an agent that dies outright takes its probe_intended
    series with it, which lands its pairs in the same fallback and preserves
    the alert's original purpose: catching an agent that stopped running or
    stopped being scraped.

This release also carries everything prepared for the never-published 2.2.0
tag (its pipeline caught two release-tooling defects before anything went
out); those changes follow below, under their original heading kept for
upgrade notes.

Carried over from the unreleased 2.2.0

Everything in this release reads the new zone-level metric family
(kconmon_ng_zone_*), and that family comes from the AGENT, not the chart:
agents below appVersion 2.2.0 do not export it (this chart pins 2.2.0, so a
default install is fine — the warning is for fleets running an older agent
image behind a newer chart). Until the fleet runs an agent image that does, the two zone
alerts are silently inert (their expressions match no series), the Zone
Heatmap dashboard renders empty, and agent.metrics.detail: zone-only
would drop the per-pair series with nothing replacing them — Prometheus
goes dark on the mesh while the console keeps working. Upgrade the agent
image first, flip the valve second. The appVersion pin is aligned when the
app release ships.

Added

  • ZoneChecksFailing and ZoneLossHigh. Two alerts on the zone plane,
    with the same per-rule knobs as the rest
    (prometheusRule.{zoneChecksFailing,zoneLossHigh}.{enabled,threshold,for,severity}).
    ZoneChecksFailing is the failure ratio of all TCP, UDP and ICMP probes
    between a zone pair, in one expression — the __name__ union keeps a
    disabled checker from blanking the ratio. ZoneLossHigh computes loss as
    (sent − received) / sent from the zone packet counters; averaging the
    per-pair loss-ratio gauges into a zone would weight an idle pair the same
    as a busy one, so the chart never does. Its default threshold is 0.1,
    lower than the per-pair UDPLossHigh at 0.5, because the zone aggregate
    dilutes any single link by the pair count: sustained loss at that level
    means the fabric, not one node. Both survive every agent.metrics.detail
    mode — that is the point of alerting on the zone family.
  • agent.metrics.detail — the cardinality valve. A scrape-time knob
    rendered as metricRelabelings on the agent ServiceMonitor:
    full (default, everything, ~70 series per directed pair),
    counters-only (drops the four per-pair histograms, ~10/pair — every pair
    alert keeps firing), zone-only (drops every series naming a
    destination_node, ~0/pair; the zone family at ~74×Z² series and the
    linear DNS/HTTP/external families remain). At 100 nodes that is ~0.7M →
    ~0.1M → practically N-independent, by configuration alone. Setting it
    without serviceMonitor.enabled is refused at render time rather than
    silently dropping nothing; plain-Prometheus equivalents are in
    docs/metrics.md.
  • controller.externalGateway — the external agent gateway, exposed by the
    chart.
    The controller's second gRPC listener (same services, but TLS with
    a bootstrap token, for agents OUTSIDE the cluster) gets a values block and
    three templates. templates/controller/service-external.yaml is a
    NodePort/LoadBalancer Service carrying the gateway port ALONE — the
    plaintext in-cluster gRPC port authenticates by network position and never
    appears on it, because a LoadBalancer in front of it would hand the whole
    mesh to anything that can reach the address. The deployment mounts two
    referenced Secrets read-only: tls.secretName (a kubernetes.io/tls
    serving pair; tls.clientCaKey names the CA bundle key in the same Secret
    and switches on client-cert identity pinning — empty is token-only mode,
    where any token holder can impersonate any agent, and NOTES.txt says so at
    install) and bootstrapToken.{secretName,key}. With
    networkPolicy.enabled, ingress on the gateway port is opened from
    networkPolicy.externalAgentCidrs toward the controller pods alone, and an
    empty list is refused at render rather than shipping a gateway no packet
    can reach; missing Secret names and a port colliding with
    config.{httpPort,grpcPort,metricsPort} are refused the same way. Two
    operational notes. Rotation: the gateway reads the certificate and token
    ONCE at startup and the chart cannot checksum content it only references,
    so rotating either Secret in place needs
    kubectl rollout restart deploy/<release>-controller. Version skew: the
    externalGateway config key is emitted only when enabled, because a
    controller image at appVersion 2.0.3 rejects the unknown key and
    crashloops — upgrade the image before flipping the switch, same rule as
    the zone family above.

Changed

  • The Zone Heatmap dashboard reads the zone family. Every panel that
    aggregated per-pair series into zones at query time now reads the
    pre-aggregated kconmon_ng_zone_* metrics, so the dashboard keeps working
    in every agent.metrics.detail mode and its queries stop scaling with the
    pair count. Loss panels are packet-weighted from the sent/received counters
    instead of averaging the per-pair ratio gauges. The one exception is the
    "MTR traces triggered" panel: MTR has no zone-level family, its counter is
    per-pair, and in zone-only mode that panel reads zero — its description
    now says so.

Performance and self-observability

  • Peer probing fans out with a bounded pool (32 in flight per round): a dead
    peer costs one timeout, not one timeout per peer in sequence, so probe
    cadence holds through partitions.
  • Reactive MTR traces are bounded by a global semaphore (4 in flight) on top
    of the existing per-pair cooldown; a mass partition trickles traces out
    instead of forking one per broken pair.
  • The agent exports self-metrics under kconmon_ng_agent_*: probe cycle
    duration and overruns per checker, controller reconnects, peer-list age,
    reactive-MTR in-flight and coalesced counters. All are fleet-size
    independent.
  • The controller coalesces peer-list broadcasts (trailing edge, 200 ms): a
    rollout's burst of registrations produces one broadcast, not one per
    change; the peer message is built once per broadcast and carries only the
    fields agents read.

v2.0.3

Choose a tag to compare

@github-actions github-actions released this 19 Aug 07:43
Immutable release. Only release title and notes can be modified.
v2.0.3
83f402b

kconmon-ng v2.0.3

Fixed

  • A config change restarts the pods that read it. The agent and the
    controller share one ConfigMap and read it once at startup, and a mounted
    ConfigMap changes under a running process without telling it — so
    controller.events.enabled: true applied to a live release updated the object
    and left the controller on the old file. It went on advertising no
    capabilities, the Console's realtime ingester retried against a stream that was
    configured but never started, and nothing anywhere reported an error: the Live
    page was simply empty. Both workloads now carry checksum/config, so a values
    change rolls them; a change the ConfigMap does not carry still does not.

v2.0.2

Choose a tag to compare

@github-actions github-actions released this 18 Aug 09:55
Immutable release. Only release title and notes can be modified.
v2.0.2
3441d98

kconmon-ng v2.0.2

Fixed

  • The matrix no longer opens at half size. The grid measured the height its
    own content had produced and fed that back into the fit, so a fresh render
    saw the container's 256px minimum, decided the grid did not fit, shrank to
    50%, and the smaller grid then held the box at 256px — a loop with no way out.
    It measures the space available instead, and a seven-node fleet opens at 100%.
  • Zooming in gives the node names back. The shared prefix every node name
    begins with is dropped to buy column width, which is right while the column is
    narrower than the names and wrong the moment it is not: at 125% a label column
    holds adm-kuber-01 with room over and still read …01. The elision is now
    decided per axis at the current scale, and the note above the grid appears only
    while an axis is actually eliding.

Added

  • A favicon. The console had none, so every tab showed the browser's blank
    square; it now wears the mark it wears in its own sidebar.

v2.0.1

Choose a tag to compare

@github-actions github-actions released this 18 Aug 08:37
Immutable release. Only release title and notes can be modified.
v2.0.1
9a1df03

kconmon-ng v2.0.1

Fixes a hole in 2.0.0: auth.mode=oidc and auth.mode=header shipped with no
way to grant anybody a role. Both modes worked, and neither was usable.

Fixed

  • An OIDC or header install can grant roles at deploy time. Role bindings
    live in the database and are created through an API that already requires
    rbac:manage, so a fresh install had nobody able to make the first binding;
    the only alternative was auth.defaultRole, which is one role for every
    authenticated subject. The way out that 2.0.0 left was to bring the console up
    in local mode, log in, create a binding by hand and only then switch — a
    workaround, published as if it were a procedure.

    console.auth.groupRoles maps a group the identity provider asserts onto a
    role this console grants, in the values file:

    console:
      auth:
        groupRoles:
          platform-oncall: admin
          everyone: viewer

    Roles resolve as the union of that map and any binding made through the API, so
    a grant by hand still adds to what the provider's groups carry. A group absent
    from the map grants nothing. What the map grants cannot be revoked through the
    API — that is what makes it declarative.

  • A role store outage no longer costs an operator their access. The store's
    half still fails closed, because an unreadable database is no evidence a
    subject holds anything; a grant that came from the claim and the config was
    never in doubt, and an outage is when the console is most needed.

v2.0.0

Choose a tag to compare

@github-actions github-actions released this 17 Aug 20:27
Immutable release. Only release title and notes can be modified.
v2.0.0
4c81710

kconmon-ng v2.0.0

A chart that installs monitoring and nothing else, a console that survives more
than one replica, and one that tells the truth about time. The chart no longer
ships a database or a cache — point it at the ones you already run. The Time
Machine moved out of the top bar and into each page's own time controls, and
the charts pin their axis to the window you asked for rather than to the data
that happened to arrive. MTR gained a Runner, path history that reads as a
timeline, and external targets.

Breaking

  • The chart no longer installs PostgreSQL or Valkey. database.mode,
    database.cnpg.* and the bundled subcharts are gone: set
    database.existingSecret to a Secret holding a postgres:// DSN and
    redis.existingSecret to one holding a redis:// DSN, and any managed
    instance works — RDS, a StatefulSet, a CloudNativePG cluster you run yourself.
    Every removed key fails the render with a message naming its replacement
    (templates/_migrations.tpl), so no old value is silently honoured.
  • console.database.* moved to the top-level database.*, and
    console.* keys that described the bundled datastores went with it.

Added

  • Chart 2.0.0, templates split per component — agent, controller, console,
    shared and observability each own their directory, with a NetworkPolicy set
    covering every component, fail-closed on external egress. Render-time guards
    refuse a port collision, an OIDC redirectURL the console would not start on,
    and more than one console replica without a shared cache.
  • The console scales past one replica — sessions, the fixed-window rate-limit
    counters and the realtime fan-out live in the Redis-compatible server, and the
    controller elects a leader so exactly one replica drives the reconcilers.
  • MTR Runner and path history — start a trace from the Explorer itself,
    with a settable cadence and duration; every distinct route the fleet has taken
    is kept, diffed and drawn on a timeline of when it changed.
  • External targets — probe a destination that is not a fleet peer, gated by
    config.checkers.external.allowedCidrs and the cluster's own egress policy.
    The console refuses at create time a target no agent could ever reach.
  • Time in the Console's result table — every figure says when it was read.

Changed

  • The Time Machine lives with the page's time filters, not in a strip across
    the top of every route. It is offered only on the pages that resolve their
    reads through ?at=, and the engaged banner stays global because writes are
    disabled console-wide.
  • Explore's axis is the window you picked — a 24h view draws 24 hours even
    when Prometheus holds less, instead of quietly redrawing three.
  • MTR Explorer is sorted by name, both destinations and their sources, with
    numbers read as numbers (m9 before m10).
  • OIDC identity is the sub claim, namespaced as oidc:<sub> — the only
    claim OIDC Core §5.7 allows as an identifier. auth.oidc.usernameClaim now
    decides the display name alone, so renaming a person no longer moves their
    roles (Grafana's CVE-2023-3128 is what the old shape risked). Group membership
    is re-read on every token refresh. Bindings made against a username stop
    granting
    ; the console names them at boot so they can be remapped.
  • The configuration bundle carries access control — custom roles and the
    grant list, but only for a caller who holds rbac:manage. Roles import;
    bindings never do, because a grant names a person in the source console's own
    identity namespace.
  • Only the chart under the cursor shows a tooltip. Its neighbours keep the
    shared crosshair and mark their own samples with a dot, instead of each
    covering its own curves with a box of numbers.

Fixed

  • WebSocket topics are authorized per topic: events:read no longer carries the
    topology and matrix snapshots that topology:read and matrix:read gate, and
    a permission taken away reaches a socket that is already open — the topics it
    may no longer have are dropped, the rest of the connection is left alone.
  • The audit row describes the mutation that happened. A body could name one
    thing for the handler and another for the audit log by spelling a key in a
    different case, and a value carrying a NUL made the whole row unwritable — in
    both directions the caller chose whether their own privileged action was
    recorded. The extraction now matches keys the way encoding/json matches
    struct fields, and is bounded before it is decoded, so a wide body on the
    public login route can no longer take the replica past its memory limit.
  • A broken alert rule no longer freezes the whole bundle. Editing a deployed
    rule into PromQL the apiserver rejects used to stop every other rule from
    being applied, while the API answered 2xx and Prometheus kept evaluating the
    stale set. The quarantine now keys on rendered content rather than rule ids,
    offers each suspect to the cluster on its own, and removes the object only
    when every rule was offered and every one refused.
  • auth.mode=anonymous is not exempt from CSRF. Any page an operator's
    browser visited could POST into a console kept off the internet; a
    cross-origin write is now refused, while a script that sends no Origin is
    unaffected.
  • The node-local HTTP checker verifies certificates. An expired certificate,
    one issued for another hostname or an interceptor's CA all used to pass, so an
    https check could not fail on the condition it was added to notice; opt out per
    target with insecureSkipVerify.
  • External metrics separate the checks on one target — the series carry
    check_type, so an icmp and a tcp check on the same target no longer average
    each other's failures away under the ExternalChecksFailing rule.
  • A check no agent could run is refused when it is written, instead of being
    dropped by every agent with nothing but a log line while the console listed it
    as enabled.
  • The MTR destination listing is complete. It is paged behind a keyset
    cursor rather than capped, so no pair is missing from the Explorer and no
    per-destination total is short.
  • A subscriber that stops reading its peer-update stream is torn down rather than
    holding a controller goroutine and its connection slot until TCP notices.
  • Shutdown finishes in-flight runs before tearing down the pipeline they publish
    onto, so a rolling update no longer logs dropped frames that were delivered.
  • Every request body is capped, so one oversized POST can no longer take a
    console replica past its memory limit.
  • The OIDC callback binds its state to the browser that started the flow.
  • A role-store failure now refuses rather than granting the default role.
  • External TCP and UDP checks probe what was asked for instead of speaking the
    agent's own protocol to something that is not an agent.
  • A user binding can no longer be resolved by a subject of another kind: role
    resolution matches the caller's kind as well as their id.
  • Revoking a role binding is auditable — the audit row names the role and the
    subject, read before the row is destroyed rather than after.
  • Path history says when it has reached the end instead of leaving a "Load
    older" button that can never be pressed, and counts the routes it is showing
    against the traces folded into them.
  • A probe tick on a diagnostics run leads with that probe — its sequence, its
    clock, its latency or its error — so two ticks on an unchanged route are no
    longer indistinguishable.