v0.1.19 AI-assisted root-cause analysis, node-to-poller visibility, generated OpenAPI contract, editable URL/DNS monitors, auth-before-availability guards, counter-threshold and IPv6 SSRF fixes
📋 Full release notes: RELEASE_NOTES.md
Ask why, and see who polls what. Two additions lead this release. AI-assisted root-cause
analysis runs from a live alert: Yagra already assembles the incident — the cascade's root cause,
the metric anomaly, the passive events and the dominant traffic around it — and now asks a language
model for the sentence a human writes at the end. It is off until you configure a provider, and
nothing in the alert path depends on it. Alongside it, node→poller visibility: every node says
which pool it belongs to and which poller is actually polling it, and pools are assignable from the
inventory tree — including on a folder, which is the only bulk assignment the tree has. Underneath
both, the northbound API now publishes its own OpenAPI 3.1 document, generated from the handlers,
with the WebUI's types generated from that — which removed 3,340 lines of hand-transcribed
TypeScript and closed a class of contract drift for good.
Breaking changes
-
An unauthenticated request now answers 401 across the whole API, not 503. Yagra answers 503
when a subsystem an endpoint needs is not configured ("skeleton mode"), and roughly a hundred
handlers ran that availability check before the permission guard — so an anonymous caller could
tell a configured deployment from an unconfigured one without holding any credential. Guard order
is now uniform API-wide: authenticate and authorize first, check availability second. Authenticated
callers see no change; only anonymous requests to an unconfigured subsystem move from 503 to 401. -
VITE_API_BASEis now an origin, not a base path. It used to default to/api/v1and be
prepended to relative paths; it now defaults to empty and the/api/v1prefix is part of every
path. Only a WebUI build that overrides it is affected, and only to drop the/api/v1suffix:
VITE_API_BASE=https://core.example.net, not…/api/v1. This also fixes the live-update streams,
which never appended/api/v1themselves and so pointed at a different host than the API client
whenever the variable was set. -
GET /api/v1/thresholdsreturns an envelope and is capped. The response changed from a bare
StoredThreshold[]to{ "items": [...], "total": <n>, "truncated": <bool> }, and the server
returns at most 500 rules per request —?limit=can narrow that, never widen it. The WebUI
ships in the same image and was updated with it, so this affects external automation only.
Anonymous requests to this endpoint now answer 401 rather than 503, matching the rest of the
API: a caller is authenticated before the server discloses whether a subsystem is configured.
Reading the rules requires ManageConfig, not View — a threshold set describes when and whom
Yagra will page — so it stays closed on a public dashboard. -
POST /api/v1/api-tokensrejects a group scope. Sending{"scope": {"Groups": [...]}}now
answers 400unsupported_scopeinstead of minting the token; omitscopeor send"All".
Nothing enforced a group scope on either surface —/mcprefuses a group-scoped token outright
and the REST endpoints never consulted the scope at all — so what was issued was a credential
that looked least-privileged and was in fact either unusable or unrestricted, depending on where
it was pointed. The WebUI never offered the field, so this affects API clients only. Group scoping
will be accepted again when the read paths actually filter by it. -
Four report and group endpoints now answer the status code the rest of the API does. The three
report deletes —DELETE /api/v1/reports/definitions/{id},/runs/{id}and/schedules/{id}—
returned200 {"ok": true}where the other 24 deletes in the API return204 No Content; they now
return 204 with no body.POST /api/v1/node-groupswas the only creator that discarded the id
it had just generated, answering 204; it now returns 201 with{"id": "<uuid>"}like the other
twenty creators.POST /api/v1/reports/schedulesreturned its{"id": …}under 200 and now
returns 201. A client that checks for an exact status code, or readsokoff a delete
response, needs updating; the WebUI ships in the same image and ignored both. -
GET /api/v1/rca/{id}andPOST /api/v1/rcanow describe what is inside a report. The
bodyfield was published as an untyped JSON blob, so a generated client gotunknownfor the
entire AI answer and its evidence. The document now carries real schemas for the answer
(summary,root_cause,dependents,next_steps,confidence,raw) and for the incident
context it was grounded in. The bytes on the wire are unchanged — only the description was
missing. Regenerate your client to pick the types up. -
Running a Troubleshoot analysis is now an Operator action on both surfaces. It required
Admin over the REST API and merely Viewer over MCP, so the on-call operator was refused in
the WebUI while the same person could run the identical analysis through an AI client. Both now
ask for the acknowledge-alerts permission (Operator and up), which is also what cancelling a run
takes. An analysis changes no configuration — the admin requirement was standing in for a rate
limit, and real admission control has done that job since it was added. Reading past runs and
their findings is unchanged and still open to Viewers. Viewer-scoped API tokens can no longer
launch analyses over/mcp. -
Four endpoints nothing called have been removed.
GET /api/v1/rca/{id},
GET /api/v1/reports/definitions/{id},PUT /api/v1/node-groups/{id}/geoand
POST /api/v1/events/alerts/closewere reachable but called by neither the WebUI, the MCP tool
surface, nor any documented automation — each answering requests, appearing in the published
contract and carrying tests for a feature that had no way in. They now404. The data each read
is still served: an RCA report comes back from thePOST /api/v1/rcathat produced it, and a
report definition fromGET /api/v1/reports/definitions. Two are a capability loss, not a
cleanup: group map coordinates can no longer be set at all — the Sites map widget reads them
and nothing writes them any more — and an event-raised alert can no longer be closed by hand; it
clears on its rule's TTL as before. Say so if either mattered to you and it can come back with a
UI attached.
New Features
-
AI-assisted root-cause analysis, on demand. Yagra already assembles the evidence for an
incident: dependency suppression attributes a cascade to its root cause, and the incident timeline
gathers the metric anomaly, the passive events and the dominant traffic around it. What no amount
of correlation produces is the sentence a human writes at the end. Active alerts gains a button
that asks a language model for that sentence, grounded in exactly that evidence, and returns a
summary, a probable root cause, the dependents it explains and suggested next steps — with a
confidence, and the model's raw answer kept beside it.
It is off until you configure it. With no provider row there is no client, no credential and no
egress. Nothing in the alert path calls into it, so hysteresis, suppression, dedup and notification
behave identically whether the provider answers, times out, or was never set up. Choose one
provider — Vertex AI, Gemini or Claude — whose credential is sealed with the same
envelope cipher as every other stored secret and is write-only once saved. Generating a report
takes Operator (the people carrying the pager are the ones who need the explanation) and is
bounded by a concurrency limit, a rate window and a context cache rather than by a narrower role;
reading one back takes View; configuring the provider takes Admin. Device output quoted into the
prompt is fenced, each provider's endpoint is a compiled-in constant rather than a settings field
so a configuration screen cannot become an exfiltration channel, and Yagra still has no way to
configure a network device. -
Every node says which poller polls it, and pools are assignable from the inventory tree.
Node→poller assignment existed but was effectively invisible: answering "which poller polls this
node?" meant runningredis-cli, and a node's pool was writable only through an API field no
screen sent. Node detail now shows Pool and Polled by — the latter a five-state answer
(assigned / pending / legacy fan-out / Meraki / unknown) read from the working set core actually
published rather than re-derived from the hash ring, so a node's answer and a poller's node list
cannot disagree. Settings ▸ Pollers drills into any poller's node set inline. Pool is now
editable on a node and on a folder, and right-clicking either in the inventory tree offers a
pool chip row — the pools that exist, plus Inherit and Custom — which is the only bulk assignment
the tree has, since it has no multi-select. A node's effective pool resolves as its own → nearest
ancestor folder →default, and every chip says whether that pool has a live poller: assigning
to a pool with none publishes its jobs to a subject nothing subscribes to, and the node silently
stops being monitored. -
A URL or DNS monitor's configuration can be edited and removed after it is created. Until now
the add-node dialog could create one and the node detail could display it, but changing a URL, a
timeout, a resolver or a record type meant deleting the node and making it again — the endpoints
existed the whole time with nothing calling them. The node's URL/DNS health card gains a ⋮ menu
with Edit and Remove monitoring; the editor covers every field including the expected
HTTP status (any 2xx, an explicit code list, or a range). Removing the configuration leaves the
node in the inventory and its recorded history intact, and simply stops probing it. Requires the
same permission as any other monitoring change (Admin). -
The API now publishes its own OpenAPI 3.1 document, at
GET /api/v1/openapi.json. It is
generated from the handlers themselves — every path, query parameter, request body, response shape
and error code — so it describes what the server actually does rather than what someone remembered
to write down. The endpoint is unauthenticated, like/api/v1/versionand/api/v1/config: it
contains no inventory, configuration or state, and is identical on every deployment. Point any
OpenAPI client generator at it. The WebUI's own types and API client are now generated from this
same document, which removes 3,340 lines of hand-transcribed TypeScript that nothing was checking.
Improvements
- Active alerts now say which device and what broke, instead of two UUIDs. A triage row read
● 550e8400-… 7c9e6679-… 2m ago. The node resolves to its name (the id is on hover, as everywhere
else), and the check's id — a one-way hash of node and metric, so it has no name to resolve to —
is replaced by what the check actually measured:icmp_rtt_ms above 100 (was 450). That detail
was already stored with every alert and shown on Alerts ▸ History; the triage screen simply never
displayed it. Both screens now render it through the same formatter, so they cannot disagree. - Settings ▸ System Health lists the flow store, and shows the server's own verdict. The
ClickHouse row was missing while the page's aggregate health counted it, so a flow-store outage
read as "everything reachable". The card now lists all five backing stores plus the bus, and
carries an "All reachable" / "Degraded" badge that comes from the server rather than being
re-derived from the rows — so the next dependency the page forgets disagrees visibly. - Poll-loop health reports working-set distribution. The widget adds pools served as a working
set versus pools falling back to per-job publish (the latter turns amber when non-zero — it means
a pool has no live registered poller), working-set snapshots versus deltas, and assignment-mirror
writes. Settings ▸ Pollers also gains a Registered column showing when each poller first
checked in. - The app icon now matches the one on the Yagra website. The browser-tab favicon and the mark in
the top bar, the mobile top bar and the sign-in panel are the topology fork that also reads as a
"Y" — one root node branching to two — replacing the older double-ring mark. The orange seal, the
colors and the sizes are unchanged; a hard refresh may be needed for the tab icon. - Event search filters and pages on when an event happened, not when Yagra ingested it. The
PostgreSQL path used the ingest timestamp while the VictoriaLogs path already used the event
timestamp, so the same search returned different rows and a different order depending on which log
store was enabled. Both now agree on the event timestamp. One consequence is inherent to
time-ordered logs: when a remote poller reconnects and replays buffered results
(store-and-forward), older events can land behind a page you have already scrolled past — which
is how the VictoriaLogs path has always behaved. - The Flapping watchlist names its nodes. The dashboard widget showed raw node and check UUIDs;
it now resolves the node's name (UUID on hover) and shows what the flapping check measures, the
same way the Active alerts list does.
Bug Fixes
- A report run in a state the WebUI didn't recognise was shown as "Failed". The status badge
ended in a catch-all that painted anything unfamiliar critical-red, so a run that had actually
succeeded could read as broken — most visibly during a rolling upgrade, where an older WebUI sees
rows written by a newer core. Run state, trigger and schedule cadence are now closed sets in the
API contract (queued/running/succeeded/failed/unknownand so on), the badge is a
per-state map with no catch-all, and a genuinely unrecognised state renders neutrally as "Unknown
state" rather than as a failure. No wire change — the same strings, now described. - A DNS monitor's failure reason was shown as a raw internal token. Node ▸ DNS rendered
nx_domain,serv_fail,depth_exceededand six others verbatim, untranslated — so a Japanese
operator got English snake_case in the resolution column. All nine now read as sentences in both
languages ("No such name (NXDOMAIN)", "名前が存在しない(NXDOMAIN)"). - An alerting rule scoped to one event stream could silently widen to all of them. An event rule
naming a source kind this build did not recognise parsed to "no kind filter", which the matcher
reads as any kind — so the rule fired on syslog, traps and webhooks alike, rather than the one
stream it was written for. Such a rule is now left out of the engine and logged until the core
understands it. - A poller went on polling nodes that had moved to another pool. The scheduler built its pool map
from node rows alone, so a pool whose last node moved away vanished from the map and was never
reconciled again — its poller kept polling the stale working set for the life of the core process,
double-polling every node that had left. Recovery took a poller restart. Every live pool is now
seeded into the map before the node pass. Until this release it took a hand-written database edit
to reach; making pool editable from the WebUI turns it into the first thing an operator does. - A filtered flow destination dropped the template datagrams its collector needed. A forwarding
destination with a filter tests the decoded flow records, and a NetFlow v9 datagram carrying only
template definitions has none — so an exporter that refreshes templates in their own record-free
datagrams left a filtered collector holding data sets it could never decode, silently and
permanently. A datagram with template definitions and no flow records now bypasses the filter. The
rule is deliberately "templates and no records", not "no records": a data set whose template is
unknown also decodes to zero records but teaches a collector nothing, so there the filter still
decides. Exporters that inline templates in every export were never affected. Found by on-metal
validation against a real collector. - A node could hold both a URL and a DNS monitor, and the DNS one would never run. The "a node is
exactly one kind" guard was enforced on the DNS writer only, so attaching a URL check to a
DNS-monitored node was accepted and stored — and the scheduler, which resolves URL first, then
never ran the DNS check. Both writers now ask the same guard, and the precedence is stated once
rather than twice. get_fleet_summaryover MCP left out node states it had not seen. The REST rollup pre-seeds all
six states; the MCP tool inserted only the ones present in the fleet, so an AI client reading
states["warning"]got a missing key where the WebUI got a zero — and reported, confidently, that
there was no warning data. Both surfaces now return the same tally. Relatedly, an empty fleet's
data coverage reads as 100% rather than 0%, so the blind-spot widget no longer lights up on every
fresh install and gets tuned out.- A broken key or an unreachable database told the operator their OIDC settings were invalid.
Saving an identity provider rendered every failure as400 invalid_providerwith the internal
error text attached, so a key-encryption problem or a failed database write was reported as bad
input — and the internal message went out on the wire. A bad submission is now a 400 that says
which field, a fault is an opaque 500 with the cause in the log, and an IdP that cannot be reached
during sign-in is a 502 rather than a 500. - A regular-expression event search was case-sensitive on VictoriaLogs, case-insensitive on
PostgreSQL. The same regex matched different rows depending on which log store was enabled; both
are now case-insensitive. A plain search term still matches whole tokens on VictoriaLogs and
substrings on PostgreSQL — an inverted word index cannot serve a leading substring without scanning
every block, measured at 30s against 0.22s on the live fleet — so that one difference is deliberate,
and the search box's regex toggle is the escape hatch when you need to reach inside a token. - The poller's store-and-forward buffer can use its disk again. The container image never
created the spill directory, and/var/libis not writable by the non-root runtime user, so the
buffer fell back to memory-only after a single startup warning. A bus outage lasting longer than
the in-memory ring therefore dropped the oldest poll results instead of spilling them, and a
poller restart mid-outage lost everything it was holding. The directory is now created in the
image with the right ownership, and the test-server deployment gives it a named volume so the
spill survives container recreation. - Adding a node twice. Nothing guarded the add-monitor dialog's submit button while the request
was in flight, so a double-click created two nodes. The dialog also kept the previous attempt's
failure message, showing a stale error above a blank form when it was reopened after a cancel. - The OpenAPI document said personal access tokens work on the REST API. They do not. The
publishedbearerscheme offered ayat_…token as an alternative to a session token, but this
API's auth edge accepts session tokens only, so a client following the contract got 401 on every
call. The description now says what is true: personal access tokens authenticate the MCP surface
at/mcp; the REST API wants a token fromPOST /api/v1/auth/login. - Dialogs keep the keyboard inside them. Tab used to walk straight out of an open dialog into
the page behind it, so a keyboard or screen-reader user could end up editing controls hidden
under the overlay whilearia-modalpromised the opposite. Tab and Shift+Tab now cycle within
the dialog, and only the frontmost one traps. - The notification bell does something. It showed the active-alert count but ignored clicks; it
now opens Active alerts, on both the desktop and mobile top bars. - Three chart colors were unreadable in dark mode. The palette's fourth, fifth and sixth
entries kept their light-theme values on the dark surface — which covered most Troubleshoot
report bodies and the passive-event and capacity dashboard widgets, not just charts. - Thresholds no longer judge raw counters. A threshold on a counter metric (
if_hc_in_octets,
errors, discards — anything the collection catalog declares a counter) compared the raw monotonic
total against the bound, so anaboverule latched permanently once the counter passed it and a
belowrule fired a phantom alert at every reboot's counter reset. Counter samples are now read
as OK — which also drains any alert such a rule had latched, through the normal recovery path —
andPOST /api/v1/thresholdson a counter metric answers 400counter_metric. Rates stay a
query-time concern (ADR-012); set thresholds on gauges. - A failed report delete no longer closes silently. Deleting a report template, saved run or
schedule swallowed the error and closed the dialog looking like success; the confirmation now
stays open and shows the message, like every other delete dialog. The report builder and schedule
dialog also disable Cancel while a save is in flight.
Security
- A URL monitor accepted an IPv6 loopback or link-local target. The edge validator blocked
SSRF-prone destinations by parsing the URL's host as an IP address — but a URL parser returns an
IPv6 literal with its brackets ([::1]), which the address parser rejects, so every IPv6 target
fell through to the hostname path and skipped the block entirely.http://[::1]/and
http://[fe80::1]/were accepted when creating or editing a URL monitor. This was a check that
silently did nothing rather than an exfiltration path — the poller refused the same addresses at
probe time, so no request was ever made to one. Three other places parse a URL host and each had
its own correct copy; all four now share one implementation.
Container images
Published to GHCR by CI for this tag:
ghcr.io/horryworks/yagra-core:v0.1.19ghcr.io/horryworks/yagra-poller:v0.1.19ghcr.io/horryworks/yagra-web:v0.1.19