The first release of Horizon UI — the next-generation web console for Apache SkyWalking. A dark, dense, information-first interface over the same OAP query protocol and MQE the previous console used, with layer-driven dashboards you configure rather than code, an AI assistant that reads your live data, and an MCP endpoint so the agent you already use can read it too.
AI assistant
Features
- Ask about your system in plain language and get answers built from real dashboard widgets, not just text. A launcher on the right edge opens a chat: describe what you want to know and the assistant reads live data, then streams back an ordered narrative with inline charts, top-N lists and tables drawn by the same components the dashboards use. Open it as a side drawer, expand it to a full page, or put it in its own tab.
- It is read-only and inherits your permissions. It can list services, read active alarms, browse each layer's metric catalog, drill a service down to its instances and endpoints, and chart any of it — never seeing more than you can, and never changing configuration, rules or dashboards.
- It embeds the real product views, scoped to the service you asked about. Ask for topology, traces, logs, browser errors, deployment, API dependencies, an instance map or a cross-layer hierarchy and the actual view mounts inside the chat, interactions intact — click a trace and its span waterfall opens. Both native SkyWalking and Zipkin tracing are covered.
- Everything it shows is a snapshot, and says so. Each block carries a replay badge and the time it was captured, and re-renders identically when you reopen the conversation — offline, with its edge sparklines and its detail views — rather than quietly re-querying and showing today's data under yesterday's question.
- It can read Kubernetes pod logs in the chat, as a result rather than a console. When a filter was applied it says so, so an empty result reads as "nothing matched" rather than a silent pod.
- It can propose profiling, and only you start it. When metrics and traces cannot localise a cause it presents a decision card explaining what it found and what profiling would reveal; nothing runs until you approve it, and only if you hold the permission. It picks the flavour that fits the target and renders the result — a flame graph, a profiled trace's waterfall beside its flame, or a network conversation graph — inline once collected.
- Guided root-cause analysis. Ask what the root cause is and it follows built-in investigation playbooks — a master method plus latency, error-rate, saturation, middleware, Kubernetes-workload and service-mesh specialisations — including following a service down into the infrastructure layer behind it, where memory, disk and connection causes live.
- It answers in each layer's own vocabulary, calling a Kubernetes instance a Pod and a mesh instance a Sidecar, and reads your configured warning thresholds rather than guessing what "healthy" means.
- An outage is reported as an outage. When Horizon cannot reach the backend, it says so and stops, instead of reporting every layer as having no metrics — which reads as "your services aren't reporting" and sends you looking for a problem in your own system.
- Your question stays in view while the answer streams, pinned to the top of the chat, with a timestamp for when you asked and when the answer finished.
- Conversations are kept per user in your browser, with a usage meter, a save toggle and a clear-all. History is stored unencrypted, and the page says so — turn it off on shared workstations. Two tabs cannot overwrite each other's chats.
- Bring your own model. Off by default, vendor-neutral, and configured by an administrator: any OpenAI-compatible endpoint — hosted, local or a gateway — or Amazon Bedrock. The assistant's instructions and the starter prompts shown in an empty chat both ship with defaults you can replace entirely. Until it is pointed at a model the panel opens read-only with a short note.
Agent access over MCP
Features
- The agent you already use can read Horizon. Point Claude Code, Codex, Claude Desktop or any Model Context Protocol client at Horizon and it gets the same tools the AI assistant uses — the metric catalog, figures, topology, traces, logs, Kubernetes, profiling proposals and the root-cause playbooks. The model stays on the caller's side, so no provider and no API key are configured here.
- It is not a new exposure. The endpoint needs the same login as every other route, a permission gates the connection, and each tool re-checks the permission its own screen needs. An agent sees exactly what the operator it authenticated as sees.
- An agent can log in through your browser instead of being handed a token. It opens Horizon's own login page — backed by whatever your organisation configured — you approve once on a consent screen, and it keeps its token from there. The screen shows the permissions the grant would really carry, filtered by what you actually hold, so it never promises access you cannot delegate. Off by default.
- A client that can draw gets the real widgets, not a picture of them: the same charts, topology graphs and trace lists the console draws, fed the captured snapshot. A terminal client reads the data itself and presents it in its own way.
- Two deployments look like two deployments. A Horizon can name itself, and that name reaches the agent — so an operator watching production and staging is never told which is which by guesswork.
- Nothing served over MCP writes anything, and every tool declares itself read-only, so a host stops asking you to approve each step of a single investigation.
Sign-in and access control
Features
- Sign in with your identity provider. Google, Okta, Entra, Keycloak or anything else speaking OpenID Connect — each configured provider becomes a button on the login page. Providers that issue only an access token are supported too. It is additive by design: password login keeps working alongside it, so a misconfigured provider never locks you out during an incident.
- The login page fits what you configured. With password login present, providers fold into one picker; where single sign-on is the only way in, up to four get their own button and the rest fold into a picker. A deployment with no password backend hides the username and password boxes entirely, rather than showing a form that cannot succeed.
- An identity provider says who you are; it does not decide what you may do here. New sign-ins are viewers unless you say otherwise, you decide which domains may sign in at all, and per-address and per-domain overrides raise individuals.
- Horizon calls you what your directory calls you, showing your display name rather than a raw address or an account id. The verified address stays the identity behind every permission check, one hover away.
- An account page, reached by clicking your own name. Who you are, how you proved it — local account, directory, single sign-on with the provider that vouched for you, break-glass, or a token — and which roles you hold and what they grant, so "why can I not see this page" has an answer that does not need an administrator.
- API tokens for callers with no browser. Scripts, CI jobs and agents authenticate on every route under exactly the permissions that route requires. A token names a user and can never carry more than that user currently holds; removing the user revokes it.
- What one session read never reaches the next person to sign in on that browser. Signing out, signing in, or having a session end mid-use discards everything cached, so a shared workstation cannot serve the previous operator's services, alarms, traces or dashboards to the next one.
- Shortening the session timeout now shortens the sessions you already have, not just new ones — which is what you want right after an incident.
- The roles board tells the truth about what each role sees. Every navigation entry is listed, permissions that gate nothing are marked as reserved, and an entry a role cannot open is not offered to it.
Fixes
-
Sign-in accepts up to 64 characters for the username and 64 for the password, and the form stops at the same limit rather than letting a longer value reach the server and come back as a generic failure. A directory that issues credentials longer than this cannot be used to sign in — see Local backend.
-
The sign-in card no longer runs off the edge of a phone-width window — it was cut off below roughly 410px.
-
Source-map upload and removal are disabled, with the reason on hover, for operators who lack
source-map:write— rather than looking available and then failing. Removing a map now asks first: it un-symbolicates stacks for everyone reading that layer.
Login audit
Features
- A durable record of who signed in, when, and from where. Optional and off by default, backed by a shared database rather than a file. An hourly summary stacked by how people signed in, then filters, then the list — statistics first, because the first question is whether anything unusual is happening and the second is which.
- It records only what a valid credential produced. Successful sign-ins, plus the two refusals that happen after authentication already succeeded. A wrong password or an unknown user stays in the application log, because those are what an anonymous caller can produce at will. Nothing that could resume a session is ever recorded.
- Signing in never waits for the database, and cannot be blocked by it. Records are written in the background; an unreachable database is invisible to the person signing in, and the page says it cannot be reached rather than showing an empty table.
- Token traffic is counted on its own tab, at its own grain. A sign-in is a person arriving; token traffic is a machine at work, and stacking them let a busy script outweigh every human sign-in beside it. One row per token per hour, grouped by hour, with each hour's totals and its busiest credentials — and a line that always says whether you are seeing all of them or the top ten.
- Reading it needs its own permission that a wildcard does not grant, because the log holds verified email addresses and client addresses. There is no write and no delete.
Fixes
- A failing statistics write no longer reports the audit store as healthy. Sign-in records, token counts and hourly statistics are written on separate schedules; one of them succeeding used to clear another's failure, and the periodic reachability check cleared a statistics failure it does not actually test. Each is now tracked on its own, so the store reads unhealthy for as long as anything is failing to write.
- An audit write that fails is no longer retried or held. Sign-in batches and hourly statistics are dropped when the database refuses them, so memory does not grow for the length of an outage and nothing is replayed afterwards — the sign-ins from that window are counted as unrecorded rather than reconstructed. Batching itself is unchanged: writes are still grouped for efficiency.
Dashboards
Features
- Layer dashboards are configuration, not code. Every layer's screens are defined by a template you edit in the console — widgets, scopes, service-list columns, thresholds and labels — and published to your backend. Forty-six bundled dashboards ship ready to use, catalogued the way the sidebar groups them.
- A layer's Service, Instance and Endpoint views can each carry more than one page. Each becomes its own row under the layer with its own URL and its own widgets, so a layer's metrics need not share one screen. A page can name the entity it lists — "Brokers" rather than "Instances" — and can narrow which services or instances it is about.
- A new tab widget packs related views into one slot. A grid tile can hold any number of named tabs, each its own small dashboard, edited right where it sits. Only the active tab is queried, so an unopened tab costs nothing.
- The dashboard editor works beside the canvas. Picking a widget kind is a menu with descriptions rather than always dropping in a card you retype; the editor pins next to the board and opens complete, wherever on the board you clicked; adding a widget scrolls it into view.
- Rows under a layer can be put in your own order, dragged in a live preview of the real menu, with a reset to the built-in order.
- Publishing refuses a template that would break the layer, naming the field at fault and writing nothing — rather than storing it and emptying that layer's screen for everyone. Work in progress still publishes: an empty expression or a half-filled section is a normal state of an unfinished draft.
- Cards can render values as coloured status chips rather than bare numbers — so the Kubernetes node status reads as
Readyin green and its pressure conditions in amber or red, instead of a raw1. - Click a latency or error point on a chart to open the matching traces, pre-filtered to that service and centred on the bucket you clicked, opening slowest-first or error-only depending on the metric. Dashboard authors turn it on per widget.
- Compare several entities on one dashboard, with one-click exit from comparison.
- Overview dashboards roll up a whole layer, with per-widget control over how that aggregation is done and how the top services are ranked.
Traces, logs and events
Features
- A trace explorer with a duration-distribution scatter, a time-positioned waterfall and a span detail modal — and Zipkin traces render with the same experience as native ones, including plain-language hints for Zipkin's annotation codes. One shareable link opens either kind.
- Logs and browser errors query on demand. Conditions stage until you press Run query, so a fresh tab prompts you rather than firing a broad query, and switching service resets rather than leaving the previous service's rows under the new name.
- Stored logs can be searched by their content where the backend supports it — the field appears only on a backend that can actually answer it, rather than silently ignoring what you typed.
- Clicking a log row opens a full payload popout with format-aware pretty-printing, the tag table, and a link to the trace.
- Cross-layer inspection for raw logs, browser errors and Kubernetes pod logs. Browser errors carry source-map upload and de-obfuscation, resolving a minified stack back to the original frames with a source snippet. Pod logs tail a container on demand and are never stored.
- A per-service events popout on every layer's service banner — agent restarts, Kubernetes events and other lifecycle records — laid out as one row per instance on a time axis, with a search box for services running hundreds of them.
- Result lists say when they were capped, and offer a next page only when there is one with rows on it. A pager reports the page and what is on it, rather than a total the backend does not provide.
- Tag fields autocomplete on theme, suggesting keys and then per-key values in a dense dropdown instead of the browser's native popup.
Fixes
-
A custom time range that cannot be read is now refused, with the reason under the control — Traces, Zipkin traces, Logs, Browser errors and both Inspect pages. A reversed, half-filled or over-wide range used to be swapped silently for a default window, so the results answered a question nobody asked. Ranges longer than six hours also carry a note that they can be slow on a large deployment; that one is advice, and the query still runs. A request made directly against the API rather than through a page is trimmed to the most recent week instead of being refused, so it still answers with the part that matters.
-
Switching service no longer leaves the previous service's endpoints, instances or profiling segments on screen. The dependent lists clear immediately and say
Reading…while the new ones load, and a slow reply for a selection you have already moved off is discarded instead of overwriting the current one. This affects the Inspect pages' Service/Instance/Endpoint pickers, all five profiling tabs' task and segment lists, the network-profiling process graph, and the Zipkin span/remote autocomplete.
Profiling
Features
- Five kinds of profiling in one place — trace sampling, async-profiler for JVM services, pprof for Go services, eBPF on/off-CPU, and network profiling — each with a task list, a create dialog that tells you upfront what it needs, and a flame graph or conversation graph for the result.
- Continuous profiling has a home. Arm a policy once and the task starts itself when a process crosses a threshold, with nobody present — which is how you catch the problem that only appears at 3 a.m. Each target lists the instances and processes actually being evaluated and how often each has fired: the difference between a policy that is stored and one that is working.
- A task a policy started is visible beside the ones you started by hand, newest first, instead of appearing nowhere.
- A profiling request that cannot be honoured is refused with the reason, rather than quietly repaired into something more expensive or trimmed to a fraction of the fleet you asked for.
- Kubernetes services gain network profiling — pick a pod and capture process-to-process conversations as a topology, the same capability the mesh layer offers.
Look and feel
Features
- Dark, dense and information-first. Horizon is built for an operator watching a system, not for a marketing page: tight tables, small type, and as much signal per screen as stays legible.
- Escape closes any dismissible panel — modals, row popouts and the topology dropdowns alike.
- Searchable, on-theme dropdowns everywhere a picker lists more than a handful of entries, replacing the browser's native controls.
- Denser Kubernetes tables, more rows without scrolling.
- The live debugger reads cleanly on tall and wide captures — the frozen first column stays pinned as you scroll sideways, clicking a source line flashes the whole matching step, and long captures scroll as one page instead of trapping the result in a fixed-height box.
- A step that dropped a record says why, in the backend's own words, instead of leaving you to reconstruct the cause from the payload.
- Serving Horizon under a path prefix is a first-class option, for a reverse proxy that strips it before forwarding.
Languages
Features
- The whole console speaks eight languages — English plus German, Spanish, French, Japanese, Korean, Portuguese and Simplified Chinese. Product, protocol and metric names stay in their original form, because those are what operators read across the docs, the source and every other SkyWalking surface.
- Dashboard text is translated in the console, per language, on a page that shows exactly what the site renders rather than the shipped defaults — with staged drafts, a diff before you publish, and a reset to bundled.
- A translation belongs to its widget, not to the widget's position, so rearranging a dashboard leaves every other widget's translation where it was — a deleted widget takes its translation with it, and a new one starts out English until you fill it in.
- Text left behind by a template edit is found and cleaned up deliberately, rather than being discarded by the next unrelated save.
Operating Horizon
Features
- The container image runs on environment variables alone — no mounted configuration file, no repackaging. The shipped configuration file doubles as the complete, self-documenting reference for every variable.
- Run against a backend whose template store you cannot write, rendering every dashboard from the bundled templates and never calling the template API. The configuration surface becomes honestly read-only, while metrics, traces, logs and topology work exactly as before. This is the supported way to run against an OAP release that has no template management endpoint.
- Cluster Status reports what is actually reachable, testing the real path each feature calls rather than inferring health from configuration being present — so a module that is loaded but broken reads as unreachable instead of a misleading green.
- Configuration hot-reloads, and a rejected reload says so out loud, naming the field at fault and continuing to serve the last valid configuration rather than silently ignoring the edit or falling back to defaults.
- Query fan-out is tunable per deployment — batch sizes, concurrency and protective caps — so a beefy backend can be pushed harder and a modest one protected. Defaults match the built-in behaviour, so the whole thing is optional.
- Responses are not cached by the browser, so metrics, traces, logs and configuration are not left behind on a shared workstation. The console's own files stay cacheable, so nothing gets slower.
- A strict content policy ships by default, permitting scripts only from Horizon's own origin, forbidding inline script, and refusing to be framed. It needs no configuration.
- Outbound documentation links are restricted to hosts you trust, and anything that is not a real web address is refused outright — both when a template is published and when one is read back.
- Warnings reach standard output by default. Break-glass logins, directory failures and rejected configuration reloads are the things you want to see without having raised the log level first.
- A duplicated dashboard record is reported, never resolved behind your back. A dashboard whose definition is ambiguous is hidden rather than rendered from whichever copy happened to win, and opening it by URL explains why and points at where to fix it. Deciding which copy survives is a deliberate cleanup, not something a restart does for you.
Fixes
-
The DSL editor asks before deleting a rule that has no bundled version, and warns before you navigate away or reload with unsaved YAML. Deleting such a rule removes the only copy.
-
The alarms and events Custom range no longer applies as soon as you open it. It takes effect on Apply, and Cancel or Escape leaves the range you were looking at alone.
-
The OAL file viewer no longer strands you on an expired session, and switching files quickly cannot leave one file's contents under another's name.
-
cli:hashcompletes when you press Enter. Typing a password interactively used to hang, because the command waited for end-of-input; the argument and piped forms were unaffected.
Source & binary releases (with signatures and checksums):
- https://dist.apache.org/repos/dist/release/skywalking/horizon-ui/1.0.0/
- KEYS: https://dist.apache.org/repos/dist/release/skywalking/KEYS
Container image: docker pull apache/skywalking-ui:horizon-1.0.0