Releases: elastic/example-mcp-app-observability
Release list
Elastic Observability MCP App v1.4.1
Full Changelog: v1.4.0...v1.4.1
Elastic Observability MCP App v1.4.0
Full Changelog: v1.3.1...v1.4.0
Elastic Observability MCP App v1.3.1
Full Changelog: v1.3.0...v1.3.1
Elastic Observability MCP App v1.3.0
Full Changelog: v1.2.0...v1.3.0
Elastic Observability MCP App v1.2.0
Elastic Observability MCP App v1.1.1
What's Changed
- observe skill: extend #8 fix to shape-2 OTel deployments (exceptions in logs) by @JM-elastic in #10
Full Changelog: v1.1.0...v1.1.1
Elastic Observability MCP App v1.1.0
A reworking of every interactive view, plus deeper integration of SLO and alert signals into the health summary so most "what's going on?" investigations stay in a single tool call.
This is the first release published as a 1.1.x — the prior public release was v1.0.16. Re-upload all six skill zips when upgrading; the routing + investigation rules in this release are skill-side and a Claude Desktop restart alone won't refresh them.
Downloads
| File | Purpose |
|---|---|
example-mcp-app-observability.mcpb |
Single-file install bundle for Claude Desktop. Double-click to install. |
apm-health-summary.zip |
Skill pack — teaches Claude when / how to call apm-health-summary. |
apm-service-dependencies.zip |
Skill pack — teaches Claude when / how to call apm-service-dependencies. |
ml-anomalies.zip |
Skill pack — teaches Claude when / how to call ml-anomalies. |
observe.zip |
Skill pack — teaches Claude when / how to call observe. |
manage-alerts.zip |
Skill pack — teaches Claude when / how to call manage-alerts. |
k8s-blast-radius.zip |
Skill pack — teaches Claude when / how to call k8s-blast-radius. |
k8s-crashloop-investigation-otel.yaml |
Optional Agent Builder workflow — automated CrashLoopBackOff / OOMKilled investigation. Import via Kibana → Workflows. |
example-mcp-app-observability-1.1.0.tgz |
npm-style tarball for Cursor / VS Code / Claude Code installs. |
Quick start: Download example-mcp-app-observability.mcpb → double-click → upload each *.zip skill via Customize → Skills → Create Skill → Upload a skill.
If upgrading from v1.0.16: the .mcpb upgrade is in-place. Re-upload all 6 skill zips to pick up the new routing + investigation discipline rules. Restarting Claude Desktop alone will not refresh skills.
Foundation
- Design-system tokens,
.ds-*utility layer, and primitives shared across all six views — colors, spacing, type scale, severity ramp, tile chrome,.ds-viewshell. - WCAG 2 AA contrast compliance across StatusBadge / ZoomControls / dim-text / banner chrome.
- Vite-based harness with per-view fixtures and an axe-core a11y runner so view changes can be validated without a live cluster.
requestDisplayModewired through the shareduseDisplayModehook — Escape key now exits fullscreen across every view.- iframe size-reporting hardened: measure
.ds-viewdirectly with ResizeObserver + MutationObserver +setTimeoutpolling, plus a 50px sanity floor against transient near-zero measurements during re-render.
apm-health-summary
- Tabbed body: Health (KPI tile rows + degraded-service triage chips + recommendation), Signals (SLOs · ML anomalies · fired alerts in one place), Resources (top pods + service throughput).
- SLO + fired-alert integration baked into the response — surfaces what's violating, what's burning, and what fired without a separate
manage-alertscall. SLO source is the authoritative.slo-observability.summary-v3*index, so it works whether or not burn-rate alerting rules are attached. - Cluster + namespace fuzzy match with explicit disambiguation: prompts like "the oteldemo cluster" route correctly without exact-name knowledge; ambiguous matches return candidate lists instead of guessing.
- Per-app filter chips with honest recomputation — service / pod / anomaly content filters client-side without re-running the tool.
- Read-only scope card with coverage-aware rendering (cluster › namespace › service / pod / node counts, plus an applications strip).
- APM + K8s KPI tile rows (throughput, p99, error rate, services on the APM side; CPU, memory, restarts, nodes on the K8s side) with sparklines, status chips, time-axis labels, and hover tooltips.
- Per-service KPIs, per-app pod rollups, pod-to-service correlation via OTel resource attributes, per-entity anomaly counts for filter recomputation.
- Anomaly entity × time heatmap replaces the legacy severity-bar layout; hover shows score + entity + time.
- K8s schema resilience: split usage / limits / restart-count queries so missing kubeletstats columns no longer 400 the whole tile row; auto-falls-back from %-utilization to cores / bytes when limits aren't populated.
- Anomaly cluster filter widened to OR
cluster.nameinfluencer match with namespace-IN-cluster (so jobs that only carryk8s.namespace.nameas influencer aren't excluded). Score floor lowered to 1; new "warning" tier (1–49) rendered in the donut + heatmap.
apm-service-dependencies
- Severity-aware edges: critical (avg latency ≥ 10s, or ≥ 5× the graph's median) render in red with thicker weight and a ⚠ glyph; warning tier in amber. Self-calibrating against the rendered graph.
- "Called slowly" floating tag on nodes whose callers are timing out — surfaces the leaf-looks-healthy / everyone-times-out-on-it pattern (the flagd / hung-feature-flag-service case).
- Top-of-graph anomalies banner names the worst critical edges so a 600s outlier can't hide between two healthy edges.
- Vertical / horizontal layout toggle.
- Click-a-node-to-pin inspection + multi-node compare strip with a hover "+" badge for adding services to comparison.
- Graph-first layout with header inline actions and an in-flow inspect strip.
ml-anomalies / anomaly-explainer
- List + detail drill-down in overview mode (paginated list on the left, click any row to inspect on the right).
- Column-grid fact panel (Function / Deviation / Detected, then Actual / Typical, then Field, then influencers).
- Byte / ms / pct unit inference shared between the tool and the view via a single helper —
metrics.k8s.pod.network.ioActual71021033499now renders as "66.2 GB" not raw bytes; chart Y-axis labels too. Network / disk byte fields specifically recognized. - Annotated time-series chart of actual vs learned-typical with proper Y-axis padding so byte labels don't clip.
- Composite-entity parsing: tolerates the
field=value; field=valueshape the tool emits on result entities being passed back as the entity arg. - Default
min_scorelowered to 1: "what anomalies do we have?" no longer silently filters to critical-only. - No auto-retry: empty results stay empty (won't pile up "Waiting…" widgets); the tool offers a wider search instead.
- Pagination + per-page selector on the overview list. Distinct empty vs waiting state in the detail pane.
observe
- Tense-based mode routing: past-tense queries →
now/table, future-tense / live →metric. "What was X for the past 60 seconds" no longer polls for the future. - Auto-charts time-series in
tablemode (date + numeric, ≥ 3 rows). - Hooks-violation fix that was unmounting the metric-mode view after polling completed (the iframe-disappears-after-60s bug).
- Skill-gap hint surfaced when ESQL fails on patterns the skill specifically warns against.
- Skill cheat sheet for the OTel kubeletstats / cluster-receiver field families with caveats for fields that don't always exist (
cpu.limit,memory.limit,restart_count). - Skill steers toward
exception.*on OTel traces (fixes #8).
manage-alerts
- Full view refresh — list / detail layout with tabs, sort, and group controls.
- Pagination at 5/page default with 5 / 10 / 25 / 50 selector.
- Source filter chip (MCP-created vs all rules).
- Tighter delete-confirm pane with explicit "reply yes" guidance.
- Default
showDetailsoff; long values wrap instead of ellipsifying. - Consistent "Alert rule" terminology in user-facing strings.
k8s-blast-radius
- Graph-first layout with header inline actions and an in-flow inspect / meta strip.
- SVG sizes to viewBox aspect; container padding cleaned up.
- Cluster-scoping parameter so multi-cluster envs can pin the analysis.
Skills
- Welcome banner / setup notices in tool views nudge users to install the matching skill pack on first run.
- Investigation discipline: one tool call per turn, narrate findings between calls, sequential offers (not "X or Y" which fires both calls in parallel).
- Routing improvements: a named-degraded service routes to
apm-service-dependenciesfirst (universally available), notml-anomalies(which depends on jobs being configured for that entity). - Disambiguation flow: ambiguous cluster / namespace returns candidates instead of guessing.
- Lookback default
1hconsistency for unqualified prompts.
Full Changelog: v1.0.16...v1.1.0
Elastic Observability MCP App v1.0.16
Full Changelog: v1.0.15...v1.0.16
Elastic Observability MCP App v1.0.15
Full Changelog: v1.0.14...v1.0.15
Elastic Observability MCP App v1.0.14
Full Changelog: v1.0.13...v1.0.14