Elastic Observability MCP App v1.0.2
v1.0.2 — Elastic Observability MCP App
Patch release. Fixes several silent-failure bugs uncovered by a post-v1.0.1 audit, and surfaces ESQL query errors in tool responses so future schema drift is diagnosable instead of invisible.
Bug fixes
apm-service-dependencies— Edge targets fromservice.target.name(gRPC FQNs likeoteldemo.AdService) were not matching theservice.namekeys used for health data (ad), so leaf nodes rendered "no traces" even when traces were present. Resolved via anrpc.service→service.namelookup against SERVER-kind spans; targets now render under their real service names and health populates for every node.apm-health-summary— The tier-1 pre-aggregated service-transaction rollup silently ignored thenamespaceargument, returning services from all namespaces regardless of the user's filter. Namespace is now threaded through.k8s-blast-radius— The downstream-services APM query had no@timestampfilter, scanning every trace ever indexed. Now scoped to a 1-hour window with the deployment cap bumped from 50 → 200 so larger clusters aren't silently truncated.ml-anomalies— The influencer-value wildcard wasn't wrapped in a nested query, so entity filters silently matched zero anomalies against modern nested mappings. Wrapped.
Error surfacing
Every tool's query wrapper previously swallowed ESQL errors and returned an empty array, which is how v1.0.0/v1.0.1 shipped with broken field names (span.status.code, span.duration.us) — queries failed silently and the UI rendered "no data" instead of a diagnosable error.
Tools now share a safeEsqlRows helper that logs failures to stderr and appends them to a per-call error collector surfaced as _query_errors in the response. Applies to apm-service-dependencies, apm-health-summary, and k8s-blast-radius.
v1.0.1 — Elastic Observability MCP App
Patch release. All tool queries now target the OpenTelemetry-native data shape in Elastic, and APM health rollups prefer Elastic's pre-aggregated service metrics with graceful fallback to raw OTel traces.
Schema requirements
All tools in this release assume OTel-native data in Elastic. If you're on classic APM agents, the pre-aggregated APM metrics path (emitted by APM Server regardless of agent type) keeps the tools working — the service maps, health rollups, and blast-radius queries all run against the normalized metrics.
- Kubernetes attributes — all tools query OTel semconv fields (
k8s.namespace.name,k8s.deployment.name,k8s.pod.name,k8s.node.name). ECS-stylekubernetes.*fields are no longer queried. - Trace fields — raw-trace queries target
duration(ns),status.code,kind,service.name. Classic APM raw-trace fields (span.duration.us,span.status.code) are no longer queried directly. - Index patterns —
traces-*.otel-*,metrics-service_*.1m.otel-*,metrics-kubeletstatsreceiver.otel-*. Legacytraces-apm*andtraces-generic.otel-*-only patterns are widened or replaced. apm-service-dependencies— call counts useSUM(response_time.count)(real call volume), notCOUNT(*)over 1-minute buckets. Prefers modernservice.target.name/service.target.typeover legacyspan.destination.service.resource. Per-service health comes frommetrics-service_summary.1m.otel-*+metrics-service_transaction.1m.otel-*when available (tier 1), otherwise falls back to rawtraces-*.otel-*with OTel-native fields (tier 2).apm-health-summary— namespace resolution and service rollups widened totraces-*.otel-*; ML influencer matching accepts both flat (k8s.namespace.name) and nested (resource.attributes.k8s.namespace.name) forms.k8s-blast-radius— migrated to OTel semconv across pods, totals, capacity, and downstream-APM queries.
v1.0.0 — Elastic Observability MCP App
Initial stable release of the Elastic Observability MCP App — an MCP App that brings interactive SRE workflows directly into Claude, Cursor, VS Code, and other MCP-compatible AI hosts. Tools return React-based UIs that render inline in the conversation, so investigations happen where the conversation is.
Tools
Six interactive SRE tools, grouped by the Elastic Observability backend they require. A logs-or-metrics-only deployment can use the Universal tools immediately; the prefixed tools (apm-*, k8s-*, ml-*) surface their requirements in both name and description.
watch (Universal) — Blocks the tool call until an ML anomaly fires or an ES|QL metric condition is met. Three modes: ML anomaly watch, live metric polling with accumulating sparkline, and single-shot "now" metric reads. Works on any numeric field in any index.
create-alert-rule (Universal) — Create a persistent Kibana custom-threshold alerting rule against any metric field in any index. Optional KQL scoping, threshold/comparator configuration, and a form UI that lets the user review and edit the rule before it's created.
ml-anomalies (ML) — Query ML anomaly-detection records and open an inline anomaly-explainer view with model_plot time series, influencers, and detail-mode drill-down. Requires ML jobs configured.
apm-health-summary (APM) — Cluster-level health rollup from APM service telemetry; fuzzy namespace matching and pod-name top_entities, with K8s and ML context layered in when available.
apm-service-dependencies (APM) — Service dependency graph showing upstream/downstream services, protocols, and call volume — rendered as an interactive graph view.
k8s-blast-radius (Kubernetes) — Assess the impact of a node going offline: full outage, degraded, unaffected, and reschedule feasibility per workload. APM context is layered in when available.
Every tool emits an investigation_actions list so the UI can surface opinionated next-step prompts — click-to-send follow-ups without forcing the user to guess the right tool name.
Installation
Multiple installation paths, depending on your AI host:
- Claude Desktop — one-click install via
.mcpbpackage - Cursor / VS Code — via
npx, local stdio, or HTTP - Claude Code — via the
claude mcp addCLI - Claude.ai — via a cloudflared tunnel
See the installation guides for step-by-step setup per target.
Skills
Six Agent Skills teach Claude when and how to use each tool from natural-language user intent — so users don't need to know tool names or deployment specifics. One skill per tool (apm-health-summary, apm-service-dependencies, create-alert-rule, k8s-blast-radius, ml-anomalies, watch) plus an anomaly-explainer view skill. Install via npx, local clone, or by uploading the individual .zip artifacts in Claude Desktop via Customize → Skills → Create Skill → Upload a skill.
Agent Builder workflow
An Agent Builder workflow ships alongside for clients that prefer Agent Builder workflows over MCP tools:
k8s-crashloop-investigation-otel.yaml— automatic CrashLoopBackOff / OOMKilled investigation for clusters on the OTel ingest path (EDOT / kube-stack). Pulls pod context, ML anomalies, upstream health, and recent changes, then synthesizes a root-cause hypothesis.
Requirements
- Node.js 22+
- Elasticsearch 8.x or 9.x
- Kibana 8.x or 9.x (for alerting rules, APM, and ML features)
- An Elasticsearch API key
Elastic Cloud users: on Elastic Cloud the same API key works for both Elasticsearch and Kibana — KIBANA_URL and KIBANA_API_KEY are optional and fall back to the Elasticsearch key.