Skip to content

Elastic Observability MCP App v1.0.2

Choose a tag to compare

@github-actions github-actions released this 18 Apr 00:03
· 133 commits to main since this release

v1.0.2 — Elastic Observability MCP App

Patch release. Fixes several silent-failure bugs uncovered by a post-v1.0.1 audit, and surfaces ESQL query errors in tool responses so future schema drift is diagnosable instead of invisible.

Bug fixes

  • apm-service-dependencies — Edge targets from service.target.name (gRPC FQNs like oteldemo.AdService) were not matching the service.name keys used for health data (ad), so leaf nodes rendered "no traces" even when traces were present. Resolved via an rpc.service → service.name lookup against SERVER-kind spans; targets now render under their real service names and health populates for every node.
  • apm-health-summary — The tier-1 pre-aggregated service-transaction rollup silently ignored the namespace argument, returning services from all namespaces regardless of the user's filter. Namespace is now threaded through.
  • k8s-blast-radius — The downstream-services APM query had no @timestamp filter, scanning every trace ever indexed. Now scoped to a 1-hour window with the deployment cap bumped from 50 → 200 so larger clusters aren't silently truncated.
  • ml-anomalies — The influencer-value wildcard wasn't wrapped in a nested query, so entity filters silently matched zero anomalies against modern nested mappings. Wrapped.

Error surfacing

Every tool's query wrapper previously swallowed ESQL errors and returned an empty array, which is how v1.0.0/v1.0.1 shipped with broken field names (span.status.code, span.duration.us) — queries failed silently and the UI rendered "no data" instead of a diagnosable error.

Tools now share a safeEsqlRows helper that logs failures to stderr and appends them to a per-call error collector surfaced as _query_errors in the response. Applies to apm-service-dependencies, apm-health-summary, and k8s-blast-radius.


v1.0.1 — Elastic Observability MCP App

Patch release. All tool queries now target the OpenTelemetry-native data shape in Elastic, and APM health rollups prefer Elastic's pre-aggregated service metrics with graceful fallback to raw OTel traces.

Schema requirements

All tools in this release assume OTel-native data in Elastic. If you're on classic APM agents, the pre-aggregated APM metrics path (emitted by APM Server regardless of agent type) keeps the tools working — the service maps, health rollups, and blast-radius queries all run against the normalized metrics.

  • Kubernetes attributes — all tools query OTel semconv fields (k8s.namespace.name, k8s.deployment.name, k8s.pod.name, k8s.node.name). ECS-style kubernetes.* fields are no longer queried.
  • Trace fields — raw-trace queries target duration (ns), status.code, kind, service.name. Classic APM raw-trace fields (span.duration.us, span.status.code) are no longer queried directly.
  • Index patterns — traces-*.otel-*, metrics-service_*.1m.otel-*, metrics-kubeletstatsreceiver.otel-*. Legacy traces-apm* and traces-generic.otel-*-only patterns are widened or replaced.
  • apm-service-dependencies — call counts use SUM(response_time.count) (real call volume), not COUNT(*) over 1-minute buckets. Prefers modern service.target.name / service.target.type over legacy span.destination.service.resource. Per-service health comes from metrics-service_summary.1m.otel-* + metrics-service_transaction.1m.otel-* when available (tier 1), otherwise falls back to raw traces-*.otel-* with OTel-native fields (tier 2).
  • apm-health-summary — namespace resolution and service rollups widened to traces-*.otel-*; ML influencer matching accepts both flat (k8s.namespace.name) and nested (resource.attributes.k8s.namespace.name) forms.
  • k8s-blast-radius — migrated to OTel semconv across pods, totals, capacity, and downstream-APM queries.

v1.0.0 — Elastic Observability MCP App

Initial stable release of the Elastic Observability MCP App — an MCP App that brings interactive SRE workflows directly into Claude, Cursor, VS Code, and other MCP-compatible AI hosts. Tools return React-based UIs that render inline in the conversation, so investigations happen where the conversation is.

Tools

Six interactive SRE tools, grouped by the Elastic Observability backend they require. A logs-or-metrics-only deployment can use the Universal tools immediately; the prefixed tools (apm-*, k8s-*, ml-*) surface their requirements in both name and description.

watch (Universal) — Blocks the tool call until an ML anomaly fires or an ES|QL metric condition is met. Three modes: ML anomaly watch, live metric polling with accumulating sparkline, and single-shot "now" metric reads. Works on any numeric field in any index.

create-alert-rule (Universal) — Create a persistent Kibana custom-threshold alerting rule against any metric field in any index. Optional KQL scoping, threshold/comparator configuration, and a form UI that lets the user review and edit the rule before it's created.

ml-anomalies (ML) — Query ML anomaly-detection records and open an inline anomaly-explainer view with model_plot time series, influencers, and detail-mode drill-down. Requires ML jobs configured.

apm-health-summary (APM) — Cluster-level health rollup from APM service telemetry; fuzzy namespace matching and pod-name top_entities, with K8s and ML context layered in when available.

apm-service-dependencies (APM) — Service dependency graph showing upstream/downstream services, protocols, and call volume — rendered as an interactive graph view.

k8s-blast-radius (Kubernetes) — Assess the impact of a node going offline: full outage, degraded, unaffected, and reschedule feasibility per workload. APM context is layered in when available.

Every tool emits an investigation_actions list so the UI can surface opinionated next-step prompts — click-to-send follow-ups without forcing the user to guess the right tool name.

Installation

Multiple installation paths, depending on your AI host:

  • Claude Desktop — one-click install via .mcpb package
  • Cursor / VS Code — via npx, local stdio, or HTTP
  • Claude Code — via the claude mcp add CLI
  • Claude.ai — via a cloudflared tunnel

See the installation guides for step-by-step setup per target.

Skills

Six Agent Skills teach Claude when and how to use each tool from natural-language user intent — so users don't need to know tool names or deployment specifics. One skill per tool (apm-health-summary, apm-service-dependencies, create-alert-rule, k8s-blast-radius, ml-anomalies, watch) plus an anomaly-explainer view skill. Install via npx, local clone, or by uploading the individual .zip artifacts in Claude Desktop via Customize → Skills → Create Skill → Upload a skill.

Agent Builder workflow

An Agent Builder workflow ships alongside for clients that prefer Agent Builder workflows over MCP tools:

  • k8s-crashloop-investigation-otel.yaml — automatic CrashLoopBackOff / OOMKilled investigation for clusters on the OTel ingest path (EDOT / kube-stack). Pulls pod context, ML anomalies, upstream health, and recent changes, then synthesizes a root-cause hypothesis.

Requirements

  • Node.js 22+
  • Elasticsearch 8.x or 9.x
  • Kibana 8.x or 9.x (for alerting rules, APM, and ML features)
  • An Elasticsearch API key

Elastic Cloud users: on Elastic Cloud the same API key works for both Elasticsearch and Kibana — KIBANA_URL and KIBANA_API_KEY are optional and fall back to the Elasticsearch key.