A read-only diagnostic MCP server for Crossplane. It gives an AI assistant (Claude, etc.) Crossplane-aware tools to debug stuck resources: it walks the Composite Resource (XR) → Managed Resource (MR) tree, pinpoints the resource that is actually blocking, and returns full condition messages, events, and provider errors — structured for an LLM, not pretty-printed for a terminal.
Status: early (
0.x). Phase 1 (diagnostics MVP) shipped; Phase 2 (discovery & schema — the package-health tools) in progress. Read-only by design — onlyget/listverbs are ever issued, so it is safe to point at a production cluster.
crossplane resource trace prints a tree to your terminal and truncates
condition messages to 64 characters. When an XR is stuck Ready: False while
every managed resource reports Ready/Synced: True, it still leaves you to
find the blocker. This server instead:
- walks the tree and ranks the deepest failing resource first — the likely root cause, not the propagated top-level symptom;
- returns full, untruncated
Ready/Synced/Healthymessages plus correlated events; - prunes noise (
managedFields, etc.) so responses stay token-light; - handles both Crossplane v2 (namespaced XRs, no Claims) and v1 / LegacyCluster (Claims) trees.
See DESIGN.md for the architecture and rationale.
A quick orientation — full detail and phased plan in ROADMAP.md.
Goals: read-only diagnosis of stuck Crossplane resources (root-cause ranking over Claim/XR→MR trees), Crossplane-aware for both v2 and v1/LegacyCluster, LLM-optimized output, safe to point at production, and discovery/schema tools.
Non-Goals: it never mutates the cluster (no create/update/delete/apply or remediation — diagnose-and-advise only); it is not a composition-authoring tool, not a general-purpose Kubernetes tool, and has no GUI.
| Tool | Purpose |
|---|---|
diagnose |
Walk the tree from a resource, rank blocking resources (deepest first) with full messages + recent events. Start here when you know the resource. |
list_unhealthy |
Triage the whole cluster: list composite resources (XRs) and claims that are not Ready/Synced — tiny rows ready to feed straight into diagnose. A resource being deleted is returned even when its conditions still say Ready (a dead reconciler freezes them) and is counted under terminating. Start here when you don't yet know what is broken. |
get_resource_tree |
The composition tree as a flat, parent-indexed node list with per-node Ready/Synced/Healthy state. Native Kubernetes resources composed by a v2 XR report Unknown — they never carry those conditions, and are never named as diagnose suspects. |
get_resource |
One resource, pruned to conditions, recent events, and spec — plus paused and, while terminating, deletionTimestamp + finalizers. |
list_providers |
Every Provider package with installed/healthy state; failing ones add full condition messages, events (e.g. the UnpackPackage registry error), per-revision health, and upgrade-skew notes. Escalate here when a managed resource's error is cryptic. |
list_functions |
Composition Function packages, same shape — a crashlooping function pod is invisible from the XR. Needs Crossplane >= 1.14 (v1beta1 Functions on 1.14–1.16 resolve too); older clusters get an explanatory note, not an unexplained empty list. |
list_configurations |
Configuration packages, same shape — the trail when Compositions/XRDs an XR needs are missing. |
list_contexts |
Available kubeconfig contexts. |
Every tool declares readOnlyHint at the MCP protocol level (so clients can
treat calls as safe), and the server publishes the recommended
list_unhealthy → diagnose → get_resource workflow — escalating to the
package-health tools when a provider/function is the suspect — as MCP
instructions.
Kind inputs are forgiving: Bucket, bucket, and buckets all resolve to the
same kind (exact kind matches always win, so nothing previously valid changes).
Homebrew (distributed as a cask)
brew install --cask briferz/tap/crossplane-mcpThe macOS binaries are unsigned; the cask strips the
com.apple.quarantineattribute on install so it runs without a Gatekeeper prompt.
Container image (GitHub Container Registry)
docker pull ghcr.io/briferz/crossplane-mcp:latestPre-built binaries — download from the latest release.
From source
go install github.com/briferz/crossplane-mcp/cmd/crossplane-mcp@latest
# or, in a clone:
make buildThe server speaks MCP over stdio. It uses your kubeconfig (honouring
KUBECONFIG and --context), falling back to in-cluster config.
crossplane-mcp --context my-cluster{
"mcpServers": {
"crossplane": {
"command": "/path/to/crossplane-mcp",
"args": ["--context", "my-cluster"]
}
}
}"The
Appclaimapp-xyzin namespaceteam-awon't become Ready — why?"
The assistant calls diagnose with {kind: "App", name: "app-xyz", namespace: "team-a"} and gets back the deepest blocking resource (say a Bucket failing
with AccessDenied: invalid credentials) instead of the unhelpful top-level
"waiting for composite resource to become Ready".
When the latest condition is a transient transport error (unexpected EOF,
connection reset, …) but a composition error keeps recurring, diagnose
surfaces that persistent root cause rather than the flake.
For provider-terraform / OpenTofu resources, diagnose also decodes the
base64+gzip error blob (the echo "…" | base64 -d | gunzip hint TF prints) and
surfaces the actionable Error: … on main.tf line NN in a decodedErrors
field — boilerplate trimmed and token-light — so the real cause is in front of
the agent without shelling out.
Each suspect also carries a lifecycle label that separates a wedged teardown
from a resource failing to come up: a resource being deleted shows
Terminating (stuck 140d) (with its deletionTimestamp and how long it has
lingered), while one still provisioning shows Creating (blocked, 5d) — so an
agent routes to "unblock the finalizer" vs "fix the create" immediately. A
terminating suspect also lists its finalizers, naming what still holds the
deletion (get_resource likewise shows deletionTimestamp + finalizers
while a resource is terminating).
A paused resource (crossplane.io/paused: "true") is flagged explicitly:
the annotation suspends reconciliation entirely — conditions go stale and a
deletion can never finish — yet nothing in status says so. Suspects carry
paused: true, a lead reason, and a Paused (blocked, 5d) /
Terminating (paused, 140d) lifecycle label; tree nodes, list_unhealthy
triage rows, and get_resource carry paused too (packages honour the same
annotation and get the same treatment in the package-health tools).
When the stuck resource's error points at the machinery itself — every MR of
one provider failing together, a cryptic gRPC/function error, a Composition
that doesn't exist — list_providers / list_functions /
list_configurations check the package layer: a healthy package costs a
tiny row, while a failing one shows its full Installed/Healthy condition
messages, the failing revision (for providers and functions its name is by
default also its runtime Deployment's name — the pivot to pod logs;
Configurations run no pods), recent events such as the UnpackPackage
registry error, and upgrade skew: an edited spec.package that never
unpacked, a Manual-policy revision waiting for approval with nothing active,
an old revision still serving while the new one is wedged (e.g. incompatible Crossplane version), or package health lagging a failing new revision.
The output stays bounded even in a mass failure (a registry outage breaking
every provider at once): only the first 10 failing packages carry full detail
(reasons, skew, revisions, events) — further rows go compact with a note;
re-call with the name filter for one package's full detail. Revision rows
are capped at 5 per package (revisionsTruncated), keeping the
current/Active/failing ones — rows are dropped whole, condition messages are
never truncated.
The server only ever issues get/list/watch, so it can run under a role
that cannot mutate anything. That invariant is enforced mechanically rather
than by convention: a forbidigo rule in .golangci.yml
fails the lint gate (a required check on main) on any write method called
through the dynamic client, and
TestHandlersIssueOnlyReadVerbs drives every tool handler against a fake API
server and asserts only get/list calls were recorded.
deploy/rbac.yaml ships two ready-made options:
- Recommended (standard Crossplane install): bind the aggregated
crossplane-viewClusterRole that Crossplane's RBAC manager maintains — it automatically covers every XRD-defined and provider-defined resource type as they are installed, and on a default install already includes read access topkg.crossplane.io(what the package-health tools list) and the events read thatdiagnose/get_resource/the package-health tools use. - Fully explicit (RBAC manager disabled): the standalone
crossplane-mcp-viewerClusterRole plus the manifest's small events-viewer role, with one rule per XR/MR API group your platform serves.
Either way, if your v2 XRs compose native Kubernetes resources directly
(Deployments, ConfigMaps, …), add explicit read rules for those types —
neither crossplane-view nor the Crossplane groups cover them, and without
read access the tree reports such a child as unreachable. The manifest shows
how (naming exact resources, never a core-group wildcard, so Secrets stay
unreadable).
For a namespace-scoped setup, bind either role with a namespaced RoleBinding
and call list_unhealthy with an explicit namespace. Note the package-health
tools are then out of reach: package types are cluster-scoped (and their events
live in the default namespace).
| Flag | Default | Description |
|---|---|---|
--kubeconfig |
$KUBECONFIG / ~/.kube/config |
Path to kubeconfig. |
--context |
current-context | Kubeconfig context to use. |
--request-timeout |
30s |
Per-request timeout for Kubernetes API calls. 0 disables it, restoring the previous behaviour where only client-go's transport defaults applied — a wedged apiserver or load balancer could then park a tool call indefinitely. Raise it if a very large cluster-wide list_unhealthy is being cut off. |
--log-file |
Append a JSONL record of each tool call to this path (or - for stderr). Also via CROSSPLANE_MCP_LOG_FILE. |
|
--log-redact |
true |
Mask sensitive values in the --log-file records; disable with --log-redact=false. Also via CROSSPLANE_MCP_LOG_REDACT=false. |
--version |
Print version and exit. |
To inspect what the server saw and returned — useful for debugging, sharing a case, or tuning — set a log file (handy when you can't pass flags through your MCP client):
export CROSSPLANE_MCP_LOG_FILE=~/crossplane-mcp.jsonl
crossplane-mcp --context my-clusterThe path expands a leading ~ and $VARS itself, so the same value works
whether set via a shell or an MCP client's JSON config (which has no shell to
expand them); an absolute path always works.
Each tool call appends one JSON line: {time, tool, durationMs, input, output, error}.
The file is created 0600 (with any missing parent directories created 0700),
so a fresh path works without a manual mkdir -p. Logging goes only to the
file/stderr — never stdout,
which is the MCP protocol channel. (- writes to stderr for ad-hoc debugging and
may interleave with other process output; use a file for clean JSONL.)
By default, three masks run before each record is written:
- key-based — scalar values under sensitive keys (
password,token,secret,credential,apikey,accesskey,privatekey,connection,dsn, …) become[redacted], so inline credentials in a resourcespecaren't written verbatim (reference structures like asecretRef's name are kept); - pair-encoded — in the
{name|key, value|val}shape that Pod-styleenvand provider-terraformvarsboth use, the value is masked when its ownname/keynames a secret ([{name: DB_PASSWORD, value: …}]). The key-based mask cannot see these: the keys it inspects are literallynameandvalue. A pair whose name isn't sensitive ({name: LOG_LEVEL, value: debug}) is left alone; - content-based — every logged string is scrubbed for a few high-precision
secret shapes (PEM private keys, AWS access-key IDs, JWTs,
Authorization: Bearertokens), which catches credential material the key-based mask misses — including in provider error text and the decoded OpenTofu blob (decodedErrors).
Disable all three with --log-redact=false or CROSSPLANE_MCP_LOG_REDACT=false.
Sensitivity: all three masks are best-effort, not a guarantee. The content scrub is deliberately high-precision — it won't catch an arbitrary or unusually-shaped secret, and it intentionally does not mask identifiers like account IDs or ARNs (often the actionable detail). Redaction applies only to the log, never to the live tool response; values that must stay hidden should be marked
sensitivein the Terraform/OpenTofu config. What gets logged is the same closed projections the tools return — built from a resource's metadata, spec, status, and events, never the raw object — and a Secret keeps itsdata/stringDataat top level, outside all of those, so Secret values never reach the log either. (A Secret composed by an XR is still fetched during a tree walk, like any other node; only its contents are withheld.) Treat the log as potentially sensitive and review it before sharing off a machine that touches production.
make test # unit tests with race detector + coverage (no cluster required)
make lint # golangci-lint
make vulncheck # govulncheck
make check # mirror all CI gates locally (fmt, vet, lint, test, vulncheck)CI (build, test, go vet, gofmt, golangci-lint, govulncheck) runs on every
push and PR via GitHub Actions.
Releases are automated with release-please driven by Conventional Commits — no manual tagging.
- Merge
feat:/fix:PRs tomainas usual. - release-please opens (and keeps updating) a release PR titled
chore(main): release X.Y.Zwith the next version and an updatedCHANGELOG.md. - Merge that release PR to cut the release: release-please creates the
vX.Y.Ztag and GitHub release, then GoReleaser attaches cross-platform binaries + checksums and publishes the Homebrew cask, and a multi-arch container image is pushed toghcr.io.
Versioning is pre-1.0 (0.x): feat: and breaking changes bump the minor,
fix: bumps the patch. This is configured in
release-please-config.json; the current version
is tracked in .release-please-manifest.json.
One-time prerequisites:
- Homebrew tap: create a public
briferz/homebrew-taprepository, and add a repo secretHOMEBREW_TAP_TOKEN(a PAT with write access to it) so GoReleaser can push the cask. - Container registry: the image publishes via the built-in
GITHUB_TOKEN; make theghcr.iopackage public in the repo's package settings if you want unauthenticated pulls. - Allow release-please: in repo Settings → Actions → General, enable "Allow GitHub Actions to create and approve pull requests" so it can open the release PR.
Contributions are welcome! See CONTRIBUTING.md for setup and guidelines, and the Code of Conduct. To report a security issue, follow SECURITY.md. File bugs and ideas via the issue templates.