Stop reading logs. Start reading diagnoses.
kdx connects to your Kubernetes cluster, collects every signal from a failing Deployment — events, pod logs, container statuses, resource limits — and produces a plain-English root cause with a copy-pasteable fix. No dashboards. No log diving. One command.
A deploy breaks in staging. You get a CrashLoopBackOff. You now have to:
- Open four tabs: build logs, deploy logs, pod logs, k8s events
- Figure out if it's your code, the config, the infra, or a transient issue
- Do all of this under pressure with no clear starting point
Existing tools show you more of the same data. kdx gives you a diagnosis.
$ kdx diagnose api-server -n production
╭─ Context ────────────────────────────────────────────────╮
│ api-server / production │
│ Cluster: docker-desktop · Pre-class: OOMKilled │
╰──────────────────────────────────────────────────────────╯
╭─ Diagnosis ──────────────────────────────────────────────╮
│ OOMKilled (high) │
│ │
│ Container 'api' is hitting its 256Mi memory limit during │
│ startup before the JVM heap is fully initialised. │
╰──────────────────────────────────────────────────────────╯
• [event] OOMKilling — pod/api-server-7d9f 14:23:01
• [log] java.lang.OutOfMemoryError: Java heap space
• [status] exit_code=137 restart_count=8 last_reason=OOMKilled
╭─ fix_command ────────────────────────────────────────────╮
│ kubectl set resources deployment/api-server │
│ -c api --limits=memory=512Mi -n production │
╰──────────────────────────────────────────────────────────╯
╭─ fix_explanation ────────────────────────────────────────╮
│ The current 256Mi limit is below the JVM's minimum heap │
│ requirement at startup. Raising to 512Mi gives the │
│ process enough headroom to initialise without OOMKilling. │
╰──────────────────────────────────────────────────────────╯
| Failure class | Typical cause | Example signal |
|---|---|---|
CrashLoopBackOff |
App exits on startup — bad config, missing dependency | exit_code=1, repeated restarts |
OOMKilled |
Container exceeds its memory limit | exit_code=137, OOMKilling event |
ImagePullBackOff |
Image doesn't exist or registry is unreachable | ErrImagePull in events |
Pending / Unschedulable |
No node satisfies the pod's constraints | FailedScheduling event |
Requirements: Python 3.12+, kubectl configured, a model provider API key (or a local model via Ollama)
# 1. Clone and install
git clone https://github.com/beejak/kdx.git
cd kdx
make venv
# 2. Set your API key
cp .env.example .env
# Edit .env and set your provider API key
# 3. Diagnose a failing deployment
.venv/bin/kdx diagnose <deployment-name> -n <namespace>No cluster? Use mock mode (built-in fixtures). You still need a working model provider (e.g. ANTHROPIC_API_KEY for hosted Claude, or KDX_PROVIDER=openai-compatible for Ollama):
.venv/bin/kdx diagnose crash-demo --mock crash_loopSee kdx diagnose --help and docs/help.md for options. The payload sent to the model is described in examples/llm_input_format.md.
git clone https://github.com/beejak/kdx.git
cd kdx
make venv # creates .venv/, installs all dependencies
make gate-phase1 # verify: kdx --version should print 0.1.0| Dependency | Version | Notes |
|---|---|---|
| Python | 3.12+ | Required — uses modern union type syntax |
| kubectl | any | Must be configured and pointing at a cluster |
| Provider API key | — | Required when using a hosted model provider |
Create a .env file in the project root (copy from .env.example):
cp .env.example .env| Variable | Required | Default | Description |
|---|---|---|---|
KDX_PROVIDER |
No | anthropic |
anthropic or openai-compatible (Ollama, LM Studio, vLLM, …) |
ANTHROPIC_API_KEY |
If provider is anthropic |
— | Anthropic API key |
KDX_MODEL |
No | Provider-specific | e.g. claude-sonnet-4-5 or llama3.1:8b |
KDX_MAX_TOKENS |
No | 1024 |
Max tokens in the model reply |
KDX_TIMEOUT |
No | 30 / 120 |
HTTP timeout seconds (hosted vs openai-compatible default) |
KDX_LOCAL_BASE_URL |
No | http://localhost:11434/v1 |
OpenAI-compatible API base URL |
KDX_LOCAL_API_KEY |
No | ollama |
API key for local server (Ollama accepts any string) |
KUBECONFIG |
No | ~/.kube/config |
Path to kubeconfig (same rules as kubectl) |
With a .env file in the current working directory, variables are loaded automatically (python-dotenv). You can still use exported shell variables instead.
.venv/bin/kdx diagnose DEPLOYMENT [OPTIONS]
.venv/bin/kdx diagnose --help # options and syntax| Option | Default | Description |
|---|---|---|
-n, --namespace TEXT |
default |
Kubernetes namespace |
--mock FIXTURE |
— | Use a fixture instead of a live cluster |
--dump-context PATH |
— | Write the collected data to a JSON file before running diagnosis |
--context TEXT |
— | Kubeconfig context name to use |
Examples:
# Diagnose a live deployment
.venv/bin/kdx diagnose api-server -n production
# Use a specific kubeconfig context
.venv/bin/kdx diagnose api-server -n staging --context my-gke-cluster
# Diagnose without a cluster (mock mode)
.venv/bin/kdx diagnose crash-demo --mock crash_loop
# Capture the raw collected data for debugging or fixture creation
.venv/bin/kdx diagnose api-server -n production --dump-context /tmp/context.json| Fixture name | Simulates |
|---|---|
crash_loop |
CrashLoopBackOff — app exits with code 1 |
oom_kill |
OOMKilled — container exceeds memory limit |
image_pull_backoff |
ImagePullBackOff — invalid image registry |
pending_unschedulable |
Pending — unsatisfiable node selector |
- Open Docker Desktop → Settings → Kubernetes → Enable Kubernetes → Apply & Restart
- Wait 2–3 minutes for the cluster to start
- Verify from your terminal:
kubectl config use-context docker-desktop kubectl get nodes # NAME STATUS ROLES AGE # docker-desktop Ready control-plane 2m
- WSL2 users: if
kdxtimes out connecting to the API server, point it at the Windows-side kubeconfig:export KUBECONFIG=/mnt/c/Users/<YourUser>/.kube/config
kdx ships with four deliberately broken Deployments to verify your setup end-to-end:
# Apply a scenario
make up SCENARIO=crash_loop
# Watch the pod enter the failure state
kubectl get pods -n kdx-test -w
# NAME READY STATUS RESTARTS
# crash-demo-... 0/1 CrashLoopBackOff 3
# Run the diagnosis
.venv/bin/kdx diagnose crash-demo -n kdx-test
# Tear everything down
make downAvailable scenarios: crash_loop, oom_kill, image_pull_backoff, pending_unschedulable
kdx uses your existing kubectl configuration. To target a remote cluster:
# Point at any context in your kubeconfig
.venv/bin/kdx diagnose my-service -n production --context my-prod-clusterRequired RBAC permissions — kdx is read-only and needs:
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: kdx-reader
rules:
- apiGroups: ["apps"]
resources: ["deployments", "replicasets"]
verbs: ["get", "list"]
- apiGroups: [""]
resources: ["pods", "pods/log", "events", "nodes"]
verbs: ["get", "list"]Apply with:
kubectl apply -f https://raw.githubusercontent.com/beejak/kdx/main/deploy/rbac.yamlkdx diagnose
│
├─► collector/k8s.py Connect to cluster via kubeconfig
│ │ Fetch: Deployment → Pods → Events → Logs
│ │ Pre-classify failure (deterministic rules)
│ ▼
│ DiagnosisContext Immutable snapshot of all signals
│ │
├─► diagnosis/engine.py Build structured prompt
│ │ Call model provider API
│ │ Parse + validate JSON response
│ ▼
│ DiagnosisResult failure_class, root_cause, evidence[], fix_command
│ │
└─► output/formatter.py Render Rich panels to terminal
Data collected per diagnosis:
- Deployment spec (replicas, image, selector, conditions)
- Up to 5 pods (statuses, restart counts, exit codes)
- Pod events from the last 30 minutes (max 50 per pod)
- Last 100 lines of logs from the failing container
- Last 50 lines of logs from the previous container instance
- Namespace events (last 30 minutes, max 50)
All data stays local except for the structured context sent to a remote model provider (if configured).
The exact user-message body (header + DiagnosisContext JSON) is documented with a full sample in examples/llm_input_format.md.
See CLAUDE.md for the full architecture, data models, import boundaries, and agent protocols.
# Run all tests (no cluster needed)
make test
# Run the full gate: lint + boundary check + tests
make gate
# Fix all auto-fixable lint issues
make fix
# Run with coverage
make coverageTests run entirely in mock mode — no live cluster. The suite patches the diagnosis engine and providers so no real model calls occur.
- docs/help.md — long-form reference: architecture, env setup (WSL2, macOS, Linux, remote clusters, in-cluster), model providers, mock fixtures, troubleshooting.
- examples/llm_input_format.md — exact shape of the user message sent to the LLM (plus how to print it from Python).
- examples/diagnosis_context.sample.json — example
DiagnosisContextJSON file (what--dump-contextwrites; same shape as mock fixtures).
- Slack integration — post diagnosis to a channel when a deploy fails
- Helm values suggestion — output a ready-to-apply Helm override for the fix
- Multi-deployment — diagnose all failing deployments in a namespace at once
- Prometheus integration — include CPU/memory usage metrics in the context
-
kubectlplugin — run askubectl diagnose - CI/CD integration — GitHub Actions step that diagnoses on deploy failure
Contributions are welcome. Before opening a PR:
- Read CONTRIBUTING.md for the development setup and workflow
- Run
make gate— it must pass - Add tests for any new behaviour
MIT — © 2026 beejak