Repository navigation
Glossary
Terms used across the code, issues, and docs. If you hit a word here that isn't defined, say so — that is a documentation defect.
Coordinator — the top-level agent that receives your question, decides which subagents to call, and assembles the final answer. The single entry point to the agent graph.
Subagent — a specialised agent the coordinator delegates to: pod, metrics, logs, events. Each owns one evidence source. If a capability can be a new tool on an existing subagent, it should not be a new subagent.
HITL (human-in-the-loop) gate — the hard stop before any mutating cluster action. Read-only queries run immediately; scale, delete, restart and apply pause for an explicit human approval. No flag or config path may skip it. See The Safety Model.
hitl_bypass — a run-config value used by scripted and autonomy-ladder paths. It does
not disable the gate for high-risk verbs — always_confirm overrides it.
always_confirm — a per-command classification that forces an approval prompt even when
hitl_bypass is set. Computed from the verb and arguments, not from configuration.
RBAC roles — readonly, operator, admin (and superadmin). Enforced before the HITL
gate: a readonly key attempting a delete gets [Permission Denied] and never reaches an
approval prompt.
Chokepoint — the single code path every mutating action must pass through, so the gate and the audit record cannot be bypassed by adding a new tool.
Episode — one complete interaction: the question, the evidence gathered, the reasoning, the answer, and any actions taken. The unit that postmortems and exports operate on.
Flight recorder — the append-only audit ledger of everything the system did, with a hash
chain so tampering is detectable. kq replay walks it; a broken chain is reported, not hidden.
Postmortem — a generated narrative of an episode: what happened, what evidence supported
it, what was changed. kq postmortem / kq export.
Detector — a rule that watches for a condition (a metric threshold, a log pattern, a Kubernetes event). Detectors can be authored in natural language.
Playbook — a YAML description of a known failure mode and how to diagnose it. Adding one is a single file and a genuinely good first contribution (#13).
Sensorium — the layer that continuously collects signals from the cluster and observability stack, feeding the detector engine.
Temporal knowledge graph (KG) — the bi-temporal store of cluster entities and their relationships over time, so the system can answer "what changed just before this broke".
Cortex — the tiered reasoning layer that decides how much reasoning a question deserves. Better diagnosis at roughly half the tokens.
Generation (v1–v4) — a self-contained re-architecture of the same product. v4 is
current; v1 is the architecture in the published paper
and is frozen. See Version Lineage.
Gate — one of the nine checks required to merge to main. Six run locally via
make setup: ruff check, mypy, the server suite, the kq suite, file modes, and syntax
warnings. Three run only in CI: the install smoke test and both suites on Python 3.13.
kq — the terminal client (kube-q on PyPI). kubeintellect — the server
distribution. ki-protocol — the shared SSE wire protocol both depend on.
Autonomy ladder — the staged model of how much the system may do without asking, from answer-only up to approval-gated action. Higher rungs are default-off.
Fabricated diagnosis — the project's cardinal sin: output that looks like a real finding but is not grounded in evidence. Any command that emits a report must fail loudly when its data source is absent rather than return something plausible.
Maintained by Mohsen Seyedkazemi Ardebili · AGPL-3.0-or-later or commercial (LICENSING) · Spotted something wrong on this wiki? Say so — wiki fixes are welcome.
Get started
Understand
Take part
- Your First Contribution
- Community & Support
- Who's using it? #51
- Vote on the roadmap #52
- Good first issues
Repo