Repository navigation
|
Evaluating this for a team that runs regulated workloads, so before I install anything I need to understand the data flow.
I would rather have an honest "no" than a marketing answer. |
Replies: 1 comment
|
Short answer: yes, cluster output is sent to whichever model provider you configure, and today that is inherent to how it works — not a setting you can turn off while keeping the diagnosis. Here is the precise version, including the parts that are weaker than you might hope. What actually leavesThe agents run real queries and the results become model context. So the provider sees:
If your threat model says pod names and application log lines must not reach a third party, then the hosted-provider configuration does not meet it. I would rather say that plainly than sell you a redaction feature that does not exist. What genuinely protects you1. Secrets and ServiceAccounts cannot be read at all. The tool layer blocks them before the Kubernetes API call, for every role including KUBECTL_BLOCKED_RESOURCES=secret,secrets,serviceaccount,serviceaccountsSince the agent can never read them, it can never put them in a prompt. That is a structural protection rather than a filter, which is why it is the one I would rely on. 2. There is a redaction layer — but know exactly what it covers. It is applied on the way into storage, not on the way out to the model. Its own docstring is explicit: "Secret/PII redaction for stored reflexion outcomes… Anything that lands in the database must pass through this redactor." The call sites are episodes, preferences, the knowledge graph and the flight recorder. So: your persisted history is redacted; your live prompts are not. Do not read "the project has a redactor" as "the project filters what it sends to OpenAI". It does not, and the redactor is heuristic anyway — its docstring says so. Keeping it on-premises
But I will not call it a supported path yet, because it is not documented, not tested in CI, and not exercised by the evaluation harness. Making it first-class is #17 — currently the top item on the roadmap-voting thread and, as of today, with zero votes on it. If self-hosted inference is what stands between "interesting" and "installable" for you, saying so on that issue is the single most useful thing you can do here. See also this discussion. Reducing exposure today
A gap this question exposedWriting this out, there is no data-handling document in the repo — nothing that states what leaves the cluster, what is persisted, what is redacted and what is not. That is a real omission for anything a platform team would evaluate, and asking here should not have been the only way to find out. I have filed it: #82. Thanks for asking the question in a form that made the gap obvious. |
Short answer: yes, cluster output is sent to whichever model provider you configure, and today that is inherent to how it works — not a setting you can turn off while keeping the diagnosis. Here is the precise version, including the parts that are weaker than you might hope.
What actually leaves
The agents run real queries and the results become model context. So the provider sees:
kubectloutput — pod, deployment and namespace names, image references, node names, resource limits, events, conditions.