Skip to content
Danny edited this page Aug 20, 2026 · 4 revisions

DProvenanceKit

Behavioral Assurance for Generative AI

Generative AI systems can change in ways that ordinary output tests do not catch.

A model update, prompt change, tool modification, framework upgrade, or seemingly harmless refactor can alter how an AI system reaches an answer without obviously changing the answer itself.

An agent may:

  • skip a required verification step,
  • call the wrong tool,
  • change the order of critical operations,
  • omit an approval or evidence check,
  • loop unexpectedly,
  • use an unregistered capability,
  • or follow a materially different decision path.

The final response may still look correct.

DProvenanceKit is designed to detect that class of failure.


Observability Is Not Assurance

AI observability tools are valuable for answering:

What happened during this run?

They provide traces, logs, dashboards, latency information, token usage, and debugging visibility.

DProvenanceKit addresses a different question:

Did the system execute the behavior we intended, and can we produce evidence of that execution?

That makes DProvenanceKit complementary to observability systems rather than a replacement for them.

Observability
     ↓
What happened?

DProvenanceKit
     ↓
Was the execution acceptable?
Did required behavior occur?
Did anything materially regress?
Can we demonstrate what happened?

Clone this wiki locally