-
Notifications
You must be signed in to change notification settings - Fork 0
Home
Generative AI systems can change in ways that ordinary output tests do not catch.
A model update, prompt change, tool modification, framework upgrade, or seemingly harmless refactor can alter how an AI system reaches an answer without obviously changing the answer itself.
An agent may:
- skip a required verification step,
- call the wrong tool,
- change the order of critical operations,
- omit an approval or evidence check,
- loop unexpectedly,
- use an unregistered capability,
- or follow a materially different decision path.
The final response may still look correct.
DProvenanceKit is designed to detect that class of failure.
AI observability tools are valuable for answering:
What happened during this run?
They provide traces, logs, dashboards, latency information, token usage, and debugging visibility.
DProvenanceKit addresses a different question:
Did the system execute the behavior we intended, and can we produce evidence of that execution?
That makes DProvenanceKit complementary to observability systems rather than a replacement for them.
Observability
↓
What happened?
DProvenanceKit
↓
Was the execution acceptable?
Did required behavior occur?
Did anything materially regress?
Can we demonstrate what happened?
Observe what happened. Define what must happen. Detect when behavior changes. Gate what should not ship. Preserve the evidence.
Home · Quickstart · Behavioral Regression Testing · Rules and Enforcement · Provenance and Audit Evidence · Architecture and Integrations · Cross-Language Conformance