Skip to content

Architecture and Integrations

Danny edited this page Aug 20, 2026 · 1 revision

Architecture and Integrations

DProvenanceKit is designed as a behavioral assurance layer that can operate alongside the rest of an AI application's infrastructure.

It does not require teams to replace their existing observability stack.

Instead, it addresses a different responsibility.


Observability and Assurance

A typical observability platform helps answer:

What happened?

DProvenanceKit is intended to help answer:

Was that behavior acceptable according to the expectations we defined?

Conceptually:

                    AI Application
                          │
            ┌─────────────┴─────────────┐
            ▼                           ▼
      Observability                DProvenanceKit
            │                           │
      logs / traces               behavioral trace
      latency / cost              rules / baselines
      debugging UI                regression analysis
            │                           │
            ▼                           ▼
     "What happened?"         "Was this acceptable?"

These concerns are complementary.

Observability helps engineers inspect a system.

Assurance helps teams evaluate the system against defined behavioral expectations.


Where DProvenanceKit Fits

A typical application may already contain several layers:

Application code
      ↓
Agent framework
      ↓
Models and tools
      ↓
Observability
      ↓
Evaluation
      ↓
Deployment pipeline

DProvenanceKit adds a behavioral assurance layer around execution:

                    AI Application
                          │
                          ▼
                   Agent / Workflow
                          │
              structured execution events
                          │
                          ▼
                  DProvenanceKit
                          │
          ┌───────────────┼───────────────┐
          ▼               ▼               ▼
       Traces          Baselines         Rules
          │               │               │
          └───────────────┼───────────────┘
                          ▼
                 Regression analysis
                          │
                          ▼
                  Pass / Fail / Review

The purpose is not to replace the application stack.

It is to add a testable layer around important execution behavior.


Local-First by Design

Execution provenance may contain sensitive information.

DProvenanceKit therefore treats local operation as a first-class capability.

The Python core can record and analyze traces locally without requiring execution data to be sent to a hosted service.

This allows teams to retain greater control over:

  • application traces,
  • customer data,
  • proprietary workflows,
  • internal reasoning artifacts,
  • test fixtures,
  • and audit evidence.

External integrations can still be added where appropriate.

Local operation remains available as a foundation.


Core and Integrations

DProvenanceKit separates its core behavioral model from framework-specific integrations.

Conceptually:

                     DProvenanceKit Core
                            │
          ┌─────────────────┼─────────────────┐
          ▼                 ▼                 ▼
      Recording          Analysis          Storage
          │                 │                 │
          └─────────────────┼─────────────────┘
                            │
                     Integration layer
                            │
      ┌─────────────┬───────┼────────┬─────────────┐
      ▼             ▼       ▼        ▼             ▼
   LangChain      OpenAI  CrewAI  LlamaIndex   Plain Python

The framework adapter captures relevant execution behavior.

The core assurance model remains independent of whichever framework generated the events.


Framework Integrations

DProvenanceKit can be used with multiple agent and application environments.

Depending on the installed extras and integration, these can include:

  • LangChain / LangGraph,
  • OpenAI Agents SDK,
  • CrewAI,
  • LlamaIndex,
  • Google GenAI,
  • FastAPI,
  • Jupyter,
  • MCP,
  • plain Python,
  • and OpenTelemetry-instrumented systems.

The objective is not to force every framework into identical runtime semantics.

The objective is to normalize enough important execution behavior that it can be recorded, compared, and evaluated consistently.


OpenTelemetry Ingestion

Some systems may already emit OpenTelemetry traces.

DProvenanceKit can ingest OTLP-style execution data so existing instrumentation can participate in behavioral analysis.

Conceptually:

Existing application
        ↓
OpenTelemetry spans
        ↓
OTLP ingestion
        ↓
DProvenanceKit trace model
        ↓
Query / Diff / Gate

This allows teams to reuse existing telemetry where appropriate rather than instrumenting every workflow from scratch.

Observability data and assurance analysis can therefore share part of the same execution foundation.


Instrumentation

For plain Python workflows, DProvenanceKit can record important execution steps directly.

For example:

from dprovenancekit import traced, traced_run

@traced
def retrieve():
    ...

@traced
def verify():
    ...

with traced_run(context_id="research-agent"):
    retrieve()
    verify()

This records behavior structurally:

research-agent
      │
      ├── retrieve
      └── verify

The purpose is to identify steps whose presence, ordering, or behavior matters.

Not every line of application code needs to become a trace event.


Storage

DProvenanceKit supports local trace storage so executions can be retained for later analysis.

A simplified architecture is:

Agent execution
      ↓
Structured events
      ↓
Trace store
      ↓
 ┌─────────────┬──────────────┬─────────────┐
 ▼             ▼              ▼
Query         Diff           Replay
                               │
                               ▼
                         Regression gate

Local SQLite storage provides a practical default for persistent trace data.

In-memory storage can support shorter-lived execution and test workflows.


CI as an Enforcement Boundary

One of the most practical places to apply behavioral assurance is the software delivery pipeline.

A typical flow is:

Developer changes
      ↓
Run agent tests
      ↓
Record candidate trace
      ↓
Compare against baseline
      ↓
Evaluate rules
      ↓
Regression detected?
   ┌───────┴───────┐
   │               │
  yes              no
   │               │
Fail PR          Continue

This places agent behavior alongside other software quality gates such as:

  • unit tests,
  • integration tests,
  • type checking,
  • security scanning,
  • API compatibility,
  • and performance regression testing.

The behavior of the AI system becomes another property that can be evaluated before deployment.


CLI Workflow

The command-line interface provides a straightforward path from trace recording to enforcement.

A typical workflow is:

python agent.py
dpk record

Later:

python agent.py
dpk compare

And when the comparison should become an enforcement boundary:

dpk gate

Conceptually:

Run
 ↓
Record
 ↓
Baseline
 ↓
Candidate
 ↓
Compare
 ↓
Gate

This makes behavioral regression testing usable without requiring a hosted control plane.


GitHub and CI Integration

A behavioral gate is especially useful during pull-request review.

Conceptually:

Pull request
     ↓
Run agent test
     ↓
Candidate trace
     ↓
Compare with accepted baseline
     ↓
Behavioral regression?
  ┌──────┴──────┐
  │             │
 yes            no
  │             │
Fail check    Continue

This allows a process regression to become visible before the code is merged.

A skipped verification step, unexpected tool, or meaningful execution change no longer has to rely solely on someone noticing it manually.


Runtime Enforcement

Some requirements cannot wait until CI.

Applications may also use DProvenanceKit primitives at configured runtime boundaries.

For example:

Incoming request
       ↓
Agent execution
       ↓
Required structural check
       ↓
Pass?
  ┌────┴────┐
  │         │
 yes        no
  │         │
continue   reject / escalate

The distinction is important:

DProvenanceKit provides mechanisms for enforcement. The application determines where those mechanisms form a required boundary.

Enabling tracing alone does not automatically make an application fail-closed.

The application architecture determines how assurance results are enforced.


A Layered AI Quality Architecture

DProvenanceKit is most useful as one part of a larger engineering system.

For example:

                   AI Application
                         │
                         ▼
                 Model / Agent Layer
                         │
       ┌─────────────────┼─────────────────┐
       ▼                 ▼                 ▼
  Output Evals      Observability     DProvenanceKit
       │                 │                 │
       ▼                 ▼                 ▼
Is the answer       What happened?     Did required
good?                                  behavior occur?
       │                 │                 │
       └─────────────────┼─────────────────┘
                         ▼
                  Quality decision

These layers answer different questions.

Output evaluation

Was the result accurate, relevant, or useful?

Observability

What happened during execution?

Behavioral assurance

Did important execution behavior remain within defined expectations?

A robust AI engineering stack may need all three.


Integration Philosophy

DProvenanceKit integrations should preserve several principles.

1. The core model remains understandable

Framework-specific adapters should feed the assurance model rather than redefine it.

2. Assurance remains explicit

Teams should be able to understand which behavior is recorded and which expectations are being evaluated.

3. Local operation remains viable

Using an integration should not automatically require execution data to leave the local environment.

4. Existing infrastructure remains useful

DProvenanceKit should complement existing telemetry and evaluation systems rather than requiring an unnecessary rewrite of the stack.


Example End-to-End Architecture

                       User Request
                            │
                            ▼
                     AI Application
                            │
                            ▼
                       Agent Layer
                            │
             ┌──────────────┴──────────────┐
             ▼                             ▼
          Tools                         Model
             │                             │
             └──────────────┬──────────────┘
                            ▼
                    Execution Events
                            │
             ┌──────────────┴──────────────┐
             ▼                             ▼
      Observability                  DProvenanceKit
             │                             │
       debug / metrics              trace / baseline
       latency / cost               rules / provenance
             │                             │
             └──────────────┬──────────────┘
                            ▼
                       QA / CI / Review
                            │
                            ▼
                     Deployment decision

The two branches serve different purposes but can operate on the same application.


Architecture Without a Hosted Dependency

A core design goal is that basic assurance does not depend on an external service.

A developer can:

record locally
      ↓
store locally
      ↓
compare locally
      ↓
gate locally

This is particularly useful for:

  • development,
  • CI,
  • privacy-sensitive environments,
  • offline testing,
  • and organizations that need stronger control over trace data.

Hosted visualization or management services can be layered on top without changing that foundation.


The Core Principle

DProvenanceKit does not need to become the center of the entire AI stack.

It needs to own one responsibility clearly:

Turn important execution behavior into something that can be recorded, compared, evaluated, and enforced.

That makes it a behavioral assurance layer rather than another general-purpose observability platform.


Next