Skip to content

Rules and Enforcement

Danny edited this page Aug 20, 2026 · 1 revision

Rules and Enforcement

Behavioral baselines are useful when a known-good execution already exists.

But some requirements should not have to be inferred from historical behavior.

They should be stated explicitly.

DProvenanceKit supports the idea that important structural requirements can become machine-checkable controls.


From Expectations to Rules

Some behaviors are important enough that they should be expressed directly.

Examples include:

A verification tool must run.

A particular operation may not repeat more than N times.

Only registered tools may execute.

A required structural event must appear.

A critical approval step must not disappear.

These are not just debugging preferences.

Depending on the application, they may represent part of the system's intended operating policy.

DProvenanceKit allows those expectations to participate in automated evaluation.


Baselines and Rules Solve Different Problems

A golden baseline answers:

How did this execution differ from the behavior we previously accepted?

A rule answers:

Did this execution violate an explicit requirement?

The two approaches complement each other.

               Agent Execution
                     │
          ┌──────────┴──────────┐
          ▼                     ▼
   Baseline comparison      Explicit rules
          │                     │
          ▼                     ▼
     What changed?          Was this allowed?
          │                     │
          └──────────┬──────────┘
                     ▼
              Assurance result

A mature assurance strategy can use both.


Required Steps

Suppose an agent workflow is expected to perform:

retrieve
   ↓
verify
   ↓
decide

A later execution becomes:

retrieve
   ↓
decide

The output might still look reasonable.

But the required verification step did not occur.

If verification is represented as an explicit requirement, the execution can be flagged regardless of whether the final answer changed.

This is the difference between testing:

What did the model say?

and testing:

Did the required process occur?


Tool Requirements

Tool-enabled agents create another class of behavioral expectations.

An application may require that certain tools are used before a decision is made.

For example:

customer request
      ↓
retrieve account
      ↓
verify policy
      ↓
generate action

If the verify policy tool disappears from a future execution, that may represent a meaningful process regression even if the generated action still appears correct.

A structural rule makes that requirement explicit.


Unexpected Tool Use

The opposite problem can also occur.

An agent may begin invoking a capability that was not part of the reviewed workflow.

For example:

EXPECTED

retrieve
   ↓
verify
   ↓
decide

becomes:

CANDIDATE

retrieve
   ↓
unknown_tool
   ↓
verify
   ↓
decide

The new tool may be legitimate.

It may also indicate:

  • a framework change,
  • unintended routing,
  • a new side effect,
  • an unapproved integration,
  • or unexpected agent behavior.

DProvenanceKit can make these changes visible through tool-aware rules and trace analysis.


Registered Tools

In higher-assurance workflows, teams may want an explicit list of capabilities an agent is allowed to invoke.

Conceptually:

Registered tools

✓ retrieve_documents
✓ verify_claim
✓ calculate_total

Unregistered execution

✗ external_lookup

If an unexpected tool appears, the execution can be flagged for review.

The goal is not to claim that a tool registry alone makes a system secure.

The goal is to make deviations from the expected capability set detectable and enforceable where configured.


Loop Detection

Agent frameworks commonly support iterative behavior.

Iteration is useful.

Unbounded or repetitive behavior may not be.

A workflow may unexpectedly become:

search
  ↓
search
  ↓
search
  ↓
search
  ↓
...

This can indicate:

  • an agent stuck in a loop,
  • a broken stopping condition,
  • a prompt regression,
  • a tool-result handling failure,
  • or an orchestration bug.

Rules can define limits around repeated behavior so these patterns become detectable.

For example:

search may repeat at most 3 times

If the candidate exceeds that expectation, the run can be flagged.


Structural Events as Policy Markers

Not every important requirement needs to map directly to a business tool.

Applications can record structural events that represent required controls.

For example:

[[RULE: Evidence Check Completed]]

[[RULE: Human Approval Required]]

[[RULE: Policy Validation Passed]]

These events can become part of the behavioral trace.

That allows an application to encode a requirement in a form that can later be:

  • recorded,
  • queried,
  • compared,
  • and gated.

Why Explicit Structural Events Matter

Consider an application where approval must occur before a final action.

A trace might be:

generate recommendation
        ↓
[[RULE: Human Approval]]
        ↓
execute action

A later candidate becomes:

generate recommendation
        ↓
execute action

The missing approval marker is immediately visible.

This creates a concrete relationship between:

Policy requirement
       ↓
Expected structural event
       ↓
Recorded execution
       ↓
Automated evaluation

Rules Should Be Narrow and Inspectable

A good assurance rule should answer a clear question.

For example:

Did verification occur?

is easier to reason about than:

Was this entire AI execution safe?

The more specific the rule, the easier it is to:

  • test,
  • audit,
  • explain,
  • and maintain.

DProvenanceKit is most useful when high-level requirements are decomposed into concrete behavioral expectations.


Enforcement Boundaries

Detecting a violation and enforcing a violation are related but separate concerns.

DProvenanceKit can provide the mechanism for checking behavior.

The application determines where those checks become mandatory.

Possible enforcement boundaries include:

Local development
       ↓
Test suite
       ↓
CI
       ↓
Pre-deployment validation
       ↓
Configured runtime boundary

This distinction matters.

Enabling tracing does not automatically make an application fail-closed.

The application must decide where a failed assurance check is allowed to stop execution or deployment.


CI Enforcement

CI is one of the most practical places to enforce behavioral rules.

A typical workflow might be:

Code change
    ↓
Run agent test
    ↓
Collect trace
    ↓
Evaluate rules
    ↓
Violation?
 ┌────┴────┐
 │         │
yes        no
 │         │
Fail PR   Continue

This turns behavioral requirements into part of the normal software quality process.


Runtime Enforcement

Some controls may need to be evaluated while an application is running.

For example:

Agent begins execution
        ↓
Critical tool invoked
        ↓
Required rule evaluated
        ↓
Continue or reject

DProvenanceKit primitives can participate in this kind of application-specific enforcement.

The boundary should be explicit.

DProvenanceKit provides the mechanisms for enforcement. The application determines where those mechanisms form a required boundary.

That is intentionally different from claiming that the library automatically guarantees application safety.


Rules and Baselines Together

The strongest behavioral assurance model often combines:

Historical expectation
      +
Explicit requirement
      ↓
Candidate execution
      ↓
Evaluation

The baseline tells you:

What changed?

The rules tell you:

Did anything required disappear or anything prohibited occur?

Together:

Golden behavior
      +
Explicit requirements
      ↓
Candidate execution
      ↓
Behavioral evaluation
      ↓
Pass / Fail / Investigate

Example

Suppose the accepted workflow is:

plan
 ↓
search
 ↓
verify
 ↓
decide

And the application also has two explicit rules:

Rule 1:
verify must occur

Rule 2:
search may not repeat more than 3 times

A candidate executes:

plan
 ↓
search
 ↓
search
 ↓
search
 ↓
search
 ↓
decide

Two separate assurance signals now appear:

Baseline comparison
→ verify disappeared
→ search behavior changed

Explicit rules
→ required verify step missing
→ search repetition exceeded

The final generated answer does not need to be obviously wrong for the execution to deserve rejection.


What Rules Do Not Prove

A passing structural rule does not prove that:

  • the model's answer is factually correct,
  • the underlying data is trustworthy,
  • the organization satisfies a regulation,
  • the application is secure,
  • or the overall system is safe.

Those require additional controls and evaluation.

DProvenanceKit rules address a narrower question:

Did the execution satisfy the behavioral requirement we defined?

That narrower claim is what makes the result testable and reviewable.


The Core Principle

Important behavioral requirements should not live only in documentation.

Where practical, they should become executable expectations.

Requirement
    ↓
Rule
    ↓
Execution
    ↓
Evaluation
    ↓
Evidence

That turns an informal process expectation into something the system can actually check.


Next