Skip to content
Sietse edited this page Sep 29, 2026 · 3 revisions

Replay

Regression testing for agents that never give the same answer twice.

Status ✅ Works
Verified 23 August 2026, A40 and H100: a model upgrade fails strict mode and passes semantic mode · bundle 1.20 MiB in the worst case, against a 10 MB target
Package merlin_replay (Python, pytest plugin)

In plain words

You cannot test an AI agent the way you test normal software, because it gives a different answer every time. So teams either do not test them, or they write tests so loose they catch nothing.

This records one run of your agent as a small file of fingerprints (kilobytes, not gigabytes) and keeps it beside the test. Later the run is played back and compared. If anything changed, it names the exact step and the exact word where it changed.

The everyday example

Someone edits one word in a prompt template. The test fails and says:

DIVERGENCE at step 3
  level:    tool call
  expected: query_db(args=revenue_2023)
  actual:   query_db(args=revenue)
  matched before: 3 steps

The mistake is caught before a customer sees it, and you are told where to look.


Two modes, and you need both

Mode Compares Use it for
strict everything, bit for bit catching an unintended change
semantic the decisions: which tools, in what order, with which arguments, how many steps, what was finally answered surviving a deliberate model upgrade

⭐ Why semantic exists. Strict mode fails on every model upgrade, and correctly so, because everything changed. Semantic mode asks whether the agent still decided the same things.

Semantic mode does not check whether two answers "mean the same". Every check is an equality, so the test itself stays deterministic.


Recording and replaying

Turn on the plugin in conftest.py:

pytest_plugins = ["merlin_replay.pytest"]

Use the replay fixture in a test:

def test_research_agent(replay):
    with replay("research-basic"):
        result = my_agent.run("What was 2023 revenue?")
    assert result.answer          # your own assertions still run

Then run pytest:

pytest --record        # capture bundles; commit tests/replays/
pytest                 # replay and compare
pytest --semantic      # compare decisions only; survives a model upgrade
pytest --re-record     # accept an intentional change (refused in CI)
pytest --strict-only   # treat a scope mismatch as a failure
pytest --replay-dir D  # where bundles live (default: tests/replays)

⚠ A missing bundle is a clear failure. It is never recorded automatically. Run pytest --record first.

⚠ --re-record is refused when a CI variable is set. Re-record on your own machine, review the change, and commit the bundle.

⭐ Tracing is off by default, and costs nothing measurable when off. Recording costs 1.05 µs per step.

Want to see it first? python -m merlin_replay.demo runs a short demo with no GPU and no model download.


Check scope before you trust a result

from merlin_replay import check_scope

result = check_scope(recorded_manifest_json, gpu_model="...", gpu_arch="...")
print(result.verdict, result.report)

From C, the same check is merlin_replay_check_scope(...).

⚠ A scope mismatch is not a test failure. A bundle recorded on different hardware is telling you something different from a bundle whose agent genuinely changed. The plugin reports a scope mismatch as a warning, not as a divergence (unless you pass --strict-only).

Where the boundary is, measured 23 August 2026:

CPU run, A40 machine vs H100 machine ⭐ identical
GPU run, A40 (sm_86) vs H100 (sm_90) ⚠ different

So strict replay is scoped to one build on one GPU architecture. The check also refuses a bundle recorded on CPU and replayed on GPU.

⚠ Passing the scope check does not guarantee a bit identical replay. It tells you the replay is meaningful here; it is not a proof.


What is in a bundle, and what is not

✅ In: a manifest, one line per step, one line per tool call, and your assertions. Fingerprints only.

Not stored:

tool responses fingerprints only; you supply the mocks
saved model memory never
model weights never

⭐ Sharing a bundle with someone who may not see the content: set GALAHAD_REPLAY_REDACT=1 when you record. The input and output token ids are replaced with a fingerprint, so the prompt and the answer cannot be read back. Without it, a bundle contains raw token ids, and the recorder warns you.

⚠ A redacted bundle needs its stored memory to still be there. If a block it uses has been removed from the store, replay stops with an error that names the step and the block.

⚠ Side effects are not undone. The email was still sent.


What it does not do

❌ it does not say whether the answer is correct that is evaluation, and is out of scope
❌ no similarity scores, no model judging a model
❌ runs with random sampling are not replayed reliably a run with randomness may replay only by luck

⚠ Random sampling is not detected for you. If your agent runs with temperature above zero, a recorded bundle may pass or fail for reasons unrelated to your change. Set temperature to zero for runs you intend to replay.


Troubleshooting

Symptom Cause Fix
every replay fails after a hardware move strict mode is per GPU architecture check scope first; use semantic mode across hardware
every replay fails after a model upgrade expected in strict mode that is what pytest --semantic is for
flaky pass and fail on the same code sampling is not pinned set temperature to zero for replayed runs
no bundle for 'name' the bundle was never recorded run pytest --record, then commit tests/replays/
--re-record refused a CI variable is set re-record locally and commit the bundle
"the replay failed" with no detail you read the status, not the report read the full report; from C, merlin_replay_last_report names step and token

Related

Clone this wiki locally