Repository navigation
Replay
Regression testing for agents that never give the same answer twice.
| Status | ✅ Works |
| Verified | 23 August 2026, A40 and H100: a model upgrade fails strict mode and passes semantic mode · bundle 1.20 MiB in the worst case, against a 10 MB target |
| Package |
merlin_replay (Python, pytest plugin) |
You cannot test an AI agent the way you test normal software, because it gives a different answer every time. So teams either do not test them, or they write tests so loose they catch nothing.
This records one run of your agent as a small file of fingerprints (kilobytes, not gigabytes) and keeps it beside the test. Later the run is played back and compared. If anything changed, it names the exact step and the exact word where it changed.
Someone edits one word in a prompt template. The test fails and says:
DIVERGENCE at step 3
level: tool call
expected: query_db(args=revenue_2023)
actual: query_db(args=revenue)
matched before: 3 steps
The mistake is caught before a customer sees it, and you are told where to look.
| Mode | Compares | Use it for |
|---|---|---|
| strict | everything, bit for bit | catching an unintended change |
| semantic | the decisions: which tools, in what order, with which arguments, how many steps, what was finally answered | surviving a deliberate model upgrade |
⭐ Why semantic exists. Strict mode fails on every model upgrade, and correctly so, because everything changed. Semantic mode asks whether the agent still decided the same things.
Semantic mode does not check whether two answers "mean the same". Every check is an equality, so the test itself stays deterministic.
Turn on the plugin in conftest.py:
pytest_plugins = ["merlin_replay.pytest"]Use the replay fixture in a test:
def test_research_agent(replay):
with replay("research-basic"):
result = my_agent.run("What was 2023 revenue?")
assert result.answer # your own assertions still runThen run pytest:
pytest --record # capture bundles; commit tests/replays/
pytest # replay and compare
pytest --semantic # compare decisions only; survives a model upgrade
pytest --re-record # accept an intentional change (refused in CI)
pytest --strict-only # treat a scope mismatch as a failure
pytest --replay-dir D # where bundles live (default: tests/replays)⚠ A missing bundle is a clear failure. It is never recorded automatically.
Run pytest --record first.
⚠ --re-record is refused when a CI variable is set. Re-record on your own
machine, review the change, and commit the bundle.
⭐ Tracing is off by default, and costs nothing measurable when off. Recording costs 1.05 µs per step.
Want to see it first? python -m merlin_replay.demo runs a short demo with no
GPU and no model download.
from merlin_replay import check_scope
result = check_scope(recorded_manifest_json, gpu_model="...", gpu_arch="...")
print(result.verdict, result.report)From C, the same check is merlin_replay_check_scope(...).
⚠ A scope mismatch is not a test failure. A bundle recorded on different
hardware is telling you something different from a bundle whose agent genuinely
changed. The plugin reports a scope mismatch as a warning, not as a divergence
(unless you pass --strict-only).
Where the boundary is, measured 23 August 2026:
| CPU run, A40 machine vs H100 machine | ⭐ identical |
| GPU run, A40 (sm_86) vs H100 (sm_90) | ⚠ different |
So strict replay is scoped to one build on one GPU architecture. The check also refuses a bundle recorded on CPU and replayed on GPU.
⚠ Passing the scope check does not guarantee a bit identical replay. It tells you the replay is meaningful here; it is not a proof.
✅ In: a manifest, one line per step, one line per tool call, and your assertions. Fingerprints only.
Not stored:
| tool responses | fingerprints only; you supply the mocks |
| saved model memory | never |
| model weights | never |
⭐ Sharing a bundle with someone who may not see the content: set
GALAHAD_REPLAY_REDACT=1 when you record. The input and output token ids are
replaced with a fingerprint, so the prompt and the answer cannot be read back.
Without it, a bundle contains raw token ids, and the recorder warns you.
⚠ A redacted bundle needs its stored memory to still be there. If a block it uses has been removed from the store, replay stops with an error that names the step and the block.
⚠ Side effects are not undone. The email was still sent.
| ❌ it does not say whether the answer is correct | that is evaluation, and is out of scope |
| ❌ no similarity scores, no model judging a model | |
| ❌ runs with random sampling are not replayed reliably | a run with randomness may replay only by luck |
⚠ Random sampling is not detected for you. If your agent runs with temperature above zero, a recorded bundle may pass or fail for reasons unrelated to your change. Set temperature to zero for runs you intend to replay.
| Symptom | Cause | Fix |
|---|---|---|
| every replay fails after a hardware move | strict mode is per GPU architecture | check scope first; use semantic mode across hardware |
| every replay fails after a model upgrade | expected in strict mode | that is what pytest --semantic is for |
| flaky pass and fail on the same code | sampling is not pinned | set temperature to zero for replayed runs |
no bundle for 'name' |
the bundle was never recorded | run pytest --record, then commit tests/replays/
|
--re-record refused |
a CI variable is set | re-record locally and commit the bundle |
| "the replay failed" with no detail | you read the status, not the report | read the full report; from C, merlin_replay_last_report names step and token |
- Memory Inspector: one moment rather than a whole run
- Agent Connectors: where the steps come from
- Install · API Reference