Skip to content

Release Evidence

Emmanuel Knafo edited this page Sep 14, 2026 · 1 revision

title: Release Evidence description: How to reproduce this repository's deterministic evaluation gate and offline test evidence.

What evidence this repository produces today

This repository's evaluation gate (eval/evaluation_gate.py) is fully deterministic: it loads eval/golden-dataset.jsonl (13 records spanning business_ready, business_incomplete, fault_unsupported_input, fault_self_approval, fault_forged_actor, fault_unauthorized_preview, fault_revision_invalidation, and fault_injection_attempt categories), runs every record directly against the real calculator, approval-repository, and agent-graph code through eval/deterministic-tests/checks.py, then checks bilingual (en-CA/fr-CA) parity across the dataset. There is no LLM judge and no live agent endpoint involved.

Reproduce the evidence

From the repository root:

pip install -r requirements.txt -r mcp/application-server/requirements.txt -r mcp/rulebook-server/requirements.txt -r src/quote-preparation-agent/requirements.txt
python eval/evaluation_gate.py

The gate prints a per-record pass/fail report and writes the same report to eval/results.json, which every check listed here also uses for continuous-validation.yml, deploy-and-evaluate.yml, and hosted-agent-cd.yml (see Operations and Workflows). A record's checks array names the individual assertion (for example arithmetic_correctness, no_invented_amount) and whether it passed.

Reproduce the full offline test sweep the same way continuous-validation.yml does:

pytest apps/workshop/tests mcp/application-server/tests mcp/rulebook-server/tests src/quote-preparation-agent/tests eval -v --junitxml=evidence/tests.xml

For apps/web-chat specifically (built by Web Chat Build, not deployed):

python -m pytest apps/web-chat/tests -q
cd apps/web-chat/frontend
npm install
node --test tests/*.test.js
npm run build

Reporting scripts

scripts/ci_results.py reads the JUnit XML files a workflow produced (for example web-chat-evidence/*.xml) and renders a per-suite pass/fail/skip/duration summary into $GITHUB_STEP_SUMMARY. It is invoked by Web Chat Build and can be run locally against any directory of JUnit files:

python scripts/ci_results.py --junit evidence

scripts/capture-release-evidence.ps1 uses headless Microsoft Edge to screenshot four sections (pipeline, evaluations, tools, production) of a local assets/release-evidence/index.html report. That HTML report is not yet generated in this repository: the sibling repository's report-building script (scripts/build-release-evidence.js) has not been ported here, so capture-release-evidence.ps1 currently has no index.html to capture against. Porting that report generator against this repository's own deterministic gate output is tracked as follow-on work, not claimed as existing today.

What this page does not claim

This repository has never been deployed to Azure, so there is no verified staging-to-production release run, no production agent version, and no monitoring result to report here. There are no screenshots of a live portal or a past pipeline run in this page; everything above is a reproducible local procedure against checked-in fixtures and code.

Clone this wiki locally