docs/release_v0.5.0.md
ersion 0.5.0 removes final compliance authority from the language model.
EvidenceShield now:
- screens provenance, scope, environment, and assessment-period validity;
- quarantines unsupported authority claims and untrusted instructions;
- sanitizes instruction-bearing comments inside verified technical exports;
- evaluates explicit control predicates as
satisfied,failed, orunresolved; - assigns the final label deterministically;
- records the model’s original advisory label for ablation analysis;
- preserves EvidenceShield v0.1 for reproducing the provenance-only defence.
Deterministic verdict policy:
Any failed predicate -> non_compliant
No failures, any unresolved -> insufficient_evidence
All predicates satisfied -> compliant
Benchmark composition
| Oracle label | Bundles |
|---|---|
| Compliant | 20 |
| Non-compliant | 15 |
| Insufficient evidence | 5 |
| Total | 40 |
Attack families: instruction injection, authority spoofing, contradiction flooding, temporal rollback, scope substitution, and evidence omission.
Benign perturbations: formatting noise, evidence reordering, irrelevant context, duplicate support, semantic paraphrasing, and metadata noise.
Installation
git clone https://github.com/anirudhnshandilya/AuditPoison.git
cd AuditPoison
python -m venv .venvWindows Command Prompt:
.venv\Scripts\activate
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"Linux and macOS:
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e ".[dev]"Validate
python scripts/validate_dataset.py
python -m pytest -qExpected: 40 bundles, 20 pairs, and 31 tests passing.
Inspect deterministic predicates
python scripts/inspect_predicates.py --bundle-id AP-IA2-005-attackedRun model experiments
Unshielded:
python scripts/run_model.py --adapter ollama --model YOUR_MODEL --defense none --output results/model_unshielded.jsonlHistorical EvidenceShield v0.1 ablation:
python scripts/run_model.py --adapter ollama --model YOUR_MODEL --defense evidenceshield-v0.1 --output results/model_v01.jsonlEvidenceShield v0.2:
python scripts/run_model.py --adapter ollama --model YOUR_MODEL --defense evidenceshield-v0.2 --output results/model_v02.jsonl--defense evidenceshield is an alias for evidenceshield-v0.2.
Important interpretation rule
The v0.2 predicate engine is explicitly designed for the frozen structured pilot. A 100% contract-test result on these 40 bundles demonstrates implementation consistency only. It must not be presented as benchmark generalisation or production effectiveness.
Core metrics
- accuracy and macro F1;
- False Assurance Rate;
- paired Attack Success Rate and robust accuracy;
- benign consistency and both-correct rate;
- citation precision, recall, and F1;
- attack-evidence detection and benign false-flag rate;
- Brier score and expected calibration error.
Repository layout
data/pilot/ Evidence bundles
docs/ Threat model, protocols, and EvidenceShield designs
prompts/ Frozen auditor and analyst prompts
results/ Local metrics and paper tables
schema/ Evidence-bundle JSON Schema
scripts/ Validation, experiments, inspection, and reporting
src/auditpoison/ Python package, screening, and predicate engine
tests/ Integrity, leakage, reproducibility, and defence tests
Versions
v0.1.0— threat model and adversarial pilotv0.2.0— balanced benchmark and evaluation pipelinev0.3.0— real-model adapters and reproducibility manifestsv0.4.0— EvidenceShield v0.1 provenance firewallv0.5.0— EvidenceShield v0.2 deterministic predicate adjudication
Licence and citation
AuditPoison is released under the MIT License. Citation metadata is provided in CITATION.cff.