ShadowShield 0.8.0
The classifier-tranche release
The semantic-pretext attack class — the documented ceiling of deterministic injection detection — is now closed.
Headline results (all measured, reproducible)
- InjecAgent (1,054 cases, enhanced): ASR 18.8% → 0.1% (−99.5%) at 76.9% utility via segment-span sanitization — 3× the utility of the redact arm at the same ~0% ASR
- AgentDojo classifier arm: ASR 0% on banking/travel/slack, 1.8% on workspace (baselines 27–62%), no abort-driven utility loss
- LLMail-Inject: 96.75% catch / 0% FPR on 2,000 real attack submissions
- Full matrix: docs/INDUSTRY_BENCHMARKS.md
What's new
TransformerDetector(segment_spans=True)— sentence-level spans redacted, legitimate tool data survives; fail-closed backstop- Chinese (Simplified) signatures — first CJK deterministic-tier coverage, first external contribution (#9, thanks @01luyicheng)
- AgentDojo adapter fix (content-block key, scan cache, block-only abort) with published disclosure of the corrected 2026-08-07/08 numbers
- Offline calibration harness (
scripts/classifier_calib.py) — predicted the live results before spending a cent of API budget
pip install -U shadowshieldFull Changelog: https://github.com/0xsl1m/shadowshield/blob/main/CHANGELOG.md