Skip to content

ShadowShield 0.8.0

Choose a tag to compare

@0xsl1m 0xsl1m released this 11 Aug 00:23
· 59 commits to main since this release
a2d4757

The classifier-tranche release

The semantic-pretext attack class — the documented ceiling of deterministic injection detection — is now closed.

Headline results (all measured, reproducible)

  • InjecAgent (1,054 cases, enhanced): ASR 18.8% → 0.1% (−99.5%) at 76.9% utility via segment-span sanitization — 3× the utility of the redact arm at the same ~0% ASR
  • AgentDojo classifier arm: ASR 0% on banking/travel/slack, 1.8% on workspace (baselines 27–62%), no abort-driven utility loss
  • LLMail-Inject: 96.75% catch / 0% FPR on 2,000 real attack submissions
  • Full matrix: docs/INDUSTRY_BENCHMARKS.md

What's new

  • TransformerDetector(segment_spans=True) — sentence-level spans redacted, legitimate tool data survives; fail-closed backstop
  • Chinese (Simplified) signatures — first CJK deterministic-tier coverage, first external contribution (#9, thanks @01luyicheng)
  • AgentDojo adapter fix (content-block key, scan cache, block-only abort) with published disclosure of the corrected 2026-08-07/08 numbers
  • Offline calibration harness (scripts/classifier_calib.py) — predicted the live results before spending a cent of API budget
pip install -U shadowshield

Full Changelog: https://github.com/0xsl1m/shadowshield/blob/main/CHANGELOG.md