umbra-eval v0.2.0 — detection benchmark
First tagged release since the head-to-head detection benchmark landed.
- 52-case, 7-language public corpus (public/OWASP, academic/CWE, crafted, hard cross-file taint, multilang, cross-file-lang) with cited provenance + safe decoys.
umbra-eval corpusscores Umbra vs competitor scanners (recall, false positives, by-language).--semgrepenables the optional layer.- Committed result: umbra-core 100% recall / 0 false positives vs Claude Opus 4.8 90% — deterministic, offline, free.
- Plus the ASR / utility-under-defense adversarial suite (
umbra-eval run).