Skip to content

umbra-eval v0.2.0 — detection benchmark

Choose a tag to compare

@bkd-dotcom bkd-dotcom released this 30 Jul 21:20
· 24 commits to main since this release
26b10d5

First tagged release since the head-to-head detection benchmark landed.

  • 52-case, 7-language public corpus (public/OWASP, academic/CWE, crafted, hard cross-file taint, multilang, cross-file-lang) with cited provenance + safe decoys.
  • umbra-eval corpus scores Umbra vs competitor scanners (recall, false positives, by-language). --semgrep enables the optional layer.
  • Committed result: umbra-core 100% recall / 0 false positives vs Claude Opus 4.8 90% — deterministic, offline, free.
  • Plus the ASR / utility-under-defense adversarial suite (umbra-eval run).