Skip to content

DataFlowBench v0.5.0

Choose a tag to compare

@DavidBakerEffendi DavidBakerEffendi released this 28 Aug 17:40
· 69 commits to main since this release
7d688c6

DataFlowBench v0.5.0

Fifth immutable release snapshot. It is two things at once: the Bifrost
v0.10.7 fix cycle
, in which every runner error in the previous freeze is
gone and Bifrost's decisive-correct count on the thirteen kernels roughly
doubles, and the first release to publish the modeling tier (#15) and the
tool-native tier (#16)
alongside the benchmark-controlled kernels. Four
analyzers — Bifrost v0.10.7, CodeQL 2.26.3, Joern 4.0.610, and Semgrep CE
1.174.0 — are bound at one fixture revision under the freeze/v1 contract
with claim scope release.

Freeze identity

  • Freeze ID (manifest SHA-256):
    43a34341ca3f818f55878bb23562da03b8fc4b1fc0c83f47b954eb22ec3f41e4
  • Manifest: reports/freeze.json
  • Benchmark revision: 7d688c6 (tag v0.5.0)
  • Fixture revision:
    sha256:9df209ed3d7723a3ee33f2b289cf2afe34a3add781bdf2a2ac445de42b8d0151
  • Case schema v2, normalized result schema v1
  • 852 frozen cases, 66 bound reports, 2884 scored case results
  • Score tiers: calibration, core, language-extension, modeling
  • Model profiles: benchmark-controlled, tool-native

Three populations, never pooled

This release publishes results under two model profiles and four score tiers.
They are separate populations and the separation is structural, not
presentational:

  • Benchmark-controlled kernels (core, calibration,
    language-extension): the thirteen expanded kernels of v0.4.0, unchanged in
    population. Each analyzer is configured by DataFlowBench to the benchmark's
    own source/sink contract.
  • Benchmark-controlled modeling (modeling tier, 12 templates in six
    balanced categories × three languages × four adapters): the same profile,
    measuring whether a tool's model declaration surface is load-bearing. A
    tool that cannot express a category takes unsupported for it, decided
    before the tool is invoked.
  • Tool-native (tool-native profile, 6 templates × three languages × four
    adapters): what each tool ships and decides on its own, with no
    DataFlowBench-supplied model.

benchmark-controlled and tool-native are never pooled and never compared
number-to-number
, on any page. A modeling row and a native row that both say
"6 templates" are not the same six templates and not the same question.

Bound evidence

Normalized report SHA-256 Analyzer Cases
reports/bifrost-c-kernel.json db4a8415f86f75c680aa52cd8e91896c70d67b6b88bab2a07f85fa0ef198a75d Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 50
reports/bifrost-cpp-kernel.json 1398c72370cd7c3d9c39fabde6c3c7230af6a376fd7c3f3571041062082abb82 Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 56
reports/bifrost-csharp-kernel.json 2fe19dbf7e5a61baeb15b9d33a2c861639f9d836dd867cce609c8d68917ceb94 Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 58
reports/bifrost-go-kernel.json 919e62b265d16b69a8288c39d5be0eff12ab7d878df3d42e116895403c4f7382 Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 58
reports/bifrost-java-kernel.json ff99ba2ee4f486ba6da1046e779581513b0a726bd0c038c9b95552c555530915 Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 58
reports/bifrost-java-modeling.json 2e6e38e4cfd5436598081866e27bbaf98c01da8c50090a513b8f53d71b483eed Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 24
reports/bifrost-java-native.json 5add3826fa45f30c92ff890e52c246843ec157761525309a34d71af24692df41 Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 — bifrost 0.10.7 built-in policy packs 12
reports/bifrost-javascript-kernel.json 58675cdd313b41add0294691d34bbafcf8c9872d60a28ddef71f44ab519b9b62 Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 58
reports/bifrost-javascript-modeling.json 751d5ae5412e87dd984de6ee687c3e8a8d3e98374de203eeceacc9987269fbfc Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 24
reports/bifrost-javascript-native.json 0a2ca366a08c571fa1d769e4f974ede005e19be2d1e4512f6bd9552aeaa3ec77 Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 — bifrost 0.10.7 built-in policy packs 12
reports/bifrost-kotlin-kernel.json db5e0633a1313a18cb248961a7a6ae5ea2bba2e31bd0a00c17b66601337705c1 Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 58
reports/bifrost-php-kernel.json ad17e035ecfabb508bf9dc9f92776cf4c6593c115cc8712380fc76c743d08d1b Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 58
reports/bifrost-python-kernel.json d0c713ed0b7f7bf24adb432443336562b53c57d9be995c26f77e9c2ae14f2419 Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 58
reports/bifrost-python-modeling.json ce470da597805006288114cd36de500bfdeeccc03a19f9569e230beb0eea840d Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 24
reports/bifrost-python-native.json ae24e34efe49f7307d571e9b3c42de1d5f829fb6d965cfe455ead4010bbaba56 Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 — bifrost 0.10.7 built-in policy packs 12
reports/bifrost-ruby-kernel.json afab27546ffb827a0244a05d6bfeb882cadf2eae26fd74f6c75051687ddcff97 Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 58
reports/bifrost-rust-kernel.json 340d6104d3ecdfba2de9429cfdfc4380bbcb30e708180644dacef588b060ca02 Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 56
reports/bifrost-scala-kernel.json 67fe1b3d9a7cfa1cd6dbed6596527f60bb25e4c1d3c467274ea57e0d5ced5bdc Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 58
reports/bifrost-smoke.json c78ed974f229ada333b621a9a325ebeca7f6aa774000482205b218e26190529c Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 118
reports/bifrost-typescript-kernel.json c4bca856809db333de1f81f8e1c5522e26a8e7577ea129c410f8abc2b1a9ed61 Bifrost 0.10.7, build 44d9a5be416432bf8ed414afd3ea0031245ebb57 58
reports/codeql-c-kernel.json 9ed3846bd073c4078f81a8702fe6d71010d9dc2a6dadf4338765daf804fb5651 CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 50
reports/codeql-cpp-kernel.json 4053f81a192b22db1c7e013af9b96e728aad6c4863ec44952f1fe02bd2d9303e CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 56
reports/codeql-csharp-kernel.json 69c0f67040a59a172858f2e29c29a8653ed0348fee81f8c11e554a4e59a0a359 CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 58
reports/codeql-go-kernel.json 715946c2e0183a20b60cf09fac97d48855e95d0eac2aaa392edf911b94d91afc CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 58
reports/codeql-java-kernel.json f13a11290939e21db527db660bf76408ec73315a3f2f1f7a9d84fd53b82c19f3 CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 58
reports/codeql-java-modeling.json c76aac2e4bf8ab05c0fa9068903c5eafb4b7407bf991031fac5e05662f992cbd CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 24
reports/codeql-java-native.json 16e2ca2ea4e7dcef4e5f4f444a33767e44c78e2ffa051e1749f3de9222037459 CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 — 2.26.3 shipped suite codeql/java-queries@1.11.9:codeql-suites/java-security-extended.qls 12
reports/codeql-javascript-kernel.json 48c18bb0c559dea6311dc34af5138e6bd0dce824581977790fc14fce5a38793f CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 58
reports/codeql-javascript-modeling.json a3e6f90dd6f068b723ebcd6e9e73c3300452eb1d9ccbf76d617f82af1c5f5525 CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 24
reports/codeql-javascript-native.json a9889691d7da7837e6edbc23b894492fe237a5dcd62d4bdd73b5b598eb4a2d9e CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 — 2.26.3 shipped suite codeql/javascript-queries@2.4.4:codeql-suites/javascript-security-extended.qls 12
reports/codeql-kotlin-kernel.json 170d5b9c7b8cd843a9deaafe291e510e483eaddb237967c2c96fb50ff42f6617 CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 58
reports/codeql-python-kernel.json 20291cc4a6a50824ca72f42f664bd580e0aff0fff769cf7c63efb8a889037717 CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 58
reports/codeql-python-modeling.json b9a3c13de153d28604d97fac51366f811a3a98dab740395a6eb0efca6811a9e2 CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 24
reports/codeql-python-native.json 21a2e4136756bb640d0f68ba6850d8315230df3f7a6622d538d858f83116ece3 CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 — 2.26.3 shipped suite codeql/python-queries@1.8.9:codeql-suites/python-security-extended.qls 12
reports/codeql-ruby-kernel.json 7aab3f0f9fe51869c73f21073b7c7c4a202c5a8e5a2e1b7c42a5059407fcf282 CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 58
reports/codeql-rust-kernel.json 02625bfd21b8609cb7800e4a55e43f57c33718b2dd8b3c07e8415404224b9066 CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 56
reports/codeql-typescript-kernel.json f702e16cc3cf3cc0e08eb038e70dda3d2e0898b66a97873d8517df704e79f5dc CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 58
reports/joern-java-kernel.json e9b98b262970b7cfa547e270b61357f5d04f6f04879a0575c39801d1eb857370 Joern 4.0.610, build joern-cli:4.0.610 58
reports/joern-java-modeling.json 062de7777d607da2225e5dfc0c76bfd71ed2c5d265a290f2d4d0a5be5b89ed1e Joern 4.0.610, build joern-cli:4.0.610 24
reports/joern-java-native.json a2bd1c79179c5dd6a70797eb7e879581ce392131cf018daf0ec36d888da3a4d3 Joern 4.0.610, build joern-cli:4.0.610 — 4.0.610 DefaultSemantics only 12
reports/joern-javascript-kernel.json 2303345687cf28bff643fe3ae57f9df01ef3fca6450e2b9f739259cdab440311 Joern 4.0.610, build joern-cli:4.0.610 58
reports/joern-javascript-modeling.json b560ae830e00453dbd7e089bbdcd005a93700274aeee0b1b1853be1820d6218c Joern 4.0.610, build joern-cli:4.0.610 24
reports/joern-javascript-native.json ca750185e576994faf9e23e7d44ddd3f1ca18134a4eae818eed2c8edd450d703 Joern 4.0.610, build joern-cli:4.0.610 — 4.0.610 DefaultSemantics only 12
reports/joern-php-kernel.json 6339d32f63c6ff1960573d7e9f3e57cc48b4b96891bebc799d3b1a1ef41a4c93 Joern 4.0.610, build joern-cli:4.0.610 58
reports/joern-python-kernel.json a47e6d4d97f660675e2781f7e3d41a6b259b89b4514ea62e7cf2bb19e6ab2f44 Joern 4.0.610, build joern-cli:4.0.610 58
reports/joern-python-modeling.json 6f7ea2ece57faf66668739ec536488c570484e6be96ec35e4a6888ad4447b011 Joern 4.0.610, build joern-cli:4.0.610 24
reports/joern-python-native.json d9612e60a43983b62bf218cd022acafc51db983a6c90a732eb4c8b59849f5444 Joern 4.0.610, build joern-cli:4.0.610 — 4.0.610 DefaultSemantics only 12
reports/joern-ruby-kernel.json 0ea9e3469ffebae67acc917035510c2804aed90b52ffc9c10013ae841a150b2e Joern 4.0.610, build joern-cli:4.0.610 58
reports/joern-rust-kernel.json 696e306951c96be64b45da6f947166a7af92c7cacf18a2195aa18ef63e0254dd Joern 4.0.610, build joern-cli:4.0.610 54
reports/semgrep-c-kernel.json 8b1978f8a41803f0341ff7778eabc0fa6665ae97edfb6468883fd93d4b42ae99 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 48
reports/semgrep-cpp-kernel.json 7a4eb1bdabedbe64555ff6f69d2c3578878aa6689a1fbdaf66742c17a32ef375 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 56
reports/semgrep-go-kernel.json 001ee7a2b34c7de1622198a27c1c95cdeb78e5d9f081cdd5ee6fd1ed1c53b2ab Semgrep CE 1.174.0, build semgrep-oss:1.174.0 58
reports/semgrep-java-kernel.json 669920574d8eb8885046cc746c1ae9f549186cd1151fc8bbac21f839c6b376f8 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 58
reports/semgrep-java-modeling.json 8af4d45c830f0f68e15474eb5f67b599e825c0e94b748cea6debb66d8d42c2c2 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 24
reports/semgrep-java-native.json ef0ec5c9bcb7ed7353cf485d1a02d30feea5820e9b0e15ef0e3eda583c79b46d Semgrep CE 1.174.0, build semgrep-oss:1.174.0 — 1.174.0 over the pinned snapshot vendored from https://github.com/semgrep/semgrep-rules into adapters/semgrep/native/java 12
reports/semgrep-javascript-kernel.json 216f093f4a5f893783771083915fc16bb2d1cfa1402ddae9de6b511fd6fa9841 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 58
reports/semgrep-javascript-modeling.json c11bf5343af35301963d064a90207ee142e06b4f5664fcdba88cedb2aa64904c Semgrep CE 1.174.0, build semgrep-oss:1.174.0 24
reports/semgrep-javascript-native.json f2c815b3b80d5c5e3f15d8c6232532a9a9afe69cd7c3fa3eea165ee8e784bd75 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 — 1.174.0 over the pinned snapshot vendored from https://github.com/semgrep/semgrep-rules into adapters/semgrep/native/javascript 12
reports/semgrep-kotlin-kernel.json 88def9a7809954ebcbfdaa174e2d518ff3e4f65f27cec2f395eb97a143905a10 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 58
reports/semgrep-php-kernel.json c265ac735c389374aebd29b7fefcce6f3b5288a868f11281c563c19a43e673e0 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 58
reports/semgrep-python-kernel.json d047262d8a19eb49bd0b0cb5f284d00f7af91af3705e2241d45fc490dc45c3e8 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 58
reports/semgrep-python-modeling.json 9c4bc7eac5d731284b8bf47f04ed3257c2271ae811a0137b912528f58eb8d58c Semgrep CE 1.174.0, build semgrep-oss:1.174.0 24
reports/semgrep-python-native.json b807832c6dfacf25489c6d05669d86dc237fb09156a5bbac4ba86cf052883151 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 — 1.174.0 over the pinned snapshot vendored from https://github.com/semgrep/semgrep-rules into adapters/semgrep/native/python 12
reports/semgrep-ruby-kernel.json 375458b3da818ef8fa52985c6718a26e68f97f93714d261fd3101ea38e12ed17 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 58
reports/semgrep-rust-kernel.json 25192435fb10edf307c5b0c5705849b5492e9d560060cc70a7fd6eaa3b4a7297 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 54
reports/semgrep-typescript-kernel.json 7ea8aaf04089d69d411a4778079513c6354b7b8a0c1fee2b73ce6bd373572ef8 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 58

Every report remains its own scorecard. The forty-two kernel and smoke reports
carry the benchmark-controlled profile on the taint track; the twelve
*-modeling reports carry the same profile on the modeling tier; the twelve
*-native reports carry the tool-native profile. The build identity
column is the witnessed identity of the binary actually invoked, extended for
the native rows with the shipped ruleset or pack that was activated — including
the rows that decide nothing. Raw evidence for every result is retained under
reports/raw/ and digest-bound in the manifest.

Results

Generated pages live in results/ and derive exclusively from the frozen
manifest; see results/index.md. Incomplete outcomes (inconclusive,
unsupported, runner-error) are capability and execution coverage and are
never counted as clean negatives.

The kernels

Correct decisions (true positives plus true negatives) on each kernel's core
tier, with incomplete outcomes shown beside them. Each language is its own
population with its own denominator, and each analyzer column is read
independently: DataFlowBench publishes no combined leaderboard, and cores of
different sizes are never pooled. n/a means the analyzer has no report for
that kernel in this freeze — no extractor, no frontend, or no adapter — which
is coverage, not a score.

Kernel core Bifrost 0.10.7 CodeQL 2.26.3 Joern 4.0.610 Semgrep CE 1.174.0
Java (29 templates, 58 assertions) 37/58 (20 inc) 48/58 47/58 12/58 (44 uns)
JavaScript (29 templates, 58 assertions) 36/58 (22 inc) 48/58 44/58 12/58 (44 uns)
TypeScript (29 templates, 58 assertions) 34/58 (24 inc) 48/58 n/a 12/58 (44 uns)
Python (29 templates, 58 assertions) 36/58 (22 inc) 48/58 48/58 12/58 (44 uns)
Kotlin (29 templates, 58 assertions) 28/58 (30 inc) 46/58 n/a 12/58 (44 uns)
Scala (29 templates, 58 assertions) 38/58 (20 inc) n/a n/a n/a
C# (29 templates, 58 assertions) 32/58 (24 inc) 47/58 n/a n/a
Go (29 templates, 58 assertions) 35/58 (22 inc) 45/58 n/a 12/58 (44 uns)
PHP (29 templates, 58 assertions) 32/58 (26 inc) n/a 48/58 12/58 (44 uns)
Ruby (29 templates, 58 assertions) 21/58 (36 inc) 49/58 40/58 12/58 (44 uns)
C++ (28 templates, 56 assertions) 30/56 (26 inc) 42/56 n/a 12/56 (42 uns)
C (24 templates, 48 assertions) 40/48 (8 inc) 41/48 n/a 12/48 (34 uns)
Rust (27 templates, 54 assertions) 36/54 (18 inc) 44/54 43/54 12/54 (40 uns)

The saturation the preregistration set out to end is still gone: no analyzer
answers a whole expanded core correctly in any of the thirteen languages, and
no column above reaches its own denominator. The population is byte-identical
to v0.4.0's, so the Bifrost column — and only the Bifrost column — is
comparable with the corresponding v0.4.0 column; CodeQL, Joern, and Semgrep are
on the same pins as v0.4.0 and reproduce their v0.4.0 numbers.

The challenge templates stay individually visible as a stratum. The same
populations, restricted to the challenge templates only:

Kernel core Bifrost 0.10.7 CodeQL 2.26.3 Joern 4.0.610 Semgrep CE 1.174.0
Java (13 templates, 26 assertions) 6/26 (20 inc) 21/26 19/26 0/26 (26 uns)
JavaScript (13 templates, 26 assertions) 4/26 (22 inc) 19/26 18/26 0/26 (26 uns)
TypeScript (13 templates, 26 assertions) 4/26 (22 inc) 19/26 n/a 0/26 (26 uns)
Python (13 templates, 26 assertions) 4/26 (22 inc) 20/26 20/26 0/26 (26 uns)
Kotlin (13 templates, 26 assertions) 6/26 (20 inc) 19/26 n/a 0/26 (26 uns)
Scala (13 templates, 26 assertions) 6/26 (20 inc) n/a n/a n/a
C# (13 templates, 26 assertions) 4/26 (22 inc) 20/26 n/a n/a
Go (13 templates, 26 assertions) 6/26 (20 inc) 19/26 n/a 0/26 (26 uns)
PHP (13 templates, 26 assertions) 6/26 (20 inc) n/a 20/26 0/26 (26 uns)
Ruby (13 templates, 26 assertions) 4/26 (22 inc) 20/26 14/26 0/26 (26 uns)
C++ (12 templates, 24 assertions) 0/24 (24 inc) 14/24 n/a 0/24 (24 uns)
C (9 templates, 18 assertions) 10/18 (8 inc) 14/18 n/a 0/18 (18 uns)
Rust (12 templates, 24 assertions) 6/24 (18 inc) 16/24 16/24 0/24 (24 uns)

The Bifrost v0.10.7 fix cycle

This is the release's largest single movement, and it is confined to one
analyzer column.

  • Runner errors: 22 → 0. The previous freeze retained 22 runner-error
    core results. This freeze retains none, anywhere in the 2884 case results.
    The two v0.4.0 defect classes moved differently, and the difference is the
    honest part:
    • Rust's ten heap and access-path cases — alias-propagation,
      array-element, nested-access-path, object-separation,
      same-object-field — previously failed with the internal_invariant
      "semantic IR gap_contract error … duplicates the same scoped fact"
      signature tracked as bifrost-dev #2638. All ten now decide, and all ten
      decide correctly. Rust's core goes 2/54 → 36/54.
    • The ten element-object cases in Java, JavaScript, Kotlin, Python, and
      TypeScript, and Rust's own element-object pair, previously failed with
      the "invalid value-flow snapshot: oracle relation does not belong to the
      required query arena and role" signature. Every one of those twelve is now
      inconclusive rather than runner-error: the crash is gone, the question
      is still declined. That is a smaller win than the Rust one and is published
      as such — none of the twelve is counted as a correct answer.
  • Decisive-correct roughly doubles: 222 → 435. Across the thirteen kernel
    cores Bifrost now produces 440 decisive outcomes of which 435 are correct
    (v0.4.0: 227 decisive, 222 correct). Inconclusive core results fall from 489
    to 298. Language by language, decided-correct moves: C 2 → 40, Scala 10 → 38,
    Rust 2 → 36, Go 14 → 35, C# 3 → 32, PHP 21 → 32, C++ 2 → 30, Kotlin 25 → 28,
    Ruby 0 → 21. Java, JavaScript, Python, and TypeScript hold their correct
    counts and each gain two inconclusive — the element-object pair that used
    to be a runner error.
  • Ruby decides for the first time. In v0.4.0 Bifrost's Ruby kernel was
    58/58 inconclusive: not one assertion was decisive. It now decides 22
    of 58, 21 of them correctly, with 36 still inconclusive. The
    analyzer-coverage gate docs/applicability-matrix.md records for Ruby, and
    which was tracked upstream as bifrost-dev #2637, has partially lifted; the
    remaining 36 declines are still coverage, and were never 58 misses.
  • Smoke: 115 → 117 of 118. The pinned breadth smoke population now decides
    117 of its 118 cases correctly. The single non-decision is the
    dfb-taint-java-modeled-external calibration case, which takes unsupported
    and is not scored — so Bifrost is 117/117 on the scored smoke partition,
    against 115/117 in v0.4.0. The two v0.4.0 non-decisions, Ruby's
    direct-propagation pair, both now decide correctly.

Two upstream fix sessions and a gap_contract fix landed in this cycle. Only
#2638 and #2637 have a delta this evidence can attribute; the other
issues closed against v0.10.7 are not cited here, because a release note that
attaches an issue number to a movement it cannot demonstrate is doing the thing
this benchmark exists to avoid.

Honest negatives

  • Four new false positives, filed rather than filtered. All four of
    Bifrost's decisive false positives in this freeze are new on v0.10.7, and
    none of them existed in v0.4.0:
    dfb-taint-csharp-infeasible-branch-negative,
    dfb-taint-csharp-loop-carried-negative,
    dfb-taint-go-loop-carried-negative, and
    dfb-taint-ruby-infeasible-branch-negative. They are the classic
    path-feasibility and loop-kill negatives — exactly the cells that punish an
    engine for deciding more. They are filed upstream as
    BrokkAi/bifrost-dev#2731 and are published here, in the tables above, in
    results/, and in the retained raw evidence. The v0.4.0 mismatches they
    replaced (Kotlin's expression false negative, Kotlin's infeasible-branch
    and loop-carried false positives, PHP's infeasible-branch false positive)
    are gone. Net: five decisive mismatches then, five now, on nearly double the
    decisive base.
  • The Java direct-propagation instability persists, unreconciled.
    dfb-taint-java-direct-positive is still reached (true positive) in
    reports/bifrost-smoke.json and not-reached (false negative) in
    reports/bifrost-java-kernel.json, at the same fixture revision and the same
    build 44d9a5be. The two are separate populations with separate scorecards,
    both raw artifacts are retained and digest-bound, and neither was re-run to
    agreement. The freeze publishes what the runs produced, for the second
    release running.
  • Where the declines still concentrate. Bifrost's remaining 298 core
    inconclusive results are not spread evenly. The heaviest kernels are Ruby
    (36 of 58), Kotlin (30 of 58), PHP and C++ (26 each), and C# and TypeScript
    (24 each); the lightest is C (8 of 48). Read as coverage, not as error: these
    are questions the engine declined with partial_discovery or
    capability_incomplete diagnostics retained, and the preregistration's
    reading rule is that a decline is not a miss.

The other three analyzers

Unchanged pins, unchanged populations, and therefore results that reproduce
v0.4.0. They are restated here because this freeze rebinds them at a new
fixture revision, not because they moved.

  • CodeQL 2.26.3 answers every one of its bound assertions definitively —
    zero incomplete outcomes across all seventeen reports. Its systematic
    character on the kernels is unchanged: loop-carried-kill is 11/22 with
    eleven false positives (every negative), reflective-invocation is 8/16 with
    all eight positives missed, chal-dispatch-table 11/22, and
    chal-callback-registration and alias-propagation-separation 12/22 each.
    Under-approximation at run-time-named callees, over-approximation at loop
    kills — approximation character, not a ranking.
  • Joern 4.0.610 covers six kernels and answers every kernel assertion
    definitively. Its widest spread is still Ruby (40/58), and infeasible-branch
    and loop-carried-kill are 6/12 each — all twelve mismatches false positives.
    chal-deep-relay-chain at 7/12 remains the predicted consequence of the
    verified maxCallDepth = 4 default.
  • Semgrep CE 1.174.0 is a bounded-profile population by construction. Of
    its 622 bound kernel-core assertions, 468 are unsupported by declared
    capability, decided from case metadata before Semgrep is invoked, and 154 are
    scored: the intraprocedural partition, identical in all eleven of its
    languages (14 scored per kernel, 12 correct). Its only two mismatched
    templates anywhere are infeasible-branch and loop-carried-kill, eleven
    false positives each. "12/58" is not a low score on a 58-assertion
    population; it is 12 correct of 14 decided, beside 44 declines.
  • language-extension tiers stay outside every core denominator. On C's
    two cases CodeQL is 2/2, and Bifrost now decides one of the two correctly
    (inconclusive on the other, against inconclusive on both in v0.4.0). On
    Rust's Result/? pair CodeQL is 1/2 (one false negative) and Bifrost is
    inconclusive on both.

First publication: the modeling tier (#15)

Twelve preregistered templates in six balanced categories — S (declared
sources and sinks), P (declared propagators), Z (declared sanitizers), O
(opaque procedure summaries), E (framework entry points), B (persistence
boundaries) — across java, javascript, python and all four adapters:
twelve reports, 24 assertions each, 288 case results in total.

The tier's contract, set before any adapter ran, is that a category is scored
for a tool only if that tool's own model-declaration surface can express it and
be made load-bearing
. A tool that cannot express a category takes
unsupported for that category, recorded in advance. The scored partitions
therefore differ per adapter, and the per-adapter denominators below are
never pooled:

Adapter Scored categories Scored templates Per language (each of java / js / python)
CodeQL 2.26.3 S, P, Z, O, E, B — 6 of 6 12 of 12 24/24 correct
Joern 4.0.610 S, Z, E, B — 4 of 6 (P and O declined by Amendment A2) 8 of 12 14/16 correct, 2 false negatives, 8 unsupported
Semgrep CE 1.174.0 S, E, and half of Z — sanitizer-selectivity declined by Amendment A3 5 of 12 10/10 correct, 14 unsupported
Bifrost 0.10.7 S, Z — 2 of 6, after Amendment A9 4 of 12 see below
  • CodeQL is the only adapter that enters all six categories, and it answers
    all 72 of its modeling assertions correctly across the three languages. That
    is the tier working as designed: the engine with the richest model surface
    has the most to be measured on, and it is measured on all of it.
  • Joern's two false negatives per language are both category B, the
    store-roundtrip and store-separation positives — the persistence boundary
    is declared, the negative is held, the positive is not reached. Its P and O
    declines are Amendment A2, recorded on a measurement that FlowSemantic is
    not load-bearing in the only direction that would decide them.
  • Semgrep's five templates are all correct. Its
    model-sanitizer-selectivity decline is Amendment A3: the mandated
    safe-function assumption makes the cell undecidable by construction, which is
    a property of the activation contract, not a failure.
  • Bifrost enters with two of six categories, and this is the first freeze
    in which its category Z is scored at all. Amendment A9 withdrew the
    preregistered unsupported for Z on a measurement that contradicted
    Bifrost's own adapter README: the analysis grammar does accept a
    (sanitizer …) stanza, the declaration suppresses on a completing run, its
    removal restores the flow with a full witness, and an undeclared
    sanitizer-shaped sibling is not suppressed. The README sentence "Sanitizer
    lowering is a future Bifrost CLI capability"
    was wrong. DataFlowBench is
    published by Bifrost's vendor, so the direction of that correction matters:
    A9 moved a category toward our own engine, on evidence, having originally
    recorded it against our own engine on the vendor's own documentation.
    Both halves of that are in the record.

Bifrost's first scored modeling run, read cell by cell:

  • Category Z decides 12 of 12 and all 12 are correctsanitizer-kill and
    sanitizer-selectivity, both polarities, in all three languages. The
    category A9 promoted is the category that answers cleanly.
  • Category S is 9 of 12 correct with 3 inconclusive and no mismatch. Java
    and python are 4/4; javascript declines the declared-sink pair and the
    declared-source positive.
  • Overall: 24 scored assertions, 21 decisive, 21 correct, 3 inconclusive, 0
    mismatches
    , beside 48 unsupported in the four categories Bifrost does not
    enter. Twenty-one of twenty-four is not a claim about Bifrost's modeling
    relative to CodeQL's 72 of 72 — the denominators are different populations by
    construction, and the whole point of publishing the scored partition is
    that "2 of 6 categories" is the load-bearing number, not the ratio inside it.

First publication: the tool-native tier (#16)

Six templates — source-sink, propagator, sanitizer, summary,
entrypoint, persistence — across java, javascript, python and all four
adapters, under the tool-native model profile: twelve reports, 12 assertions
each, 144 case results. The question is not "how good is the engine" but "what
does the shipped product decide, with nothing supplied by us".

Adapter Java JavaScript Python
CodeQL 2.26.3 (shipped *-security-extended.qls) 11/12 (1 FP) 9/12 (3 FP, 1 FN) 10/12 (2 FP)
Semgrep CE 1.174.0 (vendored semgrep-rules snapshot) 0/12 (12 uns) 0/12 (12 uns) 8/12 (4 FP)
Joern 4.0.610 (DefaultSemantics only) 0/12 (12 uns) 0/12 (12 uns) 0/12 (12 uns)
Bifrost 0.10.7 (built-in policy packs) 0/12 (12 uns) 0/12 (12 uns) 0/12 (12 uns)

Reading each row on its own terms:

  • CodeQL is the only adapter with a shipped model that decides this tier
    across all three languages.
    Its shipped suites answer all 36 assertions
    definitively and get 30 right. The one template it false-positives in every
    language is native-persistence; javascript additionally false-positives
    native-sanitizer and misses the native-persistence positive.
  • Semgrep's vendored snapshot is sharply asymmetric across languages, and
    that asymmetry is the result. Python decides all twelve, off a shipped audit
    rule that Amendment A8 promoted the column on: all six positives are
    reached, and four of six negatives are false positives — a broad rule
    behaving broadly. Java and javascript decline all twelve against the same
    pinned snapshot, recorded as Amendment A7 (java) and Amendment A6
    (javascript). One vendor, one pinned ruleset, three languages, three
    different answers.
  • Joern and Bifrost decline all twelve in every language, with their run
    identity witnessed anyway.
    Joern's DefaultSemantics ships no source or
    sink endpoints for this tier; Bifrost's built-in policy packs ship no taint
    policy and no endpoint catalog (bifrost-dev #2620, open). Amendment
    A10
    restates Bifrost's category-Z native cell on the grounds that survive
    A9: the sanitizer stanza A9 measured is reachable only through
    --policy-file, which this profile's activation contract forbids, and a
    barrier on a flow that cannot start is unobservable either way. A9 and A10
    are consistent — the same capability is present under the benchmark-controlled
    profile and out of contract under the native one.

Zero of twelve is not a score of zero. All twelve are unsupported, which
is declared coverage, and none is counted as an incorrect answer anywhere in
results/.

Amendments recorded this cycle

  • A9 (2026-08-27) — Bifrost's sanitizer category is promoted; the README's
    lowering claim was false.
    Category Z moves from preregistered unsupported
    to scored under the benchmark-controlled modeling profile, on measured
    evidence. Recorded in docs/modeling-matrix.md.
  • A10 (2026-08-28) — Bifrost's native category-Z cell is restated on the
    absent endpoint catalog.
    The tool-native Z cell keeps its unsupported
    outcome, but on grounds that survive A9: no shipped endpoint catalog, and the
    measured stanza out of the native profile's activation contract. Recorded in
    docs/native-profile.md.
  • The identity-witnessing correction. Adapters now witness the identity of
    the binary actually invoked and the ruleset or pack actually activated, and
    they do so even on a run that decides nothing. Every 0/12 native row in
    this freeze carries a full build identity in the manifest — a run that cannot
    witness its own pin has nothing truthful to write, and "we declined
    everything" is a claim that needs its provenance as much as any finding does.

Amendments A1 through A8 remain in force and are unaffected. No amendment in
this cycle invalidates a published freeze.

Reproduction

git checkout v0.5.0                # frozen benchmark revision (evidence commit)
# then check out the merge commit carrying reports/freeze.json for this release
cargo run -- validate-freeze reports/freeze.json
cargo run -- generate-results --manifest reports/freeze.json --output-directory results --check

Adapter re-execution (produces new evidence, therefore a new freeze):

cargo run -- run-bifrost-smoke --bifrost <bifrost-binary>
for kernel in java javascript typescript python kotlin scala csharp go php ruby c cpp rust; do
  cargo run -- run-bifrost-$kernel-kernel --bifrost <bifrost-binary>
done
for language in java javascript python; do
  cargo run -- run-bifrost-modeling --language $language --bifrost <bifrost-binary>
  cargo run -- run-bifrost-native --language $language --bifrost <bifrost-binary>
done

codeql pack install adapters/codeql
for pack in javascript typescript python kotlin csharp go cpp rust ruby; do
  codeql pack install adapters/codeql/$pack
done
for kernel in java javascript typescript python kotlin csharp go ruby c cpp rust; do
  cargo run -- run-codeql-$kernel-kernel --codeql <codeql-binary>
done
for language in java javascript python; do
  cargo run -- run-codeql-modeling --language $language --codeql <codeql-binary>
  cargo run -- run-codeql-native --language $language --codeql <codeql-binary>
done

for kernel in java javascript python php ruby rust; do
  cargo run -- run-joern-$kernel-kernel --joern <joern-cli-directory>
done
for language in java javascript python; do
  cargo run -- run-joern-modeling --language $language --joern <joern-cli-directory>
  cargo run -- run-joern-native --language $language --joern <joern-cli-directory>
done

for kernel in java javascript typescript python kotlin go php ruby c cpp rust; do
  cargo run -- run-semgrep-$kernel-kernel --semgrep <semgrep-binary>
done
for language in java javascript python; do
  cargo run -- run-semgrep-modeling --language $language --semgrep <semgrep-binary>
  cargo run -- run-semgrep-native --language $language --semgrep <semgrep-binary>
done

The Kotlin and Go CodeQL runners trace a real compile, so kotlinc and the Go
toolchain must be available; the Rust CodeQL runner uses the CLI's public
preview Rust extractor. Joern's php2cpg shells out to its bundled
PHP-Parser, so a host php interpreter must be on PATH, and its rust2cpg
frontend materializes each case as a minimal Cargo crate. The tool-native
Semgrep rows run against the pinned semgrep-rules snapshot vendored under
adapters/semgrep/native/, not against a live registry fetch.

Immutability

This snapshot is immutable. Corrected evidence creates a new freeze with a new
release name and digests; the v0.1.0, v0.2.0, v0.3.0, and v0.4.0 manifests and
evidence remain available at their release commits for audit. This release has
no DOI and no external review claim.