Skip to content

DataFlowBench v0.4.0

Choose a tag to compare

@DavidBakerEffendi DavidBakerEffendi released this 25 Aug 23:09
· 86 commits to main since this release
306211a

DataFlowBench v0.4.0

Fourth immutable release snapshot, and the first expanded-breadth release:
all thirteen core kernels now carry the preregistered challenge-tier templates
of the challenge-tier document, and four analyzers —
Bifrost v0.10.6, CodeQL 2.26.3, Joern 4.0.610, and Semgrep CE 1.174.0 — are
bound at one fixture revision under the freeze/v1 contract with claim scope
release.

Freeze identity

  • Freeze ID (manifest SHA-256):
    91b0008a546e6b782c1b790f174a71ce44e60039239797674eb49ebc6ac6c366
  • Manifest: reports/freeze.json
  • Benchmark revision: 306211a (tag v0.4.0)
  • Fixture revision:
    sha256:13a11ff48f26dba889f76aeb9ef60213a129abe5ebcfcb966da3a2418c12807e
  • Case schema v2, normalized result schema v1
  • 744 frozen cases, 42 bound reports, 2452 scored case results

What "expanded breadth" means here

The thirteen challenge templates were preregistered before any challenge
fixture existed, and the preregistration fixed the population decision in
advance: they carry score_tier: "core" and fold into each language's core
kernel
, with no new score tier. Each language's core denominator is therefore
its sixteen-template core (fifteen for C and Rust) plus its applicable
challenge templates: 29 templates / 58 assertions for ten languages, 28 / 56
for C++, 27 / 54 for Rust, and 24 / 48 for C.

The canonical construction of dfb-template-chal-context-pair-depth2 follows
Amendment A1 (2026-08-25) of the preregistration: helper returns its
argument and the caller sinks the result of the selected two-deep path. The
amendment was recorded before any analyzer ran against either implementing
fixture and invalidates no published freeze.

The v0.3.0 sixteen-template core and this expanded core are different
populations of the same name
. No number in this document is compared with a
v0.3.0-era number, and the frozen v0.3.0 evidence remains valid and unamended.

Bound evidence

Normalized report SHA-256 Analyzer Cases
reports/bifrost-c-kernel.json e334de9b8752daf1b0ed67bc8403232449c10805c4817be3975b2ca62643ba03 Bifrost 0.10.6, build 18d09c57d1e5044dec49acac7635d3255ea8e89c 50
reports/bifrost-cpp-kernel.json 12353998cacf0bf7c3e74574961d0eaec9204da633a6cdf87d2c36526bd71e27 Bifrost 0.10.6, build 18d09c57d1e5044dec49acac7635d3255ea8e89c 56
reports/bifrost-csharp-kernel.json a004277d218cf4793afed987394c0d6285ecd190622e23558368f2ba0ec2eeb2 Bifrost 0.10.6, build 18d09c57d1e5044dec49acac7635d3255ea8e89c 58
reports/bifrost-go-kernel.json a3597db4e3ccdb2489b720eca08173100cb3220ee5db31743bf0d7a64708c46d Bifrost 0.10.6, build 18d09c57d1e5044dec49acac7635d3255ea8e89c 58
reports/bifrost-java-kernel.json d8ad10a3cdca6d0be1e207dcbdf2b276dd46caeeb2f98d43352f003082113b0d Bifrost 0.10.6, build 18d09c57d1e5044dec49acac7635d3255ea8e89c 58
reports/bifrost-javascript-kernel.json 9a4f3c8c5a319dff7e5d5dd5f159329c2580997d742ae8f75c86176a99a1096f Bifrost 0.10.6, build 18d09c57d1e5044dec49acac7635d3255ea8e89c 58
reports/bifrost-kotlin-kernel.json dad5d772414d41644b3d8892fd3495f86974054b90a8c678d4e63364b17678f0 Bifrost 0.10.6, build 18d09c57d1e5044dec49acac7635d3255ea8e89c 58
reports/bifrost-php-kernel.json 5881b8e17f332d344f1539b79d6f1833abc1dde245302fd9ebdacb84cc14169a Bifrost 0.10.6, build 18d09c57d1e5044dec49acac7635d3255ea8e89c 58
reports/bifrost-python-kernel.json 87f9a195df465554734796412867f5b32721200ee74863fc9d1ff3b13d6ca3fe Bifrost 0.10.6, build 18d09c57d1e5044dec49acac7635d3255ea8e89c 58
reports/bifrost-ruby-kernel.json 7af673389c1146294684f2154fb8ec13c39a587a21e8b381d46119f51eb49aa1 Bifrost 0.10.6, build 18d09c57d1e5044dec49acac7635d3255ea8e89c 58
reports/bifrost-rust-kernel.json 2ad1e552b00c7c6370b26d948e2e5e850f38adeff27a87b6b1f72aefa15f4119 Bifrost 0.10.6, build 18d09c57d1e5044dec49acac7635d3255ea8e89c 56
reports/bifrost-scala-kernel.json 11c7c6462f6d7cae9d9042540014294d66909c9566d02d68c7cc2513ab5298dc Bifrost 0.10.6, build 18d09c57d1e5044dec49acac7635d3255ea8e89c 58
reports/bifrost-smoke.json bfbd71c8ea921f71eacae6983ac45361edb0264be0ef4ded17cb17449dc880f9 Bifrost 0.10.6, build 18d09c57d1e5044dec49acac7635d3255ea8e89c 118
reports/bifrost-typescript-kernel.json 6d1ec2fce985b22def1de1d32ae49049ea5be0fb6b52a8ae024adbdac9110e1c Bifrost 0.10.6, build 18d09c57d1e5044dec49acac7635d3255ea8e89c 58
reports/codeql-c-kernel.json 0b6c59ac6e4435e049a45972d297d665b55eba07fefae98535930a01543b7f0c CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 50
reports/codeql-cpp-kernel.json 6e772fe6740133ee4a0b20a9683d145b832f3ede088c535936ecdaf2f6a802ff CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 56
reports/codeql-csharp-kernel.json 3436137a9aa293bbf0efb62e263a90b4c16e47ede83cb45602cf7c229fb31fed CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 58
reports/codeql-go-kernel.json 7812a935ee53a26fab3a7f3b1d74c169e5ae98a4c3d1e0c6b02a09496612aa93 CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 58
reports/codeql-java-kernel.json bad0469701d0f3c45825cc4ee8d0448bdbec40e9006cf78935112e09e77e08db CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 58
reports/codeql-javascript-kernel.json 98d8064493ddfbbee98e63f855cba6c6dc3eb7ae85ee6405ad0e0ab0167fc045 CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 58
reports/codeql-kotlin-kernel.json bada2d8ddf781ff12b569fd240f4b94014f647d71f5fb1d653308d5a5995858f CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 58
reports/codeql-python-kernel.json 8cad62ca0ae9206f172ab9cdbdbe95a8dc72d5af62d9624a97e7afb491eafea7 CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 58
reports/codeql-ruby-kernel.json 09358d58fb97df1bc2024545ec56438d23989ad4c3a2c13a4a0e7c407a450858 CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 58
reports/codeql-rust-kernel.json 30882a1c1aab9919ff484a89166f10ded6050e9b3ffef37f1568c6cce605195f CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 56
reports/codeql-typescript-kernel.json ffe51480b9c3ab103e67484e5a86e4ba1c35c15d2ffaca6ff11831d7694a501e CodeQL 2.26.3, build codeql-cli:7d097a43199effe04ecd9c6bd3ad9bb02a45b3d7 58
reports/joern-java-kernel.json 7e3c4cb6adbe7325bb5f9f12a62f0d0176faacd5825ce5c9f32c7f67a03c1e85 Joern 4.0.610, build joern-cli:4.0.610 58
reports/joern-javascript-kernel.json d794fbb8d72cb5d836ffe45d8caf421113931b9bc376424c62b57f753b820dfb Joern 4.0.610, build joern-cli:4.0.610 58
reports/joern-php-kernel.json 1638b81a764de42658700e1f3deeeebf808055b2582db4e21c65c8ed0b7ddab5 Joern 4.0.610, build joern-cli:4.0.610 58
reports/joern-python-kernel.json 0facb3aa7e4e2855f04c356c103beb9f1a84b4879c41f7b826747d7d25c637f4 Joern 4.0.610, build joern-cli:4.0.610 58
reports/joern-ruby-kernel.json 0d1bd231bae8ffec048e0cd7cbe83d128e7cf3c0cb49a3a1faafea3a10960b3c Joern 4.0.610, build joern-cli:4.0.610 58
reports/joern-rust-kernel.json cf346a6c0bda2dc3b6cd6af1f24bc2c4b895fdea0cce7c85f4b2c7998ff54351 Joern 4.0.610, build joern-cli:4.0.610 54
reports/semgrep-c-kernel.json a75b5f37004166d20de264ee95ba7c6f4905ab0cc9e82c47aeed69ee35f4e4c5 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 48
reports/semgrep-cpp-kernel.json ef9b23c9a1fa0764ffbe0db62e80c9bd5986374a24ecf7db702b221661d5445d Semgrep CE 1.174.0, build semgrep-oss:1.174.0 56
reports/semgrep-go-kernel.json bab1bcdeef91fc8fbaece22857db61efac141793573046d13a06ee9921889426 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 58
reports/semgrep-java-kernel.json 60ccdc90635606d8eaa150d0823c626ea5b3cf69963f6575f324fb020287e2bb Semgrep CE 1.174.0, build semgrep-oss:1.174.0 58
reports/semgrep-javascript-kernel.json df957a71497cec253003f0c0a1941d9bc13a8424d0d9bb9058e079f11259d0c9 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 58
reports/semgrep-kotlin-kernel.json e1e23bfcf109af1e7345649be232de413e7f1e655513947ecb0224d9bbaefba6 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 58
reports/semgrep-php-kernel.json 22bc4cf3c0d3192edf1391eef4e2c9e0e4b3e69fbe9769737df397dd8c58e695 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 58
reports/semgrep-python-kernel.json 9eb2778656f6f46c95181fca43832e0c1f8ed03ce7e85b3e7ded51a80bc45137 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 58
reports/semgrep-ruby-kernel.json 6f3b58ae53d16816b90076d449361e75728b20af8f9866c0e1f37ce12874b279 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 58
reports/semgrep-rust-kernel.json 1739e7212692082488ac9ef6f3661a46dce3f9536f708b41113cb0b7574a3e89 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 54
reports/semgrep-typescript-kernel.json 50a0635d83a7d24c6ae71b12b1e0d066ff1262b35d1934d5bc3827d8af97fba3 Semgrep CE 1.174.0, build semgrep-oss:1.174.0 58

Every report uses the benchmark-controlled model profile on the taint track
and remains its own scorecard. The Bifrost smoke population is the pinned
118-case breadth slice — challenge templates are excluded from the smoke
selection outright, so the smoke population did not grow — and every other
report is a single-language kernel population. The C and Rust kernel reports
each carry two language-extension cases in addition to their core tier;
those stay on their own tier and are never folded into a core denominator.
Raw evidence for every result is retained under reports/raw/ and
digest-bound in the manifest.

Results

Generated pages live in results/ and derive exclusively from the frozen
manifest; see results/index.md. Incomplete outcomes (inconclusive,
unsupported, runner-error) are capability and execution coverage and are
never counted as clean negatives.

Correct decisions (true positives plus true negatives) on each kernel's core
tier, with incomplete outcomes shown beside them. Each language is its own
population with its own denominator, and each analyzer column is read
independently: DataFlowBench publishes no combined leaderboard, and cores of
different sizes are never pooled. n/a means the analyzer has no report for
that kernel in this freeze — no extractor, no frontend, or no adapter — which
is coverage, not a score.

Kernel core Bifrost 0.10.6 CodeQL 2.26.3 Joern 4.0.610 Semgrep CE 1.174.0
Java (29 templates, 58 assertions) 37/58 (18 inc, 2 err) 48/58 47/58 12/58 (44 uns)
JavaScript (29 templates, 58 assertions) 36/58 (20 inc, 2 err) 48/58 44/58 12/58 (44 uns)
TypeScript (29 templates, 58 assertions) 34/58 (22 inc, 2 err) 48/58 n/a 12/58 (44 uns)
Python (29 templates, 58 assertions) 36/58 (20 inc, 2 err) 48/58 48/58 12/58 (44 uns)
Kotlin (29 templates, 58 assertions) 25/58 (28 inc, 2 err) 46/58 n/a 12/58 (44 uns)
Scala (29 templates, 58 assertions) 10/58 (48 inc) n/a n/a n/a
C# (29 templates, 58 assertions) 3/58 (55 inc) 47/58 n/a n/a
Go (29 templates, 58 assertions) 14/58 (44 inc) 45/58 n/a 12/58 (44 uns)
PHP (29 templates, 58 assertions) 21/58 (36 inc) n/a 48/58 12/58 (44 uns)
Ruby (29 templates, 58 assertions) 0/58 (58 inc) 49/58 40/58 12/58 (44 uns)
C++ (28 templates, 56 assertions) 2/56 (54 inc) 42/56 n/a 12/56 (42 uns)
C (24 templates, 48 assertions) 2/48 (46 inc) 41/48 n/a 12/48 (34 uns)
Rust (27 templates, 54 assertions) 2/54 (40 inc, 12 err) 44/54 43/54 12/54 (40 uns)

Nobody is perfect on the expanded kernels. The saturation the preregistration
set out to end — a top scorer answering every question correctly — is gone: no
analyzer answers a whole expanded core correctly in any of the thirteen
languages, and no column above reaches its own denominator.

The challenge templates fold into those core denominators, but they stay
individually visible as a stratum, exactly as the preregistration requires.
The same populations, restricted to the challenge templates only (13 templates
/ 26 assertions, 12 / 24 for C++ and Rust, 9 / 18 for C):

Kernel core Bifrost 0.10.6 CodeQL 2.26.3 Joern 4.0.610 Semgrep CE 1.174.0
Java (13 templates, 26 assertions) 6/26 (18 inc, 2 err) 21/26 19/26 0/26 (26 uns)
JavaScript (13 templates, 26 assertions) 4/26 (20 inc, 2 err) 19/26 18/26 0/26 (26 uns)
TypeScript (13 templates, 26 assertions) 4/26 (20 inc, 2 err) 19/26 n/a 0/26 (26 uns)
Python (13 templates, 26 assertions) 4/26 (20 inc, 2 err) 20/26 20/26 0/26 (26 uns)
Kotlin (13 templates, 26 assertions) 6/26 (18 inc, 2 err) 19/26 n/a 0/26 (26 uns)
Scala (13 templates, 26 assertions) 0/26 (26 inc) n/a n/a n/a
C# (13 templates, 26 assertions) 1/26 (25 inc) 20/26 n/a n/a
Go (13 templates, 26 assertions) 4/26 (22 inc) 19/26 n/a 0/26 (26 uns)
PHP (13 templates, 26 assertions) 4/26 (22 inc) n/a 20/26 0/26 (26 uns)
Ruby (13 templates, 26 assertions) 0/26 (26 inc) 20/26 14/26 0/26 (26 uns)
C++ (12 templates, 24 assertions) 0/24 (24 inc) 14/24 n/a 0/24 (24 uns)
C (9 templates, 18 assertions) 0/18 (18 inc) 14/18 n/a 0/18 (18 uns)
Rust (12 templates, 24 assertions) 0/24 (20 inc, 4 err) 16/24 16/24 0/24 (24 uns)

Reading the four populations, each on its own terms:

  • Bifrost 0.10.6 decides 115 of the 118 cases in the pinned breadth
    smoke population correctly; the three non-decisive results are Ruby's
    direct-propagation pair (inconclusive) and the modeled-external Java
    calibration case (unsupported, and not scored). Across the thirteen
    expanded kernels it produces 227 decisive outcomes, of which 222 are
    correct
    : the five decisive mismatches are one false negative on Java's
    direct-propagation positive, three on Kotlin (an expression false
    negative plus infeasible-branch and loop-carried false positives), and
    one PHP infeasible-branch false positive. Everything else it does not
    answer it declines: 489 inconclusive core results retaining
    partial_discovery or capability_incomplete diagnostics. On the challenge
    strata specifically it decides 33 of 326 assertions — the depth-6 relay, the
    two-level context pair, and Java's and Kotlin's recursive carry — and all
    33 are correct
    . Per the preregistration's own reading rule, correct
    stratum-D negatives beside undecided positives describe a bound, not
    precision.
  • Bifrost's two published defects. 22 core results are runner-error and
    are retained verbatim as execution coverage. Ten of them are the
    element-object pair in Java, JavaScript, Kotlin, Python and TypeScript,
    where the run fails with internal_invariant and "invalid value-flow
    snapshot: oracle relation does not belong to the required query arena and
    role". The other twelve are Rust's heap and access-path cases, failing with
    a different signature — internal_invariant, "semantic IR gap_contract
    error … duplicates the same scoped fact". They are tracked upstream as
    bifrost-dev #2639 (element-object) and #2638 (Rust gap_contract). Bifrost's Ruby kernel is 58/58 inconclusive: no
    assertion is decisive, which is the analyzer-coverage gate
    docs/applicability-matrix.md already records for Ruby, now measured over
    the whole expanded population; it is tracked upstream as bifrost-dev #2637
    and is never counted as 58 misses.
  • An observed instability, published as observed. The case
    dfb-taint-java-direct-positive is reached (true positive) in
    reports/bifrost-smoke.json and not-reached (false negative) in
    reports/bifrost-java-kernel.json, at the same fixture revision and the
    same build. The two are separate populations with separate scorecards, and
    both raw artifacts are retained and digest-bound. Neither result was
    re-run to agreement: the freeze publishes what the runs produced.
  • CodeQL 2.26.3 brings the first challenge-tier evidence across eleven
    languages, with zero incomplete outcomes anywhere: every one of its 626
    bound assertions gets a definitive answer, 509 of them correct (506 of the
    622 core assertions, plus 3 of the 4 language-extension assertions). That
    freeze-wide tally is a count of bound evidence, not a score — the eleven
    populations behind it have different denominators and are never pooled.
    Per kernel the results run from 42/56 on C++ and 41/48 on C to 49/58 on
    Ruby, and on the challenge strata from 14/24 (C++) and 14/18 (C) to 21/26
    (Java) — read one denominator at a time, never as a sequence. Its systematic miss is dynamic dispatch: across all
    eleven languages the reflective-invocation pairs score 8/16 — every
    negative correct, every one of the eight positives missed — and the
    dispatch-table pairs 11/22, where ten of the eleven positives are missed.
    That is an under-approximating refusal to follow a callee named at run time,
    which stratum A was written to make visible as approximation character
    rather than as a ranking. At the other end, deep-relay-chain and
    recursive-carry are 22/22 each.
  • Joern 4.0.610 covers six kernels and answers every assertion
    definitively, 270 of 344 correctly. Its per-language character is the story
    rather than any single number: Python and PHP at 48/58, Java at 47/58,
    JavaScript at 44/58 (6 false positives, tied with Ruby), Rust at 43/54,
    and Ruby at 40/58, where 12 false negatives and 6 false positives make it
    the frontend with the widest approximation spread. Stratum D behaves as the
    preregistration predicted from the verified maxCallDepth = 4 default:
    five of the six kernels miss the depth-6 relay positive while answering its
    negative correctly. Ruby departs from that pattern — it is the one
    frontend that reaches the depth-6 positive, and it is also the one that
    false-positives the recursive-carry negative. Both readings belong
    together: a bound that does not bind here comes with a widening that does
    not kill.
  • Semgrep CE 1.174.0 is a bounded-profile population by construction. Of
    its 622 bound core assertions, 468 are unsupported by declared
    capability
    , decided from case metadata before Semgrep is invoked, and
    154 are scored — the intraprocedural partition, identical in every one
    of its eleven languages (14 scored per kernel, 12 correct, and the same two
    false positives everywhere: the infeasible-branch and loop-carried-kill
    negatives). All 274 challenge assertions in its populations are declined. The preregistration said in
    advance that this would happen and why: an engine that documents a construct
    as out of scope takes unsupported, and that is correct behavior, not a
    gap. "12/58" is not a low score on a 58-assertion population; it is 12
    correct of 14 decided, beside 44 declines.
  • language-extension tiers stay outside every core denominator. On C's
    two cases CodeQL is 2/2 and Bifrost is inconclusive on both; on Rust's
    Result/? pair CodeQL is 1/2 (one false negative) and Bifrost is
    inconclusive on both.

No analyzer is declared a winner. The populations, model profiles, coverage,
and denominators differ per language and per analyzer; incomplete outcomes are
coverage evidence rather than incorrect answers; and stratum A results are
approximation character, not a ranking.

Reproduction

git checkout v0.4.0                # frozen benchmark revision (evidence commit)
# then check out the merge commit carrying reports/freeze.json for this release
cargo run -- validate-freeze reports/freeze.json
cargo run -- generate-results --manifest reports/freeze.json --output-directory results --check

Adapter re-execution (produces new evidence, therefore a new freeze):

cargo run -- run-bifrost-smoke --bifrost <bifrost-binary>
for kernel in java javascript typescript python kotlin scala csharp go php ruby c cpp rust; do
  cargo run -- run-bifrost-$kernel-kernel --bifrost <bifrost-binary>
done

codeql pack install adapters/codeql
for pack in javascript typescript python kotlin csharp go cpp rust ruby; do
  codeql pack install adapters/codeql/$pack
done
for kernel in java javascript typescript python kotlin csharp go ruby c cpp rust; do
  cargo run -- run-codeql-$kernel-kernel --codeql <codeql-binary>
done

for kernel in java javascript python php ruby rust; do
  cargo run -- run-joern-$kernel-kernel --joern <joern-cli-directory>
done

for kernel in java javascript typescript python kotlin go php ruby c cpp rust; do
  cargo run -- run-semgrep-$kernel-kernel --semgrep <semgrep-binary>
done

The Kotlin and Go CodeQL runners trace a real compile, so kotlinc and the Go
toolchain must be available; the Rust CodeQL runner uses the CLI's public
preview Rust extractor. Joern's php2cpg shells out to its bundled
PHP-Parser, so a host php interpreter must be on PATH, and its rust2cpg
frontend materializes each case as a minimal Cargo crate.

Immutability

This snapshot is immutable. Corrected evidence creates a new freeze with a new
release name and digests; the v0.1.0, v0.2.0, and v0.3.0 manifests and evidence
remain available at their release commits for audit. This release has no DOI
and no external review claim.