Skip to content

Balanced DOOP corpus: 21 measured applications

Pre-release
Pre-release

Choose a tag to compare

Balanced, measured DOOP application corpus

21 distinct applications: 6 small / 5 medium / 6 large / 4 xlarge, versus the original 5/2/4/1. All thresholds are unchanged. The catalog now consistently uses canonical SRDatalog CPU VarPointsTo counts, not mixed upstream counts or input size.

New applications: Bloat 11,227,250; Chart 17,221,612; Clojure 16,591,860; javac 24,163,325; JRuby 60,512,570; PDFBox 69,315,812; Soot 412,802,921; Scala 612,741,889; Kotlin 1,076,040,872 VPT rows. Nine were selected from sixteen newly measured application families; no duplicated versions, multiplied facts, sampled rows or synthetic roots.

Assets include all original raw fact relations for each new application, per-application extraction provenance, full CPU count/report and input-fingerprint evidence in measurements.json, candidate-selection audit, extraction logs, preparation-verification.json, and SHA 256SUMS. The original twelve public archives remain at their unchanged pinned HuggingFace URLs in the catalog.

Every selected dataset completed the unchanged canonical CPU fixedpoint, with 74 relation counts and 37 full IDB exports. Zero warmups/one repeat: cardinality and completion evidence, NOT comparative timing or GPU-equivalence evidence. The portable raw archive preparation is checked against the immutable CPU-validated input hashes rather than rewriting old CPU report hashes.

DOOP 4.24.9/java_8, pinned image and full recorded application/dependency artifacts. Static extraction has reflection and phantom limitations: Kotlin retains 23 phantom methods and 3 phantom-based methods after dependency completion; Scala also retains phantom-based methods. See provenance and extraction logs; no ignored-error flags or synthetic stand-ins were used. Failed extraction attempts are not datasets.

Recorded absolute paths identify original evidence and require mount-path remapping for regeneration. Upstream applications and libraries retain their own licenses; no relicensing of third-party inputs is asserted. Facts/binaries are release assets, not Git source files.