Sciev v0.2.0 — aligned scientific decision pipeline
Sciev v0.2.0 — aligned scientific decision pipeline
Typed decision heads (choice/noul/score) over frozen LLaDA-8B, trained
and evaluated under the systemone-v2 contract: group-disjoint splits,
explicit overflow handling (max_ctx=960, no silent truncation), shared
R2 preprocessing across train/calibrate/eval/serve, dev-only temperature
fitting, and provenance manifests for every file.
Highlights
- Matched frozen/adapted comparison (identical recipe + seed + data):
mean±sd over seeds 7331–7333. Domain adaptation robustly improves
out-of-distribution verification (SciFact noul 0.709 vs 0.483, +22.6pt)
at a noisy cost to in-distribution rubric scoring — task-dependent, not
a uniform win/loss. - Evidence controls (empty/shuffle, agreement-with-reference): expose
option-side leakage — choice is ~80–87% solvable without reading the
passage; score is the task whose accuracy most directly reflects
evidence use. - Hard-distractor battery (
--hard-choice): strict same-passage +
numeric-perturb pool reduces leakage ~6–8pt but does not remove it.
Contents
heads/release/— fr_matched_s7331 heads (recommended; frozen backbone)heads/experimental/— da_matched_s7331 heads (DAPT arm; documented caveats)evals/— 32 evaluation JSONs (arms × tasks × benchmarks × controls)manifests/— dataset manifests incl. exclusion provenance and input hashessciev-v0.2-paper.pdf,MODEL_CARD.md,LICENSE,SHA256SUMS,
release-manifest.json
Checkpoints are self-describing (inference config, training-data
contract, input fingerprints, recipe, seed). Controls are scored as
agreement-with-reference, never relabelled.