Skip to content

Sciev v0.2.0 — aligned scientific decision pipeline

Choose a tag to compare

@alrobles alrobles released this 27 Sep 01:24
· 20 commits to main since this release

Sciev v0.2.0 — aligned scientific decision pipeline

Typed decision heads (choice/noul/score) over frozen LLaDA-8B, trained
and evaluated under the systemone-v2 contract: group-disjoint splits,
explicit overflow handling (max_ctx=960, no silent truncation), shared
R2 preprocessing across train/calibrate/eval/serve, dev-only temperature
fitting, and provenance manifests for every file.

Highlights

  • Matched frozen/adapted comparison (identical recipe + seed + data):
    mean±sd over seeds 7331–7333. Domain adaptation robustly improves
    out-of-distribution verification (SciFact noul 0.709 vs 0.483, +22.6pt)
    at a noisy cost to in-distribution rubric scoring — task-dependent, not
    a uniform win/loss.
  • Evidence controls (empty/shuffle, agreement-with-reference): expose
    option-side leakage — choice is ~80–87% solvable without reading the
    passage; score is the task whose accuracy most directly reflects
    evidence use.
  • Hard-distractor battery (--hard-choice): strict same-passage +
    numeric-perturb pool reduces leakage ~6–8pt but does not remove it.

Contents

  • heads/release/ — fr_matched_s7331 heads (recommended; frozen backbone)
  • heads/experimental/ — da_matched_s7331 heads (DAPT arm; documented caveats)
  • evals/ — 32 evaluation JSONs (arms × tasks × benchmarks × controls)
  • manifests/ — dataset manifests incl. exclusion provenance and input hashes
  • sciev-v0.2-paper.pdf, MODEL_CARD.md, LICENSE, SHA256SUMS,
    release-manifest.json

Checkpoints are self-describing (inference config, training-data
contract, input fingerprints, recipe, seed). Controls are scored as
agreement-with-reference, never relabelled.