Skip to content

Sciev v0.2.3 — docs and eval-infra alignment

Latest

Choose a tag to compare

@alrobles alrobles released this 28 Sep 02:51
· 2 commits to main since this release

Sciev v0.2.3 — documentation and evaluation-infra alignment.

Weights, datasets (systemone-v2) and all reported metrics are identical
to v0.2.2; head checksums were re-verified byte-for-byte against the
published v0.2.2 SHA256SUMS. This release aligns the package and public
documentation with the verified matched study (three head seeds,
evidence controls) and the manuscript prepared for arXiv submission.

Package changes since v0.2.2:

  • sciev.eval: --predictions emits per-item prediction records
  • sciev.train: --steps overrides the recipe's default training budget
  • new eval/baseline scripts under scripts/ (r2_*, eval_cbjev.py,
    eval_llada_prompted.py, train_encoder_scifact.py, fit_dev_temp.py)
  • MODEL_CARD rewritten for fr_* with evidence controls and the
    elite-v1 truncation-defect disclosure; README aligns public claims

Contents:

  • fr_{choice,noul,score}.pt — frozen-backbone release heads (v0.2 bits)
  • da_{choice,noul,score}.pt — experimental arm (caveats in the paper)
  • evals.tar.gz — archived 32-report evaluation set (same as v0.2.2)
  • evidence.json — compact ledger of the complete 60-report study with
    SHA-256 fingerprints, per-seed metrics, dataset and checkpoint metadata
  • all-manifests.txt, hard-controls-manifests.txt — dataset manifests
  • sciev-0.2.3-{py3-none-any.whl,tar.gz} — same bits as PyPI
  • sciev-v0.2.3-paper.pdf — current manuscript (in preparation for arXiv;
    supersedes the archived v0.2.2 paper PDF)
  • MODEL_CARD.md, LICENSE, release-notes.txt, SHA256SUMS

Known caveats (unchanged): the adapted arm used one archival LoRA
adapter with a non-native loss normalization — it is not a causal DAPT
estimate; evidence controls report agreement with the original
reference, not relabeled truth; selective-coverage metrics are
retrospective diagnostics, not deployment error guarantees. See the
paper and MODEL_CARD for the full scope of each claim.