Sciev v0.2.3 — documentation and evaluation-infra alignment.
Weights, datasets (systemone-v2) and all reported metrics are identical
to v0.2.2; head checksums were re-verified byte-for-byte against the
published v0.2.2 SHA256SUMS. This release aligns the package and public
documentation with the verified matched study (three head seeds,
evidence controls) and the manuscript prepared for arXiv submission.
Package changes since v0.2.2:
- sciev.eval: --predictions emits per-item prediction records
- sciev.train: --steps overrides the recipe's default training budget
- new eval/baseline scripts under scripts/ (r2_*, eval_cbjev.py,
eval_llada_prompted.py, train_encoder_scifact.py, fit_dev_temp.py) - MODEL_CARD rewritten for fr_* with evidence controls and the
elite-v1 truncation-defect disclosure; README aligns public claims
Contents:
- fr_{choice,noul,score}.pt — frozen-backbone release heads (v0.2 bits)
- da_{choice,noul,score}.pt — experimental arm (caveats in the paper)
- evals.tar.gz — archived 32-report evaluation set (same as v0.2.2)
- evidence.json — compact ledger of the complete 60-report study with
SHA-256 fingerprints, per-seed metrics, dataset and checkpoint metadata - all-manifests.txt, hard-controls-manifests.txt — dataset manifests
- sciev-0.2.3-{py3-none-any.whl,tar.gz} — same bits as PyPI
- sciev-v0.2.3-paper.pdf — current manuscript (in preparation for arXiv;
supersedes the archived v0.2.2 paper PDF) - MODEL_CARD.md, LICENSE, release-notes.txt, SHA256SUMS
Known caveats (unchanged): the adapted arm used one archival LoRA
adapter with a non-native loss normalization — it is not a causal DAPT
estimate; evidence controls report agreement with the original
reference, not relabeled truth; selective-coverage metrics are
retrospective diagnostics, not deployment error guarantees. See the
paper and MODEL_CARD for the full scope of each claim.