Releases: alrobles/sciev
Release list
Sciev v0.2.3 — docs and eval-infra alignment
Sciev v0.2.3 — documentation and evaluation-infra alignment.
Weights, datasets (systemone-v2) and all reported metrics are identical
to v0.2.2; head checksums were re-verified byte-for-byte against the
published v0.2.2 SHA256SUMS. This release aligns the package and public
documentation with the verified matched study (three head seeds,
evidence controls) and the manuscript prepared for arXiv submission.
Package changes since v0.2.2:
- sciev.eval: --predictions emits per-item prediction records
- sciev.train: --steps overrides the recipe's default training budget
- new eval/baseline scripts under scripts/ (r2_*, eval_cbjev.py,
eval_llada_prompted.py, train_encoder_scifact.py, fit_dev_temp.py) - MODEL_CARD rewritten for fr_* with evidence controls and the
elite-v1 truncation-defect disclosure; README aligns public claims
Contents:
- fr_{choice,noul,score}.pt — frozen-backbone release heads (v0.2 bits)
- da_{choice,noul,score}.pt — experimental arm (caveats in the paper)
- evals.tar.gz — archived 32-report evaluation set (same as v0.2.2)
- evidence.json — compact ledger of the complete 60-report study with
SHA-256 fingerprints, per-seed metrics, dataset and checkpoint metadata - all-manifests.txt, hard-controls-manifests.txt — dataset manifests
- sciev-0.2.3-{py3-none-any.whl,tar.gz} — same bits as PyPI
- sciev-v0.2.3-paper.pdf — current manuscript (in preparation for arXiv;
supersedes the archived v0.2.2 paper PDF) - MODEL_CARD.md, LICENSE, release-notes.txt, SHA256SUMS
Known caveats (unchanged): the adapted arm used one archival LoRA
adapter with a non-native loss normalization — it is not a causal DAPT
estimate; evidence controls report agreement with the original
reference, not relabeled truth; selective-coverage metrics are
retrospective diagnostics, not deployment error guarantees. See the
paper and MODEL_CARD for the full scope of each claim.
Sciev v0.2.2 — PyPI release, public URLs
Sciev v0.2.2 — PyPI release with public repository URLs.
Identical to v0.2.1 except packaging metadata: project URLs now point
at the public repo github.com/alrobles/sciev. Package is installable
via pip install sciev (PyPI). Weights, datasets (systemone-v2) and
all reported metrics are unchanged.
Contents:
- heads/release/fr_{choice,noul,score}.pt — frozen-backbone release heads
- heads/experimental/da_{choice,noul,score}.pt — DAPT arm (caveats in paper)
- evals.tar.gz — 32 evaluation JSONs (matched arms, controls, multi-seed)
- manifests/ — dataset manifests for sci_battery, controls, benchmarks
- sciev-0.2.2-{py3-none-any.whl,tar.gz} — same bits as PyPI
- sciev-v0.2.2-paper.pdf, MODEL_CARD.md, LICENSE, SHA256SUMS
Sciev v0.2.1 — sciev rebrand
Sciev v0.2.1 — namespace rebrand release.
The Python package was renamed reverse_jev -> sciev (import sciev,
python -m sciev.train / sciev.eval). Checkout paths moved to
~/GitHub/sciev and /beegfs/a474r867/sciev.
This is a namespace-only release: the decision heads, datasets
(systemone-v2), calibration artifacts and all reported metrics are
bit-identical to v0.2.0. Head SHA-256s match the v0.2.0 assets.
Contents:
- heads/release/fr_{choice,noul,score}.pt — frozen-backbone release heads
- heads/experimental/da_{choice,noul,score}.pt — DAPT arm (documented caveats)
- evals.tar.gz — 32 evaluation JSONs (matched arms, controls, multi-seed)
- manifests/ — dataset manifests for sci_battery, controls, benchmarks
- sciev-0.2.1-py3-none-any.whl — installable package (pip install sciev)
- sciev-v0.2.1-paper.pdf — updated draft
- MODEL_CARD.md, LICENSE, SHA256SUMS, release-manifest.json
Upgrade note: replace reverse_jev with sciev in imports. Checkpoints
are pure state_dicts and load unchanged.
Sciev v0.2.0 — aligned scientific decision pipeline
Sciev v0.2.0 — aligned scientific decision pipeline
Typed decision heads (choice/noul/score) over frozen LLaDA-8B, trained
and evaluated under the systemone-v2 contract: group-disjoint splits,
explicit overflow handling (max_ctx=960, no silent truncation), shared
R2 preprocessing across train/calibrate/eval/serve, dev-only temperature
fitting, and provenance manifests for every file.
Highlights
- Matched frozen/adapted comparison (identical recipe + seed + data):
mean±sd over seeds 7331–7333. Domain adaptation robustly improves
out-of-distribution verification (SciFact noul 0.709 vs 0.483, +22.6pt)
at a noisy cost to in-distribution rubric scoring — task-dependent, not
a uniform win/loss. - Evidence controls (empty/shuffle, agreement-with-reference): expose
option-side leakage — choice is ~80–87% solvable without reading the
passage; score is the task whose accuracy most directly reflects
evidence use. - Hard-distractor battery (
--hard-choice): strict same-passage +
numeric-perturb pool reduces leakage ~6–8pt but does not remove it.
Contents
heads/release/— fr_matched_s7331 heads (recommended; frozen backbone)heads/experimental/— da_matched_s7331 heads (DAPT arm; documented caveats)evals/— 32 evaluation JSONs (arms × tasks × benchmarks × controls)manifests/— dataset manifests incl. exclusion provenance and input hashessciev-v0.2-paper.pdf,MODEL_CARD.md,LICENSE,SHA256SUMS,
release-manifest.json
Checkpoints are self-describing (inference config, training-data
contract, input fingerprints, recipe, seed). Controls are scored as
agreement-with-reference, never relabelled.
Sciev v0.1.1 — audited paper and evaluation patch
Sciev v0.1.1 — audited paper and evaluation patch
The three frozen-LLaDA c_* head checkpoints are byte-for-byte identical to v0.1. No DAPT adapter is selected. The existing v0.1 tag and release are unchanged.
Paper and evidence
- Updated Table 3 with sample counts, accuracy, ECE, and retrospective automation from the archived evaluation JSONs.
- Added Table 4 comparing the frozen release, DAPT g2000/g5000 candidates, and the early bw1_sr control.
- Includes the compiled six-page paper (
main.pdf),results.jsonwith source hashes and 12 head-training configurations, andsciev-v0.1.1-evidence.tar.gzwith the 32 raw evaluation reports, training configurations, and bw1_sr training log. - Corrected earlier metric transcriptions: GPQA main/diamond 0.3125/0.3182; SciFact noul/score 0.8500/0.6294. These are archived runs, not new GPU evaluations.
Metric correction and interpretation
- Fixed option-order flip measurement to compare original option identities instead of positions after permutation. Added regression tests for both an invariant predictor and a position-biased predictor.
- Withdrawn legacy non-canonical flip comparisons pending re-evaluation; checkpoint weights, accuracy, and calibration calculations are unchanged.
- Head optimization differs between
c_*and the DAPT candidates; Table 4 is not a controlled DAPT-only ablation. - bw1_sr is a 20-step checkpoint with a shorter context window, not a fully trained 1.4B baseline.
- DAPT g5000 warm-started from g2000 without restoring optimizer/RNG state. The nominal 81.9M token-slot budget is not a measured unique-token count.
- Automation@5% is retrospective, not a deployment risk guarantee. External benchmarks informed candidate selection.
Assets and use
| asset | head | parameters | recorded dev temperature |
|---|---|---|---|
sciev-0.1-choice.pt |
AttnPool, four layers | 67,166,209 | 0.8274 |
sciev-0.1-noul.pt |
final-layer MLP | 16,793,601 | 2.7127 |
sciev-0.1-score.pt |
final-layer MLP, ordinal training | 16,793,601 | 0.9367 |
The 0.1 filenames identify unchanged weights. The backbone is fetched separately from GSAI-ML/LLaDA-8B-Instruct. Use v0.1.1 code, --r2-mode spanpool, --canonical-order, and the matching --r2-temp. Choice also needs --r2-layers=-1,-9,-17,-25. See MODEL_CARD.md and the source README for input formatting and reproducibility.
The manifest records artifact hashes, sizes, architecture, and temperatures. Verify downloads against SHA256SUMS. Package metadata now matches the repository's Apache-2.0 license; backbones and datasets retain their own licenses.
Verification
- 12 tests passed, including regression tests and exact checks of Tables 3/4 against
results.json. - Paper compiled twice without LaTeX warnings; tables visually checked.
- All three checkpoint hashes match both HPC originals and v0.1 GitHub assets.
- Weights-only deserialization, strict state loading, finite tensors, and CPU head-forward smoke checks passed.
- Wheel built offline, installed into an isolated target, and checked for version 0.1.1, Apache-2.0 metadata, and CLI imports. The 12 tests also passed from a clean archive of the release commit.
Research preview. The paper remains a draft and has not been submitted to arXiv. Repository visibility is unchanged (private).
Generated with Devin
Sciev v0.1
Sciev (sci+ev) — open System-One-style decision model: typed questions over a state → calibrated probabilities in one forward pass. No generation, nothing to parse.
Frozen LLaDA-8B (masked-diffusion, bidirectional) + specialist heads (~2M params). The only open entry in the ecosystem on a diffusion backbone, and the only one with exact option-order invariance (flip rate 0.00 by construction via canonical ordering).
Head checkpoints (assets)
| asset | type | head |
|---|---|---|
sciev-v0.1-c_choice.pt |
choice | AttnPool multi-layer {8,16,24,32} |
sciev-v0.1-c_noul.pt |
noul | MLP |
sciev-v0.1-c_score.pt |
score | MLP + ordinal aux loss |
Backbone downloads from HF automatically (GSAI-ML/LLaDA-8B-Instruct).
Headline numbers
| benchmark | acc | note |
|---|---|---|
| sci battery choice | 0.870 | auto@5% = 0.81, flip = 0 |
| SciFact noul | 0.853 | ~4pts from open SOTA |
| SST-2 / AG News (zero-shot) | 0.930 / 0.854 | general transfer |
| GPQA | 0.315 | = backbone ceiling (~0.33) |
Model card: MODEL_CARD.md · Paper draft: paper/main.tex · Status: docs/PROJECT-STATUS.md
Apache-2.0. Research preview — K>4 options and score remain limited (see model card).