Skip to content

Releases: alrobles/sciev

Sciev v0.2.3 — docs and eval-infra alignment

Choose a tag to compare

@alrobles alrobles released this 28 Sep 02:51

Sciev v0.2.3 — documentation and evaluation-infra alignment.

Weights, datasets (systemone-v2) and all reported metrics are identical
to v0.2.2; head checksums were re-verified byte-for-byte against the
published v0.2.2 SHA256SUMS. This release aligns the package and public
documentation with the verified matched study (three head seeds,
evidence controls) and the manuscript prepared for arXiv submission.

Package changes since v0.2.2:

  • sciev.eval: --predictions emits per-item prediction records
  • sciev.train: --steps overrides the recipe's default training budget
  • new eval/baseline scripts under scripts/ (r2_*, eval_cbjev.py,
    eval_llada_prompted.py, train_encoder_scifact.py, fit_dev_temp.py)
  • MODEL_CARD rewritten for fr_* with evidence controls and the
    elite-v1 truncation-defect disclosure; README aligns public claims

Contents:

  • fr_{choice,noul,score}.pt — frozen-backbone release heads (v0.2 bits)
  • da_{choice,noul,score}.pt — experimental arm (caveats in the paper)
  • evals.tar.gz — archived 32-report evaluation set (same as v0.2.2)
  • evidence.json — compact ledger of the complete 60-report study with
    SHA-256 fingerprints, per-seed metrics, dataset and checkpoint metadata
  • all-manifests.txt, hard-controls-manifests.txt — dataset manifests
  • sciev-0.2.3-{py3-none-any.whl,tar.gz} — same bits as PyPI
  • sciev-v0.2.3-paper.pdf — current manuscript (in preparation for arXiv;
    supersedes the archived v0.2.2 paper PDF)
  • MODEL_CARD.md, LICENSE, release-notes.txt, SHA256SUMS

Known caveats (unchanged): the adapted arm used one archival LoRA
adapter with a non-native loss normalization — it is not a causal DAPT
estimate; evidence controls report agreement with the original
reference, not relabeled truth; selective-coverage metrics are
retrospective diagnostics, not deployment error guarantees. See the
paper and MODEL_CARD for the full scope of each claim.

Sciev v0.2.2 — PyPI release, public URLs

Choose a tag to compare

@alrobles alrobles released this 27 Sep 02:10

Sciev v0.2.2 — PyPI release with public repository URLs.

Identical to v0.2.1 except packaging metadata: project URLs now point
at the public repo github.com/alrobles/sciev. Package is installable
via pip install sciev (PyPI). Weights, datasets (systemone-v2) and
all reported metrics are unchanged.

Contents:

  • heads/release/fr_{choice,noul,score}.pt — frozen-backbone release heads
  • heads/experimental/da_{choice,noul,score}.pt — DAPT arm (caveats in paper)
  • evals.tar.gz — 32 evaluation JSONs (matched arms, controls, multi-seed)
  • manifests/ — dataset manifests for sci_battery, controls, benchmarks
  • sciev-0.2.2-{py3-none-any.whl,tar.gz} — same bits as PyPI
  • sciev-v0.2.2-paper.pdf, MODEL_CARD.md, LICENSE, SHA256SUMS

Sciev v0.2.1 — sciev rebrand

Choose a tag to compare

@alrobles alrobles released this 27 Sep 01:27

Sciev v0.2.1 — namespace rebrand release.

The Python package was renamed reverse_jev -> sciev (import sciev,
python -m sciev.train / sciev.eval). Checkout paths moved to
~/GitHub/sciev and /beegfs/a474r867/sciev.

This is a namespace-only release: the decision heads, datasets
(systemone-v2), calibration artifacts and all reported metrics are
bit-identical to v0.2.0. Head SHA-256s match the v0.2.0 assets.

Contents:

  • heads/release/fr_{choice,noul,score}.pt — frozen-backbone release heads
  • heads/experimental/da_{choice,noul,score}.pt — DAPT arm (documented caveats)
  • evals.tar.gz — 32 evaluation JSONs (matched arms, controls, multi-seed)
  • manifests/ — dataset manifests for sci_battery, controls, benchmarks
  • sciev-0.2.1-py3-none-any.whl — installable package (pip install sciev)
  • sciev-v0.2.1-paper.pdf — updated draft
  • MODEL_CARD.md, LICENSE, SHA256SUMS, release-manifest.json

Upgrade note: replace reverse_jev with sciev in imports. Checkpoints
are pure state_dicts and load unchanged.

Sciev v0.2.0 — aligned scientific decision pipeline

Choose a tag to compare

@alrobles alrobles released this 27 Sep 01:24

Sciev v0.2.0 — aligned scientific decision pipeline

Typed decision heads (choice/noul/score) over frozen LLaDA-8B, trained
and evaluated under the systemone-v2 contract: group-disjoint splits,
explicit overflow handling (max_ctx=960, no silent truncation), shared
R2 preprocessing across train/calibrate/eval/serve, dev-only temperature
fitting, and provenance manifests for every file.

Highlights

  • Matched frozen/adapted comparison (identical recipe + seed + data):
    mean±sd over seeds 7331–7333. Domain adaptation robustly improves
    out-of-distribution verification (SciFact noul 0.709 vs 0.483, +22.6pt)
    at a noisy cost to in-distribution rubric scoring — task-dependent, not
    a uniform win/loss.
  • Evidence controls (empty/shuffle, agreement-with-reference): expose
    option-side leakage — choice is ~80–87% solvable without reading the
    passage; score is the task whose accuracy most directly reflects
    evidence use.
  • Hard-distractor battery (--hard-choice): strict same-passage +
    numeric-perturb pool reduces leakage ~6–8pt but does not remove it.

Contents

  • heads/release/ — fr_matched_s7331 heads (recommended; frozen backbone)
  • heads/experimental/ — da_matched_s7331 heads (DAPT arm; documented caveats)
  • evals/ — 32 evaluation JSONs (arms × tasks × benchmarks × controls)
  • manifests/ — dataset manifests incl. exclusion provenance and input hashes
  • sciev-v0.2-paper.pdf, MODEL_CARD.md, LICENSE, SHA256SUMS,
    release-manifest.json

Checkpoints are self-describing (inference config, training-data
contract, input fingerprints, recipe, seed). Controls are scored as
agreement-with-reference, never relabelled.

Sciev v0.1.1 — audited paper and evaluation patch

Choose a tag to compare

@alrobles alrobles released this 27 Sep 01:23

Sciev v0.1.1 — audited paper and evaluation patch

The three frozen-LLaDA c_* head checkpoints are byte-for-byte identical to v0.1. No DAPT adapter is selected. The existing v0.1 tag and release are unchanged.

Paper and evidence

  • Updated Table 3 with sample counts, accuracy, ECE, and retrospective automation from the archived evaluation JSONs.
  • Added Table 4 comparing the frozen release, DAPT g2000/g5000 candidates, and the early bw1_sr control.
  • Includes the compiled six-page paper (main.pdf), results.json with source hashes and 12 head-training configurations, and sciev-v0.1.1-evidence.tar.gz with the 32 raw evaluation reports, training configurations, and bw1_sr training log.
  • Corrected earlier metric transcriptions: GPQA main/diamond 0.3125/0.3182; SciFact noul/score 0.8500/0.6294. These are archived runs, not new GPU evaluations.

Metric correction and interpretation

  • Fixed option-order flip measurement to compare original option identities instead of positions after permutation. Added regression tests for both an invariant predictor and a position-biased predictor.
  • Withdrawn legacy non-canonical flip comparisons pending re-evaluation; checkpoint weights, accuracy, and calibration calculations are unchanged.
  • Head optimization differs between c_* and the DAPT candidates; Table 4 is not a controlled DAPT-only ablation.
  • bw1_sr is a 20-step checkpoint with a shorter context window, not a fully trained 1.4B baseline.
  • DAPT g5000 warm-started from g2000 without restoring optimizer/RNG state. The nominal 81.9M token-slot budget is not a measured unique-token count.
  • Automation@5% is retrospective, not a deployment risk guarantee. External benchmarks informed candidate selection.

Assets and use

asset head parameters recorded dev temperature
sciev-0.1-choice.pt AttnPool, four layers 67,166,209 0.8274
sciev-0.1-noul.pt final-layer MLP 16,793,601 2.7127
sciev-0.1-score.pt final-layer MLP, ordinal training 16,793,601 0.9367

The 0.1 filenames identify unchanged weights. The backbone is fetched separately from GSAI-ML/LLaDA-8B-Instruct. Use v0.1.1 code, --r2-mode spanpool, --canonical-order, and the matching --r2-temp. Choice also needs --r2-layers=-1,-9,-17,-25. See MODEL_CARD.md and the source README for input formatting and reproducibility.

The manifest records artifact hashes, sizes, architecture, and temperatures. Verify downloads against SHA256SUMS. Package metadata now matches the repository's Apache-2.0 license; backbones and datasets retain their own licenses.

Verification

  • 12 tests passed, including regression tests and exact checks of Tables 3/4 against results.json.
  • Paper compiled twice without LaTeX warnings; tables visually checked.
  • All three checkpoint hashes match both HPC originals and v0.1 GitHub assets.
  • Weights-only deserialization, strict state loading, finite tensors, and CPU head-forward smoke checks passed.
  • Wheel built offline, installed into an isolated target, and checked for version 0.1.1, Apache-2.0 metadata, and CLI imports. The 12 tests also passed from a clean archive of the release commit.

Research preview. The paper remains a draft and has not been submitted to arXiv. Repository visibility is unchanged (private).

Generated with Devin

Sciev v0.1

Choose a tag to compare

@alrobles alrobles released this 24 Sep 20:28

Sciev (sci+ev) — open System-One-style decision model: typed questions over a state → calibrated probabilities in one forward pass. No generation, nothing to parse.

Frozen LLaDA-8B (masked-diffusion, bidirectional) + specialist heads (~2M params). The only open entry in the ecosystem on a diffusion backbone, and the only one with exact option-order invariance (flip rate 0.00 by construction via canonical ordering).

Head checkpoints (assets)

asset type head
sciev-v0.1-c_choice.pt choice AttnPool multi-layer {8,16,24,32}
sciev-v0.1-c_noul.pt noul MLP
sciev-v0.1-c_score.pt score MLP + ordinal aux loss

Backbone downloads from HF automatically (GSAI-ML/LLaDA-8B-Instruct).

Headline numbers

benchmark acc note
sci battery choice 0.870 auto@5% = 0.81, flip = 0
SciFact noul 0.853 ~4pts from open SOTA
SST-2 / AG News (zero-shot) 0.930 / 0.854 general transfer
GPQA 0.315 = backbone ceiling (~0.33)

Model card: MODEL_CARD.md · Paper draft: paper/main.tex · Status: docs/PROJECT-STATUS.md

Apache-2.0. Research preview — K>4 options and score remain limited (see model card).