Skip to content

S1Q v0.1.0: five-model quantization study

Pre-release
Pre-release

Choose a tag to compare

@CYMCharming CYMCharming released this 01 Oct 08:01

S1Q v0.1.0 — five-model research release

S1Q: Low-Bit Quantization for System One Decision Models studies Kev-0.8B, Kev-4B, Kev-9B, NanoJev and Laya with revision-pinned native adapters and real A100/A800 GPU experiments.

This release includes code, per-question paired predictions, unquantized and matched RTN baselines, development ablations, temperature-calibrated probability metrics, source-transfer/game OOD evaluation, a frozen 85-task public JevBench cohort, measured packed memory, and exportable SVG/PDF research figures.

Model Selected format Main native accuracy Matched RTN S1Q
Kev-0.8B W4, Fisher-weighted 82.93% 80.63% 81.18%
Kev-4B W4, activation-aware 85.45% 84.79% 84.79%
Kev-9B W4A8 simulation 86.98% 84.35% 86.54%
NanoJev W4, activation-aware 79.86% 79.47% 80.45%
Laya W4, activation-aware 66.85% 63.68% 64.55%

Kev/Laya use 914 decisions; NanoJev uses 1,023 game decisions. These accuracies are not a cross-dataset model leaderboard. Main cohorts are exploratory following inspected pilots. The separate external cohort uses profiles frozen before inference; all five external S1Q-minus-matched-RTN intervals include zero. S1Q does not consistently outperform RTN.

Weight assets

Seven ordered parts reconstruct the exact evaluated packed linears for the three Kev checkpoints and Laya. Download with python scripts/fetch_artifacts.py --model kev-0.8b, using the checked-in configs/release-artifacts.json. Every part and reconstructed artifact is SHA256-verified. The same pinned original checkpoint/tokenizer/head is required; these are not standalone models.

NanoJev quantization and its evaluated local artifact are reproducible through the published recipe. Its audited model card does not separately state an explicit fine-tuned weight redistribution grant, so S1Q publishes the source/recipe/results without redistributing that weight derivative.

The attached license bundle retains exact upstream Apache/MIT notices and a quantization modification notice. It does not grant a new license for any upstream model. Native matrix multiplications remain floating point, activation quantization is simulated, and no native INT4 speedup is claimed. The release does not claim to be the first quantization of Jev-like/System One models; direct prior Laya/Kev work is credited.

Full results and intervals · Method · Artifact usage · Upstream licenses

源码、五个模型的量化实验与负面结果均已公开;当前推荐以 W4 权重量化为起点,Kev-9B 使用开发集选出的 W4A8 模拟方案。本文实现强调可复现决策评测与存储压缩,不宣称普遍超越 RTN 或原生 INT4 加速。