v0.50.0 — 50 correctness beats (PMAT-827..876)
[0.50.0] - 2026-06-21
Fixed
Provable-correctness wave — fifty shipped-green correctness defects (PMAT-827..876),
each fixed with a named proof-obligation + a RED-on-bug / GREEN-on-fix falsifier + a
pv-validated contract. Spans all four pillars (replace+beat scikit-learn / PyTorch /
Unsloth / Ollama) plus eval/format/export and CI determinism. The first fifteen:
stats::incomplete_betaextra/a(PMAT-827, Pillar-1) — the regularized
incomplete beta was wrong fora != 1, so every t-test (df ≤ 30) and ANOVA F-test
p-value was too small (falsely significant). e.g. a one-sample t-test reported p=0.115
when scipy gives 0.230. Now matchesscipy.special.betainc.- rsLoRA adapter scale dropped on load (PMAT-828, Pillar-3) —
LoRAAdapter::to_layer
recomputed Standardalpha/rankand discarded the serialized rsLoRAalpha/sqrt(rank)
scale, silently re-scaling a saved adapter bysqrt(rank)(e.g. 4× at rank 16). --grad-clipsilent no-op on the CPU trainer (PMAT-829, Pillar-2) —clip_and_step
computed the clip coefficient then discarded it (let _ = scale); the optimizer stepped
on raw, unclipped gradients (divergence risk), while the WGPU path clipped correctly.apr prune --sparsityover-pruned (PMAT-830) —sparsity.max(target_ratio)raised any
--sparsitybelow the 0.5--target-ratiodefault, so--sparsity 0.3zeroed 50% of
weights (not 30%) and the output metadata misreported the sparsity actually applied.GradientBoostingClassifier::predict_probasaturated (PMAT-831, Pillar-1) — the weak
learner fit a classification tree tosign(residual)and added a fixed ±1 step instead of a
regression tree to the continuous residuals, so probabilities saturated to 0/1 (50/164 →
P=0.99998 vs the correct 0.75). Now uses aDecisionTreeRegressor(Friedman gradient step).- Q3_K GGUF dequant corrupted weights on import (PMAT-832) — the 6-bit super-block scales
were unpacked as 4-bit (offset −8 instead of −32) with the wrong quant/high-bit layout, so
~252/256 elements were wrong on any Q3_K_S/Q3_K_M model. Ported the correct GGML algorithm. - MoE /
head_dimdropped on SafeTensors import (PMAT-833) —load_model_config_from_json
hardcodednum_experts/num_experts_per_tok/moe_intermediate_size/head_dimtoNone,
so a MoE model (Mixtral/Qwen3-MoE/DeepSeek) silently converted to a DENSE.apr, and an
explicithead_dimwas lost (wrong RoPE/attention dims for Qwen3/Gemma2/Phi3). - ARIMA forecast wrong for
d >= 2(PMAT-834, Pillar-1) — reverse-differencing re-seeded
every un-differencing pass withy[n]instead of the matching intermediate difference, so
every forecast with two or more differencing orders overshot (e.g. 165 vs the correct 110). apr evalpass@k inflated under single greedy sampling (PMAT-835) — the Chen et al.
estimator was fed the problem-count/solved-count in its per-sample(n, c)slots, so a model
solving 50/164 HumanEval reported pass@10=98% / pass@100=100% (correct: 30% for every k under
one deterministic sample) in the CI-consumed JSON. Now collapses to pass@1.- User
__metadata__dropped on everyapr export(PMAT-836) —extract_user_metadata
read a fabricated APR v2 header layout (length @ byte 8, JSON @ 16) instead of the real
64-byte header (metadata_offset@ 12, JSON @metadata_offset), always returning empty —
so the user's SafeTensors__metadata__was silently lost on re-export. - GPT-2 byte-level BPE decode produced mojibake (PMAT-837, Pillar-4) —
gpt2_char_to_byte
used a linearcode − 0x100offset instead of the GPT-2byte_encoderstaircase, so 129/256
bytes failed round-trip and all non-ASCII serve output was garbled (中 →ä¸Ń). Now delegates
to the correct unicode→byte map. - GLM IRLS swapped the link / inverse-link derivative (PMAT-838, Pillar-1) — the IRLS working
response and weights usedLink::derivative(the inverse-link derivativedμ/dη) where the
link derivativedη/dμis required, so coefficients were wrong for every non-identity link
(logistic slope 1.033 vs the correct 1.127). Now inverts it. - Gradient accumulation stepped on the SUM not the MEAN (PMAT-839, Pillar-2) — backward ops
accumulate into shared grad cells, but the trainer stepped without dividing by the accumulation
window, inflating the effective learning rate ×window (K-fold LR inflation / divergence). Now
scales grads by1/windowat the accumulation boundary. cargo install aprenderbroke on macOS (PMAT-840) —configure_parent_death_signalused
libc::prctl(PR_SET_PDEATHSIG)under#[cfg(unix)], but that prctl form is Linux-only, so
aprender-orchestrate(a dependency ofapr-cli) failed to compile on*-apple-darwin,
breaking the published binary for every macOS user. Now gated to#[cfg(target_os = "linux")].- Batched-GPU serving crashed on every GQA model (PMAT-841, Pillar-4) —
batch_generate_gpu
dispatched ≥32-prompt batches into an MHA-only path that assumesQKV = 3 × hidden_dim, so
every grouped-query-attention model (Qwen2 / Llama-3 / Mistral) crashed with a CUDA GEMM size
mismatch (B expected 3·hidden·hidden). Now routes GQA through the per-prompt path.
The remaining thirty-five (PMAT-842..876), each with a falsifier + pv-validated contract:
Pillar-1 — scikit-learn parity: macro precision/recall/f1/jaccard/fbeta averaged over
max(label)+1 instead of present labels (844); silhouette_score scored singleton clusters
+1.0 instead of 0 (845); FastICA whitening matrix transposed → Cov(X_white) ≠ I (847);
Lasso/ElasticNet alpha ignored the 1/(2·n) loss normalization (848); Ward linkage used the
wrong Lance-Williams coefficient (849); tree/RandomForest feature_importances used raw sample
counts not impurity decrease/MDI (851); train_test_split used round not ceil for float
test_size (852); two-tailed t-test used a normal approximation for df>30 (853); Brandes
betweenness counted the source's own dependency (860); TfidfVectorizer omitted L2 row
normalization (861); ARIMA AR coefficients estimated on uncentered data (862); Bayesian-logistic
MAP converged to precision n·λ not λ (864); KNN tie-break used randomized HashMap order not
smallest-label (865); StratifiedKFold dumped every class remainder into the low folds (866);
isotonic regression interpolated inside pooled PAV blocks (870); Calinski-Harabasz/Davies-Bouldin
counted phantom empty clusters / not relabel-invariant (871).