Skip to content

v0.50.0 — 50 correctness beats (PMAT-827..876)

Choose a tag to compare

@noahgift noahgift released this 21 Jun 01:11
· 537 commits to main since this release
e3a28d6

[0.50.0] - 2026-06-21

Fixed

Provable-correctness wave — fifty shipped-green correctness defects (PMAT-827..876),
each fixed with a named proof-obligation + a RED-on-bug / GREEN-on-fix falsifier + a
pv-validated contract. Spans all four pillars (replace+beat scikit-learn / PyTorch /
Unsloth / Ollama) plus eval/format/export and CI determinism. The first fifteen:

  • stats::incomplete_beta extra /a (PMAT-827, Pillar-1) — the regularized
    incomplete beta was wrong for a != 1, so every t-test (df ≤ 30) and ANOVA F-test
    p-value was too small (falsely significant). e.g. a one-sample t-test reported p=0.115
    when scipy gives 0.230. Now matches scipy.special.betainc.
  • rsLoRA adapter scale dropped on load (PMAT-828, Pillar-3) — LoRAAdapter::to_layer
    recomputed Standard alpha/rank and discarded the serialized rsLoRA alpha/sqrt(rank)
    scale, silently re-scaling a saved adapter by sqrt(rank) (e.g. 4× at rank 16).
  • --grad-clip silent no-op on the CPU trainer (PMAT-829, Pillar-2) — clip_and_step
    computed the clip coefficient then discarded it (let _ = scale); the optimizer stepped
    on raw, unclipped gradients (divergence risk), while the WGPU path clipped correctly.
  • apr prune --sparsity over-pruned (PMAT-830) — sparsity.max(target_ratio) raised any
    --sparsity below the 0.5 --target-ratio default, so --sparsity 0.3 zeroed 50% of
    weights (not 30%) and the output metadata misreported the sparsity actually applied.
  • GradientBoostingClassifier::predict_proba saturated (PMAT-831, Pillar-1) — the weak
    learner fit a classification tree to sign(residual) and added a fixed ±1 step instead of a
    regression tree to the continuous residuals, so probabilities saturated to 0/1 (50/164 →
    P=0.99998 vs the correct 0.75). Now uses a DecisionTreeRegressor (Friedman gradient step).
  • Q3_K GGUF dequant corrupted weights on import (PMAT-832) — the 6-bit super-block scales
    were unpacked as 4-bit (offset −8 instead of −32) with the wrong quant/high-bit layout, so
    ~252/256 elements were wrong on any Q3_K_S/Q3_K_M model. Ported the correct GGML algorithm.
  • MoE / head_dim dropped on SafeTensors import (PMAT-833) — load_model_config_from_json
    hardcoded num_experts/num_experts_per_tok/moe_intermediate_size/head_dim to None,
    so a MoE model (Mixtral/Qwen3-MoE/DeepSeek) silently converted to a DENSE .apr, and an
    explicit head_dim was lost (wrong RoPE/attention dims for Qwen3/Gemma2/Phi3).
  • ARIMA forecast wrong for d >= 2 (PMAT-834, Pillar-1) — reverse-differencing re-seeded
    every un-differencing pass with y[n] instead of the matching intermediate difference, so
    every forecast with two or more differencing orders overshot (e.g. 165 vs the correct 110).
  • apr eval pass@k inflated under single greedy sampling (PMAT-835) — the Chen et al.
    estimator was fed the problem-count/solved-count in its per-sample (n, c) slots, so a model
    solving 50/164 HumanEval reported pass@10=98% / pass@100=100% (correct: 30% for every k under
    one deterministic sample) in the CI-consumed JSON. Now collapses to pass@1.
  • User __metadata__ dropped on every apr export (PMAT-836) — extract_user_metadata
    read a fabricated APR v2 header layout (length @ byte 8, JSON @ 16) instead of the real
    64-byte header (metadata_offset @ 12, JSON @ metadata_offset), always returning empty —
    so the user's SafeTensors __metadata__ was silently lost on re-export.
  • GPT-2 byte-level BPE decode produced mojibake (PMAT-837, Pillar-4) — gpt2_char_to_byte
    used a linear code − 0x100 offset instead of the GPT-2 byte_encoder staircase, so 129/256
    bytes failed round-trip and all non-ASCII serve output was garbled (中 → ä¸Ń). Now delegates
    to the correct unicode→byte map.
  • GLM IRLS swapped the link / inverse-link derivative (PMAT-838, Pillar-1) — the IRLS working
    response and weights used Link::derivative (the inverse-link derivative dμ/dη) where the
    link derivative dη/dμ is required, so coefficients were wrong for every non-identity link
    (logistic slope 1.033 vs the correct 1.127). Now inverts it.
  • Gradient accumulation stepped on the SUM not the MEAN (PMAT-839, Pillar-2) — backward ops
    accumulate into shared grad cells, but the trainer stepped without dividing by the accumulation
    window, inflating the effective learning rate ×window (K-fold LR inflation / divergence). Now
    scales grads by 1/window at the accumulation boundary.
  • cargo install aprender broke on macOS (PMAT-840) — configure_parent_death_signal used
    libc::prctl(PR_SET_PDEATHSIG) under #[cfg(unix)], but that prctl form is Linux-only, so
    aprender-orchestrate (a dependency of apr-cli) failed to compile on *-apple-darwin,
    breaking the published binary for every macOS user. Now gated to #[cfg(target_os = "linux")].
  • Batched-GPU serving crashed on every GQA model (PMAT-841, Pillar-4) — batch_generate_gpu
    dispatched ≥32-prompt batches into an MHA-only path that assumes QKV = 3 × hidden_dim, so
    every grouped-query-attention model (Qwen2 / Llama-3 / Mistral) crashed with a CUDA GEMM size
    mismatch (B expected 3·hidden·hidden). Now routes GQA through the per-prompt path.

The remaining thirty-five (PMAT-842..876), each with a falsifier + pv-validated contract:

Pillar-1 — scikit-learn parity: macro precision/recall/f1/jaccard/fbeta averaged over
max(label)+1 instead of present labels (844); silhouette_score scored singleton clusters
+1.0 instead of 0 (845); FastICA whitening matrix transposed → Cov(X_white) ≠ I (847);
Lasso/ElasticNet alpha ignored the 1/(2·n) loss normalization (848); Ward linkage used the
wrong Lance-Williams coefficient (849); tree/RandomForest feature_importances used raw sample
counts not impurity decrease/MDI (851); train_test_split used round not ceil for float
test_size (852); two-tailed t-test used a normal approximation for df>30 (853); Brandes
betweenness counted the source's own dependency (860); TfidfVectorizer omitted L2 row
normalization (861); ARIMA AR coefficients estimated on uncentered data (862); Bayesian-logistic
MAP converged to precision n·λ not λ (864); KNN tie-break used randomized HashMap order not
smallest-label (865); StratifiedKFold dumped every class remainder into the low folds (866);
isotonic regression interpolated inside pooled PAV blocks (870); Calinski-Harabasz/Davies-Bouldin
counted phantom empty clusters / not relabel-invariant (871).