Skip to content

v0.101.0

Choose a tag to compare

@github-actions github-actions released this 21 Aug 18:21
· 522 commits to main since this release

Changes since v0.100.1:

Features

  • sampled admission (T>0) — true rejection-sampling verify for the dspark spec route
  • port the z-lab DFlash2 drafter into the dspark spec route

Fixes

  • boundaries never leave a sub-PRIME_MIN_T prompt remainder — closes the W1 tokenwise door for captures
  • prime boundaries land on the GDN grid — stop-vs-mono byte divergence at long ctx goes to zero (B3 disposition)
  • lo-clipped windowed SDPA for the round attention — the >12k-ctx crash (B2)
  • default id prompt must clear prime_cache's T>=16 floor

Documentation

  • wgmma dedup recorded; receipts push note
  • rows for the v0.100.1 samplat doors (MEMRA_FILTER_COOP, MEMRA_NVFP4_FUSED3B) — check-flags census green on the merged train
  • the ratification train ships as v0.101.0 (v0.100.0 was taken mid-train by the moeprime prime-wall release)
  • v0.100.0 — the DFlash2 drafter x round-cost engine train (family-keyed harvest, sampled admission, >12k longctx fix, prime-grid law); measured serve-class numbers with receipts
  • v0.100 train doors — family-then-strategy harvest keying on the MEMRA_DSPARK_HARVEST row; rows for MEMRA_DEBUG_PRIMESEG and MEMRA_DSPARK_FULLG_DEBUG (check-flags census green on the merged tree)
  • MEMRA_DFLASH2_SDPA_CLIP row (lo-clipped DFlash2 round SDPA, default ON)
  • rows for the dspark_sample_gate knobs (MEMRA_SG_*) + the MEMRA_DSPARK_SPEC route truth (greedy OR sampled) + backfill MEMRA_SPEC_ANATOMY
  • FLAGS rows for the dspark verify revert seams (STATE_COPY_BATCH slice-1 drift + slice-4 FA_ROWS / VG_AUTOFREE) — check-flags census green

Other

  • refactor(wgmma): dedup the triplicated wgmma helpers into wgmma_common.cuh — SASS-identical 12/12
  • boundary: remediate all 66 unremediated allowlist findings — migrate to darklanes, prune grants; tomllib fallback for 3.10 boxes
  • release: v0.101.0 — the ratification train: DFlash2 drafter x rank-1..3 round-cost engine (family-keyed harvest, sampled admission, >12k longctx fix, prime-grid law, dsv4 flip policy); version + internal pins
  • train integration: frspec stable-boundary test inherits the prime-grid contract
  • diag(primeseg): MEMRA_DEBUG_PRIMESEG prints the prime CALL SEQUENCE and the W1 tokenwise door
  • gate(primegrid): the prime-grid law on the qwen hybrid, gated both directions; flags + serving-contract docs
  • RECEIPTS: Server-window default-env battery DISCHARGED (box6 Server Edition, PASS 14/14 x3+control, banked flip verbatim, Server<->WS dumps byte-identical) — no open items on the ws-ref-thresholds derivation
  • ci journal: perf-quick rows for lane/dspark-trunk-kernels-20260820 close (correctness GREEN; 31b cells flat: plain 42.26/39.75, spec 107.94@0.798 / 102.20@0.817 — the norm ILP twins are shared kernels and the 31b cells hold; dual/group doors are dspark-only)
  • dspark(q38) trunk-kernels: per-door engagement receipts — the shared Once suppressed FA group3's ENGAGED print
  • ci: perf-quick journal rows on the lane tip (rc=0, perf 0 fail / 0 warn; 31b cells flat vs baseline)
  • dspark(q38) trunk-kernels slice D: FA q/k/v ride the group kernel (group3 = group4 with n3=0)
  • RECEIPTS: lane ws-ref-thresholds — (REF variant x bf16 arm) compressor_kv flip policy derivation banked (characterization: one-step e4m3 flip, cross-class bit-equivalence, no Server default-env receipt existed; calibration rows: exceeders 1/4096 x3 boots + swap, worst-case flippable census 2.44%/3.12% vs budget 5%; OWED: one Server-window default-env battery, prediction PASS)
  • dspark(q38) trunk-kernels slice C: group-4 GDN-tuple batched matvec — one launch for wqkv/gate/beta/alpha
  • dspark(q38) trunk-kernels slice B: qwen35 t-parallel FFN rides the proven dual gate+up chain
  • dspark(q38) trunk-kernels slice A: norm ILP twins — rms_norm_f32_v2 / add_rms_norm_f32_v2
  • dsv4-gpu-gate: REF-variant compressor_kv flip policy on the bf16-dequant arm — derived, (variant,arm)-keyed, never card-keyed
  • dsv4-gpu-gate: MEMRA_DSV4_GATE_DUMP calibration instrument — dumps every compared GPU array verbatim (LE f32) for offline per-element error measurement vs the fixture npz; dump-only, verdicts/thresholds/compare path byte-unchanged (WS REF-threshold lane characterization)
  • ci journal: perf-quick rows for lane/dspark-fa-execupdate-20260820 close (correctness green; 31b cells flat: plain 42.23/39.79, spec 107.89@0.798 / 102.17@0.817 — the batched fa/append rows + graphs ctx are dspark-only, MTP untouched)
  • dspark(q38): MEMRA_DSPARK_VERIFY_GRAPH measured disposition updated (slice 4a-4c receipts)
  • ci: perf-quick receipt on the pushed tip 8aa76d7 (rc=0, perf 0 fail / 0 warn) — discharges the §10.1 OWED item; earlier attempts died to rig CI contention, this run had the rig solo
  • dspark(q38): slice 4c fix — full-graph gate is WALK COVERAGE, not slot arithmetic
  • dspark(q38): MEMRA_DSPARK_FULLG_DEBUG=1 one-shot guard print (full-verify graphs did not engage on the s4c battery — zero 'full' census lines; diagnose which gate refuses)
  • dflash2 parity gate: require NVIDIA_TF32_OVERRIDE=0 (instrument setting)
  • dspark(q38): slice 4c — full-verify single graph per (vt, rung) under the opt-in
  • dspark(q38): slice 4 — batched fa/append rows + PRIORITY-instantiated segment graphs

Boards + reproduction artifacts: https://huggingface.co/Avifenesh/memra-bench · full experiment log in research/tune-data/