Skip to content

LSE v0.4.14

Choose a tag to compare

@Geramy Geramy released this 29 Sep 10:32

LSE v0.4.14

Source: c11f103c86627e66c36d1b109fb02281c40ad686.

Improvements

  • c1ee42d: M1024 Q4 down projection shares activation fragments and reuses weight groups across four M16 row blocks.
  • 67a0b4e: WMMA paged attention skips fully masked 256-key windows with a workgroup-uniform predicate. Contributing windows retain their arithmetic and ordering.
  • de3afcf and 8150263: M1024 Q4 gate/up and the measured QKV, GDN and attention projection shapes stage activation fragments once in LDS for column waves to share. Inputs share a producer; the central shape table checks barrier support and the LDS budget.
  • c11f103: Two M8 Q4 shapes consume independent activation row pairs while reusing decoded weights and wider activation loads. The per-row chunk/FMA and split-reduction order is preserved. Accepted down-projection bodies and unrelated shapes remain unchanged.

Floating-point accumulation remains FP32. Compact Q4 weights are unchanged. Complete output bits, independent arithmetic reference, codec, fused residual, replay, readonly-input and buffer-guard checks pass. Measured component kernels report zero private scratch. Focused suites passed: activation panel 9/9, matrix panel 10/10 and dispatch 10/10. No additional perplexity run was needed for these bit-identical schedules.

Final same-binary measurements

macOS/R9700 gfx1201, Qwen Q4 target, Q8 auxiliary/draft weights, BF16 KV, temperature 0.6, top-k 20, top-p 0.95, seed 1234, batch/ubatch 1024 and configured KV capacity 262100. Each 1024-token coding request generates 384 tokens (383 timed decode tokens). MTP uses depth 3; DFlash2 uses all seven proposals in block 8.

Each mode starts with an empty private kernel cache. The resident request retains compiled code and reuses zero prompt KV. Compilation is included. Requests execute with zero host groups or fallbacks, without profiling or a competing workload. One pair per mode is not a statistical estimate or a long-context Pi replay; rates vary with prompt, sampling and live context.

Mode Cold prefill tokens/s Cold decode tokens/s Resident prefill tokens/s Resident decode tokens/s
Baseline 442.18 24.14 624.10 24.73
MTP=3 403.95 37.96 608.19 48.62
DFlash2, seven proposals 417.14 31.39 616.78 42.41

All three controlled resident requests exceed 600 prefill tokens/s. The 29 TPS baseline, 49 TPS MTP and 103 TPS DFlash2 goals remain unmet.

The matched projection change improved resident DFlash2 prefill from 568.03 to 619.08 tokens/s (+8.99%), with decode effectively unchanged. The subsequent M8 row-pair change improved matched resident decode from 40.68 to 42.41 tokens/s (+4.27%), with prefill effectively unchanged. Each scoped comparison preserved exact responses and acceptance/probability statistics. Across different generation modes, RNG execution can produce different answers.

Final mode comparison, M8 row-pair methods, prefill projection methods.

Packages

Both archives are built from the exact source above and include a runtime/source manifest and SHA256 file.

  • Linux x86_64 bundles its selected HRX runtime and patched Loom compiler. Root and bin/ launchers select them. Compatible Linux C/C++ libraries, ROCm 7.x, HSA and the GPU driver remain external requirements.
  • macOS arm64 bundles HSA, HRX, patched Loom and its runtime dependency closure. Use bin/lse or bin/lse-server; install and activate MacAMDGPU separately. Native binaries target Apple Silicon macOS 15 or newer; GPU execution requires the driver-supported macOS version.

K/V defaults to model-declared BF16 for BF16 checkpoints, and FP16 otherwise, with explicit overrides. Seven-proposal DFlash2 and batch/ubatch 1024 remain defaults. --temperature sets the server default; request values take precedence and launch examples use 0.6.

Qualification

  • Linux final archive build: all 86 tests, selected compiler/runtime verification, live engine/HTTP smoke and relocated launcher/loader checks passed. Selected targets: gfx942, gfx1150, gfx1151, gfx1200 and gfx1201.
  • macOS final archive build: 53 CPU-only tests, four HRX-linked host fixtures, 89 CLI cases, native gfx1201 compilation and runtime relocation passed. The hosted runner has no external AMD GPU; actual execution was qualified locally for the accepted source changes.
  • Both archives were checked against their SHA256 files, exact source/compiler/patch manifests and bundled runtime hashes. Relocated macOS launchers, signatures and HSA loading passed without opening a GPU. Uploaded asset bytes and tag identity were verified before publication.

Verify downloads with the accompanying .sha256 files.