Releases: Geramy/LSE
Release list
LSE v0.4.23
- Faster long-context prefill: FlashPrefill V2 reached 632.1 pp/s at 16K and 604.9 pp/s at 32K on Qwen3.8-27B Q4, R9700/HRX/LOOM. Default on for supported configurations; use
--FlashPrefillV2=offfor dense prefill. - MTP and DFlash2 support: at 16K, prompt throughput reached 582.6 pp/s with MTP3 and 615.2 pp/s with DFlash2. Both on/off pairs matched the 64-token output and aggregate acceptance. Drafting and verification stay dense. Includes approximately 5× faster block selection and Q4 SwiGLU bias-tail unrolling.
- Previously observed DFlash2 peaks: 67.6 tok/s decode and 96% acceptance. These are separate workload peaks; the new sparse tests establish prefill gains, not a decode speedup. Perplexity is not yet measured.
LSE v0.4.22 — faster Q4 prefill
- Faster Q4 prefill with optimized nibble expansion and paired activation staging: 597.9 pp/s peak on a warm full prompt with BF16 KV and no prompt cache reuse.
- Fused eligible single-token Q4 SwiGLU gate/up projections and activation; updated dispatch regression coverage.
- Fixed RDNA4 prefetch address/span lowering in bundled Loom. Previously observed live DFlash2 peaks: 67.6 tok/s decode and 96% acceptance on Qwen3.8-27B Q4, R9700/macOS. These are separate workload peaks.
LSE v0.4.21 — joint decode attention and automatic LICM
- Faster attention: share K/V reads across query heads and token rows during speculative verification.
- Automatic Loom optimization: enable loop-invariant code motion, with cooperative staging for FP16, BF16, FP8 and BF8.
- Real-world peak performance: 67.6 output tokens/s, 390.2 prompt tokens/s, and 96% draft acceptance with Qwen3.8-27B Q4 + Q8 DFlash2 on AMD R9700, macOS HRX/Loom. Peaks from individual requests before the reported tool-call loop.
LSE v0.4.20 — K/V memory management and automatic split attention
K/V memory management and automatic split attention
- Share 256 MiB K/V arenas across layers and sessions, with 256 KiB fragments.
- Reuse available slots before allocating another arena.
- Grow K/V without copying existing cache data or changing its addresses.
- Retain fragment storage while a submitted or cached graph still uses it; release an arena after its final fragment is released.
- Bundle the matching HRX native-address API and Loom compiler support for macOS and Linux.
This change is scoped to paged attention K/V on the Loom backend. Weights, activations, general tensor allocation, K/V precision, and FP32 attention accumulation retain their existing policies. It does not offload active K/V to host RAM.
Validation and performance
On the local R9700, a Q4 target with Q8 DFlash2 and BF16 K/V reached 68,301 total tokens without an allocation failure. Peak VRAM was 27.83 GB, leaving at least 6.11 GB free in sampled counters. The 61,452-token request measured 165.88 PP/s and 26.21 TPS; the 68,173-token request measured 139.31 PP/s and 10.10 TPS. These repeated prompts had unusually high acceptance and are not representative coding workloads.
Fragment addressing added 0.7–8.9% to isolated attention at 16K context and 2.5% for eight queries at 68K. Outputs matched exactly. No matched small-context HTTP throughput comparison was made. These K/V growth measurements do not establish a throughput change.
Short-query split attention now derives its partition count from actual K/V table capacity and checks the device's LDS requirements. It no longer stops at 65,536 keys. Four-query tiling and empty-partition skipping cover compatible capacities automatically.
A same-executable attention check at 65,656 live keys measured 8.12 ms with fragmented split attention versus 71.79 ms with forced unsplit attention. A cold 64K HTTP comparison measured 13.87 TPS versus 5.72 TPS from the preserved earlier executable, with byte-identical response text. Prefill stayed near 146 PP/s. The executables include other revision differences, so the entire HTTP gain cannot be attributed solely to this dispatch change. Automatic split-attention evidence.
Vector staging and a shape-gated cache hint improve long-context Flash WMMA prefill. In a matched cold 64K HTTP pair, prefill rose from 146.06 to 224.02 tokens/s (+53.4%), and prompt time fell from 448.70 to 292.54 seconds. Decode measured 13.85 versus 13.79 tokens/s. At 1K, prefill measured 413.16 versus 408.43 tokens/s and decode 27.75 versus 27.55 tokens/s. Both pairs produced identical response text and acceptance counts with zero CPU fallback; each is a single comparison. Vector-staging evidence.
The bundled Loom compiler now performs cooperative matrix operand staging automatically for eligible FP16/BF16 loads and FP8/BF8 loads converted to FP16. Selection follows the access pattern, alignment, barrier safety, and shared-memory budget; it does not depend on model or kernel names. LSE no longer authors the staging operations directly. The long-prefill cache hint remains selected by LSE’s architecture and shape rules. In isolated 64K attention checks, the packed-load compiler path reduced FP8 kernel time from 559.6 to 457.4 ms (18.3%) and BF8 from 557.7 to 458.6 ms (17.8%), with bit-identical full outputs and zero spills. These are kernel measurements, not end-to-end throughput claims. Attention accumulation remains FP32.
The configured 262,100-token context limit is not a tested usable capacity. The final arena can be partially filled, and each K/V tensor can have less than one fragment of unused tail space.
K/V storage design · Measurement details
Correctness
HIP kernel caches now distinguish operand wiring and shared view ownership, and fused phases preserve in-place aliasing. Temporary-buffer lifetime accounting counts distinct consumers consistently when an operand is repeated. If a phase cannot execute and the graph is repartitioned, pending outputs receive independent storage and view bindings are refreshed; completed work, including a partly submitted joined run, remains covered during replay. This prevents the new execution order from reusing storage based on the abandoned schedule. HIP rejects fragmented K/V bindings explicitly; the fragmented storage layout uses Loom. Native F16 regression coverage checks K/V writes, attention, and prefill-to-decode state against CPU references.
Downloads
macOS ARM64 and Linux x86-64 archives include the matching HRX runtime and Loom compiler. Use the packaged launcher to select the bundled libraries. Checksums accompany each archive. macOS requires the installed MacAMDGPU driver; Linux requires compatible ROCm/HSA and GPU drivers.
Release verification
Built from 497b56f53e51b3121cd15f3ccb6eb49f1e7d3ae3. The Linux release workflow passed all 88 CTest executables plus native GPU engine and HTTP smoke checks. The macOS archive is checked for host behavior and relocation on an ARM64 runner; native GPU correctness and performance were separately measured on the local R9700.
LSE v0.4.19
LSE v0.4.19
Built from 3d9addd on the GPU families below.
What this build runs on
- Architectures: CDNA3, RDNA3.5, RDNA4
- GPU targets:
gfx942,gfx1150,gfx1151,gfx1200,gfx1201 - Platform: Linux x86_64
A GPU target that is not listed here has no ahead-of-time kernels in this build. The engine compiles its kernels at run time, so an unlisted device still runs, paying the compile once per kernel and caching it on disk.
Requirements
The archive bundles the selected HRX runtime and patched Loom compiler. Use the root or bin/ launchers to select the bundled libraries. Compatible Linux C/C++ libraries, ROCm 7.x, HSA and the GPU driver remain external requirements.
Make the installed ROCm libraries available to the dynamic loader. If they are not in its search path, set LD_LIBRARY_PATH to their installation directory, for example /opt/rocm/lib. An external HRX installation is not required for this archive.
Building from source needs a compiler with C++26 reflection support; see the README for the supported toolchain.
Changes since v0.4.18
Fixes
- Fix long-context KV growth and unify split attention
Measured performance (gfx1201 / R9700, Qwen3.8-27B-Q4)
Long-context split attention
The separate single-token and short-query split kernels are consolidated into
one partial and one merge family. Four verifier query rows share KV loads at
capacities through 65K keys. In matched cold 65,126-token HTTP requests on the
local R9700, Q4+DFlash2 decode rose from 8.51 to 14.57 tokens/s. Prefill was
147.7 versus 147.0 tokens/s, and the 64-token output matched. The prompts
repeat a short sequence; this is not a general chat throughput claim.
Method and limits.
Long-context KV growth
The macOS R9700 with a Q4 target, Q8 DFlash2 draft and BF16 KV completed
successive 30,726, 33,126, 37,126 and 41,126-token prompts in one HTTP
process. It then completed a 65,354-token prompt with 41,394 cached tokens.
Peak reserved device memory during that extension was 29.72 GB, leaving at
least 4.22 GB free in sampled counters. A 32-token continuation at 65K
completed without device allocation growth. The shorter seeded 8,192-token
request generated exactly the same 128-token response as v0.4.18.
The long prompts repeat a short synthetic sequence. Its decode rates are not
representative of general conversations. The configured 262,100-token KV limit
is not a measured usable capacity. Method and limits.
Weight slab VRAM reduction
The macOS HRX weight allocator now packs the Q4 target and Q8 DFlash2 draft
into 512 MiB slabs. In matched local R9700 runs, reserved GPU memory after
loading fell from 20.583 to 18.870 GB and after a request from 24.694 to
22.978 GB. Warm repeated requests measured 545.5 versus 548.3 prompt
tokens/s and 49.9 decode tokens/s in both builds. Generated output hashes and
80% DFlash2 acceptance matched. These numbers use decimal GB and driver
reserved-memory counters; timing does not show a throughput regression, but
does not establish identical performance for every workload.
Earlier prefill workspace memory fixes
Completed target graphs are released before the next request's prefill and
when the prefill chunk width changes. Consecutive full-size chunks keep replay.
DFlash2 releases completed large context-projection programs while preserving
narrow decode programs. Live recurrent state, KV pools, weights and compiled
kernels remain reusable.
The matched local macOS/R9700 test uses Qwen3.8-27B Q4, Q8 DFlash2, BF16 KV,
FP32 accumulation, batch/ubatch 1024, temperature 0.6, top-k 20, top-p 0.95,
seed 1234 and KV capacity 262100. A 5120-token request precedes a 6143-token
request with a 1023-token remainder. Each generates 128 tokens.
| Metric | v0.4.16 control | Memory fix |
|---|---|---|
| Reserved VRAM after ragged request | 29.77 GB | 28.26 GB |
| Sampled peak reserved VRAM | 30.09 GB | 28.87 GB |
| Ragged request prefill | 457.82 tok/s | 458.49 tok/s |
| Ragged request decode | 68.21 tok/s | 67.88 tok/s |
Responses and acceptance results match exactly. The synthetic decode request
has 100% proposal acceptance; its rate is not a general DFlash2 rate. The small
timing differences do not establish a throughput improvement or regression.
An eight-turn cached-follow-up comparison also matches all outputs and cache
lengths. Reserved VRAM grows only 12.14 MB over those turns. Aggregate decode
is 42.89 tok/s for control and 42.94 tok/s for the memory fix.
Measurements use memory-fix source ec71e14; v0.4.17 adds version, documentation
and packaging changes. Numbers use decimal GB and driver reserved-memory
counters sampled every 0.5 seconds. Sampling can miss instantaneous peaks.
The reported allocation failure has not been replayed. The result
establishes reduced workspace retention, not elimination of every possible OOM.
Full memory method and limits.
The earlier combined M8 gate/up kernels and all accepted typed attention,
activation panels, buffer views and speculative decoding paths remain active.
Earlier all-mode throughput
uses a different 1024-token workload and is not a new v0.4.17 measurement.
Verification
- Test suite: 100% tests passed, 0 tests failed out of 87
- End to end: greedy decode and the
/v1HTTP surface were exercised against a real model on the build machine - Built with: g++ 16.2.0, CMake 3.31.6, ROCm rocm-7.2.1
- Commit:
3d9addd7cbbd1816f9ded365403237b2c0384ad4 - Artifact:
lse-v0.4.19-linux-x86_64.tar.gz(25.8 MiB) - Verify the download with the accompanying
.sha256file:sha256sum -c <file>.sha256
LSE v0.4.18
LSE v0.4.18
Built from 83719e7 on the GPU families below.
What this build runs on
- Architectures: CDNA3, RDNA3.5, RDNA4
- GPU targets:
gfx942,gfx1150,gfx1151,gfx1200,gfx1201 - Platform: Linux x86_64
A GPU target that is not listed here has no ahead-of-time kernels in this build. The engine compiles its kernels at run time, so an unlisted device still runs, paying the compile once per kernel and caching it on disk.
Requirements
The archive bundles the selected HRX runtime and patched Loom compiler. Use the root or bin/ launchers to select the bundled libraries. Compatible Linux C/C++ libraries, ROCm 7.x, HSA and the GPU driver remain external requirements.
Make the installed ROCm libraries available to the dynamic loader. If they are not in its search path, set LD_LIBRARY_PATH to their installation directory, for example /opt/rocm/lib. An external HRX installation is not required for this archive.
Building from source needs a compiler with C++26 reflection support; see the README for the supported toolchain.
Changes since v0.4.17
Other
- Include weight memory report in macOS package
- Reduce unused VRAM in weight slabs
Measured performance (gfx1201 / R9700, Qwen3.8-27B-Q4)
Weight slab VRAM reduction
The macOS HRX weight allocator now packs the Q4 target and Q8 DFlash2 draft
into 512 MiB slabs. In matched local R9700 runs, reserved GPU memory after
loading fell from 20.583 to 18.870 GB and after a request from 24.694 to
22.978 GB. Warm repeated requests measured 545.5 versus 548.3 prompt
tokens/s and 49.9 decode tokens/s in both builds. Generated output hashes and
80% DFlash2 acceptance matched. These numbers use decimal GB and driver
reserved-memory counters; timing does not show a throughput regression, but
does not establish identical performance for every workload.
Earlier prefill workspace memory fixes
Completed target graphs are released before the next request's prefill and
when the prefill chunk width changes. Consecutive full-size chunks keep replay.
DFlash2 releases completed large context-projection programs while preserving
narrow decode programs. Live recurrent state, KV pools, weights and compiled
kernels remain reusable.
The matched local macOS/R9700 test uses Qwen3.8-27B Q4, Q8 DFlash2, BF16 KV,
FP32 accumulation, batch/ubatch 1024, temperature 0.6, top-k 20, top-p 0.95,
seed 1234 and KV capacity 262100. A 5120-token request precedes a 6143-token
request with a 1023-token remainder. Each generates 128 tokens.
| Metric | v0.4.16 control | Memory fix |
|---|---|---|
| Reserved VRAM after ragged request | 29.77 GB | 28.26 GB |
| Sampled peak reserved VRAM | 30.09 GB | 28.87 GB |
| Ragged request prefill | 457.82 tok/s | 458.49 tok/s |
| Ragged request decode | 68.21 tok/s | 67.88 tok/s |
Responses and acceptance results match exactly. The synthetic decode request
has 100% proposal acceptance; its rate is not a general DFlash2 rate. The small
timing differences do not establish a throughput improvement or regression.
An eight-turn cached-follow-up comparison also matches all outputs and cache
lengths. Reserved VRAM grows only 12.14 MB over those turns. Aggregate decode
is 42.89 tok/s for control and 42.94 tok/s for the memory fix.
Measurements use memory-fix source ec71e14; v0.4.17 adds version, documentation
and packaging changes. Numbers use decimal GB and driver reserved-memory
counters sampled every 0.5 seconds. Sampling can miss instantaneous peaks.
The reported allocation failure has not been replayed. The result
establishes reduced workspace retention, not elimination of every possible OOM.
Full memory method and limits.
The earlier combined M8 gate/up kernels and all accepted typed attention,
activation panels, buffer views and speculative decoding paths remain active.
Earlier all-mode throughput
uses a different 1024-token workload and is not a new v0.4.17 measurement.
Verification
- Test suite: 100% tests passed, 0 tests failed out of 87
- End to end: greedy decode and the
/v1HTTP surface were exercised against a real model on the build machine - Built with: g++ 16.2.0, CMake 3.31.6, ROCm rocm-7.2.1
- Commit:
83719e7789ee5d7cfb871d5c61ff0a39bd017b7c - Artifact:
lse-v0.4.18-linux-x86_64.tar.gz(25.6 MiB) - Verify the download with the accompanying
.sha256file:sha256sum -c <file>.sha256
Known long-context issue
A separate Pi session reported a GPU allocation failure on the request after
32,170 total tokens, despite the smaller weight slabs. v0.4.18 does not claim
to fix KV pool growth, retained graph memory, or this allocation failure.
The failure is under investigation.
LSE v0.4.17 — prefill memory bug fixes
LSE v0.4.17 — prefill memory bug fixes
Bug fixes
- Release the previous request's completed verifier graphs before a new prefill.
- Release completed target activation workspaces when prefill changes chunk width. Repeated full-size chunks preserve replay.
- Bound DFlash2's wide context-projection programs and release them after prefill. Preserve narrow context and draft programs used during decoding.
- Preserve materialized recurrent state and live KV allocations while releasing their obsolete graph owners.
- Include the memory comparison report in both platform archives and clarify which earlier release produced the retained throughput measurements.
Prefill workspace memory fixes
Completed target graphs are released before the next request's prefill and
when the prefill chunk width changes. Consecutive full-size chunks keep replay.
DFlash2 releases completed large context-projection programs while preserving
narrow decode programs. Live recurrent state, KV pools, weights and compiled
kernels remain reusable.
The matched local macOS/R9700 test uses Qwen3.8-27B Q4, Q8 DFlash2, BF16 KV,
FP32 accumulation, batch/ubatch 1024, temperature 0.6, top-k 20, top-p 0.95,
seed 1234 and KV capacity 262100. A 5120-token request precedes a 6143-token
request with a 1023-token remainder. Each generates 128 tokens.
| Metric | v0.4.16 control | Memory fix |
|---|---|---|
| Reserved VRAM after ragged request | 29.77 GB | 28.26 GB |
| Sampled peak reserved VRAM | 30.09 GB | 28.87 GB |
| Ragged request prefill | 457.82 tok/s | 458.49 tok/s |
| Ragged request decode | 68.21 tok/s | 67.88 tok/s |
Responses and acceptance results match exactly. The synthetic decode request
has 100% proposal acceptance; its rate is not a general DFlash2 rate. The small
timing differences do not establish a throughput improvement or regression.
An eight-turn cached-follow-up comparison also matches all outputs and cache
lengths. Reserved VRAM grows only 12.14 MB over those turns. Aggregate decode
is 42.89 tok/s for control and 42.94 tok/s for the memory fix.
Measurements use memory-fix source ec71e14; v0.4.17 adds version, documentation
and packaging changes. Numbers use decimal GB and driver reserved-memory
counters sampled every 0.5 seconds. Sampling can miss instantaneous peaks.
The reported allocation failure has not been replayed. The result
establishes reduced workspace retention, not elimination of every possible OOM.
Full memory method and limits.
The earlier combined M8 gate/up kernels and all accepted typed attention,
activation panels, buffer views and speculative decoding paths remain active.
Earlier all-mode throughput
uses a different 1024-token workload and is not a new v0.4.17 measurement.
Packaging and compatibility
Both macOS arm64 and Linux x86_64 archives are built from tag v0.4.17, source
e4a8eb1d1b7f59f2d20b831d94ed5c703f1ea6a4. Each has a SHA256 checksum and a
source/runtime manifest. Use the archive's bin/lse-server or bin/lse launcher.
macOS includes HSA, HRX, Loom and runtime dependencies. The installed and activated
MacAMDGPU driver is required separately. Linux includes HRX and Loom; compatible
ROCm/HSA and GPU drivers remain external requirements. Linux targets are gfx942,
gfx1150, gfx1151, gfx1200 and gfx1201.
Kernel-cache ownership advances to 0.4.17. Startup removes complete older
LSE-owned artifact families. The default remains ~/.lse/cache/; --cache-dir
overrides it. First use of the new version can include recompilation.
FP32 floating-point accumulation, model generation defaults, thinking output,
tool calls, OpenAI-compatible HTTP routes, MTP=3 and full-width DFlash2 remain
available. No kernel arithmetic or sampling change is introduced by this fix.
Verification
The local runtime and DFlash2 suites pass. Matched local GPU requests preserve
complete generated responses, cached-prefix lengths and acceptance statistics.
No additional perplexity run was made for these ownership changes.
Both platform workflows passed their build, test, packaging, and artifact-upload
steps. The downloaded macOS arm64 and Linux x86_64 archives passed checksum,
tag/source-manifest, patched-compiler, bundled-library, and relocated-launcher
checks. Archive SHA256 values are:
- macOS arm64:
0d1203a8039dd330ceeada8f08b2122fb408a5c30048ff0cd6d4b953d7396005 - Linux x86_64:
a80dcd830b5a516804141d7d79cd33517d8f1731e8632bbea1d90cf008448b52
The extracted macOS release archive also initialized the local R9700 and
completed a short Q4 + Q8 DFlash2 HTTP request (12 prompt tokens, 29 generated
tokens). This smoke check verifies package loading and execution, not
long-context stability or throughput.
One separate local long-context Pi session lost the GPU driver service at about
18K context tokens and returned hsa_amd_signal_wait_any index 4294967295.
That failure is not claimed fixed by this release. A replay using the saved
conversation structure and an adjusted system-prompt token budget completed at
21.8K tokens, including two further requests, with reserved VRAM stable near
31.14 GB. The original session's full expanded system prompt was unavailable,
so the replay does not rule out an intermittent device or runtime fault.
The hosted macOS runner verifies host behavior, native gfx1201 compilation,
bundled libraries and relocated launchers. It has no external AMD GPU; live GPU
execution was measured locally in the comparisons linked above.
LSE v0.4.16
LSE v0.4.16
Changes
- Activate the measured combined M8 Q4 gate/up kernel through the graph optimizer and central shape table.
- Remove one gate/up launch and the gate temporary while preserving FP32 accumulation order and exact outputs.
- Advance kernel cache ownership to 0.4.16 and refresh the README download and launch examples.
- Include the measured gate/up report in both platform archives.
Combined M8 Q4 gate/up execution
The optimizer selects one shared-panel SwiGLU consumer for the measured gfx1201
M8/N17408/K5120 shape. It removes one launch and the gate temporary. Independent
FP32 projection accumulation order and the existing SiLU intrinsic are preserved.
The architecture and shape rule is in the central dispatch table.
Full-size native component checks match all 139,264 output bits. Paired component
GPU time falls 27.69%, and wall time falls 20.50%. The combined kernel uses 120 VGPRs
versus 88 for each separate kernel, with zero private scratch and zero LDS. The
measured gains include the higher register pressure; occupied waves were not measured.
Measured workload and results
Measurements use source 257cc97f3bbc4fa1ae1877f068103415ea944857 and local server
SHA256 30aa8dd5651d7fd541a48f3a7fa0c35c782a13e6a1dcfcbe8cb8296c734e6582.
The release adds a version bump, documentation and packaging changes. These rates
were measured locally on macOS/R9700 gfx1201. They are not Linux throughput results.
Qwen3.8-27B Q4 target, Q8 auxiliary/draft weights, BF16 KV, FP32 floating
accumulation, temperature 0.6, top-k 20, top-p 0.95, seed 1234, batch/ubatch 1024 and
KV capacity 262100. Two 1024-token coding requests generate 384 tokens each, with
383 timed decode tokens. MTP uses depth 3; DFlash2 uses all seven proposals in block 8.
Each mode starts a new process with an empty private kernel cache. The resident
request reuses compiled code and zero prompt KV tokens. Compilation is included.
No profiler or competing GPU workload runs during throughput timing.
| Mode | Cold PP/s | Cold TPS | Resident PP/s | Resident TPS |
|---|---|---|---|---|
| Baseline | 446.28 | 24.23 | 634.34 | 24.86 |
| MTP=3 | 403.89 | 37.44 | 607.88 | 48.42 |
| DFlash2, seven proposals | 418.73 | 33.42 | 629.80 | 45.64 |
Matched resident DFlash2 improves 43.44 → 45.64 TPS (+5.07%). Both requests preserve
exact generated choices, acceptance counts and speculative pass counts. Verifier
time falls 70.7413 → 66.5850 ms per pass; draft time remains about 14.8 ms. All modes
record zero host groups and zero host fallbacks. MTP=3 remains faster on this prompt.
One cold/resident pair per mode does not establish long-context Pi throughput.
The 29/49/103 TPS targets remain unmet.
Full method and resource results.
Retained optimizations and compatibility
All accepted M8 down WMMA, M1024 prefill panels, GDN panel reuse, typed attention,
conditional DFlash2 sampling, memory retirement, buffer views, model generation
configuration and thinking/tool-call HTTP behavior remain active. No additional
perplexity run is needed for the bit-identical gate/up change.
Kernel cache ownership advances to 0.4.16. Startup removes only complete,
identifiable older LSE artifact families from the selected cache directory.
The default is ~/.lse/cache/; --cache-dir overrides it. Foreign, partial,
symlink and newer-version entries are preserved.
Use bin/lse or bin/lse-server in the archives. The macOS arm64 package bundles
HSA/HRX/Loom and the dependency closure; MacAMDGPU must be installed and activated
separately. The Linux x86_64 package bundles HRX/Loom and uses compatible installed
ROCm/HSA. Each archive contains source/runtime manifests and a SHA256 checksum.
The macOS build runner checks host behavior, native gfx1201 compilation and
relocated launchers. It has no external AMD GPU; actual GPU execution was checked
locally in the qualified component and HTTP runs above.
Release identity
Release source: 2e705a0c45953e2126825c95398e7b1455d00ccf. Performance measurements identify their exact earlier binary above.
Downloads and qualification
| Platform | Archive | SHA256 |
|---|---|---|
| macOS arm64 | Download | 644a4324eef06362437ec2e7c29ce9623d0ab8838ddce795f599505369492b13 |
| Linux x86_64 | Download | eda7a019b8a4adaadb5460f348d24a14782f2ad9339eb49651db587c0bd3061e |
Both downloaded build archives passed checksum, exact source, runtime/compiler
manifest, compiled version, launcher and README/report checks. Linux passed 87/87
tests and live GPU/HTTP checks. macOS passed 54 host suites and 4 HRX-linked checks,
native gfx1201 compilation, relocated launchers and Mach-O signature checks.
The macOS hosted runner has no AMD GPU. A second independent local relocation
check verified bundled library mappings and signatures without GPU initialization.
LSE v0.4.15
v0.4.15
Source: cb285b136fbb7d45b23ce4d0ffc7f0dfb0a4d665. All accepted v0.4.14 optimization schedules remain active.
Changes
- Share the existing M8 Q4 activation panel between GDN alpha and beta projections. QKV and the gate use the same producer. No new kernel body or persistent weight copy is added.
- Include the engine release version in compiled kernel cache identities, ownership metadata and artifact names. Updates use a new release namespace. Startup removes only complete identifiable older LSE artifact families from the selected cache directory. Same/newer releases, unrelated files, incomplete entries and symlinks remain.
- Keep
~/.lse/cache/as the automatic default;--cache-dirselects another directory. Retain compiler, device and exact-source validation and same-release reuse. - Serialize cache publication and restrict generated-source cleanup to its own identifiable files.
Qualification and measurements
The GDN component check includes activation preparation in both arms. Twenty paired rounds measured combined alpha/beta GPU time 0.050042→0.024935 ms (−50.17%), and wall time 0.237459→0.187337 ms (−21.11%). Complete fused output bits, panel codec, padding, readonly inputs, guards and replay pass. Real graph tests prove one producer shared by four projections. Registered native integration passes 24 device dispatches with zero host groups or fallbacks.
The cache suite passes 21/21 focused cases, including ownership cleanup, same-release reuse, directory selection, concurrent publication and codec separation. Independent review found no remaining blocking issue.
One matching DFlash2 HTTP cold/resident pair was run locally with the final source and bundled runtime. Requests have 1,024 input tokens, 384 generated tokens and 383 timed decode tokens; temperature 0.6, top-k 20, top-p 0.95, seed 1234, BF16 KV, batch/ubatch 1,024 and KV capacity 262,100. Each process starts with an empty private disk cache. The second request retains compiled code and reuses zero prompt KV. Compilation is included. No profiler, competing workload, host groups or fallbacks were recorded.
| DFlash2, seven proposals | v0.4.14 PP/s | v0.4.15 PP/s | v0.4.14 TPS | v0.4.15 TPS |
|---|---|---|---|---|
| Cold | 417.14 | 413.74 | 31.39 | 32.01 |
| Resident | 616.78 | 616.38 | 42.41 | 43.05 |
Both complete responses and acceptance statistics match v0.4.14 exactly. The resident decode difference is about +1.5% in this single pair; it is not a statistical or long-context Pi guarantee. The live GPU cache uses the new release namespace. No additional perplexity run is added for the bit-identical schedule change.
Baseline and MTP=3 were not rerun for this addition. Their preceding v0.4.14 resident results remain 624.10 PP/s / 24.73 TPS and 608.19 PP/s / 48.62 TPS, respectively. These historical numbers are not new v0.4.15 measurements. Baseline 29 TPS, MTP=3 49 TPS and DFlash2 103 TPS targets remain unmet.
A private paired-load WMMA layout probe was correct but measured about 1.1% more GPU time and 1.4% more wall time. It was never selected in production. The accepted M8 down WMMA and other qualified paths remain active.
Archives and requirements
Use the macOS arm64 or Linux x86_64 archive and its SHA256 checksum. Both archives are built from the exact tagged source. Linux passed 86/86 tests, live GPU/HTTP smoke and relocated loader checks. macOS passed 53 host suites, four HRX fixtures, 89 CLI cases, native gfx1201 compilation and relocated launcher/signature checks. Both source manifests and bundled runtime identities are verified.
| Archive | Verified SHA256 |
|---|---|
lse-v0.4.15-linux-x86_64.tar.gz |
f9656a7ab07720d992dc243e06018a5ab7d696287d764d0682f4efa734de5790 |
lse-v0.4.15-macos-arm64.tar.gz |
e6af6ddf4402c62012a4ca4fa6c6f464df9bfc482607b5d27e8b876e986ea92b |
The macOS archive bundles HSA, HRX, Loom and runtime dependencies. It requires an activated compatible MacAMDGPU DriverKit extension, which the archive does not install. The macOS binaries target macOS 15 or later; GPU execution also requires a macOS version supported by MacAMDGPU.
The Linux archive bundles its selected HRX runtime and patched Loom compiler. Compatible ROCm 7.x, HSA, the GPU driver and Linux C/C++ runtime remain required. The root and bin/ launchers select bundled libraries and preserve CLI parameters. Each BUILD.json records source pins, patches and library identities.
Remote macOS workflow checks do not establish GPU inference throughput. Local gfx1201 HTTP/native qualification is reported separately above.
HTTP launch
Use lse-server with your local Q4 model, --pool hrx:0 --dialect loom --batch-size 1024 --ubatch-size 1024 --temperature 0.6 --kv-cache-dtype bf16 --kv-len 262100. For DFlash2, add --dflash2=on --dflash2-model PATH. For MTP=3, use --mtp PATH --mtp-depth 3. Explicit request sampling settings override launch defaults.
LSE v0.4.14
LSE v0.4.14
Source: c11f103c86627e66c36d1b109fb02281c40ad686.
Improvements
c1ee42d: M1024 Q4 down projection shares activation fragments and reuses weight groups across four M16 row blocks.67a0b4e: WMMA paged attention skips fully masked 256-key windows with a workgroup-uniform predicate. Contributing windows retain their arithmetic and ordering.de3afcfand8150263: M1024 Q4 gate/up and the measured QKV, GDN and attention projection shapes stage activation fragments once in LDS for column waves to share. Inputs share a producer; the central shape table checks barrier support and the LDS budget.c11f103: Two M8 Q4 shapes consume independent activation row pairs while reusing decoded weights and wider activation loads. The per-row chunk/FMA and split-reduction order is preserved. Accepted down-projection bodies and unrelated shapes remain unchanged.
Floating-point accumulation remains FP32. Compact Q4 weights are unchanged. Complete output bits, independent arithmetic reference, codec, fused residual, replay, readonly-input and buffer-guard checks pass. Measured component kernels report zero private scratch. Focused suites passed: activation panel 9/9, matrix panel 10/10 and dispatch 10/10. No additional perplexity run was needed for these bit-identical schedules.
Final same-binary measurements
macOS/R9700 gfx1201, Qwen Q4 target, Q8 auxiliary/draft weights, BF16 KV, temperature 0.6, top-k 20, top-p 0.95, seed 1234, batch/ubatch 1024 and configured KV capacity 262100. Each 1024-token coding request generates 384 tokens (383 timed decode tokens). MTP uses depth 3; DFlash2 uses all seven proposals in block 8.
Each mode starts with an empty private kernel cache. The resident request retains compiled code and reuses zero prompt KV. Compilation is included. Requests execute with zero host groups or fallbacks, without profiling or a competing workload. One pair per mode is not a statistical estimate or a long-context Pi replay; rates vary with prompt, sampling and live context.
| Mode | Cold prefill tokens/s | Cold decode tokens/s | Resident prefill tokens/s | Resident decode tokens/s |
|---|---|---|---|---|
| Baseline | 442.18 | 24.14 | 624.10 | 24.73 |
| MTP=3 | 403.95 | 37.96 | 608.19 | 48.62 |
| DFlash2, seven proposals | 417.14 | 31.39 | 616.78 | 42.41 |
All three controlled resident requests exceed 600 prefill tokens/s. The 29 TPS baseline, 49 TPS MTP and 103 TPS DFlash2 goals remain unmet.
The matched projection change improved resident DFlash2 prefill from 568.03 to 619.08 tokens/s (+8.99%), with decode effectively unchanged. The subsequent M8 row-pair change improved matched resident decode from 40.68 to 42.41 tokens/s (+4.27%), with prefill effectively unchanged. Each scoped comparison preserved exact responses and acceptance/probability statistics. Across different generation modes, RNG execution can produce different answers.
Final mode comparison, M8 row-pair methods, prefill projection methods.
Packages
Both archives are built from the exact source above and include a runtime/source manifest and SHA256 file.
- Linux x86_64 bundles its selected HRX runtime and patched Loom compiler. Root and
bin/launchers select them. Compatible Linux C/C++ libraries, ROCm 7.x, HSA and the GPU driver remain external requirements. - macOS arm64 bundles HSA, HRX, patched Loom and its runtime dependency closure. Use
bin/lseorbin/lse-server; install and activate MacAMDGPU separately. Native binaries target Apple Silicon macOS 15 or newer; GPU execution requires the driver-supported macOS version.
K/V defaults to model-declared BF16 for BF16 checkpoints, and FP16 otherwise, with explicit overrides. Seven-proposal DFlash2 and batch/ubatch 1024 remain defaults. --temperature sets the server default; request values take precedence and launch examples use 0.6.
Qualification
- Linux final archive build: all 86 tests, selected compiler/runtime verification, live engine/HTTP smoke and relocated launcher/loader checks passed. Selected targets: gfx942, gfx1150, gfx1151, gfx1200 and gfx1201.
- macOS final archive build: 53 CPU-only tests, four HRX-linked host fixtures, 89 CLI cases, native gfx1201 compilation and runtime relocation passed. The hosted runner has no external AMD GPU; actual execution was qualified locally for the accepted source changes.
- Both archives were checked against their SHA256 files, exact source/compiler/patch manifests and bundled runtime hashes. Relocated macOS launchers, signatures and HSA loading passed without opening a GPU. Uploaded asset bytes and tag identity were verified before publication.
Verify downloads with the accompanying .sha256 files.