Skip to content

LSE v0.4.15

Choose a tag to compare

@Geramy Geramy released this 29 Sep 11:35

v0.4.15

Source: cb285b136fbb7d45b23ce4d0ffc7f0dfb0a4d665. All accepted v0.4.14 optimization schedules remain active.

Changes

  • Share the existing M8 Q4 activation panel between GDN alpha and beta projections. QKV and the gate use the same producer. No new kernel body or persistent weight copy is added.
  • Include the engine release version in compiled kernel cache identities, ownership metadata and artifact names. Updates use a new release namespace. Startup removes only complete identifiable older LSE artifact families from the selected cache directory. Same/newer releases, unrelated files, incomplete entries and symlinks remain.
  • Keep ~/.lse/cache/ as the automatic default; --cache-dir selects another directory. Retain compiler, device and exact-source validation and same-release reuse.
  • Serialize cache publication and restrict generated-source cleanup to its own identifiable files.

Qualification and measurements

The GDN component check includes activation preparation in both arms. Twenty paired rounds measured combined alpha/beta GPU time 0.050042→0.024935 ms (−50.17%), and wall time 0.237459→0.187337 ms (−21.11%). Complete fused output bits, panel codec, padding, readonly inputs, guards and replay pass. Real graph tests prove one producer shared by four projections. Registered native integration passes 24 device dispatches with zero host groups or fallbacks.

The cache suite passes 21/21 focused cases, including ownership cleanup, same-release reuse, directory selection, concurrent publication and codec separation. Independent review found no remaining blocking issue.

One matching DFlash2 HTTP cold/resident pair was run locally with the final source and bundled runtime. Requests have 1,024 input tokens, 384 generated tokens and 383 timed decode tokens; temperature 0.6, top-k 20, top-p 0.95, seed 1234, BF16 KV, batch/ubatch 1,024 and KV capacity 262,100. Each process starts with an empty private disk cache. The second request retains compiled code and reuses zero prompt KV. Compilation is included. No profiler, competing workload, host groups or fallbacks were recorded.

DFlash2, seven proposals v0.4.14 PP/s v0.4.15 PP/s v0.4.14 TPS v0.4.15 TPS
Cold 417.14 413.74 31.39 32.01
Resident 616.78 616.38 42.41 43.05

Both complete responses and acceptance statistics match v0.4.14 exactly. The resident decode difference is about +1.5% in this single pair; it is not a statistical or long-context Pi guarantee. The live GPU cache uses the new release namespace. No additional perplexity run is added for the bit-identical schedule change.

Baseline and MTP=3 were not rerun for this addition. Their preceding v0.4.14 resident results remain 624.10 PP/s / 24.73 TPS and 608.19 PP/s / 48.62 TPS, respectively. These historical numbers are not new v0.4.15 measurements. Baseline 29 TPS, MTP=3 49 TPS and DFlash2 103 TPS targets remain unmet.

A private paired-load WMMA layout probe was correct but measured about 1.1% more GPU time and 1.4% more wall time. It was never selected in production. The accepted M8 down WMMA and other qualified paths remain active.

Archives and requirements

Use the macOS arm64 or Linux x86_64 archive and its SHA256 checksum. Both archives are built from the exact tagged source. Linux passed 86/86 tests, live GPU/HTTP smoke and relocated loader checks. macOS passed 53 host suites, four HRX fixtures, 89 CLI cases, native gfx1201 compilation and relocated launcher/signature checks. Both source manifests and bundled runtime identities are verified.

Archive Verified SHA256
lse-v0.4.15-linux-x86_64.tar.gz f9656a7ab07720d992dc243e06018a5ab7d696287d764d0682f4efa734de5790
lse-v0.4.15-macos-arm64.tar.gz e6af6ddf4402c62012a4ca4fa6c6f464df9bfc482607b5d27e8b876e986ea92b

The macOS archive bundles HSA, HRX, Loom and runtime dependencies. It requires an activated compatible MacAMDGPU DriverKit extension, which the archive does not install. The macOS binaries target macOS 15 or later; GPU execution also requires a macOS version supported by MacAMDGPU.

The Linux archive bundles its selected HRX runtime and patched Loom compiler. Compatible ROCm 7.x, HSA, the GPU driver and Linux C/C++ runtime remain required. The root and bin/ launchers select bundled libraries and preserve CLI parameters. Each BUILD.json records source pins, patches and library identities.

Remote macOS workflow checks do not establish GPU inference throughput. Local gfx1201 HTTP/native qualification is reported separately above.

HTTP launch

Use lse-server with your local Q4 model, --pool hrx:0 --dialect loom --batch-size 1024 --ubatch-size 1024 --temperature 0.6 --kv-cache-dtype bf16 --kv-len 262100. For DFlash2, add --dflash2=on --dflash2-model PATH. For MTP=3, use --mtp PATH --mtp-depth 3. Explicit request sampling settings override launch defaults.