LSE v0.4.17 — prefill memory bug fixes
LSE v0.4.17 — prefill memory bug fixes
Bug fixes
- Release the previous request's completed verifier graphs before a new prefill.
- Release completed target activation workspaces when prefill changes chunk width. Repeated full-size chunks preserve replay.
- Bound DFlash2's wide context-projection programs and release them after prefill. Preserve narrow context and draft programs used during decoding.
- Preserve materialized recurrent state and live KV allocations while releasing their obsolete graph owners.
- Include the memory comparison report in both platform archives and clarify which earlier release produced the retained throughput measurements.
Prefill workspace memory fixes
Completed target graphs are released before the next request's prefill and
when the prefill chunk width changes. Consecutive full-size chunks keep replay.
DFlash2 releases completed large context-projection programs while preserving
narrow decode programs. Live recurrent state, KV pools, weights and compiled
kernels remain reusable.
The matched local macOS/R9700 test uses Qwen3.8-27B Q4, Q8 DFlash2, BF16 KV,
FP32 accumulation, batch/ubatch 1024, temperature 0.6, top-k 20, top-p 0.95,
seed 1234 and KV capacity 262100. A 5120-token request precedes a 6143-token
request with a 1023-token remainder. Each generates 128 tokens.
| Metric | v0.4.16 control | Memory fix |
|---|---|---|
| Reserved VRAM after ragged request | 29.77 GB | 28.26 GB |
| Sampled peak reserved VRAM | 30.09 GB | 28.87 GB |
| Ragged request prefill | 457.82 tok/s | 458.49 tok/s |
| Ragged request decode | 68.21 tok/s | 67.88 tok/s |
Responses and acceptance results match exactly. The synthetic decode request
has 100% proposal acceptance; its rate is not a general DFlash2 rate. The small
timing differences do not establish a throughput improvement or regression.
An eight-turn cached-follow-up comparison also matches all outputs and cache
lengths. Reserved VRAM grows only 12.14 MB over those turns. Aggregate decode
is 42.89 tok/s for control and 42.94 tok/s for the memory fix.
Measurements use memory-fix source ec71e14; v0.4.17 adds version, documentation
and packaging changes. Numbers use decimal GB and driver reserved-memory
counters sampled every 0.5 seconds. Sampling can miss instantaneous peaks.
The reported allocation failure has not been replayed. The result
establishes reduced workspace retention, not elimination of every possible OOM.
Full memory method and limits.
The earlier combined M8 gate/up kernels and all accepted typed attention,
activation panels, buffer views and speculative decoding paths remain active.
Earlier all-mode throughput
uses a different 1024-token workload and is not a new v0.4.17 measurement.
Packaging and compatibility
Both macOS arm64 and Linux x86_64 archives are built from tag v0.4.17, source
e4a8eb1d1b7f59f2d20b831d94ed5c703f1ea6a4. Each has a SHA256 checksum and a
source/runtime manifest. Use the archive's bin/lse-server or bin/lse launcher.
macOS includes HSA, HRX, Loom and runtime dependencies. The installed and activated
MacAMDGPU driver is required separately. Linux includes HRX and Loom; compatible
ROCm/HSA and GPU drivers remain external requirements. Linux targets are gfx942,
gfx1150, gfx1151, gfx1200 and gfx1201.
Kernel-cache ownership advances to 0.4.17. Startup removes complete older
LSE-owned artifact families. The default remains ~/.lse/cache/; --cache-dir
overrides it. First use of the new version can include recompilation.
FP32 floating-point accumulation, model generation defaults, thinking output,
tool calls, OpenAI-compatible HTTP routes, MTP=3 and full-width DFlash2 remain
available. No kernel arithmetic or sampling change is introduced by this fix.
Verification
The local runtime and DFlash2 suites pass. Matched local GPU requests preserve
complete generated responses, cached-prefix lengths and acceptance statistics.
No additional perplexity run was made for these ownership changes.
Both platform workflows passed their build, test, packaging, and artifact-upload
steps. The downloaded macOS arm64 and Linux x86_64 archives passed checksum,
tag/source-manifest, patched-compiler, bundled-library, and relocated-launcher
checks. Archive SHA256 values are:
- macOS arm64:
0d1203a8039dd330ceeada8f08b2122fb408a5c30048ff0cd6d4b953d7396005 - Linux x86_64:
a80dcd830b5a516804141d7d79cd33517d8f1731e8632bbea1d90cf008448b52
The extracted macOS release archive also initialized the local R9700 and
completed a short Q4 + Q8 DFlash2 HTTP request (12 prompt tokens, 29 generated
tokens). This smoke check verifies package loading and execution, not
long-context stability or throughput.
One separate local long-context Pi session lost the GPU driver service at about
18K context tokens and returned hsa_amd_signal_wait_any index 4294967295.
That failure is not claimed fixed by this release. A replay using the saved
conversation structure and an adjusted system-prompt token budget completed at
21.8K tokens, including two further requests, with reserved VRAM stable near
31.14 GB. The original session's full expanded system prompt was unavailable,
so the replay does not rule out an intermittent device or runtime fault.
The hosted macOS runner verifies host behavior, native gfx1201 compilation,
bundled libraries and relocated launchers. It has no external AMD GPU; live GPU
execution was measured locally in the comparisons linked above.