LSE v0.4.20 — K/V memory management and automatic split attention
K/V memory management and automatic split attention
- Share 256 MiB K/V arenas across layers and sessions, with 256 KiB fragments.
- Reuse available slots before allocating another arena.
- Grow K/V without copying existing cache data or changing its addresses.
- Retain fragment storage while a submitted or cached graph still uses it; release an arena after its final fragment is released.
- Bundle the matching HRX native-address API and Loom compiler support for macOS and Linux.
This change is scoped to paged attention K/V on the Loom backend. Weights, activations, general tensor allocation, K/V precision, and FP32 attention accumulation retain their existing policies. It does not offload active K/V to host RAM.
Validation and performance
On the local R9700, a Q4 target with Q8 DFlash2 and BF16 K/V reached 68,301 total tokens without an allocation failure. Peak VRAM was 27.83 GB, leaving at least 6.11 GB free in sampled counters. The 61,452-token request measured 165.88 PP/s and 26.21 TPS; the 68,173-token request measured 139.31 PP/s and 10.10 TPS. These repeated prompts had unusually high acceptance and are not representative coding workloads.
Fragment addressing added 0.7–8.9% to isolated attention at 16K context and 2.5% for eight queries at 68K. Outputs matched exactly. No matched small-context HTTP throughput comparison was made. These K/V growth measurements do not establish a throughput change.
Short-query split attention now derives its partition count from actual K/V table capacity and checks the device's LDS requirements. It no longer stops at 65,536 keys. Four-query tiling and empty-partition skipping cover compatible capacities automatically.
A same-executable attention check at 65,656 live keys measured 8.12 ms with fragmented split attention versus 71.79 ms with forced unsplit attention. A cold 64K HTTP comparison measured 13.87 TPS versus 5.72 TPS from the preserved earlier executable, with byte-identical response text. Prefill stayed near 146 PP/s. The executables include other revision differences, so the entire HTTP gain cannot be attributed solely to this dispatch change. Automatic split-attention evidence.
Vector staging and a shape-gated cache hint improve long-context Flash WMMA prefill. In a matched cold 64K HTTP pair, prefill rose from 146.06 to 224.02 tokens/s (+53.4%), and prompt time fell from 448.70 to 292.54 seconds. Decode measured 13.85 versus 13.79 tokens/s. At 1K, prefill measured 413.16 versus 408.43 tokens/s and decode 27.75 versus 27.55 tokens/s. Both pairs produced identical response text and acceptance counts with zero CPU fallback; each is a single comparison. Vector-staging evidence.
The bundled Loom compiler now performs cooperative matrix operand staging automatically for eligible FP16/BF16 loads and FP8/BF8 loads converted to FP16. Selection follows the access pattern, alignment, barrier safety, and shared-memory budget; it does not depend on model or kernel names. LSE no longer authors the staging operations directly. The long-prefill cache hint remains selected by LSE’s architecture and shape rules. In isolated 64K attention checks, the packed-load compiler path reduced FP8 kernel time from 559.6 to 457.4 ms (18.3%) and BF8 from 557.7 to 458.6 ms (17.8%), with bit-identical full outputs and zero spills. These are kernel measurements, not end-to-end throughput claims. Attention accumulation remains FP32.
The configured 262,100-token context limit is not a tested usable capacity. The final arena can be partially filled, and each K/V tensor can have less than one fragment of unused tail space.
K/V storage design · Measurement details
Correctness
HIP kernel caches now distinguish operand wiring and shared view ownership, and fused phases preserve in-place aliasing. Temporary-buffer lifetime accounting counts distinct consumers consistently when an operand is repeated. If a phase cannot execute and the graph is repartitioned, pending outputs receive independent storage and view bindings are refreshed; completed work, including a partly submitted joined run, remains covered during replay. This prevents the new execution order from reusing storage based on the abandoned schedule. HIP rejects fragmented K/V bindings explicitly; the fragmented storage layout uses Loom. Native F16 regression coverage checks K/V writes, attention, and prefill-to-decode state against CPU references.
Downloads
macOS ARM64 and Linux x86-64 archives include the matching HRX runtime and Loom compiler. Use the packaged launcher to select the bundled libraries. Checksums accompany each archive. macOS requires the installed MacAMDGPU driver; Linux requires compatible ROCm/HSA and GPU drivers.
Release verification
Built from 497b56f53e51b3121cd15f3ccb6eb49f1e7d3ae3. The Linux release workflow passed all 88 CTest executables plus native GPU engine and HTTP smoke checks. The macOS archive is checked for host behavior and relocation on an ARM64 runner; native GPU correctness and performance were separately measured on the local R9700.