Version-aligned with SKaiNET 0.31.0. Completes the eager board-decode path for
FunctionGemma on the Astra Machina SL2610.
Added
- maxInferenceLen on GemmaNetworkLoader.load() — optional cap on the context the
eager network sizes its KV cache + RoPE tables for (default
min(contextLength, 4096)), threaded through applyWeightsToNetwork ->
gemmaNetwork. A constrained-device consumer can pass a small value (e.g. 32 for
a short tool-call prompt) to shrink the KV cache ~100x, which otherwise
allocates ~0.4 GB at the first forward and OOMs the 1.9 GB board. (#180)
Fixed
- Tied Q8_0 lm_head stays packed in the eager NATIVE_OPTIMIZED Gemma path.
FunctionGemma's token_embd is Q8_0 and tied, so convertGemmaWeightsPacked was
dequantizing BOTH token_embd and output to FP32 (2x~0.67 GB) -> OOM on the
board. output/lm_head now packs as Q8_0 (packGemmaKQuant gained a Q8_0 case;
the row-major->block-major relayout is generalized with a blockSize param) and
runs on the (NEON) Q8_0 kernel; token_embd stays FP32 (it is gathered) but is
wrapped no-copy. Tied embed/lm_head footprint ~1.34 GB -> ~0.76 GB. Verified
byte-identical decode parity (GemmaQ5KPackedParityTest) and a stable ~1.06 GB
load on the SL2610. Requires engine 0.31.0's ops.transpose fix for packed
Q8_0/Q4_0 weights. (#179)
Changed
- skainet pin 0.30.0 -> 0.31.0; json-schema-validator -> 3.0.4 (#175).
- CI: opt out of AGP's config-time dependency-resolution check (gradle#31483) in
the project gradle.properties so the Kotlin/JS + Wasm npm resolution doesn't
fail assemble/allTests.
See CHANGELOG.md [0.31.0].