Skip to content

0.30.0

@michalharakal michalharakal tagged this 14 Jun 19:03
Version-aligned with SKaiNET 0.30.0 (Q5_K packed matmul, NEON native
kernels, Kotlin/Native cinterop). Skips 0.29.x — tracked internally without
a tagged release.

Highlights
- Q5_K stays packed in the eager Gemma runtime. GemmaMemSegConverter now
  relayouts GGUF bytes to block-major and wraps them as Q5_KBlockTensorData
  (176 B/block) instead of dequantizing to FP32 on load, reaching the engine's
  new Q5KMatmulKernel via DefaultCpuOps. FunctionGemma-270M (Q5_K_M) decodes
  byte-identically to the FP32 baseline (GemmaQ5KPackedParityTest).
- Gemma NATIVE_OPTIMIZED path is Kotlin/Native-ready. The layout + packing
  helpers (GemmaQuantLayout.kt, GemmaPackedWeights.kt) moved to commonMain and
  GemmaNetworkLoader.load() runs convertGemmaWeightsPacked — the board binary
  keeps K-quant weights packed with no java.lang.foreign MemSeg dependency.
  Verified on JVM and linuxX64.

Fixes
- Kernel-less quant types under NATIVE_OPTIMIZED dequant to FP32 [out, in]
  instead of crashing on a rank-1 transpose.
- DecoderGgufMemSegConverter dequantizes Q4_1 and every other non-packed quant
  type instead of passing raw bytes through to a matmul crash (#654).

Build / engine
- skainet pin 0.28.1 -> 0.30.0; VERSION_NAME 0.28.1 -> 0.30.0.
- mavenLocal()-first dev shim reverted; release resolves the engine from
  Maven Central.

Validation
- ./gradlew build green (compile + unit tests + apiCheck).
- Full integration suite green: gemma jvmTest -PincludeIntegration, 87 tests,
  6 skipped, 0 failures (GemmaQ5KPackedParityTest + the RealGemma* real-model
  tests run).

See CHANGELOG.md [0.30.0] for the full entry.
Assets 2
Loading