Highlights
- Apertus end-to-end. Real-GGUF loading on top of skainet 0.23.x's
block-major Q4_K TensorData wiring, routed through OptimizedLLMRuntime
+ apertusNetwork(). Chat template, tool calling, and integration tests
against Apertus-8B-Q4_K_S. See APERTUS_ROLLOUT.md.
- Gemma 4 chat-model JVM facade (Gemma4ChatModel) for embedded text-only
deployments; close() propagates to the mmap arena; PLE mmap path now
consumes upstream loadTensorStorageMapped.
- Multi-id EOS / stop-token support in the chat layer.
- Tokenizer auto-detect for SentencePiece in fromTokenizerJson.
- New end-to-end smoke test in llm-test/llm-test-java that wires LEAF
(mdbr-leaf-mt via KBertJava) and Llama 3.2-1B (KLlamaJava) in one JVM,
gated on env vars / cache fallbacks.
- Apertus tool calling as a first-class family alongside Llama 3, Gemma 4,
Qwen, and ChatML/Hermes.
- kllama-cli + skainet-cli shadow-jar ServiceLoader fix-up so the
priority-100 skainet-backend-native-cpu provider is picked up at runtime.
Fixes
- fix(apertus): force-dequant token_embd under NATIVE_OPTIMIZED.
- fix(tokenizer): auto-detect SentencePiece marker in fromTokenizerJson.
- fix(gemma4): produce coherent text on real SafeTensors checkpoint.
- fix(apertus): route through OptimizedLLMRuntime + apertusNetwork().
Build / version
- VERSION_NAME 0.21.1 → 0.23.1; skainet pin 0.23.0 → 0.23.1.
- llm-test/llm-test-java maxHeapSize 8g → 16g (Llama 3.2-1B + LEAF in one JVM).
- No 0.22.x transformers release was tagged; the version line jumps to
re-sync with the engine.
See CHANGELOG.md for the full list of changes.