Repository navigation
Release 1.0.0 (09-aug-2026)
[1.0.0] — 2026-08-09
First public release of the nano-vllm-java CPU inference library
(JPMS module com.igormaznitsa.nanollvm, Maven coordinates com.igormaznitsa:nano-vllm-java:1.0.0).
Added
- Pure Java 21+ offline LLM inference: continuous batching, paged KV cache, Hugging Face safetensors and GGUF weights (Qwen3, Gemma3, LFM2).
- Shared models via
LlmModelFactory.makeand per-engineLLM.Builder/LLM(chat, one-shot, completion, cancel, timeout, generation stats). LlmModelisAutoCloseable: close engines first, then the shared model, to release weight resources; closed models and engines reject further use (isClosed()on both).- Chat sessions with history limits, listeners, optional advisors and mixers, plus lexical BM25 RAG (
RagFactory/RagSession). - Process-wide
ResourceLimitsfor file, PDF, corpus, JSON, GGUF, safetensors, and history budgets (overridable per process or per corpus). - Configuration knobs including
kvHeapFractionfor heap-based KV auto-sizing, CPU matmul thread control, andnanollvm.*/NANOLLVM_*system properties and environment variables. - Documented RAG/advisor trust boundary and optional suppression of prepared-prompt debug events (
emitDebugPrompts(false)). - Samples: minimal Gemma
HelloWorld, log-triage demo, interactiveExample, andBench.
Changed
- Explicit
.cpuThreads(N)/.disableMultiCpu()wins over-Dnanollvm.cpu.threads; sequential mode creates no matmul executor.