Repository navigation
Release 1.2.0 (22-aug-2026)
[1.2.0] — 2026-08-22
Public release of nano-vllm-java 1.2.0 (Maven coordinates com.igormaznitsa:nano-vllm-java:1.2.0).
SentencePiece and extra tokenizer families, RAG load tuners, engine-owned matmul pool, deterministic sampling, checkpoint modalities, and close() resource reclaim.
Added
-
LlmModel.modalities()(andinputModalities()/outputModalities()): input and output
content types asLlmModality(TEXT,IMAGE,AUDIO,VIDEO,EMBEDDING). Values follow
the checkpoint config (Gemma 4 QAT mobile declares text+image+audio+video in). Embedding
encoders are text→embedding.LlmModel.usableModalities()is what this library actually runs
(text→text chat, or text→embedding); vision/audio towers are still skipped at load.
LlmModel.toString()includesmodalities=andusable=when they differ. TheExample
demo prints checkpoint modalities after load, plus the runtime line when they are not the same. -
RAG user turns with retrieved passages now start with
RagPrompts.GROUNDINGso small chat models
are told to answer from those lines only and not invent a book or play. -
HybridRagIndex.of(RagIndex…)fuses any two or more indexes with RRF (nested hybrids flatten).
BM25+dense viaRagFactory.withEmbeddingsis unchanged. To mix disk folders, classpath files,
and inline strings, add them on oneRagFactory.builder()—HybridRagIndexdoes not concatenate
corpora.Builder.addFolderswalks several directories in one call. -
LLM.Builder.deterministic()(alsoSamplingParams/ChatSession/RagSession): same prompt
always picks the highest-logit token (topK = 1, nucleus off).temperature(0)stays rejected. -
LLM.Builder.dedicatedMatmulPool(): a bounded matmul thread pool owned by that engine and shut
down onclose(), so a server need not join the process-widenanollvm-matmul-*pool or hand
the library a foreignExecutorService. Combining it withmatmulExecutorfails atbuild(). -
Optional download scripts for intfloat/multilingual-e5-small
(models/download-multilingual-e5-small.sh/.ps1/.cmd→models/multilingual-e5-small/, ONNX fp32 ~470 MB).
Hugging Face Unigram SentencePiece tokenizers load fromtokenizer.json; embedding wrap accepts XLM-R
<s>/</s>as well as BERT[CLS]/[SEP]. Optimum-style BERT ONNX (anonymousonnx::MatMul_*
weights) remaps those MatMul aliases through the BERT schema, same as named HF tensors.
EmbeddingsHelloWorlddefaults to this folder and prefixes non-retrieval text withquery:.
Unigram now applies the Hugging Face precompiled charsmap (needed for many non-ASCII XLM-R / E5
strings). Folders withouttokenizer.jsonload SentencePiecetokenizer.model. WordPiece,
WordLevel, and character models are recognized fromtokenizer.jsonmodel.type.
Tokenizer.fromSentencePiecebuilds a tokenizer from protobuf bytes. -
Load-time RAG tuners on
RagFactory.Builder.addProcessor: skip files, supply custom document
text, or rewrite extracted strings before chunking. Several tuners run in order.
SampleRagTunerHelloWorldextracts a bundled Project Gutenberg EPUB of Karel Čapek's R.U.R.
(JDK zip + StAX, no epub4j / xpp3), keeps the short OPF title (before a subtitle slash) as a
Markdown heading so later chunks still name the play, indexes with BM25, and asks questions from that play with
LLM.Builder.deterministic()so repeats pick the same tokens.DenseRagIndex.of/HybridRagIndex.ofaccept
a per-passage embed callback and an optional callerExecutorfor parallel indexing (sequential
when omitted).
Removed
- Built-in
PdfTextExtractor. Folder walks no longer pick up.pdfby default. Index PDFs (or
other binaries) withRagTuner.extracting, same pattern as the EPUB sample.ResourceLimits
drops PDF inflate / ToUnicode cmap caps (maxPdfInflateBytes,maxCmapRangeSpan,
maxCmapEntries).
Fixed
- Closing a model, engine, or weight reader now drops file buffers, KV pages, rotary tables, and
the last shared matmul pool instead of pinning them until process exit. GGUF and safetensors files
≤ 2 GiB are copied into heap soclose()can reclaim them; shards larger than that still use a
positioned file channel. Late unpack releases packed bytes when no other engine is using the model.
Changed
- Dense RAG query embedding stays concurrent (each
LlmModel.embeduses a fresh step context).
Index-time passage embedding is sequential unless the caller supplies anExecutor. - Extra CPU cores now work on long prompts, not only GEMM: independent attention heads, rotary
embeddings, token embedding gathers, and Gemma QAT activation scaling run on the same matmul
pool as linear layers whencpuThreads > 1. Causal attention splits that work head-major so
each worker covers a full query range (later tokens attend more keys; query-major chunks
left the last worker with most of the prefill). LlmModelis a sealed API type. The factory still returnsLlmModel; the transformer graph
and engine lease live on a hidden implementation (no static access registry).