Skip to content

Release 1.2.0 (22-aug-2026)

Choose a tag to compare

@raydac raydac released this 22 Aug 11:38
· 34 commits to main since this release

[1.2.0] — 2026-08-22

Public release of nano-vllm-java 1.2.0 (Maven coordinates com.igormaznitsa:nano-vllm-java:1.2.0).
SentencePiece and extra tokenizer families, RAG load tuners, engine-owned matmul pool, deterministic sampling, checkpoint modalities, and close() resource reclaim.

Added

  • LlmModel.modalities() (and inputModalities() / outputModalities()): input and output
    content types as LlmModality (TEXT, IMAGE, AUDIO, VIDEO, EMBEDDING). Values follow
    the checkpoint config (Gemma 4 QAT mobile declares text+image+audio+video in). Embedding
    encoders are text→embedding. LlmModel.usableModalities() is what this library actually runs
    (text→text chat, or text→embedding); vision/audio towers are still skipped at load.
    LlmModel.toString() includes modalities= and usable= when they differ. The Example
    demo prints checkpoint modalities after load, plus the runtime line when they are not the same.

  • RAG user turns with retrieved passages now start with RagPrompts.GROUNDING so small chat models
    are told to answer from those lines only and not invent a book or play.

  • HybridRagIndex.of(RagIndex…) fuses any two or more indexes with RRF (nested hybrids flatten).
    BM25+dense via RagFactory.withEmbeddings is unchanged. To mix disk folders, classpath files,
    and inline strings, add them on one RagFactory.builder() — HybridRagIndex does not concatenate
    corpora. Builder.addFolders walks several directories in one call.

  • LLM.Builder.deterministic() (also SamplingParams / ChatSession / RagSession): same prompt
    always picks the highest-logit token (topK = 1, nucleus off). temperature(0) stays rejected.

  • LLM.Builder.dedicatedMatmulPool(): a bounded matmul thread pool owned by that engine and shut
    down on close(), so a server need not join the process-wide nanollvm-matmul-* pool or hand
    the library a foreign ExecutorService. Combining it with matmulExecutor fails at build().

  • Optional download scripts for intfloat/multilingual-e5-small
    (models/download-multilingual-e5-small.sh / .ps1 / .cmd → models/multilingual-e5-small/, ONNX fp32 ~470 MB).
    Hugging Face Unigram SentencePiece tokenizers load from tokenizer.json; embedding wrap accepts XLM-R
    <s> / </s> as well as BERT [CLS] / [SEP]. Optimum-style BERT ONNX (anonymous onnx::MatMul_*
    weights) remaps those MatMul aliases through the BERT schema, same as named HF tensors.
    EmbeddingsHelloWorld defaults to this folder and prefixes non-retrieval text with query: .
    Unigram now applies the Hugging Face precompiled charsmap (needed for many non-ASCII XLM-R / E5
    strings). Folders without tokenizer.json load SentencePiece tokenizer.model. WordPiece,
    WordLevel, and character models are recognized from tokenizer.json model.type.
    Tokenizer.fromSentencePiece builds a tokenizer from protobuf bytes.

  • Load-time RAG tuners on RagFactory.Builder.addProcessor: skip files, supply custom document
    text, or rewrite extracted strings before chunking. Several tuners run in order.
    Sample RagTunerHelloWorld extracts a bundled Project Gutenberg EPUB of Karel Čapek's R.U.R.
    (JDK zip + StAX, no epub4j / xpp3), keeps the short OPF title (before a subtitle slash) as a
    Markdown heading so later chunks still name the play, indexes with BM25, and asks questions from that play with
    LLM.Builder.deterministic() so repeats pick the same tokens. DenseRagIndex.of / HybridRagIndex.of accept
    a per-passage embed callback and an optional caller Executor for parallel indexing (sequential
    when omitted).

Removed

  • Built-in PdfTextExtractor. Folder walks no longer pick up .pdf by default. Index PDFs (or
    other binaries) with RagTuner.extracting, same pattern as the EPUB sample. ResourceLimits
    drops PDF inflate / ToUnicode cmap caps (maxPdfInflateBytes, maxCmapRangeSpan,
    maxCmapEntries).

Fixed

  • Closing a model, engine, or weight reader now drops file buffers, KV pages, rotary tables, and
    the last shared matmul pool instead of pinning them until process exit. GGUF and safetensors files
    ≤ 2 GiB are copied into heap so close() can reclaim them; shards larger than that still use a
    positioned file channel. Late unpack releases packed bytes when no other engine is using the model.

Changed

  • Dense RAG query embedding stays concurrent (each LlmModel.embed uses a fresh step context).
    Index-time passage embedding is sequential unless the caller supplies an Executor.
  • Extra CPU cores now work on long prompts, not only GEMM: independent attention heads, rotary
    embeddings, token embedding gathers, and Gemma QAT activation scaling run on the same matmul
    pool as linear layers when cpuThreads > 1. Causal attention splits that work head-major so
    each worker covers a full query range (later tokens attend more keys; query-major chunks
    left the last worker with most of the prefill).
  • LlmModel is a sealed API type. The factory still returns LlmModel; the transformer graph
    and engine lease live on a hidden implementation (no static access registry).