Repository navigation
Release 1.1.0 (16-aug-2026)
[1.1.0] — 2026-08-16
Public release of nano-vllm-java 1.1.0 (Maven coordinates com.igormaznitsa:nano-vllm-java:1.1.0).
ONNX weight import, Llama and Gemma 4 text chat, BERT embeddings, Qwen3 GGUF, dense/hybrid RAG, and related API cleanup.
Changed
- Prepared-prompt
TEXT_DEBUGevents are off unless you callChatSession.emitDebugPrompts(true)
orRagSession.emitDebugPrompts(true)(the previous default dumped mixed RAG/advisor prompts to
any listener /streamTosink). - CPU inference kernels are faster on the decode path: linear layers reuse the activation
vector across outputs (GEMV), Vector-API dots use independent accumulators, and residual /
MLP / RMSNorm / attention value mix run through SIMD instead of scalar Java loops. Paged KV
attention reads cache slots in place instead of copying each page into a dense tensor. - Construct sampling knobs only with
SamplingParams.builder()(orSamplingDefaults/ withers).
Conveniencenew SamplingParams(), two-arg, and three-arg constructors are removed so they cannot
skip named knobs or disagree with builder defaults (those shortcuts used top-p 0.9).
SamplingParamsis a class with a private constructor, same Builder contract asLLM/
LlmAdvisor. - Construct an engine only with
LLM.builder(model).build(). Thenew LLM(model)shortcut is
removed so closed and embedding checkpoints cannot skip builder checks. - Interactive
Exampledemo is a line-oriented terminal app: model menu, RAG mode
(none / BM25 / dense / hybrid), then advisor count (0–3). Enter selects the first item
(downloaded models first; Qwen3-0.6B preferred for chat quality). If no checkpoint is on disk, the demo exits
with download instructions. Embedding models skip to an embed REPL. Prepared-prompt debug
(debug>on stderr) is off unless you pass--debug. Readsamples.Exampletop to bottom —
each mode is a named method. - GGUF and ONNX weight loads now redraw the same in-place percent/ETA bar as Hugging Face
safetensors (current tensor on one line) instead of a tensor list or a late “assembled” summary. - Library chat/sampling/RAG defaults are architecture-marker driven, not product-tuned:
Tokenizer.ChatFormat
(ChatML / turn-based / plain), neutralSamplingDefaults, no Qwen EOS id fallback, no default
unknown-arch→Qwen3, no Gemma-only session retry or RAG isolate. Demo policies (system prompts,
turn-based top-k, unusable-answer recovery, advisor setup-boilerplate filter) live in
nano-vllm-java-samples(SampleChatPrompts,Example,HelloWorld). - Advisor demo role text, shared advisor instructions, Greek name catalog, and advisor-aware
system add-on are samples-only (SampleAdvisorPrompts). The library uses caller-supplied
LlmAdvisorname/prompt only, plus structural note-mixing helpers. - Demo advisor role strings and advisor-aware system add-on moved to samples
(SampleAdvisorPrompts); library no longer appends advisor prose to system prompts. - Library chat defaults no longer inject model-family system prose (Qwen
<think>rules, plain
assistant text).ChatPrompts.systemForis always empty; demos set policy via
SampleChatPromptsinnano-vllm-java-samples. 1.0 boolean chat shims
(ChatMessages.newConversation(boolean),scrubSetupBoilerplateTurns,
ChatPrompts.systemFor(boolean)/systemFor(Tokenizer, boolean)) are removed in this
release rather than kept as@Deprecated— usenewConversation(String),
scrubMatchingAssistantTurns, andsystemFor(Tokenizer)/LLM.Builder.systemPrompt. - Maven layout is now multi-module: parent
nano-vllm-java-pom, librarynano-vllm-java, demos
nano-vllm-java-samples. Sample mains are no longer packaged in the library JAR; run demos with
mvn -pl nano-vllm-java-samples exec:javafrom the repository root.
Added
- Fluent load, sampling, and call shortcuts:
LlmModelFactory.open(path).listen(…).unpackParameters().thinkTags(…).chatSpecials(…).make()
(existingmake/fromClasspath*remain);SamplingParams.builder()and withers;SamplingDefaults.neutral();
LLM.Builder.sampling/stopTokenIds;chatOnce/completemax-token overloads;generate(…, Duration)/
seq-awareLLM.TokenEventcallbacks;ChatSession.seed/maxTokens/send(text, params);
ChatReply.parse(raw, llm)using the model's think tags.ChatSessionstreaming also emits
TEXT_RAW(unparsed tokenizer decode, think tags and chat specials kept) alongside parsed
thinking/answer. ModelSupport/UnsupportedModelException: exact architecture detection (no substringqwen→Qwen3),
a user-facing support catalog, and fail-fast load errors for look-alike families (Qwen2, Qwen3.5/Fara,
Gemma 2, Gemma 4 vision/audio-only variants, Mistral, other VLMs, GGUF Llama/Gemma, HF BERT safetensors).-Dnanollvm.archcannot override a
different family.LLM.builder/LlmModel.embedmisuse messages include the same catalog. Gemma 4 text (including QAT mobile) is a supported chat graph.Tokenizer.ChatFormat/isTurnBasedChat()/skipSpecialTokensOnChatDecode()(product-named
chat helpers such as{@code isGemmaChat}removed; use format/architecture APIs).- RMSNorm offset scale flag renamed to
{@code onePlusWeight}(math convention, not a product name). ChatSession/RagSessionopt-inrecoverUnusableAnswers/unusableAnswer/
unusableAnswerFallback;RagSessionalso exposesmaxHistoryMessagesandemitDebugPrompts
so RAG chats do not have to drop through.chat().LLM.Builder.advisorNoteFilterso apps can drop demo setup fillers before advisor mix.- ONNX folder weight import (Tier A):
LlmModelFactory.make(folder)loadsconfig.json+ tokenizer +.onnx(root oronnx/) like safetensors — Qwen3 / Gemma3 / Llama chat and BERT embeddings; no ONNX Runtime. - Optional download scripts for Gemma 4 E2B QAT mobile (
models/download-gemma4-e2b-qat-mobile.sh/.ps1/.cmd→models/Gemma4-E2B-IT-QAT-Mobile/, ~2.3 GB). Hugging Face folders withmodel_typegemma4/gemma4_textnow load as text-only chat (packed QAT int2/4/8, per-layer embeddings, KV sharing). Vision and audio towers in the same checkpoint are skipped. Safetensors shards larger than 2 GiB are read viaFileChannel(Java mmap stays limited to 2 GiB). - Llama causal architecture (
LlamaForCausalLM) for HF safetensors and ONNX (Tiny-LLM-ONNX base demo;
SmolLM2-135M-Instruct-ONNX chat demo viamodels/download-smollm2-135m-instruct-onnx.sh).
SampleNextTokenHelloWorldencodes a seed and prints the next sampled tokens plus the continued
text (default Tiny-LLM-ONNX). - Transparent GGUF BERT embedding support (e.g. GTE-small): load via
LlmModelFactory.makeand callLlmModel.embed(...)with text or token ids;LLM.builderrejects embedding-only models;Examplemenu option runs an embedding REPL. SampleEmbeddingsHelloWorldprints the vector dim, a preview, and cosine vs the same / related / unrelated text (defaultmodels/gte-small.Q2_K.gguf). - Qwen3 GGUF chat:
LlmModelFactory.make(pathToQwen3.gguf)loads a self-containedgeneral.architecture=qwen3file (embedded tokenizer + packed weights). Load is split into container transport (GGUF / HF safetensors / ONNX) and a per-family architecture processor that binds config/schema, fills weights, and builds the graph. Qwen2, MoE, VL, and Gemma/Llama GGUF stay rejected. - Dense / hybrid RAG:
DenseRagIndex,HybridRagIndex, andRagFactory.withEmbeddings(PreparedRag, LlmModel)(BM25 + embedding cosine via RRF).Exampleoffers a RAG-mode menu (none / BM25 / dense / hybrid) after choosing a chat model, then an advisor-count menu (0–3). SampleAdvisorRagHelloWorldshows one custom advisor (Alex) plus BM25 overrag/on Gemma3-270M. - RAG classpath documents:
RagFactory.makeResource/Builder.addResource(absolute ClassLoader path orClass.getResourceAsStreamresolution); source labels useclasspath:…. - GGML dequant for the remaining GGUF weight types: K-quants (
Q2_K,Q5_K,Q8_K), legacy (Q4_1,Q5_0,Q5_1,Q8_1,Q1_0,Q2_0), IQ (IQ1_S/M,IQ2_*,IQ3_*,IQ4_XS), ternary (TQ1_0/TQ2_0), MXFP4 / NVFP4, and integer/F64 tensors.Q3_K/IQ4_NLwere already present. File recipes likeQ4_K_Mstill mix those GGML types (not a separate dtype). - Stream / classpath model load:
ModelFileId+ModelFileSource,LlmModelFactory.make(source), andfromClasspath/fromClasspathGgufhelpers (bytes stay in heap; no disk cache). Filesystemmake(Path)unchanged. - Custom chat scratchpad markers via
ThinkTagsinLlmModelFactory.make(…, Map)under
LlmModel.OPTION_THINK_TAGS(default remains<think>/</think>), and chat-markup search
strings viaChatSpecialsunderLlmModel.OPTION_CHAT_SPECIALS(defaults cover ChatML, Gemma
turn markers, Llama stops, and the default think pair). Omitted keys are filled with those
defaults on the frozen options map.ChatSession.thinkTags/RagSession.thinkTagsoverride the
scratchpad pair for one conversation. Parse, ChatML skip-seed, and history truncation use the
same think pair when both markers are in vocab.
Fixed
- Weightless RMSNorm (no affine scale, used on Gemma 4 shared-KV V) no longer NPEs on the fused
residual pathforward(x, residual). - Windows model download scripts (
.ps1/.cmd) resolve their install dir via$PSScriptRoot,
write with absolute paths (no longer change the caller’s working directory), prefercurl.exe,
stay ASCII-only for Windows PowerShell 5.1 encoding, and the Gemma script also honors
HF_HOME\token. - Lexical RAG off-topic detection no longer treats conversational fillers (
what/think/
about/ …) as topic evidence, so queries like “what do you think about BMW?” stay outside
fairy-tale corpora even when those glue words appear in the documents. - ChatML models without
<think>vocab tokens no longer setTokenizer.invitesThinking(); think
invitation is vocab-gated for HF and GGUF loads alike. Library system prompts stay empty
regardless (demo policies live in samples). - Generation stops at
maxModelLen(and clampsmaxTokensto remaining context) so short-context models such as Tiny-LLM-ONNX no longer crash RoPE pastmax_position_embeddings. - ONNX load skips non-float graph constants (e.g. INT64) and scalar initializers so transformers.js exports like SmolLM2 Instruct load cleanly; unknown / float8 / nibble weight types fail with an explicit error instead of silent ignore.
- Decode stops early on degenerate token loops (exact repeated blocks, long same-token streaks, or overused n-grams) so tiny models cannot fill the whole
maxTokensbudget with the same paragraph; Example caps compact ONNX demos (SmolLM2 / Tiny) to 256 new tokens in chat and RAG. BundledModels.findaccepts absolute filesystem paths (it no longer strips a leading/before the absolute-path check).