Repository navigation
Releases: raydac/nano-vllm-java
Release list
Release 1.4.0 (06-sep-2026)
[1.4.0] — 2026-09-06
Public release of nano-vllm-java 1.4.0 (Maven coordinates com.igormaznitsa:nano-vllm-java:1.4.0).
Meta fastText text classification, optional TornadoVM GEMV acceleration, and typed generate results
(LlmModality.resultType / cast — no caller cast).
Added
- Meta fastText supervised text classification (
*.bin/*.ftz), including the official
language-identification models
(lid.176.binpreferred /lid.176.ftz, 176 languages). Load a file or folder with
LlmModelFactory.make, thenLLM.generate(LlmInText, LABELS)→ rankedLlmOutLabels
(__label__xxcodes with probabilities). Pure Java (no JNI); hierarchical softmax,
softmax, and one-vs-all. Download:models/download-fasttext-lid-176.sh(or.ps1/.cmd)
fetches the denserlid.176.bin(~126 MB). Sample:LanguageIdHelloWorld. - Optional TornadoVM acceleration compiled into the main library (
tornado-api/
tornado-runtimeoptional Maven deps at6.0.0-jdk22plus, not transitive). When TornadoVM is
on the module path and reports at least one device,-Dnanollvm.kernels=auto(default) prefers
TornadoVM for large dense GEMV; elementwise kernels stay on the Vector/scalar CPU backend.
Explicit modes:tornado/gpu,vector/simd,scalar/plain. Tornado GEMV prefers
the Kernel API (KernelContext+WorkerGrid1D, one thread per output row) per TornadoVM’s
SGEMV guidance, with Loop Parallel (@Parallelover0..outCount)
as fallback; reuses compiled plans (LRU), keeps weights on-device for matching buffers/shape,
and runs as one full-range launch (no CPU tile sharding against the device execute lock).
Hybrid cuBLAS needs a CUDA Tornado SDK (not the OpenCL-only path). Samples: Maven profile
-Ptornadolaunches via thetornadoCLI with-Dnanollvm.kernels=tornado.
Changed
LlmModalitycarries the concretegenerateresult class (TEXT→LlmOutText,
AUDIO→LlmOutSoundData,EMBEDDING→LlmOutEmbedding,LABELS→LlmOutLabels).
LLM.generate/LlmModel.generatereturn that type viaLlmModality.cast, so callers no
longer need an explicit cast.
Release 1.3.0 (29-aug-2026)
[1.3.0] — 2026-08-29
Public release of nano-vllm-java 1.3.0 (Maven coordinates com.igormaznitsa:nano-vllm-java:1.3.0).
Whisper speech-to-text, Piper text-to-speech, typed generate(LlmInput, LlmModality) for every graph
kind, XLM-RoBERTa embeddings, and optionalData load extras.
Added
-
Load-time extras via typed keys:
LlmModelFactory.open(path).optionalData(key, value).make().
Unknown keys are kept and ignored by graphs that do not read them. Empty extras stay off
LlmModel.options()so existing think-tag / chat-special maps are unchanged. Piper voices
useLlmOptionalData.ESPEAK_DATAfor the espeak-ng-data directory (default
{model}/espeak-ng-data). A missing or incomplete folder is ignored; Piper still
synthesizes. When the folder includesdictsource(*_list/*_rules) for the
voice language, Piper reads those files for G2P instead of the built-in letter tables.
Suffix and prefixS/Prules retranslate the stem; digits use the list's number
fragments (_1,_2X,_0C, …). Russian*_ruleshonor the A/B/C/F/G/H/Y letter
groups (soиafter a hard consonant isы, not after every consonant) and then
palatalize beforeе/и, reduce unstressedо/е, and place stress from
compiled{lang}_dict/$1–$7list flags (otherwise espeak's Russian
syllable-count guess, so two-syllable words likeэто/мамаare not
end-stressed). That stops Irina stressing every unknown word on the final
syllable, which sounded Slavic but not Russian. Ifdictsourceis missing but compiled
phontaband{lang}_dictare present, listed words still come from those files. -
Piper text-to-speech from a voice folder (
*.onnx+*.onnx.json). Load with
LlmModelFactory.make, thenLLM.builder(model).build()and
generate(LlmInText, AUDIO)→LlmOutSoundData(WAV PCM16 LE mono + sample rate) —
writesound.wav()only if you need a file.LlmModel.generateremains a sequential shortcut.
Official Piper ONNX exports (including ru_RU-irina-medium) load: the phoneme table exported as
sid, WaveNet kernels namedonnx::Conv_*, and HiFi-GAN vocoder residual blocks. Conv and
ConvTranspose geometry (stride, padding, dilation, groups, output padding) comes from the ONNX
graph instead of vocoder-family guesses, so voices are not locked to Lessac/Irina kernel sizes.
Irina uses ResBlock2 (convs.0/convs.1) with those graph dilations, so output is speech rather
than a metallic buzz. Reverse synthesis follows VITS: all residual couplings with channel flips,
WaveNet skip paths, and stochastic duration (rational-quadratic conv flows).
espeak-ng-data is optional (compiledphontab/{lang}_dictare used when
dictsourceis absent). Download scripts
installlang/plusdictsource/(*_list/*_rules) and, when a system
espeak-ng-data is present, compiledphontab/phondata/{lang}_dictso Russian
lexical stress is available. The dictionary is the G2P when
present. Otherwise a letter-to-sound fallback is used: Russian uses espeak phones for ш/ж/ы (ʃ/ʒ/y), so those
consonants stay audible (academic IPAʂ/ʐ/ɨare unused by this voice). Piper voices map the voiced velar to
IPAɡ(not ASCIIg); G2P emitsɡso Russianгis spoken instead of dropped. English Lessac G2P maps
espeak@2/@5to schwa (no leaked digit2), and an en-us post-pass turns British-leaning phones into Piper
rhotic IPA (ɹ/ɚ/ɑː) with content-word stress so pangrams stay close to official Piper.
Russian letter G2P also destresses word-final obstruents, assimilates voicing across clusters
(including prepositions), and applies common orthoepic rewrites (жи/ши, -ого/-его, что, -ться)
while still emitting Piper espeak phones. Latin letter G2P uses a small English lexicon plus
letter/digraph rules so a US English Piper voice still works without dictionaries. Optional
models/download-piper-en-lessac-medium.sh/.ps1/.cmd(Lessac medium + espeak-ng-data)
andmodels/download-piper-ru-irina-medium.sh/.ps1/.cmd(Irina medium + espeak-ng-data).
SampleSynthesizeHelloWorldprefers Lessac when both folders exist;VoiceReplyHelloWorld
speaks the chat reply with that same voice and plays it through Java Sound
(javax.sound.sampled) when a mixer is present. With no input WAV it records 15
seconds from the default microphone (TargetDataLine), trims leading and trailing
silence, and falls back to a canned
Piper line if that input is missing;Examplelists both
voices and the TTS session suggestsHello world/Привет, мир.
LLM.builderis the runtime for every graph kind: Piper uses the samecpuThreads/ matmul
pool as chat, and 1-D conv / conv-transpose split independent output channels across that pool.
Non-chat engines skip KV paging (numKvcacheBlocksis 0), so Piper no longer fails to build
on a large heap (the chat auto-sizer treated a missing transformer as 1-byte pages and overflowed).
Synthesis skips identity attention masks, norms channels in place (no layout ping-pong), and uses
SIMD add/dot paths on the VITS hot path. -
Whisper speech-to-text from Hugging Face safetensors (
openai/whisper-*). Load with
LlmModelFactory.make, thenLLM.builder(model).build()and
generate(LlmInSound, TEXT)→LlmOutText(WAV bytes viaLlmInSound.ofWav, or PCM via
ofPcm; read files yourself). Optional language is aLocaleonLlmInSound
(null/Locale.ROOT= auto; region ignored; ISO aliases such asjvmap to Whisper's
jwtoken).LlmModel.generateremains a sequential shortcut. The builder
uses the same CPU matmul pool as chat (Linear, attention, and stem convs). Decoder steps keep a
growing self-attention KV cache (one step per new token). Log-mel STFT uses Bluestein FFT with
sparse mel filters and optional frame parallelism on that pool. CTranslate2 /
faster-whispermodel.binfolders, Whisper GGUF, and Whisper ONNX are refused. Optional
models/download-whisper-base.sh/.ps1/.cmd(~290 MB) anddownload-whisper-tiny.sh
(~150 MB). SampleTranscribeHelloWorld;VoiceReplyHelloWorldchains Whisper → few-shot mood
of the transcript → chat → Piper;Examplelists Whisper in the menu and opens a WAV
transcribe session. -
Cross-kind typed facade:
LlmOutput generate(LlmInput, LlmModality)onLlmModelandLLM
is the single inference entry for embeddings, Whisper, Piper, and raw text completion.
Sealed inputs (LlmInText,LlmInSound,LlmInTokenIds) and outputs (LlmOutText,
LlmOutSoundDatawith WAV + sample rate,LlmOutEmbedding). Chat stays on
ChatSession/chatOnce, and batched token generation stays ongenerate(List, …).
Text completion via the facade needs anLLMengine.
LLM.Builder.random(Random)injects the engine-owned RNG used for chat
Gumbel sampling and Piper noise (default remains unseeded, or seed1Lfor synthesis). -
Optional download scripts for FacebookAI/xlm-roberta-base
(models/download-xlm-roberta-base.sh/.ps1/.cmd→models/xlm-roberta-base/, ONNX fp32
saved asonnx/model.onnx, ~1.9 GB). BERT-encoder embeddings (bert/roberta/xlm-roberta,
and the same graph under those family names) load from GGUF or ONNX viaLLM.builderthen
generate(LlmInText, EMBEDDING)(orLlmModel.generateas a sequential shortcut) — the
library keys off architecture, not a named Hub checkpoint. ONNX BERT weight names drop whatever
module sits in front ofembeddings./encoder.(not a fixed prefix list).
ModelSupport.isEmbeddingCheckpointclassifies a folder or GGUF fromconfig.json/ metadata
without loading weights. TheExampledemo lists the optional xlm-roberta-base download in its
menu, picks any BERT encoder undermodels/for dense/hybrid RAG (smallest by weight file), and
can few-shot classify on encoder vectors (centered prototypes: teachlabel | text, then predict).
RagFactory.withEmbeddingscan report per-passage embed progress (and run those embeds on a
callerExecutor). DistilBERT / ALBERT / DeBERTa / ELECTRA stay unsupported; safetensors
BERT-family folders are still rejected.
Removed
- Public
embed/transcribe/synthesize/completeshortcuts. Call typed
generate(LlmInput, LlmModality)instead (LlmInText→EMBEDDING/AUDIO/TEXT,
LlmInSound→TEXT). BERT embeddings now go throughLLM.builderlike every other kind.
Fixed
-
Chat answers no longer keep a typed ChatML lookalike
<|im_ended|>at the end of a turn
(ChatSpecials.DEFAULT). Generation still stops only on real EOS ids such as<|im_end|>. -
Model download scripts treat HTTP 416 on resume as "already complete", so re-running a script
after a sidecar such asconfig.jsonfinished no longer fails (curl: (22) ... 416).
Release 1.2.0 (22-aug-2026)
[1.2.0] — 2026-08-22
Public release of nano-vllm-java 1.2.0 (Maven coordinates com.igormaznitsa:nano-vllm-java:1.2.0).
SentencePiece and extra tokenizer families, RAG load tuners, engine-owned matmul pool, deterministic sampling, checkpoint modalities, and close() resource reclaim.
Added
-
LlmModel.modalities()(andinputModalities()/outputModalities()): input and output
content types asLlmModality(TEXT,IMAGE,AUDIO,VIDEO,EMBEDDING). Values follow
the checkpoint config (Gemma 4 QAT mobile declares text+image+audio+video in). Embedding
encoders are text→embedding.LlmModel.usableModalities()is what this library actually runs
(text→text chat, or text→embedding); vision/audio towers are still skipped at load.
LlmModel.toString()includesmodalities=andusable=when they differ. TheExample
demo prints checkpoint modalities after load, plus the runtime line when they are not the same. -
RAG user turns with retrieved passages now start with
RagPrompts.GROUNDINGso small chat models
are told to answer from those lines only and not invent a book or play. -
HybridRagIndex.of(RagIndex…)fuses any two or more indexes with RRF (nested hybrids flatten).
BM25+dense viaRagFactory.withEmbeddingsis unchanged. To mix disk folders, classpath files,
and inline strings, add them on oneRagFactory.builder()—HybridRagIndexdoes not concatenate
corpora.Builder.addFolderswalks several directories in one call. -
LLM.Builder.deterministic()(alsoSamplingParams/ChatSession/RagSession): same prompt
always picks the highest-logit token (topK = 1, nucleus off).temperature(0)stays rejected. -
LLM.Builder.dedicatedMatmulPool(): a bounded matmul thread pool owned by that engine and shut
down onclose(), so a server need not join the process-widenanollvm-matmul-*pool or hand
the library a foreignExecutorService. Combining it withmatmulExecutorfails atbuild(). -
Optional download scripts for intfloat/multilingual-e5-small
(models/download-multilingual-e5-small.sh/.ps1/.cmd→models/multilingual-e5-small/, ONNX fp32 ~470 MB).
Hugging Face Unigram SentencePiece tokenizers load fromtokenizer.json; embedding wrap accepts XLM-R
<s>/</s>as well as BERT[CLS]/[SEP]. Optimum-style BERT ONNX (anonymousonnx::MatMul_*
weights) remaps those MatMul aliases through the BERT schema, same as named HF tensors.
EmbeddingsHelloWorlddefaults to this folder and prefixes non-retrieval text withquery:.
Unigram now applies the Hugging Face precompiled charsmap (needed for many non-ASCII XLM-R / E5
strings). Folders withouttokenizer.jsonload SentencePiecetokenizer.model. WordPiece,
WordLevel, and character models are recognized fromtokenizer.jsonmodel.type.
Tokenizer.fromSentencePiecebuilds a tokenizer from protobuf bytes. -
Load-time RAG tuners on
RagFactory.Builder.addProcessor: skip files, supply custom document
text, or rewrite extracted strings before chunking. Several tuners run in order.
SampleRagTunerHelloWorldextracts a bundled Project Gutenberg EPUB of Karel Čapek's R.U.R.
(JDK zip + StAX, no epub4j / xpp3), keeps the short OPF title (before a subtitle slash) as a
Markdown heading so later chunks still name the play, indexes with BM25, and asks questions from that play with
LLM.Builder.deterministic()so repeats pick the same tokens.DenseRagIndex.of/HybridRagIndex.ofaccept
a per-passage embed callback and an optional callerExecutorfor parallel indexing (sequential
when omitted).
Removed
- Built-in
PdfTextExtractor. Folder walks no longer pick up.pdfby default. Index PDFs (or
other binaries) withRagTuner.extracting, same pattern as the EPUB sample.ResourceLimits
drops PDF inflate / ToUnicode cmap caps (maxPdfInflateBytes,maxCmapRangeSpan,
maxCmapEntries).
Fixed
- Closing a model, engine, or weight reader now drops file buffers, KV pages, rotary tables, and
the last shared matmul pool instead of pinning them until process exit. GGUF and safetensors files
≤ 2 GiB are copied into heap soclose()can reclaim them; shards larger than that still use a
positioned file channel. Late unpack releases packed bytes when no other engine is using the model.
Changed
- Dense RAG query embedding stays concurrent (each
LlmModel.embeduses a fresh step context).
Index-time passage embedding is sequential unless the caller supplies anExecutor. - Extra CPU cores now work on long prompts, not only GEMM: independent attention heads, rotary
embeddings, token embedding gathers, and Gemma QAT activation scaling run on the same matmul
pool as linear layers whencpuThreads > 1. Causal attention splits that work head-major so
each worker covers a full query range (later tokens attend more keys; query-major chunks
left the last worker with most of the prefill). LlmModelis a sealed API type. The factory still returnsLlmModel; the transformer graph
and engine lease live on a hidden implementation (no static access registry).
Release 1.1.0 (16-aug-2026)
[1.1.0] — 2026-08-16
Public release of nano-vllm-java 1.1.0 (Maven coordinates com.igormaznitsa:nano-vllm-java:1.1.0).
ONNX weight import, Llama and Gemma 4 text chat, BERT embeddings, Qwen3 GGUF, dense/hybrid RAG, and related API cleanup.
Changed
- Prepared-prompt
TEXT_DEBUGevents are off unless you callChatSession.emitDebugPrompts(true)
orRagSession.emitDebugPrompts(true)(the previous default dumped mixed RAG/advisor prompts to
any listener /streamTosink). - CPU inference kernels are faster on the decode path: linear layers reuse the activation
vector across outputs (GEMV), Vector-API dots use independent accumulators, and residual /
MLP / RMSNorm / attention value mix run through SIMD instead of scalar Java loops. Paged KV
attention reads cache slots in place instead of copying each page into a dense tensor. - Construct sampling knobs only with
SamplingParams.builder()(orSamplingDefaults/ withers).
Conveniencenew SamplingParams(), two-arg, and three-arg constructors are removed so they cannot
skip named knobs or disagree with builder defaults (those shortcuts used top-p 0.9).
SamplingParamsis a class with a private constructor, same Builder contract asLLM/
LlmAdvisor. - Construct an engine only with
LLM.builder(model).build(). Thenew LLM(model)shortcut is
removed so closed and embedding checkpoints cannot skip builder checks. - Interactive
Exampledemo is a line-oriented terminal app: model menu, RAG mode
(none / BM25 / dense / hybrid), then advisor count (0–3). Enter selects the first item
(downloaded models first; Qwen3-0.6B preferred for chat quality). If no checkpoint is on disk, the demo exits
with download instructions. Embedding models skip to an embed REPL. Prepared-prompt debug
(debug>on stderr) is off unless you pass--debug. Readsamples.Exampletop to bottom —
each mode is a named method. - GGUF and ONNX weight loads now redraw the same in-place percent/ETA bar as Hugging Face
safetensors (current tensor on one line) instead of a tensor list or a late “assembled” summary. - Library chat/sampling/RAG defaults are architecture-marker driven, not product-tuned:
Tokenizer.ChatFormat
(ChatML / turn-based / plain), neutralSamplingDefaults, no Qwen EOS id fallback, no default
unknown-arch→Qwen3, no Gemma-only session retry or RAG isolate. Demo policies (system prompts,
turn-based top-k, unusable-answer recovery, advisor setup-boilerplate filter) live in
nano-vllm-java-samples(SampleChatPrompts,Example,HelloWorld). - Advisor demo role text, shared advisor instructions, Greek name catalog, and advisor-aware
system add-on are samples-only (SampleAdvisorPrompts). The library uses caller-supplied
LlmAdvisorname/prompt only, plus structural note-mixing helpers. - Demo advisor role strings and advisor-aware system add-on moved to samples
(SampleAdvisorPrompts); library no longer appends advisor prose to system prompts. - Library chat defaults no longer inject model-family system prose (Qwen
<think>rules, plain
assistant text).ChatPrompts.systemForis always empty; demos set policy via
SampleChatPromptsinnano-vllm-java-samples. 1.0 boolean chat shims
(ChatMessages.newConversation(boolean),scrubSetupBoilerplateTurns,
ChatPrompts.systemFor(boolean)/systemFor(Tokenizer, boolean)) are removed in this
release rather than kept as@Deprecated— usenewConversation(String),
scrubMatchingAssistantTurns, andsystemFor(Tokenizer)/LLM.Builder.systemPrompt. - Maven layout is now multi-module: parent
nano-vllm-java-pom, librarynano-vllm-java, demos
nano-vllm-java-samples. Sample mains are no longer packaged in the library JAR; run demos with
mvn -pl nano-vllm-java-samples exec:javafrom the repository root.
Added
- Fluent load, sampling, and call shortcuts:
LlmModelFactory.open(path).listen(…).unpackParameters().thinkTags(…).chatSpecials(…).make()
(existingmake/fromClasspath*remain);SamplingParams.builder()and withers;SamplingDefaults.neutral();
LLM.Builder.sampling/stopTokenIds;chatOnce/completemax-token overloads;generate(…, Duration)/
seq-awareLLM.TokenEventcallbacks;ChatSession.seed/maxTokens/send(text, params);
ChatReply.parse(raw, llm)using the model's think tags.ChatSessionstreaming also emits
TEXT_RAW(unparsed tokenizer decode, think tags and chat specials kept) alongside parsed
thinking/answer. ModelSupport/UnsupportedModelException: exact architecture detection (no substringqwen→Qwen3),
a user-facing support catalog, and fail-fast load errors for look-alike families (Qwen2, Qwen3.5/Fara,
Gemma 2, Gemma 4 vision/audio-only variants, Mistral, other VLMs, GGUF Llama/Gemma, HF BERT safetensors).-Dnanollvm.archcannot override a
different family.LLM.builder/LlmModel.embedmisuse messages include the same catalog. Gemma 4 text (including QAT mobile) is a supported chat graph.Tokenizer.ChatFormat/isTurnBasedChat()/skipSpecialTokensOnChatDecode()(product-named
chat helpers such as{@code isGemmaChat}removed; use format/architecture APIs).- RMSNorm offset scale flag renamed to
{@code onePlusWeight}(math convention, not a product name). ChatSession/RagSessionopt-inrecoverUnusableAnswers/unusableAnswer/
unusableAnswerFallback;RagSessionalso exposesmaxHistoryMessagesandemitDebugPrompts
so RAG chats do not have to drop through.chat().LLM.Builder.advisorNoteFilterso apps can drop demo setup fillers before advisor mix.- ONNX folder weight import (Tier A):
LlmModelFactory.make(folder)loadsconfig.json+ tokenizer +.onnx(root oronnx/) like safetensors — Qwen3 / Gemma3 / Llama chat and BERT embeddings; no ONNX Runtime. - Optional download scripts for Gemma 4 E2B QAT mobile (
models/download-gemma4-e2b-qat-mobile.sh/.ps1/.cmd→models/Gemma4-E2B-IT-QAT-Mobile/, ~2.3 GB). Hugging Face folders withmodel_typegemma4/gemma4_textnow load as text-only chat (packed QAT int2/4/8, per-layer embeddings, KV sharing). Vision and audio towers in the same checkpoint are skipped. Safetensors shards larger than 2 GiB are read viaFileChannel(Java mmap stays limited to 2 GiB). - Llama causal architecture (
LlamaForCausalLM) for HF safetensors and ONNX (Tiny-LLM-ONNX base demo;
SmolLM2-135M-Instruct-ONNX chat demo viamodels/download-smollm2-135m-instruct-onnx.sh).
SampleNextTokenHelloWorldencodes a seed and prints the next sampled tokens plus the continued
text (default Tiny-LLM-ONNX). - Transparent GGUF BERT embedding support (e.g. GTE-small): load via
LlmModelFactory.makeand callLlmModel.embed(...)with text or token ids;LLM.builderrejects embedding-only models;Examplemenu option runs an embedding REPL. SampleEmbeddingsHelloWorldprints the vector dim, a preview, and cosine vs the same / related / unrelated text (defaultmodels/gte-small.Q2_K.gguf). - Qwen3 GGUF chat:
LlmModelFactory.make(pathToQwen3.gguf)loads a self-containedgeneral.architecture=qwen3file (embedded tokenizer + packed weights). Load is split into container transport (GGUF / HF safetensors / ONNX) and a per-family architecture processor that binds config/schema, fills weights, and builds the graph. Qwen2, MoE, VL, and Gemma/Llama GGUF stay rejected. - Dense / hybrid RAG:
DenseRagIndex,HybridRagIndex, andRagFactory.withEmbeddings(PreparedRag, LlmModel)(BM25 + embedding cosine via RRF).Exampleoffers a RAG-mode menu (none / BM25 / dense / hybrid) after choosing a chat model, then an advisor-count menu (0–3). SampleAdvisorRagHelloWorldshows one custom advisor (Alex) plus BM25 overrag/on Gemma3-270M. - RAG classpath documents:
RagFactory.makeResource/Builder.addResource(absolute ClassLoader path orClass.getResourceAsStreamresolution); source labels useclasspath:…. - GGML dequant for the remaining GGUF weight types: K-quants (
Q2_K,Q5_K,Q8_K), legacy (Q4_1,Q5_0,Q5_1,Q8_1,Q1_0,Q2_0), IQ (IQ1_S/M,IQ2_*,IQ3_*,IQ4_XS), ternary (TQ1_0/TQ2_0), MXFP4 / NVFP4, and integer/F64 tensors.Q3_K/IQ4_NLwere already present. File recipes likeQ4_K_Mstill mix those GGML types (not a separate dtype). - Stream / classpath model load:
ModelFileId+ModelFileSource,LlmModelFactory.make(source), andfromClasspath/fromClasspathGgufhelpers (bytes stay in heap; no disk cache). Filesystemmake(Path)unchanged. - Custom chat scratchpad markers via
ThinkTagsinLlmModelFactory.make(…, Map)under
LlmModel.OPTION_THINK_TAGS(default remains<think>/</think>), and chat-markup search
strings viaChatSpecialsunderLlmModel.OPTION_CHAT_SPECIALS(defaults cover ChatML, Gemma
turn markers, Llama stops, and the default think pair). Omitted keys are filled with those
defaults on the frozen options map.ChatSession.thinkTags/RagSession.thinkTagsoverride the
scratchpad pair for one conversation. Parse, ChatML skip-seed, and history truncation use the
same think pair when both markers are in vocab.
Fixed
- Weightless RMSNorm (no affine scale, used on Gemma 4 shared-KV V) no longer NPEs on the fused
residual pathforward(x, residual). - Windows model download scripts (
.ps1/.cmd) resolve their install dir via$PSScriptRoot,
write with absolute paths (no longer change the caller’s working directory), prefercurl.exe,
stay ASCII-only for Windows PowerShell 5.1 encoding, and the Gemma script also honors
HF_HOME\token. - Lexical RAG off-topic detection no longer treats conversational fillers (
what/think/
about/ …) as topic evidence, so queries like “what do you think about BMW?” stay outside
fai...
Release 1.0.0 (09-aug-2026)
[1.0.0] — 2026-08-09
First public release of the nano-vllm-java CPU inference library
(JPMS module com.igormaznitsa.nanollvm, Maven coordinates com.igormaznitsa:nano-vllm-java:1.0.0).
Added
- Pure Java 21+ offline LLM inference: continuous batching, paged KV cache, Hugging Face safetensors and GGUF weights (Qwen3, Gemma3, LFM2).
- Shared models via
LlmModelFactory.makeand per-engineLLM.Builder/LLM(chat, one-shot, completion, cancel, timeout, generation stats). LlmModelisAutoCloseable: close engines first, then the shared model, to release weight resources; closed models and engines reject further use (isClosed()on both).- Chat sessions with history limits, listeners, optional advisors and mixers, plus lexical BM25 RAG (
RagFactory/RagSession). - Process-wide
ResourceLimitsfor file, PDF, corpus, JSON, GGUF, safetensors, and history budgets (overridable per process or per corpus). - Configuration knobs including
kvHeapFractionfor heap-based KV auto-sizing, CPU matmul thread control, andnanollvm.*/NANOLLVM_*system properties and environment variables. - Documented RAG/advisor trust boundary and optional suppression of prepared-prompt debug events (
emitDebugPrompts(false)). - Samples: minimal Gemma
HelloWorld, log-triage demo, interactiveExample, andBench.
Changed
- Explicit
.cpuThreads(N)/.disableMultiCpu()wins over-Dnanollvm.cpu.threads; sequential mode creates no matmul executor.