Repository navigation
Release 1.3.0 (29-aug-2026)
[1.3.0] — 2026-08-29
Public release of nano-vllm-java 1.3.0 (Maven coordinates com.igormaznitsa:nano-vllm-java:1.3.0).
Whisper speech-to-text, Piper text-to-speech, typed generate(LlmInput, LlmModality) for every graph
kind, XLM-RoBERTa embeddings, and optionalData load extras.
Added
-
Load-time extras via typed keys:
LlmModelFactory.open(path).optionalData(key, value).make().
Unknown keys are kept and ignored by graphs that do not read them. Empty extras stay off
LlmModel.options()so existing think-tag / chat-special maps are unchanged. Piper voices
useLlmOptionalData.ESPEAK_DATAfor the espeak-ng-data directory (default
{model}/espeak-ng-data). A missing or incomplete folder is ignored; Piper still
synthesizes. When the folder includesdictsource(*_list/*_rules) for the
voice language, Piper reads those files for G2P instead of the built-in letter tables.
Suffix and prefixS/Prules retranslate the stem; digits use the list's number
fragments (_1,_2X,_0C, …). Russian*_ruleshonor the A/B/C/F/G/H/Y letter
groups (soиafter a hard consonant isы, not after every consonant) and then
palatalize beforeе/и, reduce unstressedо/е, and place stress from
compiled{lang}_dict/$1–$7list flags (otherwise espeak's Russian
syllable-count guess, so two-syllable words likeэто/мамаare not
end-stressed). That stops Irina stressing every unknown word on the final
syllable, which sounded Slavic but not Russian. Ifdictsourceis missing but compiled
phontaband{lang}_dictare present, listed words still come from those files. -
Piper text-to-speech from a voice folder (
*.onnx+*.onnx.json). Load with
LlmModelFactory.make, thenLLM.builder(model).build()and
generate(LlmInText, AUDIO)→LlmOutSoundData(WAV PCM16 LE mono + sample rate) —
writesound.wav()only if you need a file.LlmModel.generateremains a sequential shortcut.
Official Piper ONNX exports (including ru_RU-irina-medium) load: the phoneme table exported as
sid, WaveNet kernels namedonnx::Conv_*, and HiFi-GAN vocoder residual blocks. Conv and
ConvTranspose geometry (stride, padding, dilation, groups, output padding) comes from the ONNX
graph instead of vocoder-family guesses, so voices are not locked to Lessac/Irina kernel sizes.
Irina uses ResBlock2 (convs.0/convs.1) with those graph dilations, so output is speech rather
than a metallic buzz. Reverse synthesis follows VITS: all residual couplings with channel flips,
WaveNet skip paths, and stochastic duration (rational-quadratic conv flows).
espeak-ng-data is optional (compiledphontab/{lang}_dictare used when
dictsourceis absent). Download scripts
installlang/plusdictsource/(*_list/*_rules) and, when a system
espeak-ng-data is present, compiledphontab/phondata/{lang}_dictso Russian
lexical stress is available. The dictionary is the G2P when
present. Otherwise a letter-to-sound fallback is used: Russian uses espeak phones for ш/ж/ы (ʃ/ʒ/y), so those
consonants stay audible (academic IPAʂ/ʐ/ɨare unused by this voice). Piper voices map the voiced velar to
IPAɡ(not ASCIIg); G2P emitsɡso Russianгis spoken instead of dropped. English Lessac G2P maps
espeak@2/@5to schwa (no leaked digit2), and an en-us post-pass turns British-leaning phones into Piper
rhotic IPA (ɹ/ɚ/ɑː) with content-word stress so pangrams stay close to official Piper.
Russian letter G2P also destresses word-final obstruents, assimilates voicing across clusters
(including prepositions), and applies common orthoepic rewrites (жи/ши, -ого/-его, что, -ться)
while still emitting Piper espeak phones. Latin letter G2P uses a small English lexicon plus
letter/digraph rules so a US English Piper voice still works without dictionaries. Optional
models/download-piper-en-lessac-medium.sh/.ps1/.cmd(Lessac medium + espeak-ng-data)
andmodels/download-piper-ru-irina-medium.sh/.ps1/.cmd(Irina medium + espeak-ng-data).
SampleSynthesizeHelloWorldprefers Lessac when both folders exist;VoiceReplyHelloWorld
speaks the chat reply with that same voice and plays it through Java Sound
(javax.sound.sampled) when a mixer is present. With no input WAV it records 15
seconds from the default microphone (TargetDataLine), trims leading and trailing
silence, and falls back to a canned
Piper line if that input is missing;Examplelists both
voices and the TTS session suggestsHello world/Привет, мир.
LLM.builderis the runtime for every graph kind: Piper uses the samecpuThreads/ matmul
pool as chat, and 1-D conv / conv-transpose split independent output channels across that pool.
Non-chat engines skip KV paging (numKvcacheBlocksis 0), so Piper no longer fails to build
on a large heap (the chat auto-sizer treated a missing transformer as 1-byte pages and overflowed).
Synthesis skips identity attention masks, norms channels in place (no layout ping-pong), and uses
SIMD add/dot paths on the VITS hot path. -
Whisper speech-to-text from Hugging Face safetensors (
openai/whisper-*). Load with
LlmModelFactory.make, thenLLM.builder(model).build()and
generate(LlmInSound, TEXT)→LlmOutText(WAV bytes viaLlmInSound.ofWav, or PCM via
ofPcm; read files yourself). Optional language is aLocaleonLlmInSound
(null/Locale.ROOT= auto; region ignored; ISO aliases such asjvmap to Whisper's
jwtoken).LlmModel.generateremains a sequential shortcut. The builder
uses the same CPU matmul pool as chat (Linear, attention, and stem convs). Decoder steps keep a
growing self-attention KV cache (one step per new token). Log-mel STFT uses Bluestein FFT with
sparse mel filters and optional frame parallelism on that pool. CTranslate2 /
faster-whispermodel.binfolders, Whisper GGUF, and Whisper ONNX are refused. Optional
models/download-whisper-base.sh/.ps1/.cmd(~290 MB) anddownload-whisper-tiny.sh
(~150 MB). SampleTranscribeHelloWorld;VoiceReplyHelloWorldchains Whisper → few-shot mood
of the transcript → chat → Piper;Examplelists Whisper in the menu and opens a WAV
transcribe session. -
Cross-kind typed facade:
LlmOutput generate(LlmInput, LlmModality)onLlmModelandLLM
is the single inference entry for embeddings, Whisper, Piper, and raw text completion.
Sealed inputs (LlmInText,LlmInSound,LlmInTokenIds) and outputs (LlmOutText,
LlmOutSoundDatawith WAV + sample rate,LlmOutEmbedding). Chat stays on
ChatSession/chatOnce, and batched token generation stays ongenerate(List, …).
Text completion via the facade needs anLLMengine.
LLM.Builder.random(Random)injects the engine-owned RNG used for chat
Gumbel sampling and Piper noise (default remains unseeded, or seed1Lfor synthesis). -
Optional download scripts for FacebookAI/xlm-roberta-base
(models/download-xlm-roberta-base.sh/.ps1/.cmd→models/xlm-roberta-base/, ONNX fp32
saved asonnx/model.onnx, ~1.9 GB). BERT-encoder embeddings (bert/roberta/xlm-roberta,
and the same graph under those family names) load from GGUF or ONNX viaLLM.builderthen
generate(LlmInText, EMBEDDING)(orLlmModel.generateas a sequential shortcut) — the
library keys off architecture, not a named Hub checkpoint. ONNX BERT weight names drop whatever
module sits in front ofembeddings./encoder.(not a fixed prefix list).
ModelSupport.isEmbeddingCheckpointclassifies a folder or GGUF fromconfig.json/ metadata
without loading weights. TheExampledemo lists the optional xlm-roberta-base download in its
menu, picks any BERT encoder undermodels/for dense/hybrid RAG (smallest by weight file), and
can few-shot classify on encoder vectors (centered prototypes: teachlabel | text, then predict).
RagFactory.withEmbeddingscan report per-passage embed progress (and run those embeds on a
callerExecutor). DistilBERT / ALBERT / DeBERTa / ELECTRA stay unsupported; safetensors
BERT-family folders are still rejected.
Removed
- Public
embed/transcribe/synthesize/completeshortcuts. Call typed
generate(LlmInput, LlmModality)instead (LlmInText→EMBEDDING/AUDIO/TEXT,
LlmInSound→TEXT). BERT embeddings now go throughLLM.builderlike every other kind.
Fixed
-
Chat answers no longer keep a typed ChatML lookalike
<|im_ended|>at the end of a turn
(ChatSpecials.DEFAULT). Generation still stops only on real EOS ids such as<|im_end|>. -
Model download scripts treat HTTP 416 on resume as "already complete", so re-running a script
after a sidecar such asconfig.jsonfinished no longer fails (curl: (22) ... 416).