Repository navigation
[1.4.0] — 2026-09-06
Public release of nano-vllm-java 1.4.0 (Maven coordinates com.igormaznitsa:nano-vllm-java:1.4.0).
Meta fastText text classification, optional TornadoVM GEMV acceleration, and typed generate results
(LlmModality.resultType / cast — no caller cast).
Added
- Meta fastText supervised text classification (
*.bin/*.ftz), including the official
language-identification models
(lid.176.binpreferred /lid.176.ftz, 176 languages). Load a file or folder with
LlmModelFactory.make, thenLLM.generate(LlmInText, LABELS)→ rankedLlmOutLabels
(__label__xxcodes with probabilities). Pure Java (no JNI); hierarchical softmax,
softmax, and one-vs-all. Download:models/download-fasttext-lid-176.sh(or.ps1/.cmd)
fetches the denserlid.176.bin(~126 MB). Sample:LanguageIdHelloWorld. - Optional TornadoVM acceleration compiled into the main library (
tornado-api/
tornado-runtimeoptional Maven deps at6.0.0-jdk22plus, not transitive). When TornadoVM is
on the module path and reports at least one device,-Dnanollvm.kernels=auto(default) prefers
TornadoVM for large dense GEMV; elementwise kernels stay on the Vector/scalar CPU backend.
Explicit modes:tornado/gpu,vector/simd,scalar/plain. Tornado GEMV prefers
the Kernel API (KernelContext+WorkerGrid1D, one thread per output row) per TornadoVM’s
SGEMV guidance, with Loop Parallel (@Parallelover0..outCount)
as fallback; reuses compiled plans (LRU), keeps weights on-device for matching buffers/shape,
and runs as one full-range launch (no CPU tile sharding against the device execute lock).
Hybrid cuBLAS needs a CUDA Tornado SDK (not the OpenCL-only path). Samples: Maven profile
-Ptornadolaunches via thetornadoCLI with-Dnanollvm.kernels=tornado.
Changed
LlmModalitycarries the concretegenerateresult class (TEXT→LlmOutText,
AUDIO→LlmOutSoundData,EMBEDDING→LlmOutEmbedding,LABELS→LlmOutLabels).
LLM.generate/LlmModel.generatereturn that type viaLlmModality.cast, so callers no
longer need an explicit cast.