Skip to content

Release 1.0.0 (09-aug-2026)

Choose a tag to compare

@raydac raydac released this 09 Aug 08:39
· 96 commits to main since this release

[1.0.0] — 2026-08-09

First public release of the nano-vllm-java CPU inference library
(JPMS module com.igormaznitsa.nanollvm, Maven coordinates com.igormaznitsa:nano-vllm-java:1.0.0).

Added

  • Pure Java 21+ offline LLM inference: continuous batching, paged KV cache, Hugging Face safetensors and GGUF weights (Qwen3, Gemma3, LFM2).
  • Shared models via LlmModelFactory.make and per-engine LLM.Builder / LLM (chat, one-shot, completion, cancel, timeout, generation stats).
  • LlmModel is AutoCloseable: close engines first, then the shared model, to release weight resources; closed models and engines reject further use (isClosed() on both).
  • Chat sessions with history limits, listeners, optional advisors and mixers, plus lexical BM25 RAG (RagFactory / RagSession).
  • Process-wide ResourceLimits for file, PDF, corpus, JSON, GGUF, safetensors, and history budgets (overridable per process or per corpus).
  • Configuration knobs including kvHeapFraction for heap-based KV auto-sizing, CPU matmul thread control, and nanollvm.* / NANOLLVM_* system properties and environment variables.
  • Documented RAG/advisor trust boundary and optional suppression of prepared-prompt debug events (emitDebugPrompts(false)).
  • Samples: minimal Gemma HelloWorld, log-triage demo, interactive Example, and Bench.

Changed

  • Explicit .cpuThreads(N) / .disableMultiCpu() wins over -Dnanollvm.cpu.threads; sequential mode creates no matmul executor.