Skip to content

NVMAI 3.5

Choose a tag to compare

@Pummelchen Pummelchen released this 13 Aug 10:03
· 1046 commits to main since this release

NVMAI 3.5

Production-readiness release: full-codebase audit fixes, the complete test-suite green-up, and benchmark tooling.

  • Tool-call loop. Responses API streaming announces every tool call (previously only the first) with unique item IDs and matching function_call_arguments.done / output_item.done events; function_call input items require a call_id; store: true is rejected as unsupported.
  • Decode-service IPC. The Mac app's response router uses per-request signaling (fixing a lost-update race and a shared-semaphore wakeup hazard) and only sends a cancel when the consumer actually cancels.
  • Defaults aligned. Chat-completions temperature default 0.6 -> 0.2 and top_k default 20 -> 64, matching the documented/app defaults; the launchers' nothink choice now actually disables reasoning.
  • Robustness. Prompt-cache snapshot ceiling corrected to 4 GiB; the CLI-strip heuristic drops empty assistant turns and matches reminder tags case-insensitively; installer manifest/receipt writes are fsync-durable and verified against the format's 16 KB alignment contract; expert-pread bookkeeping uses a single lock; shader NaN guards in softmax and decode attention; dead code removed.
  • Test suite green. All 658 tests / 120 suites pass. Fixes include: ServerCoordinator admits up to queueLimit; the HTTP handler cancels an active generation on a client I/O error; the trusted-receipt policy is strict; the detokenizer recognizes ByteLevel byte tokens and never emits U+FFFD; PromptSubmissionPolicy defers to the editor for .return newline mode.
  • Tooling. benchmark/combos.sh kills only its own servers, records failed runs explicitly, no longer edits the user's global OpenCode config, and gains a LIMIT=N fastest-first shortcut; tools/responses_bridge.py announces output items before streaming and maps max_output_tokens: 0 to "no cap".

See the wiki changelog.