Skip to content

Release 5.18.0

Latest

Choose a tag to compare

@vasqu vasqu released this 30 Sep 16:46
· 35 commits to main since this release

New Model additions

Nemotron 3 Diarization

image

Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio. It supports both streaming and offline inference, handles up to eight speakers, and orders speaker outputs by each speaker's first arrival in the input audio.

The model uses the Arrival-Order Speaker Cache (AOSC) 1 and FIFO queue introduced for Streaming Sortformer 1, 2. A single checkpoint supports configurable latency profiles, from an 80 ms input buffer to a 30.4 s offline-style buffer, and configurable output frame resolution in multiples of 10 ms. With chunked inference, the maximum audio duration is not limited.

Links: Documentation

NemotronH Omni

NemotronH Omni is a multimodal reasoning model from NVIDIA that pairs the NemotronH hybrid
Mamba-Transformer language model with a RADIO vision encoder and an optional Parakeet-based sound encoder.
Image (and video) patches are projected through a RADIO tower and a pixel-shuffle MLP into the language model's
embedding space at the <image> / <video> context-token positions; audio clips are projected in the same way at
<audio> positions. The result is a single autoregressive model that reasons jointly over text, images, video and
sound.

Links: Documentation

HyperCLOVAX Vision V2

HyperCLOVAX Vision V2 is a multimodal vision-language model developed by NAVER. It combines the HyperClovaX language model backbone with a Qwen2.5-VL vision encoder. The model supports text, image, and video inputs and is capable of chain-of-thought reasoning via built-in thinking tokens (<think>...</think>).

Links: Documentation

GTE

GTE was proposed in mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval by Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, Meishan Zhang, Wenjie Li and Min Zhang.

GTE is a BERT-style bidirectional encoder that replaces absolute position embeddings with RoPE, uses a gated MLP, and applies layer normalization after each residual connection. The same architecture backs Alibaba's gte-*-v1.5, gte-multilingual-* and gte-en-mlm-* checkpoints as well as Snowflake's snowflake-arctic-embed-m-v2.0.

Links: Documentation

Breaking changes

Bugfixes and improvements

Significant community contributions

The following contributors have made significant changes to the library over the last release:

  • @harshaljanjani
    • model: Add GTE to Transformers (#48416)
    • [vLLM] Fix video token counting for Transformers backend video inputs (Part 2) (#48900)
    • 🚨 [vLLM] Fix video token counting for Transformers backend video inputs (Part 1) (#48894)
  • @eustlb
    • [Nemotron3Diarization] fix streaming last stft frame dropped (#49167)
    • [Nemotron3Diarization] nit: hub pr merged to main (#49065)
    • Add Nemotron3Diarization (#49056)
    • [apply_chat_template] pass sampling_rate to call (#48794)
  • @tarekziade
    • QA: restore masking_utils export comment with a noqa (#49209)
    • QA: fix noisy comment checker (#49184)
    • QA: Fix llama4 leak (#49042)
    • Shorten noisy OpenVINO SDPA comment (#49033)
    • Added more usage of MemoryCleanupMixin (#49011)
    • fix noisy comments (#49013)
    • updated GraniteMoeHybrid expectations (CPU) (#49016)
    • QA: applied ruff rule PLW1514 (#48990)
    • Contract the mamba2 chunk scan with einsum instead of broadcast-then-sum (#48978)
    • QA: deactivate rule 41 (#48852)
    • QA: Added a MemoryCleanupMixin class for tests (#48681)
    • Synthetic test assets (#48589)
  • @ydshieh
    • [CacheHardIntegrationTest] use dedicated safetensors repo to fix Xet bucket cache corruption (#49182)
    • [imagegpt/vilt/trocr] fix cache fixture tests: use hf_hub_download instead of load_dataset (#49178)
    • [PerceptionLM] Restore test_inputs_embeds overrides to fix flaky test (#49165)
    • [CI] ssh-runner: add optional cache_type input to switch between bucket and EFS runners (#49159)
    • [Gemma] Fix test_model_7b_fp16_static_cache expected value for cuda 8 after #49084 (#49133)
    • [gemma3n] fix audio test fixture: use hf_hub_download instead of load_dataset (#49128)
    • Fix exporters import on torch < 2.9 (is_contiguous_or_false) (#49124)
    • [NemotronH-Omni] Fix device mismatch in test tensor creation (#49102)
    • [Zamba] Fix associative scan breaking ONNX export and OOM in integration test (#49092)
    • Fix Qwen3OmniMoeIntegrationTest OOM (#48987)
    • Reduce peak memory in examples_torch CI job (OOM fix) (#48983)
    • Fix AXK2 integration test: update CUDA expected text and rename class (#48941)
    • Clean up some MoE models' integration tests (#48833)
    • Mark test_generate_with_static_cache as flaky for olmo and bigbird_pegasus (#48856)
    • Fix GLM rope_parameters dict shared mutation across test instances (#48895)
    • Fix Glm4MoeIntegrationTest: split into 3 tests, switch to GLM-4.5-Air (#48820)
    • Fix flaky RfDetr test_save_load (#48869)
    • docs: fix broken #combining-with-fsdp2 anchor in expert_parallelism.md (#48854)
    • Relax static-cache tolerance in test_generate_with_static_cache (1e-5 → 5e-5) (#48815)
    • Add integration tests for MuseGlimmerAssistantModel (#48796)
    • Fix Glm4vMoeIntegrationTest: offload_folder + MemoryCleanupMixin (#48776)
    • Switch daily CI to torch 2.14 — update expected outputs (#48750)
    • [DeepseekV3] OOM cascade root-cause investigation (generator ref leak in conversion_mapping) (#48720)
    • [CI] Replace hardcoded username allowlists with author_association check in workflow triggers (#48712)
    • [MusicgenMelody] Fix conditioning silently dropped at generation step 0 (#48679)
    • [fix] Fix GlmOcr integration tests: wrong token IDs and image token decode bug (#48650)
    • [tests] Fix NougatModelIntegrationTest: pin artifact revision and update golden values (#48638)
    • [CI] Deduplicate Nvidia/AMD CI reply comments and add headers (#48655)
  • @molbap
    • Fix lost call (#49163)
    • 🚨 🚨Bring some dinos to modern standards (#46266)
    • [Fix] yolos offload issue (#48688)
  • @remi-or
    • [CB] Fix failing tests discovered when using the B200 (#49171)
    • [CB] [Major] Upgrade the cache to support different attention types (#47809)
    • [Improvement] Rework the docstring and comments of mHC (#48888)
    • [CB] Add pause mechanism (#48462)
    • [Fix] Clean-up ternaries in the DeepSeek family (#48447)
  • @jiqing-feng
    • Keep attention_mask as None in OPT's causal mask creation (#49002)
    • Register the remaining mamba-ssm kernel layers on XPU (#49035)
    • Only seed numpy in BigBird block-sparse attention during training (#49037)
    • Fix mask creation not being skipped under torch.compile (#48975)
    • Fix StaticCache for Mllama and enable torch.compile (#48141)
    • Fix reset on the dynamic cache layers (#48809)
    • Cast pixel values to the patch embedding dtype in DeepSeek-OCR-2 (#48632)
    • Fix TextToAudioPipeline crash for tokenizer-only models (#48505)
    • Fix slow integration tests on XPU (#48611)
    • Register activation kernel layers on XPU (#47858)
  • @Wauplin
    • Bump huggingface_hub upper bound to <3.0 (transformers) (#49083)
    • [docs] Fix legacy hf CLI references (transformers) (#48988)
    • Resolve the Hub revision once per load instead of passing a private _commit_hash around (#47611)
    • Use huggingface_hub httpx export (#48685)
  • @meatybobby
    • Add support for Nemotron Omni (#46509)
  • @Sainava
    • Add pose estimation keypoint preprocessing to Sapiens2ImageProcessor (#47199)
  • @jp1924
    • add HyperClovaX Vision (#44314)