Skip to content

NVIDIA Neural Modules 1.18.0

Choose a tag to compare

@ericharper ericharper released this 12 May 17:49
· 4210 commits to main since this release

Highlights

Models

NeMo ASR

  • Hybrid Autoregressive Transducer (HAT) #6260
  • Apple MPS Support for ASR Inference #6289
  • InterCTC Support for Hybrid ASR Models #6215
  • RNNT N-Gram Fusion with mAES algo #6118
  • ASR + Apple M2 CPU/GPU MPS #6289

NeMo TTS

  • TTS directory structure refactor
  • User-set symbol vocabulary #6172

NeMo Megatron

  • Model parallelism from Megatron Core #6393
  • Continued training for P-tuning #6273
  • SFT for GPT-3 #6210
  • Tensor and pipeline model parallel conversion #6218
  • Megatron NMT Export to Riva

NeMo Core

Detailed Changelogs

ASR

Changelog

TTS

Changelog

NLP / NMT

Changelog
  • [Core] return_config=True now extracts just config, not full tarfile by @titu1994 :: PR: #6346
  • restore path for p-tuning by @arendu :: PR: #6273
  • taskname and early stopping for adapters by @arendu :: PR: #6366
  • Adapter tuning accepts expanded language model dir by @arendu :: PR: #6376
  • Update gpt_training.rst by @blisc :: PR: #6378
  • Megatron GPT model finetuning by @MaximumEntropy :: PR: #6210
  • [NeMo Megatron] Cleanup configs to infer the models TP PP config automatically by @titu1994 :: PR: #6368
  • Fix prompt template unescaping by @MaximumEntropy :: PR: #6399
  • Add support for Megatron GPT Untied Embd TP PP Change by @titu1994 :: PR: #6388
  • Move Parallelism usage from Apex -> Megatron Core by @aklife97 :: PR: #6393
  • Add ability to enable/disable act ckpt and seq parallelism in GPT by @markelsanz14 :: PR: #6327
  • Refactor PP conversion + add support for TP only conversion by @titu1994 :: PR: #6419
  • fix CPU overheads of GPT synthetic dataset by @xrennvidia :: PR: #6427
  • check if grad is none before calling all_reduce by @arendu :: PR: #6428
  • Fix replace_bos_with_pad not found by @aklife97 :: PR: #6443
  • Support Swiglu in TP PP Conversion by @titu1994 :: PR: #6437
  • BERT pre-training mp fork to spawn by @aklife97 :: PR: #6442
  • Meagtron encoder decoder fix for empty validation outputs by @michalivne :: PR: #6459
  • Reduce workers on NMT CI by @aklife97 :: PR: #6472
  • Switch to NVIDIA Megatron repo by @aklife97 :: PR: #6465
  • Megatron KERPLE positional embeddings by @michalivne :: PR: #6478
  • Support in external sample mapping for Megatron datasets by @michalivne :: PR: #6462
  • Fix custom by @aklife97 :: PR: #6512
  • GPT fp16 inference fix by @MaximumEntropy :: PR: #6543
  • Fix for T5 FT model by @aklife97 :: PR: #6529
  • Pass instead of scaler object to core by @aklife97 :: PR: #6545
  • Change Megatron Enc Dec model to use persistent_workers by @aklife97 :: PR: #6548
  • Turn autocast off when precision is fp32 by @aklife97 :: PR: #6554
  • Fix batch size reconf for T5 FT for multi-validation by @aklife97 :: PR: #6582
  • Make tensor split contiguous for qkv and kv in attention by @aklife97 :: PR: #6580
  • Patches from main to r1.18.0 for Virtual Parallel by @titu1994 :: PR: #6592
  • Create dummy iters to satisy iter type len checks in core + update core commit by @aklife97 :: PR: #6600
  • Restore GPT support for interleaved pipeline parallelism by @timmoon10 :: PR: #6528
  • Add megatron_core to requirements by @ericharper :: PR: #6639

Export

Changelog

Bugfixes

Changelog
  • Fix the GPT SFT datasets loss mask bug by @yidong72 :: PR: #6409
  • [BugFix] Fix multi-processing bug in data simulator by @tango4j :: PR: #6310
  • Fix cache aware hybrid bugs by @VahidooX :: PR: #6466
  • [BugFix] Force _get_batch_preds() to keep logits in decoder timestamp… by @tango4j :: PR: #6500
  • Fixing bug in unsort_tensor by @borisfom :: PR: #6320
  • Bugfix for BF16 grad reductions with distopt by @timmoon10 :: PR: #6340
  • Limit urllib3 version to patch issue with RTD by @aklife97 :: PR: #6568

General improvements

Changelog