Skip to content

v1.12.0

Choose a tag to compare

@RanoRaraku RanoRaraku released this 15 Aug 10:54
· 4 commits to main since this release

Release notes

This release reduces model latencies significantly with the addition of delay-penalty training and speeds up training by up to 2.0x (base model on an 8xA100 node). It also speeds up greedy decoding by up to 40x (depending on the dataset, model and machine).

This release:

  • Reduces the solution's latency. For a breakdown of the system latencies see user-perceived latency docs. Specifically this release:
    • Adds support to train with a delay-penalty, either fixed or on a linear schedule, to encourage early emissions of tokens. The algorithm is turned on by default, see docs here
    • Adds support to calculate emission latency during training and validation. Turn on with --calculate_emission_latency
  • Speeds up training, (see training times docs) via:
    • Reducing RNN-T loss' memory consumption by performing log_softmax in-place
    • Running data preparation on GPU. Default is now --dali_train_device='gpu'
    • Support to train on CPUs with heterogenous cores, more information found here
  • Adds support for a batched greedy decoder. This should work by default (use decode='greedy' and --val_batch_size=<INT_VALUE>)
  • Adds option to floor gradient scaler to help stabilize training. Turn on with --scaler_lbl2=<FLOAT_VALUE>
  • Makes various documentation improvements