Skip to content

v0.77.0-dev20260812

Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 12 Aug 03:21
· 701 commits to main since this release
Immutable release. Only release title and notes can be modified.
ea042c4

Note

If you are installing from a release, please refer to the README, INSTALLATION instructions, and any other documentation packaged with the release, not on the main branch. There may be differences between the latest main and the previous release.

The changelog will now follow, showing the changes from last release.

This release was generated by the CI workflow https://github.com/tenstorrent/tt-metal/actions/runs/31550319955

LLK (low-level kernels)

  • LLK Test Infra transpose dest tests with Dest bank switching PR 51673
  • SFPU edge cases initial pass ( expand ranges on ops) PR 52172
  • fix: improve tanh bf16 performance PR 52732
  • Add static asserts to dest dvalid sync functions on llk api level PR 52843
  • binary op sfpu (add) in parallel with matmul PR 47469
  • #49944: adding RNE rounding to SFPU ADD, SUB, RSUB llk API PR 51060

Metalium (tt-metal core)

  • Fix tensor prefetcher stop() deadlocking under the Tracy device profiler PR 52529
  • Make runtime_noc_debugging run and pass in CI PR 52672
  • Metal 2.0: record dataflow buffers in graph-capture CB accounting PR 52757
  • Fixing Producer Injection Counter Overflow PR 52396
  • emule: resolve a fabric connection's direction per worker, not per chip PR 52888
  • Implement vararg CTA for metal 2.0 PR 52391
  • Automated UMD Bump 07.08.2026 PR 52447
  • Consolidate mailbox tests; cover DM->TRISC for all compute threads PR 52791
  • Weight-load + JIT-compile observability in logs PR 52258
  • Graduate runtime tensor out of experimental PR 51640
  • #49944: adding RNE rounding to SFPU ADD, SUB, RSUB llk API PR 51060

TT-NN

  • Fix Blackhole hang on RM reshape to [N,1] PR 50967
  • Avoid use of hard-coded bfloat16 intermediate CBs in Layernorm PR 52081
  • Bug #48267 Pair printed shards with their own coordinates in to_string PR 52708
  • fix device_id remap skips buffer_pages_by_address snapshot format and buffer_pages PR 51678
  • Metal 2.0: record dataflow buffers in graph-capture CB accounting PR 52757
  • Fix AllGather BH CI machine requirement PR 52862
  • Fused CCL + matmul: fabric-bound minimal-matmul, strided AGMM and MMRS PR 52513
  • Expose forwarding link indices to Python PR 52642
  • Add KDA reference semantics and test utilities PR 52781
  • Sparse MLA: move gathers to high-bandwidth all-gather PR 52606
  • #47644: Extend universal input/output support for ttnn::fold PR 50385
  • reduce_scatter "direct" algorithm for small shapes PR 51741
  • Graduate runtime tensor out of experimental PR 51640
  • #49944: adding RNE rounding to SFPU ADD, SUB, RSUB llk API PR 51060
  • Neighbor-pad halo exchange and halo-mode conv3d PR 52514

tt-train

  • Improve training stability and introduce new training (scheduler/optimizer) knobs PR 48716
  • Generate and upload loss plot summaries to CI PR 51051

Models

  • Fix MoE reduce score-channel shape inference PR 52435
  • SDXL VAE pcc relaxed PR 50930
  • Add Wan2.2 Models to QB2 PR 52772
  • Fused CCL + matmul: fabric-bound minimal-matmul, strided AGMM and MMRS PR 52513
  • Add KDA reference semantics and test utilities PR 52781
  • Sparse MLA: move gathers to high-bandwidth all-gather PR 52606
  • Kimi-K3 MLA: NoPE, output gate, 96 heads (MLA layer only, random weights) PR 52068
  • prefill: remove standalone serving mode (default 11-chunk KV, fix rank claims) PR 52213
  • Weight-load + JIT-compile observability in logs PR 52258
  • Clean up/document TT-DiT test organization PR 52504
  • #47644: Extend universal input/output support for ttnn::fold PR 50385

Infrastructure & CI

  • Bug #48267 Pair printed shards with their own coordinates in to_string PR 52708
  • Make runtime_noc_debugging run and pass in CI PR 52672
  • Retune MiniMax-M3 prefill perf gate to 4500 +/- 7% PR 52722
  • Remove passing tests from ttsim skip list PR 52871
  • Add Wan2.2 Models to QB2 PR 52772
  • Add KDA reference semantics and test utilities PR 52781
  • Kimi-K3 MLA: NoPE, output gate, 96 heads (MLA layer only, random weights) PR 52068
  • Grant actions: write so the stale-PR state cache can refresh PR 52765
  • Use new exabox runners and remove old ones PR 52803
  • Remove ownership requirement for .github/deprecations.json PR 52848
  • Clean up/document TT-DiT test organization PR 52504
  • Stop ttnn-core owning ttnn/cpp/ttnn/kernel/ PR 52846
  • Graduate runtime tensor out of experimental PR 51640
  • Generate and upload loss plot summaries to CI PR 51051

Other

  • Fused CCL + matmul: fabric-bound minimal-matmul, strided AGMM and MMRS PR 52513
  • Kimi-K3 MLA: NoPE, output gate, 96 heads (MLA layer only, random weights) PR 52068
  • #47644: Extend universal input/output support for ttnn::fold PR 50385
  • reduce_scatter "direct" algorithm for small shapes PR 51741
  • Graduate runtime tensor out of experimental PR 51640
  • Neighbor-pad halo exchange and halo-mode conv3d PR 52514