Skip to content

v0.78.0-dev20260828

Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 28 Aug 05:32
· 50 commits to main since this release
Immutable release. Only release title and notes can be modified.
f9bf473

Note

If you are installing from a release, please refer to the README, INSTALLATION instructions, and any other documentation packaged with the release, not on the main branch. There may be differences between the latest main and the previous release.

The changelog will now follow, showing the changes from last release.

This release was generated by the CI workflow https://github.com/tenstorrent/tt-metal/actions/runs/33137911692

LLK (low-level kernels)

  • llk perf: qualify SfpuType globally in sfpu_binop_scalar_perf (fixes LLK perf nightly) PR 54560
  • One Parquet writer per run, named from the run tag PR 53928
  • Added STALLWAIT on MATH for SFPU path on moe_gate_test to fix race PR 54483

Metalium (tt-metal core)

  • Skip WH perf check for One to All Multicast Directed Ideal PR 54593
  • Fix Ethernet firmware teardown in multichip simulation PR 54484
  • Validate the Quasar FDS go/done wiring between dispatch engines and worker cores PR 54341
  • Expand PCH coverage: shared test PCHs and distributed.hpp in PCHFull (~11% less clean-build CPU) PR 53733
  • Extend the zeroing API to support CoreLocalMem/Scratchpad/LocalTensorAccessor PR 54122
  • Automated UMD Bump 27.08.2026 PR 54501
  • Metal 2.0: Implement variadic tensor parameters PR 54480
  • Fix Galaxy multi-allocator resize and reset PR 53809
  • Bring up tt-metal on the Quasar 8x4 grid in craq-sim PR 54374
  • Tensor prefetcher: mcast-in0 on K-row-major weights, and stop requiring a page-multiple GCB size PR 52384
  • Reducing MeshDevice::get_devices usage PR 54336
  • Use correct MetalContext in device, dispatch, program, and kernel internals PR 54618
  • add web viewer section to dm readme PR 54652
  • Update DM readme with web viewer file location PR 54663
  • Delete HAL_MEM_L1_* and DEBUG_VALID_L1_* macros PR 54500

TT-NN

  • Fix expand right-align when output rank > input rank PR 53517
  • ttnn.topk: validate the preallocated output indices width PR 54287
  • #52651: Fix TILE slice cache-hit all-zero output on divergent-partition hit PR 53017
  • embedding_bw: zero the trailing partial tile row when num_embeddings is not tile-aligned PR 54556
  • Use up to 8 tiles per reduction and remove STALLWAIT workaround PR 54284
  • Test Only: cut nightly fused test runtime via grid de-crossing and module-scoped device PR 54319
  • Lower Galaxy high_bw_all_gather 55K floor PR 54555
  • Expand PCH coverage: shared test PCHs and distributed.hpp in PCHFull (~11% less clean-build CPU) PR 53733
  • repeat_interleave: generated DeviceOperation + ProgramFactory PR 50700
  • refactor: deduplicate bfloat16 ULP bit-manipulation helpers in eltwise tests PR 54497
  • Improve topk_large_indices with kernel-only fused reduction PR 53953
  • Metal 2.0 port: normalization/layernorm PR 54363
  • MMRS: enforce the windowed L1 MM-output handoff, wire all models onto it, adopt windowed-swept block configs PR 54239
  • Add chunk recurrence preparation PR 52797
  • Pass1 Cleanup: Call compute API, drop packer SFPU_ACTIVATION, unguard fast_tilize PR 54585
  • Add program factory identity to graph tracing reports (#54158) PR 54304
  • Port Paged Cache to Device 2.0 PR 54598
  • Tensor prefetcher: mcast-in0 on K-row-major weights, and stop requiring a page-multiple GCB size PR 52384
  • SDPA windowed mode with no masking and variable window size PR 54492
  • Correct a topk ttsim skip comment: SFPCONFIG, not SFPLOADMACRO PR 54586
  • ccache: force depend mode for GCC PCH providers to stop stale .gch reuse PR 54645

tt-train

  • Unpack P/dS to dest at full FP32 in sdpa_bw kv kernel PR 53315
  • Fix AdamW state-restore beta ordering; validate optimizer state shapes; per-buffer SGD dampening PR 53222
  • pair zeros-fills with write-zeros barriers in variable_matmul PR 53440
  • Expand PCH coverage: shared test PCHs and distributed.hpp in PCHFull (~11% less clean-build CPU) PR 53733
  • Expose GELU approximation variant in ttml.ops.unary.gelu PR 53959
  • ccache: force depend mode for GCC PCH providers to stop stale .gch reuse PR 54645

Models

  • Fix for Qwen3-32B and Llama-3.3-70B sampling test failures on Galaxy PR 54292
  • Use up to 8 tiles per reduction and remove STALLWAIT workaround PR 54284
  • #54604: add qwen3_vl ops tests for quasar PR 54625
  • Prune blaze-models-prefill-tests.yaml PR 53530
  • Remove debug env hooks and investigation labels from the sampling path PR 54567
  • Migrate the Galaxy demo tests to tier 1 Models e2e and retire the pipeline PR 54306
  • Migrate the Galaxy perf tests to tier 1 e2e and device perf, and retire the pipeline PR 54324
  • Rename Blackhole e2e job display names PR 54558
  • MMRS: enforce the windowed L1 MM-output handoff, wire all models onto it, adopt windowed-swept block configs PR 54239
  • AGMM performance: block-size heuristics for out-of-the-box performance of unswept AGMM shapes PR 54477
  • #54583: add qwen3 copy for quasar work PR 54588
  • Migrate the Galaxy TTTv2 2D module tests to tier 1 Models unit CI PR 54443
  • Gemma4 optimizations [Performance] PR 50648
  • #48144: Migrate the tt_dit encoder tests to the tiered Models CI as dit-encoders PR 54465
  • Create DiT/VAE ping-pong buffers on device PR 51496

Infrastructure & CI

  • One Parquet writer per run, named from the run tag PR 53928
  • Validate the Quasar FDS go/done wiring between dispatch engines and worker cores PR 54341
  • Prune blaze-models-prefill-tests.yaml PR 53530
  • Enable tt-split-pr-by-codeowners in tenstorrent-skills-reviewer PR 54511
  • repeat_interleave: generated DeviceOperation + ProgramFactory PR 50700
  • ci: give the Copilot cloud agent an environment and instructions to compile tt-metal PR 54520
  • Migrate the Galaxy demo tests to tier 1 Models e2e and retire the pipeline PR 54306
  • Migrate the Galaxy perf tests to tier 1 e2e and device perf, and retire the pipeline PR 54324
  • Rename Blackhole e2e job display names PR 54558
  • Remove tdowdallTT and roseli-TT as CODEOWNERs for tests/pipeline_reorg/ PR 54589
  • Increase BH Sanity frequency to every 2 hours for more feedback PR 54591
  • Drop the retired Galaxy demo and perf pipelines from /test-command PR 54569
  • Raise the TT-Transformers common unit tests timeout to 23 min on wh_llmbox PR 54608
  • Migrate the Galaxy TTTv2 2D module tests to tier 1 Models unit CI PR 54443
  • Gemma4 optimizations [Performance] PR 50648
  • Add program factory identity to graph tracing reports (#54158) PR 54304
  • #48144: Migrate the tt_dit encoder tests to the tiered Models CI as dit-encoders PR 54465
  • Migrate the Galaxy Wan2.2 integration tests to tier 1 Models unit CI PR 54462
  • Restrict skills reviewers to PRs targeting main PR 54656
  • ccache: force depend mode for GCC PCH providers to stop stale .gch reuse PR 54645

Documentation

  • Extend the zeroing API to support CoreLocalMem/Scratchpad/LocalTensorAccessor PR 54122
  • Add program factory identity to graph tracing reports (#54158) PR 54304

Tooling

  • hc: surface failed self-heal reboot and its error on JIRA ticket + comment PR 54419
  • Add SQLite serializer to tt-triage PR 54394

Other

  • Clear program cache after NO_DISPATCH memory capture PR 54262
  • Expand PCH coverage: shared test PCHs and distributed.hpp in PCHFull (~11% less clean-build CPU) PR 53733
  • ci: give the Copilot cloud agent an environment and instructions to compile tt-metal PR 54520
  • Migrate the Galaxy demo tests to tier 1 Models e2e and retire the pipeline PR 54306
  • Migrate the Galaxy perf tests to tier 1 e2e and device perf, and retire the pipeline PR 54324
  • MMRS: enforce the windowed L1 MM-output handoff, wire all models onto it, adopt windowed-swept block configs PR 54239
  • Add program factory identity to graph tracing reports (#54158) PR 54304
  • ccache: force depend mode for GCC PCH providers to stop stale .gch reuse PR 54645