v0.77.0-dev20260805
Pre-release
Pre-release
·
182 commits
to main
since this release
Immutable
release. Only release title and notes can be modified.
Note
If you are installing from a release, please refer to the README, INSTALLATION instructions, and any other documentation packaged with the release, not on the main branch. There may be differences between the latest main and the previous release.
The changelog will now follow, showing the changes from last release.
This release was generated by the CI workflow https://github.com/tenstorrent/tt-metal/actions/runs/30963951219
LLK (low-level kernels)
- Add back the sfpu_math value so the perf report headers match again PR 52048
- fix(perf): regenerate catalog for formats.sfpu_math — unbreak main gate PR 52058
- Exempt reduce_block_max_row negative controls from the bit-exact check PR 51677
- Add golden perf-CSV header catalog and drift gate PR 51484
- add unpack tilize to fuser PR 51978
- Enable Quasar selection in LLK ttsim regression script PR 51649
- cleanup(llk): remove dead llk_pack_untilize_hw_configure PR 51838
- unpack_tilize tiny tiles compute api bringup PR 51405
- Add blaze-authored Blackhole LLKs to experimental PR 51361
- add pack untilize kernel to fuser PR 52041
Metalium (tt-metal core)
- jit_build: fix kernel-cache temp-path collision across forked processes PR 51570
- Fix Tensix hang in unary broadcast compute API PR 51772
- Quasar - Fixing watcher to always analyze unmapped TCs PR 52087
- Updating experimental::quasar DM CreateKernel to skip DM0 & DM1 PR 52095
- Per-fiber my_logical_x_/y_ globals + build_core_map stale-SWEmuleChip cache invalidation PR 52071
- Tensor prefetcher: take a MeshCommandQueue, not a cq_id PR 51635
- Automated UMD Bump 03.08.2026 PR 51947
- Automated UMD Bump 04.08.2026 PR 52055
- Collapse fabric stability intensity forks into one YAML (gtest-style name filters) PR 50270
- Collapse BH-glx torus/superpod YAMLs into test_fabric_2d_torus_stability.yaml PR 50282
- Preallocate vector capacity and move vectors into structured return values PR 51988
- unpack_tilize tiny tiles compute api bringup PR 51405
- Performance: deepen the fast dispatch prefetcher's relay_paged scratch to a 3-buffer ring PR 51272
- Add blaze-authored Blackhole LLKs to experimental PR 51361
- Collapse per-core-type command queue dispatch layouts into a single layout PR 52105
- Break DispatchCoreManager/ServiceCoreManager MetalContext::instance() cycle PR 52023
TT-NN
- Avoid optimized sharded tilize for row-major tiny-tile inputs PR 51517
- Fix moe_routing_remap expert_parallel_size vs mesh-axis validation PR 51686
- Tensor prefetcher: take a MeshCommandQueue, not a cq_id PR 51635
- #47500: Traceable chunked prefill — trace layer + on-device MoE padding config PR 51624
- Fix nonzero row-major block-sharded L1 reads PR 47148
- Port
experimental/transformer/rotary_embedding_llamafactories PR 51490 - Fix full ND sharded TILE test cases to use tile-aligned shard shapes PR 49576
- Preallocate vector capacity and move vectors into structured return values PR 51988
- Revert "[Bug fix] Make sampling and top-k deterministic on sub-core g… PR 52088
- Add support for sharded joint on the sequence dimension to ring_joint_scaled_dot_product_sdpa PR 48677
- DS Prefill :: Improve fabric links usage for dispatch operation PR 51019
tt-train
- Preallocate vector capacity and move vectors into structured return values PR 51988
Models
- Fix moe_routing_remap expert_parallel_size vs mesh-axis validation PR 51686
- Add LTX 2.3 Blackhole T2V and I2V tests to CI PR 51767
- #47500: Traceable chunked prefill — trace layer + on-device MoE padding config PR 51624
- Fix merged KV chunk table stage layout PR 51996
- SP-only merged KVPE+indexer KV chunk address table PR 51917
- multirank external runner pcc test PR 51269
- MiniMax-M3 prefill zone profiler PR 51776
- remove _trace_prefill_supported_seq_lens PR 52032
- tt_transformers: remove unreachable per_core_M arm in prefill QKV config PR 52044
- Warm the SP ring cache-read path in prefill compile() PR 52022
- Migrate vLLM nightly to "vLLM Model Tests" (tiered-format, single cadence) PR 50910
- Revert "[Bug fix] Make sampling and top-k deterministic on sub-core g… PR 52088
- Add support for sharded joint on the sequence dimension to ring_joint_scaled_dot_product_sdpa PR 48677
- DS Prefill :: Improve fabric links usage for dispatch operation PR 51019
- Replace transformers fx-availability shim with direct torch.fx import PR 52067
TT-STL
- Preallocate vector capacity and move vectors into structured return values PR 51988
Infrastructure & CI
- Batch lead-model sweep jobs by device key: 71 -> 0 test failures on Galaxy PR 51817
- Add LTX 2.3 Blackhole T2V and I2V tests to CI PR 51767
- repo-assist: gate confident bug root-cause claims behind verification PR 51791
- #47500: Traceable chunked prefill — trace layer + on-device MoE padding config PR 51624
- SP-only merged KVPE+indexer KV chunk address table PR 51917
- Collapse fabric stability intensity forks into one YAML (gtest-style name filters) PR 50270
- Collapse BH-glx torus/superpod YAMLs into test_fabric_2d_torus_stability.yaml PR 50282
- Enable sc16 deepseek blitz tests PR 51980
- ttop: call the await loop via $GITHUB_ACTION_PATH, not a nested uses PR 52028
- Migrate vLLM nightly to "vLLM Model Tests" (tiered-format, single cadence) PR 50910
Tooling
- Collapse BH-glx torus/superpod YAMLs into test_fabric_2d_torus_stability.yaml PR 50282