Skip to content

1.4.0

Choose a tag to compare

@tenstorrent-github-bot tenstorrent-github-bot released this 30 Jul 11:17
· 2 commits to main since this release
ffd8ff1

Installation

Via PyPI

pip install tt-forge==1.4.0 --extra-index-url https://pypi.eng.aws.tenstorrent.com/

Via Docker

docker pull ghcr.io/tenstorrent/tt-forge-slim:1.4.0

Dependency commits

tt-xla commit: 941a6b866a9284cadd4b33e3c4b0dc973b3e8ba5
tt-mlir commit: 6b2d83bbee7bc60e8521b814368d32785df6cd73
tt-metal commit: f1f4ff75579ebd7a69c7da52d45368f273026d85

What's Changed

  • Add workflow to test installations steps using AI by @vmilosevic in #921
  • Uplift third_party/tt_forge_models to 37332bc2fa41c6d8ff7ec994ad19b84ab266429f 2026-03-28 by @vvukomanTT in #922
  • Improve claude instructions testing by @vmilosevic in #923
  • Bump release version to 1.1.0 by @vvukomanTT in #924
  • Add a number of skills for the tenstorrent ecosystem by @zoecarver in #925
  • Update claude workflow by @vmilosevic in #926
  • Uplift third_party/tt_forge_models to e7ba96f22de7ced01b151393f4d4e1bc312fac23 2026-04-04 by @vvukomanTT in #930
  • Remove benchmark infrastructure from tt-forge by @chandrasekaranpradeep in #928
  • Upgrade checkout from v4 to v5, Add venv activate by @nsumrakTT in #929
  • Uplift third_party/tt_forge_models to 752832c839d59bb6c3fde43d7b9c03ebc820fa57 2026-04-11 by @vvukomanTT in #931
  • Add tt-xla llama demo by @ddilbazTT in #932
  • Add tiny llama demo by @ddilbazTT in #933
  • Uplift third_party/tt_forge_models to 534f4821827f4e92af6a1ad58e8c836c8ccd38fd 2026-04-18 by @vvukomanTT in #935
  • Stop tt-forge-onnx nightly releases by @vvukomanTT in #936
  • Uplift third_party/tt_forge_models to d11fe6671e44d286c2e8c6dcd9bd46a3b84f9ea8 2026-04-25 by @vvukomanTT in #937
  • Uplift third_party/tt_forge_models to cae9ccbc67a318736f656bee9a9ea776eb73e69c 2026-05-02 by @vvukomanTT in #940
  • Remove depreciated files by @nsumrakTT in #941
  • Add qwen3 demo by @abrown in #934
  • Add filter test matrix file by @vvukomanTT in #942
  • Refactor into call claude by @vmilosevic in #946
  • Update requirements to uplift PyTorch by @mmanzoorTT in #944
  • Uplift third_party/tt_forge_models to f224af305a10d38acb9fbd72c0c3514b26ec4544 2026-05-09 by @vvukomanTT in #980
  • Add ai model bringup workflow by @vmilosevic in #986
  • Ai bringup update by @vmilosevic in #987
  • Update ai-bringup workflow by @vmilosevic in #988
  • Update ai workflow docker image by @vmilosevic in #989
  • Update ai workflow by @vmilosevic in #990
  • Update ai workflow by @vmilosevic in #992
  • Uplift third_party/tt_forge_models to a64a98131c35b010895198f489355d0e6306934f 2026-05-16 by @vvukomanTT in #993
  • Run cpu bringup on builder by @vmilosevic in #994
  • Add master ai bringup by @vmilosevic in #1002
  • Fix push problem in ai workflows by @vmilosevic in #1003
  • Uplift third_party/tt_forge_models to 7201811e7020d0e35e908df47a9e57926ba0aa1c 2026-05-23 by @vvukomanTT in #1005
  • Uplift third_party/tt_forge_models to b845dfd5fefc65d5b7aff0515243b8c520758ba5 2026-05-30 by @vvukomanTT in #1006
  • Uplift third_party/tt_forge_models to 156da1137c7fde3afeea2801cf7e8f25f0d3000d 2026-06-03 by @vvukomanTT in #1009
  • Uplift third_party/tt_forge_models to 6b4b47a7c419cdc2713ddfc6e3179f61012c12f4 2026-06-06 by @vvukomanTT in #1011
  • Uplift third_party/tt_forge_models to 09239ae98eb4f0b03abe5240aca9418eb3131717 2026-06-13 by @vvukomanTT in #1013
  • Move ai bringup scripts to tt-forge-ai-bringup repo by @vmilosevic in #1004
  • Uplift third_party/tt_forge_models to 32d5c2e4a8cfd55b0f2ec99b3ec8d1b217fcb742 2026-06-20 by @vvukomanTT in #1014
  • Uplift third_party/tt_forge_models to 9af516c8daaa0634c10f00b1cd745042883845ea 2026-06-27 by @vvukomanTT in #1015
  • Uplift third_party/tt_forge_models to 2e59f9b5f72a6bfacace309d4a809300f4ad4eb7 2026-07-04 by @vvukomanTT in #1016
  • [PyTorch v2.11][Uplift] Update requirements for demo examples by @mmanzoorTT in #1018
  • Uplift third_party/tt_forge_models to 7f1ab80e38802437070ca6a102ace98d5ae49267 2026-07-11 by @vvukomanTT in #1019
  • Uplift third_party/tt_forge_models to 64b77e1dd83e2d2516cc216ef90c143553b362bf 2026-07-18 by @vvukomanTT in #1020
  • Uplift third_party/tt_forge_models to a18b58d353826ee96f696ce3069993c2a1edc165 2026-07-25 by @vvukomanTT in #1022

New Contributors

Full Changelog: 1.0.0...1.4.0


LLM Performance

Model Token/sec/user Batch Token/sec ttft (ms) Hardware
pytorch_DeepSeek-V3.2_deepseek_v3_2_exp_modified_nlp_causal_lm_custom 3.0 128 384.0 8216.11 n150
pytorch_Falcon_3_10B_Base_nlp_causal_lm_huggingface 40.0 32 1280.0 835.41 p150
pytorch_Falcon_3_1B_Base_nlp_causal_lm_huggingface 54.0 32 1728.0 658.97 n150
pytorch_Falcon_3_1B_Base_nlp_causal_lm_huggingface 105.0 32 3360.0 298.38 p150
pytorch_Falcon_3_3B_Base_nlp_causal_lm_huggingface 36.0 32 1152.0 820.12 n150
pytorch_Falcon_3_3B_Base_nlp_causal_lm_huggingface 64.0 32 2048.0 376.11 p150
pytorch_Falcon_3_7B_Base_nlp_causal_lm_huggingface 18.0 32 576.0 1153.17 n150
pytorch_Falcon_3_7B_Base_nlp_causal_lm_huggingface 36.0 32 1152.0 500.88 p150
pytorch_GLM_4.7_nlp_causal_lm_huggingface 7.0 128 896.0 2731.29 n150
pytorch_GPT-OSS_20B_nlp_causal_lm_huggingface 8.0 64 512.0 6253.4 n150
pytorch_GPT-OSS_20B_nlp_causal_lm_huggingface 21.0 1 21.0 333.23 p150
pytorch_Gemma_1.1_2B_IT_nlp_causal_lm_huggingface 38.0 32 1216.0 669.59 n150
pytorch_Gemma_1.1_2B_IT_nlp_causal_lm_huggingface 77.0 32 2464.0 243.57 p150
pytorch_Kimi-K2.6_kimi_k2_6_modified_nlp_causal_lm_custom 6.0 64 384.0 3556.51 n150
pytorch_Kimi-K2_kimi_k2_instruct_modified_nlp_causal_lm_custom 6.0 64 384.0 3446.21 n150
pytorch_Llama_3.1_70B_Instruct_nlp_causal_lm_huggingface 6.0 32 192.0 11018.66 n150
pytorch_Llama_3.1_70B_Instruct_nlp_causal_lm_huggingface 12.0 32 384.0 1780.34 p150
pytorch_Llama_3.1_8B_Instruct_nlp_causal_lm_huggingface 22.0 32 704.0 1246.98 n150
pytorch_Llama_3.1_8B_Instruct_nlp_causal_lm_huggingface 38.0 32 1216.0 582.74 p150
pytorch_Llama_3.2_1B_Instruct_nlp_causal_lm_huggingface 64.0 32 2048.0 560.88 n150
pytorch_Llama_3.2_1B_Instruct_nlp_causal_lm_huggingface 123.0 32 3936.0 236.32 p150
pytorch_Llama_3.2_3B_Instruct_nlp_causal_lm_huggingface 30.0 32 960.0 589.99 n150
pytorch_Llama_3.2_3B_Instruct_nlp_causal_lm_huggingface 55.0 32 1760.0 284.05 p150
pytorch_Mistral_7B_INSTRUCT_v03_nlp_causal_lm_huggingface 20.0 32 640.0 1224.87 n150
pytorch_Mistral_7B_INSTRUCT_v03_nlp_causal_lm_huggingface 34.0 32 1088.0 557.78 p150
pytorch_Mistral_Small_24B_INSTRUCT_2501_nlp_causal_lm_huggingface 29.0 32 928.0 894.37 p150
pytorch_Phi-1.5_Phi_1_5_nlp_causal_lm_huggingface 20.0 32 640.0 628.66 n150
pytorch_Phi-1.5_Phi_1_5_nlp_causal_lm_huggingface 37.0 32 1184.0 334.16 p150
pytorch_Phi-1_Phi_1_nlp_causal_lm_huggingface 20.0 32 640.0 629.87 n150
pytorch_Phi-1_Phi_1_nlp_causal_lm_huggingface 37.0 32 1184.0 330.3 p150
pytorch_Phi-2_Phi_2_nlp_causal_lm_huggingface 8.0 32 256.0 1412.12 n150
pytorch_Phi-2_Phi_2_nlp_causal_lm_huggingface 20.0 32 640.0 694.87 p150
pytorch_Qwen 2.5 Coder_32B_Instruct_nlp_causal_lm_huggingface 17.0 32 544.0 1467.01 p150
pytorch_Qwen 2.5_0.5B_Instruct_nlp_causal_lm_huggingface 69.0 32 2208.0 405.43 n150
pytorch_Qwen 2.5_0.5B_Instruct_nlp_causal_lm_huggingface 126.0 32 4032.0 183.54 p150
pytorch_Qwen 2.5_1.5B_Instruct_nlp_causal_lm_huggingface 36.0 32 1152.0 483.81 n150
pytorch_Qwen 2.5_1.5B_Instruct_nlp_causal_lm_huggingface 63.0 32 2016.0 211.96 p150
pytorch_Qwen 2.5_3B_Instruct_nlp_causal_lm_huggingface 30.0 32 960.0 670.98 n150
pytorch_Qwen 2.5_3B_Instruct_nlp_causal_lm_huggingface 58.0 32 1856.0 293.38 p150
pytorch_Qwen 2.5_7B_Instruct_nlp_causal_lm_huggingface 16.0 32 512.0 801.72 n150
pytorch_Qwen 2.5_7B_Instruct_nlp_causal_lm_huggingface 27.0 32 864.0 354.84 p150
pytorch_Qwen 3_0_6B_nlp_causal_lm_huggingface 27.0 32 864.0 1154.41 n150
pytorch_Qwen 3_0_6B_nlp_causal_lm_huggingface 59.0 32 1888.0 528.65 p150
pytorch_Qwen 3_1_7B_nlp_causal_lm_huggingface 22.0 32 704.0 695.31 n150
pytorch_Qwen 3_1_7B_nlp_causal_lm_huggingface 46.0 32 1472.0 327.83 p150
pytorch_Qwen 3_32B_nlp_causal_lm_huggingface 14.0 32 448.0 1827.03 p150
pytorch_Qwen 3_4B_nlp_causal_lm_huggingface 13.0 32 416.0 964.19 n150
pytorch_Qwen 3_4B_nlp_causal_lm_huggingface 27.0 32 864.0 462.53 p150
pytorch_Qwen 3_8B_nlp_causal_lm_huggingface 10.0 32 320.0 1552.55 n150
pytorch_Qwen 3_8B_nlp_causal_lm_huggingface 21.0 32 672.0 718.31 p150

Non-LLM Performance

Model Batch Sample/sec Hardware
Wan2.2-I2V-A14B-DiT 1 0.0 p150
Wan2.2-I2V-A14B-UMT5-Text-Encoder 1 12.0 p150
Wan2.2-I2V-A14B-VAE-Decoder 1 0.0 p150
Wan2.2-I2V-A14B-VAE-Encoder 1 2.0 p150
flux1-dev 1 0.0 p150
flux2 1 0.0 p150
glm-image 1 0.0 p150
janus-pro-1b 1 0.0 n150
janus-pro-1b 1 0.0 p150
janus-pro-7b 1 0.0 p150
playground-v2.5 1 0.0 n150
playground-v2.5 1 0.0 p150
pytorch_BERT_emrecan/bert-base-turkish-cased-mean-nli-stsb-tr_nlp_embed_gen_huggingface 8 159.0 n150
pytorch_BGE-M3_Base_nlp_embed_gen_custom 4 9.0 n150
pytorch_BGE-M3_Base_nlp_embed_gen_custom 4 18.0 p150
pytorch_EfficientNet_Timm_B0_cv_image_cls_timm 8 352.0 n150
pytorch_EfficientNet_Timm_B0_cv_image_cls_timm 8 793.0 p150
pytorch_MNIST_Cnn_Dropout_cv_image_cls_custom 32 13915.0 n150
pytorch_MNIST_Cnn_Dropout_cv_image_cls_custom 32 27582.0 p150
pytorch_MobileNetV2_Mobilenet_v2_cv_image_cls_torch_hub 12 1241.0 n150
pytorch_MobileNetV2_Mobilenet_v2_cv_image_cls_torch_hub 12 2917.0 p150
pytorch_Qwen 3_Embedding_4B_nlp_embed_gen_huggingface 32 49.0 n150
pytorch_Qwen 3_Embedding_4B_nlp_embed_gen_huggingface 32 109.0 p150
pytorch_ResNet_ResNet50_HuggingFace_cv_image_cls_huggingface 8 1352.0 n150
pytorch_ResNet_ResNet50_HuggingFace_cv_image_cls_huggingface 8 2750.0 p150
pytorch_SegFormer_B0_Finetuned_Ade_512_512_cv_image_seg_huggingface 1 37.0 n150
pytorch_SegFormer_B0_Finetuned_Ade_512_512_cv_image_seg_huggingface 1 79.0 p150
pytorch_Swin_S_cv_image_cls_torchvision 1 10.0 n150
pytorch_Swin_S_cv_image_cls_torchvision 1 22.0 p150
pytorch_U-Net for Conditional Generation_Base_conditional_generation_huggingface 1 5.0 n150
pytorch_U-Net for Conditional Generation_Base_conditional_generation_huggingface 1 9.0 p150
pytorch_Ultra-Fast Lane Detection v2_TuSimple_ResNet34_Backbone_cv_image_seg_github 1 136.0 n150
pytorch_Ultra-Fast Lane Detection v2_TuSimple_ResNet34_Backbone_cv_image_seg_github 1 236.0 p150
pytorch_VGG19-UNet_base_cv_image_seg_custom 1 153.0 n150
pytorch_VGG19-UNet_base_cv_image_seg_custom 1 309.0 p150
pytorch_ViT_Base_cv_image_cls_huggingface 8 235.0 n150
pytorch_ViT_Base_cv_image_cls_huggingface 8 566.0 p150
pytorch_VoVNet_Ese_Vovnet19b_Dw.ra_In1k_cv_image_cls_timm 8 742.0 n150
pytorch_VoVNet_Ese_Vovnet19b_Dw.ra_In1k_cv_image_cls_timm 8 1601.0 p150
sdxl-lightning 1 0.0 n150
sdxl-lightning 1 0.0 p150
zimage 1 0.0 p150

Model coverage

Info: Full list of supported models is available in the assets section.

Model task Model architecture Model variant Model framework Inference Training n150 n300 p150 Single device Data parallel Tensor parallel Model source
cv image cls DINOv2 Small pytorch View Source
nlp causal lm Phi-1 Phi 1 jax View Source
mm image text similarity CLIP Base Patch16 pytorch View Source
cv image cls EfficientNet B0 pytorch View Source
nlp causal lm Phi-4 Phi 4 pytorch View Source
nlp causal lm Phi-1 Phi 1 pytorch View Source
cv image cls MobileNetV1 Mobilenet v1 pytorch View Source
nlp causal lm Qwen 2.5 7B Instruct pytorch View Source
nlp causal lm Phi-2 Phi 2 jax View Source
nlp causal lm Phi-1 LoRA Phi 1 pytorch View Source
nlp causal lm GPT-OSS 20B pytorch View Source
nlp embed gen Qwen 3 Embedding 8B pytorch View Source
mm image text similarity SigLIP Base Patch16 224 pytorch View Source
cv image cls VGG HF Vgg19 pytorch View Source
nlp causal lm Qwen 2.5 1.5B Instruct jax View Source
cv image cls VoVNet Ese Vovnet19b Dw.ra In1k pytorch View Source
nlp causal lm Qwen 2.5 0.5B Instruct jax View Source
nlp causal lm Falcon 3 10B Base pytorch View Source
nlp causal lm Qwen 3 4B jax View Source
nlp causal lm Qwen 2.5 Coder 1.5B Instruct jax View Source
nlp causal lm Phi-3 Mini Instruct pytorch View Source
nlp causal lm Gemma 1.1 2B IT pytorch View Source
mm visual qa Mistral base pytorch View Source
nlp causal lm Gemma 2 2B IT pytorch View Source
cv object det TransFuser None pytorch View Source
nlp causal lm Falcon 3 1B Base pytorch View Source
nlp token cls BiLSTM-CRF Default pytorch View Source
cv image cls MobileNetV2 Mobilenet v2 pytorch View Source
cv img to img Autoencoder linear pytorch View Source
nlp causal lm olmo_3 3 7b instruct pytorch View Source
nlp causal lm GPT-2 Xl jax View Source
nlp causal lm Qwen 3 1 7B jax View Source
mm visual qa Llama 3.2 11B Vision pytorch View Source
cv object det OWL-ViT Base Patch32 pytorch View Source
nlp causal lm olmo_3 3 32b think pytorch View Source
nlp causal lm olmo_3 3 7b think pytorch View Source
nlp causal lm Phi-3 Mini 128K Instruct pytorch View Source
nlp causal lm Qwen 3 1 7B pytorch View Source
cv object det DETR ResNet50 Backbone pytorch View Source
cv image seg VGG19-UNet base pytorch View Source
nlp causal lm Qwen 2.5 Coder 3B Instruct jax View Source
cv image cls MNIST Cnn Dropout jax View Source
cv image cls MNIST Mlp Custom jax View Source
nlp causal lm Llama 3.2 1B pytorch View Source
cv image seg MaskFormer Swin-B Swin Base Coco pytorch View Source
cv object det YOLOv9 T pytorch View Source
nlp causal lm Qwen 2.5 Coder 3B jax View Source
nlp causal lm Falcon 3 7B Base pytorch View Source
nlp causal lm Llama 3.1 70B pytorch View Source
nlp causal lm Qwen 2.5 0.5B Instruct pytorch View Source
nlp causal lm Phi-1.5 Phi 1 5 jax View Source
nlp causal lm Qwen 2.5 0.5B jax View Source
nlp causal lm Qwen 2.5 72B Instruct pytorch View Source
nlp causal lm olmo_3 3 1025 7b pytorch View Source
nlp causal lm Phi-1 Phi 1 pytorch View Source
cv object det PointPillars pointpillars pytorch View Source
nlp causal lm Qwen 2.5 Coder 32B Instruct pytorch View Source
nlp causal lm GPT-2 Base jax View Source
nlp causal lm Gemma 1.1 7B IT pytorch View Source
cv image cls MNIST Cnn Nodropout pytorch View Source
cv image cls MNIST Mlp Custom 1x2 jax View Source
cv object det YOLOS Small Small pytorch View Source
nlp causal lm Qwen 2.5 Coder 0.5B jax View Source
nlp causal lm Qwen 3 8B pytorch View Source
conditional generation U-Net for Conditional Generation Base pytorch View Source
nlp causal lm Qwen 3 14B pytorch View Source
nlp causal lm Gemma 2 27B IT pytorch View Source
nlp causal lm Qwen 3 0 6B pytorch View Source
nlp causal lm Qwen 2.5 1.5B Instruct pytorch View Source
nlp causal lm Gemma 2 9B IT pytorch View Source
nlp causal lm Mistral 7B INSTRUCT v03 pytorch View Source
nlp causal lm Mistral Magistral Small 2506 pytorch View Source
nlp causal lm Qwen 3 4B pytorch View Source
cv image cls Swin S pytorch View Source
cv image cls MNIST Cnn Batchnorm jax View Source
nlp causal lm Phi-2 Phi 2 pytorch View Source
nlp embed gen Qwen 3 Embedding 4B pytorch View Source
nlp causal lm Llama 3.1 8B Instruct pytorch View Source
cv object det EfficientDet D0 pytorch View Source
nlp causal lm Mistral Ministral 8B Instruct pytorch View Source