Skip to content

1.5.0.dev20260807012312

Choose a tag to compare

@tenstorrent-github-bot tenstorrent-github-bot released this 07 Aug 02:03
· 1 commit to main since this release
923fe85

Installation

Via PyPI

pip install tt-forge==1.5.0.dev20260807012312 --extra-index-url https://pypi.eng.aws.tenstorrent.com/

Via Docker

docker pull ghcr.io/tenstorrent/tt-forge-slim:1.5.0.dev20260807012312

Dependency commits

tt-xla commit: 65fa938d0d761da579998e50eb2bb485af3c7449
tt-mlir commit: 724d2ff7b48da2fa73977e85476f6da1c835434b
tt-metal commit: f1f4ff75579ebd7a69c7da52d45368f273026d85

What's Changed

  • Uplift third_party/tt_forge_models to 8734a136b439fc3261145a411acbc038ca108e04 2026-08-01 by @vvukomanTT in #1023

Full Changelog: 1.5.0.dev20260801001900...1.5.0.dev20260807012312


LLM Performance

Model Token/sec/user Batch Token/sec ttft (ms) Hardware
pytorch_DeepSeek-V3.1_deepseek_v3_1_modified_nlp_causal_lm_custom 2.0 128 256.0 5398.34 n150
pytorch_DeepSeek-V3.2_deepseek_v3_2_exp_modified_nlp_causal_lm_custom 3.0 128 384.0 8208.69 n150
pytorch_Falcon_3_10B_Base_nlp_causal_lm_huggingface 40.0 32 1280.0 863.44 p150
pytorch_Falcon_3_1B_Base_nlp_causal_lm_huggingface 54.0 32 1728.0 631.18 n150
pytorch_Falcon_3_1B_Base_nlp_causal_lm_huggingface 101.0 32 3232.0 293.55 p150
pytorch_Falcon_3_3B_Base_nlp_causal_lm_huggingface 36.0 32 1152.0 802.59 n150
pytorch_Falcon_3_3B_Base_nlp_causal_lm_huggingface 66.0 32 2112.0 373.16 p150
pytorch_Falcon_3_7B_Base_nlp_causal_lm_huggingface 18.0 32 576.0 1166.21 n150
pytorch_Falcon_3_7B_Base_nlp_causal_lm_huggingface 35.0 32 1120.0 491.43 p150
pytorch_GLM_4.7_nlp_causal_lm_huggingface 6.0 128 768.0 2784.66 n150
pytorch_GPT-OSS_20B_nlp_causal_lm_huggingface 21.0 1 21.0 343.05 p150
pytorch_Gemma_1.1_2B_IT_nlp_causal_lm_huggingface 38.0 32 1216.0 621.01 n150
pytorch_Gemma_1.1_2B_IT_nlp_causal_lm_huggingface 78.0 32 2496.0 238.13 p150
pytorch_Kimi-K2.6_kimi_k2_6_modified_nlp_causal_lm_custom 6.0 64 384.0 3552.42 n150
pytorch_Kimi-K2_kimi_k2_instruct_modified_nlp_causal_lm_custom 6.0 64 384.0 3409.06 n150
pytorch_Llama_3.1_70B_Instruct_nlp_causal_lm_huggingface 4.0 32 128.0 12838.83 n150
pytorch_Llama_3.1_70B_Instruct_nlp_causal_lm_huggingface 12.0 32 384.0 1806.1 p150
pytorch_Llama_3.1_8B_Instruct_nlp_causal_lm_huggingface 21.0 32 672.0 1205.14 n150
pytorch_Llama_3.1_8B_Instruct_nlp_causal_lm_huggingface 50.0 32 1600.0 643.93 p150
pytorch_Llama_3.2_1B_Instruct_nlp_causal_lm_huggingface 64.0 32 2048.0 555.42 n150
pytorch_Llama_3.2_1B_Instruct_nlp_causal_lm_huggingface 120.0 32 3840.0 231.19 p150
pytorch_Llama_3.2_3B_Instruct_nlp_causal_lm_huggingface 30.0 32 960.0 579.22 n150
pytorch_Llama_3.2_3B_Instruct_nlp_causal_lm_huggingface 54.0 32 1728.0 280.05 p150
pytorch_Mistral_7B_INSTRUCT_v03_nlp_causal_lm_huggingface 20.0 32 640.0 1216.73 n150
pytorch_Mistral_7B_INSTRUCT_v03_nlp_causal_lm_huggingface 35.0 32 1120.0 555.48 p150
pytorch_Mistral_Small_24B_INSTRUCT_2501_nlp_causal_lm_huggingface 29.0 32 928.0 885.96 p150
pytorch_Phi-1.5_Phi_1_5_nlp_causal_lm_huggingface 20.0 32 640.0 660.73 n150
pytorch_Phi-1.5_Phi_1_5_nlp_causal_lm_huggingface 36.0 32 1152.0 326.49 p150
pytorch_Phi-1_Phi_1_nlp_causal_lm_huggingface 20.0 32 640.0 641.19 n150
pytorch_Phi-1_Phi_1_nlp_causal_lm_huggingface 37.0 32 1184.0 322.75 p150
pytorch_Phi-2_Phi_2_nlp_causal_lm_huggingface 8.0 32 256.0 1402.2 n150
pytorch_Phi-2_Phi_2_nlp_causal_lm_huggingface 20.0 32 640.0 699.18 p150
pytorch_Qwen 2.5 Coder_32B_Instruct_nlp_causal_lm_huggingface 17.0 32 544.0 1493.27 p150
pytorch_Qwen 2.5_0.5B_Instruct_nlp_causal_lm_huggingface 69.0 32 2208.0 408.89 n150
pytorch_Qwen 2.5_0.5B_Instruct_nlp_causal_lm_huggingface 127.0 32 4064.0 182.03 p150
pytorch_Qwen 2.5_1.5B_Instruct_nlp_causal_lm_huggingface 35.0 32 1120.0 473.71 n150
pytorch_Qwen 2.5_1.5B_Instruct_nlp_causal_lm_huggingface 63.0 32 2016.0 211.37 p150
pytorch_Qwen 2.5_3B_Instruct_nlp_causal_lm_huggingface 30.0 32 960.0 695.97 n150
pytorch_Qwen 2.5_3B_Instruct_nlp_causal_lm_huggingface 56.0 32 1792.0 295.71 p150
pytorch_Qwen 2.5_7B_Instruct_nlp_causal_lm_huggingface 16.0 32 512.0 809.05 n150
pytorch_Qwen 2.5_7B_Instruct_nlp_causal_lm_huggingface 27.0 32 864.0 353.96 p150
pytorch_Qwen 3_0_6B_nlp_causal_lm_huggingface 50.0 32 1600.0 1107.48 n150
pytorch_Qwen 3_0_6B_nlp_causal_lm_huggingface 98.0 32 3136.0 519.51 p150
pytorch_Qwen 3_1_7B_nlp_causal_lm_huggingface 35.0 32 1120.0 693.7 n150
pytorch_Qwen 3_1_7B_nlp_causal_lm_huggingface 67.0 32 2144.0 324.73 p150
pytorch_Qwen 3_32B_nlp_causal_lm_huggingface 17.0 32 544.0 1818.0 p150
pytorch_Qwen 3_4B_nlp_causal_lm_huggingface 23.0 32 736.0 924.64 n150
pytorch_Qwen 3_4B_nlp_causal_lm_huggingface 40.0 32 1280.0 460.98 p150
pytorch_Qwen 3_8B_nlp_causal_lm_huggingface 16.0 32 512.0 1542.21 n150
pytorch_Qwen 3_8B_nlp_causal_lm_huggingface 31.0 32 992.0 720.89 p150

Non-LLM Performance

Model Batch Sample/sec Hardware
Wan2.2-I2V-A14B-DiT 1 0.0 p150
Wan2.2-I2V-A14B-UMT5-Text-Encoder 1 13.0 p150
Wan2.2-I2V-A14B-VAE-Decoder 1 0.0 p150
Wan2.2-I2V-A14B-VAE-Encoder 1 1.0 p150
flux1-dev 1 0.0 p150
flux2 1 0.0 p150
glm-image 1 0.0 p150
hunyuan-image-2.1 1 0.0 p150
janus-pro-1b 1 0.0 n150
janus-pro-1b 1 0.0 p150
janus-pro-7b 1 0.0 p150
playground-v2.5 1 0.0 n150
playground-v2.5 1 0.0 p150
pytorch_BERT_emrecan/bert-base-turkish-cased-mean-nli-stsb-tr_nlp_embed_gen_huggingface 8 161.0 n150
pytorch_BGE-M3_Base_nlp_embed_gen_custom 4 9.0 n150
pytorch_BGE-M3_Base_nlp_embed_gen_custom 4 18.0 p150
pytorch_EfficientNet_Timm_B0_cv_image_cls_timm 8 352.0 n150
pytorch_EfficientNet_Timm_B0_cv_image_cls_timm 8 797.0 p150
pytorch_MNIST_Cnn_Dropout_cv_image_cls_custom 32 14068.0 n150
pytorch_MNIST_Cnn_Dropout_cv_image_cls_custom 32 28317.0 p150
pytorch_MobileNetV2_Mobilenet_v2_cv_image_cls_torch_hub 12 1241.0 n150
pytorch_MobileNetV2_Mobilenet_v2_cv_image_cls_torch_hub 12 2915.0 p150
pytorch_Qwen 3_Embedding_4B_nlp_embed_gen_huggingface 32 48.0 n150
pytorch_Qwen 3_Embedding_4B_nlp_embed_gen_huggingface 32 109.0 p150
pytorch_ResNet_ResNet50_HuggingFace_cv_image_cls_huggingface 8 1355.0 n150
pytorch_ResNet_ResNet50_HuggingFace_cv_image_cls_huggingface 8 2751.0 p150
pytorch_SegFormer_B0_Finetuned_Ade_512_512_cv_image_seg_huggingface 1 37.0 n150
pytorch_SegFormer_B0_Finetuned_Ade_512_512_cv_image_seg_huggingface 1 81.0 p150
pytorch_Swin_S_cv_image_cls_torchvision 1 10.0 n150
pytorch_Swin_S_cv_image_cls_torchvision 1 23.0 p150
pytorch_U-Net for Conditional Generation_Base_conditional_generation_huggingface 1 5.0 n150
pytorch_U-Net for Conditional Generation_Base_conditional_generation_huggingface 1 9.0 p150
pytorch_Ultra-Fast Lane Detection v2_TuSimple_ResNet34_Backbone_cv_image_seg_github 1 136.0 n150
pytorch_Ultra-Fast Lane Detection v2_TuSimple_ResNet34_Backbone_cv_image_seg_github 1 233.0 p150
pytorch_VGG19-UNet_base_cv_image_seg_custom 1 150.0 n150
pytorch_VGG19-UNet_base_cv_image_seg_custom 1 309.0 p150
pytorch_ViT_Base_cv_image_cls_huggingface 8 236.0 n150
pytorch_ViT_Base_cv_image_cls_huggingface 8 566.0 p150
pytorch_VoVNet_Ese_Vovnet19b_Dw.ra_In1k_cv_image_cls_timm 8 741.0 n150
pytorch_VoVNet_Ese_Vovnet19b_Dw.ra_In1k_cv_image_cls_timm 8 1565.0 p150
sdxl-lightning 1 0.0 n150
sdxl-lightning 1 0.0 p150
zimage 1 0.0 p150

Model coverage

Info: Full list of supported models is available in the assets section.

Model task Model architecture Model variant Model framework Inference Training n150 n300 p150 Single device Data parallel Tensor parallel Model source