1.5.0.dev20260807012312
·
1 commit
to main
since this release
Installation
Via PyPI
pip install tt-forge==1.5.0.dev20260807012312 --extra-index-url https://pypi.eng.aws.tenstorrent.com/Via Docker
docker pull ghcr.io/tenstorrent/tt-forge-slim:1.5.0.dev20260807012312Dependency commits
tt-xla commit: 65fa938d0d761da579998e50eb2bb485af3c7449
tt-mlir commit: 724d2ff7b48da2fa73977e85476f6da1c835434b
tt-metal commit: f1f4ff75579ebd7a69c7da52d45368f273026d85
What's Changed
- Uplift third_party/tt_forge_models to 8734a136b439fc3261145a411acbc038ca108e04 2026-08-01 by @vvukomanTT in #1023
Full Changelog: 1.5.0.dev20260801001900...1.5.0.dev20260807012312
LLM Performance
| Model | Token/sec/user | Batch | Token/sec | ttft (ms) | Hardware |
|---|---|---|---|---|---|
| pytorch_DeepSeek-V3.1_deepseek_v3_1_modified_nlp_causal_lm_custom | 2.0 | 128 | 256.0 | 5398.34 | n150 |
| pytorch_DeepSeek-V3.2_deepseek_v3_2_exp_modified_nlp_causal_lm_custom | 3.0 | 128 | 384.0 | 8208.69 | n150 |
| pytorch_Falcon_3_10B_Base_nlp_causal_lm_huggingface | 40.0 | 32 | 1280.0 | 863.44 | p150 |
| pytorch_Falcon_3_1B_Base_nlp_causal_lm_huggingface | 54.0 | 32 | 1728.0 | 631.18 | n150 |
| pytorch_Falcon_3_1B_Base_nlp_causal_lm_huggingface | 101.0 | 32 | 3232.0 | 293.55 | p150 |
| pytorch_Falcon_3_3B_Base_nlp_causal_lm_huggingface | 36.0 | 32 | 1152.0 | 802.59 | n150 |
| pytorch_Falcon_3_3B_Base_nlp_causal_lm_huggingface | 66.0 | 32 | 2112.0 | 373.16 | p150 |
| pytorch_Falcon_3_7B_Base_nlp_causal_lm_huggingface | 18.0 | 32 | 576.0 | 1166.21 | n150 |
| pytorch_Falcon_3_7B_Base_nlp_causal_lm_huggingface | 35.0 | 32 | 1120.0 | 491.43 | p150 |
| pytorch_GLM_4.7_nlp_causal_lm_huggingface | 6.0 | 128 | 768.0 | 2784.66 | n150 |
| pytorch_GPT-OSS_20B_nlp_causal_lm_huggingface | 21.0 | 1 | 21.0 | 343.05 | p150 |
| pytorch_Gemma_1.1_2B_IT_nlp_causal_lm_huggingface | 38.0 | 32 | 1216.0 | 621.01 | n150 |
| pytorch_Gemma_1.1_2B_IT_nlp_causal_lm_huggingface | 78.0 | 32 | 2496.0 | 238.13 | p150 |
| pytorch_Kimi-K2.6_kimi_k2_6_modified_nlp_causal_lm_custom | 6.0 | 64 | 384.0 | 3552.42 | n150 |
| pytorch_Kimi-K2_kimi_k2_instruct_modified_nlp_causal_lm_custom | 6.0 | 64 | 384.0 | 3409.06 | n150 |
| pytorch_Llama_3.1_70B_Instruct_nlp_causal_lm_huggingface | 4.0 | 32 | 128.0 | 12838.83 | n150 |
| pytorch_Llama_3.1_70B_Instruct_nlp_causal_lm_huggingface | 12.0 | 32 | 384.0 | 1806.1 | p150 |
| pytorch_Llama_3.1_8B_Instruct_nlp_causal_lm_huggingface | 21.0 | 32 | 672.0 | 1205.14 | n150 |
| pytorch_Llama_3.1_8B_Instruct_nlp_causal_lm_huggingface | 50.0 | 32 | 1600.0 | 643.93 | p150 |
| pytorch_Llama_3.2_1B_Instruct_nlp_causal_lm_huggingface | 64.0 | 32 | 2048.0 | 555.42 | n150 |
| pytorch_Llama_3.2_1B_Instruct_nlp_causal_lm_huggingface | 120.0 | 32 | 3840.0 | 231.19 | p150 |
| pytorch_Llama_3.2_3B_Instruct_nlp_causal_lm_huggingface | 30.0 | 32 | 960.0 | 579.22 | n150 |
| pytorch_Llama_3.2_3B_Instruct_nlp_causal_lm_huggingface | 54.0 | 32 | 1728.0 | 280.05 | p150 |
| pytorch_Mistral_7B_INSTRUCT_v03_nlp_causal_lm_huggingface | 20.0 | 32 | 640.0 | 1216.73 | n150 |
| pytorch_Mistral_7B_INSTRUCT_v03_nlp_causal_lm_huggingface | 35.0 | 32 | 1120.0 | 555.48 | p150 |
| pytorch_Mistral_Small_24B_INSTRUCT_2501_nlp_causal_lm_huggingface | 29.0 | 32 | 928.0 | 885.96 | p150 |
| pytorch_Phi-1.5_Phi_1_5_nlp_causal_lm_huggingface | 20.0 | 32 | 640.0 | 660.73 | n150 |
| pytorch_Phi-1.5_Phi_1_5_nlp_causal_lm_huggingface | 36.0 | 32 | 1152.0 | 326.49 | p150 |
| pytorch_Phi-1_Phi_1_nlp_causal_lm_huggingface | 20.0 | 32 | 640.0 | 641.19 | n150 |
| pytorch_Phi-1_Phi_1_nlp_causal_lm_huggingface | 37.0 | 32 | 1184.0 | 322.75 | p150 |
| pytorch_Phi-2_Phi_2_nlp_causal_lm_huggingface | 8.0 | 32 | 256.0 | 1402.2 | n150 |
| pytorch_Phi-2_Phi_2_nlp_causal_lm_huggingface | 20.0 | 32 | 640.0 | 699.18 | p150 |
| pytorch_Qwen 2.5 Coder_32B_Instruct_nlp_causal_lm_huggingface | 17.0 | 32 | 544.0 | 1493.27 | p150 |
| pytorch_Qwen 2.5_0.5B_Instruct_nlp_causal_lm_huggingface | 69.0 | 32 | 2208.0 | 408.89 | n150 |
| pytorch_Qwen 2.5_0.5B_Instruct_nlp_causal_lm_huggingface | 127.0 | 32 | 4064.0 | 182.03 | p150 |
| pytorch_Qwen 2.5_1.5B_Instruct_nlp_causal_lm_huggingface | 35.0 | 32 | 1120.0 | 473.71 | n150 |
| pytorch_Qwen 2.5_1.5B_Instruct_nlp_causal_lm_huggingface | 63.0 | 32 | 2016.0 | 211.37 | p150 |
| pytorch_Qwen 2.5_3B_Instruct_nlp_causal_lm_huggingface | 30.0 | 32 | 960.0 | 695.97 | n150 |
| pytorch_Qwen 2.5_3B_Instruct_nlp_causal_lm_huggingface | 56.0 | 32 | 1792.0 | 295.71 | p150 |
| pytorch_Qwen 2.5_7B_Instruct_nlp_causal_lm_huggingface | 16.0 | 32 | 512.0 | 809.05 | n150 |
| pytorch_Qwen 2.5_7B_Instruct_nlp_causal_lm_huggingface | 27.0 | 32 | 864.0 | 353.96 | p150 |
| pytorch_Qwen 3_0_6B_nlp_causal_lm_huggingface | 50.0 | 32 | 1600.0 | 1107.48 | n150 |
| pytorch_Qwen 3_0_6B_nlp_causal_lm_huggingface | 98.0 | 32 | 3136.0 | 519.51 | p150 |
| pytorch_Qwen 3_1_7B_nlp_causal_lm_huggingface | 35.0 | 32 | 1120.0 | 693.7 | n150 |
| pytorch_Qwen 3_1_7B_nlp_causal_lm_huggingface | 67.0 | 32 | 2144.0 | 324.73 | p150 |
| pytorch_Qwen 3_32B_nlp_causal_lm_huggingface | 17.0 | 32 | 544.0 | 1818.0 | p150 |
| pytorch_Qwen 3_4B_nlp_causal_lm_huggingface | 23.0 | 32 | 736.0 | 924.64 | n150 |
| pytorch_Qwen 3_4B_nlp_causal_lm_huggingface | 40.0 | 32 | 1280.0 | 460.98 | p150 |
| pytorch_Qwen 3_8B_nlp_causal_lm_huggingface | 16.0 | 32 | 512.0 | 1542.21 | n150 |
| pytorch_Qwen 3_8B_nlp_causal_lm_huggingface | 31.0 | 32 | 992.0 | 720.89 | p150 |
Non-LLM Performance
| Model | Batch | Sample/sec | Hardware |
|---|---|---|---|
| Wan2.2-I2V-A14B-DiT | 1 | 0.0 | p150 |
| Wan2.2-I2V-A14B-UMT5-Text-Encoder | 1 | 13.0 | p150 |
| Wan2.2-I2V-A14B-VAE-Decoder | 1 | 0.0 | p150 |
| Wan2.2-I2V-A14B-VAE-Encoder | 1 | 1.0 | p150 |
| flux1-dev | 1 | 0.0 | p150 |
| flux2 | 1 | 0.0 | p150 |
| glm-image | 1 | 0.0 | p150 |
| hunyuan-image-2.1 | 1 | 0.0 | p150 |
| janus-pro-1b | 1 | 0.0 | n150 |
| janus-pro-1b | 1 | 0.0 | p150 |
| janus-pro-7b | 1 | 0.0 | p150 |
| playground-v2.5 | 1 | 0.0 | n150 |
| playground-v2.5 | 1 | 0.0 | p150 |
| pytorch_BERT_emrecan/bert-base-turkish-cased-mean-nli-stsb-tr_nlp_embed_gen_huggingface | 8 | 161.0 | n150 |
| pytorch_BGE-M3_Base_nlp_embed_gen_custom | 4 | 9.0 | n150 |
| pytorch_BGE-M3_Base_nlp_embed_gen_custom | 4 | 18.0 | p150 |
| pytorch_EfficientNet_Timm_B0_cv_image_cls_timm | 8 | 352.0 | n150 |
| pytorch_EfficientNet_Timm_B0_cv_image_cls_timm | 8 | 797.0 | p150 |
| pytorch_MNIST_Cnn_Dropout_cv_image_cls_custom | 32 | 14068.0 | n150 |
| pytorch_MNIST_Cnn_Dropout_cv_image_cls_custom | 32 | 28317.0 | p150 |
| pytorch_MobileNetV2_Mobilenet_v2_cv_image_cls_torch_hub | 12 | 1241.0 | n150 |
| pytorch_MobileNetV2_Mobilenet_v2_cv_image_cls_torch_hub | 12 | 2915.0 | p150 |
| pytorch_Qwen 3_Embedding_4B_nlp_embed_gen_huggingface | 32 | 48.0 | n150 |
| pytorch_Qwen 3_Embedding_4B_nlp_embed_gen_huggingface | 32 | 109.0 | p150 |
| pytorch_ResNet_ResNet50_HuggingFace_cv_image_cls_huggingface | 8 | 1355.0 | n150 |
| pytorch_ResNet_ResNet50_HuggingFace_cv_image_cls_huggingface | 8 | 2751.0 | p150 |
| pytorch_SegFormer_B0_Finetuned_Ade_512_512_cv_image_seg_huggingface | 1 | 37.0 | n150 |
| pytorch_SegFormer_B0_Finetuned_Ade_512_512_cv_image_seg_huggingface | 1 | 81.0 | p150 |
| pytorch_Swin_S_cv_image_cls_torchvision | 1 | 10.0 | n150 |
| pytorch_Swin_S_cv_image_cls_torchvision | 1 | 23.0 | p150 |
| pytorch_U-Net for Conditional Generation_Base_conditional_generation_huggingface | 1 | 5.0 | n150 |
| pytorch_U-Net for Conditional Generation_Base_conditional_generation_huggingface | 1 | 9.0 | p150 |
| pytorch_Ultra-Fast Lane Detection v2_TuSimple_ResNet34_Backbone_cv_image_seg_github | 1 | 136.0 | n150 |
| pytorch_Ultra-Fast Lane Detection v2_TuSimple_ResNet34_Backbone_cv_image_seg_github | 1 | 233.0 | p150 |
| pytorch_VGG19-UNet_base_cv_image_seg_custom | 1 | 150.0 | n150 |
| pytorch_VGG19-UNet_base_cv_image_seg_custom | 1 | 309.0 | p150 |
| pytorch_ViT_Base_cv_image_cls_huggingface | 8 | 236.0 | n150 |
| pytorch_ViT_Base_cv_image_cls_huggingface | 8 | 566.0 | p150 |
| pytorch_VoVNet_Ese_Vovnet19b_Dw.ra_In1k_cv_image_cls_timm | 8 | 741.0 | n150 |
| pytorch_VoVNet_Ese_Vovnet19b_Dw.ra_In1k_cv_image_cls_timm | 8 | 1565.0 | p150 |
| sdxl-lightning | 1 | 0.0 | n150 |
| sdxl-lightning | 1 | 0.0 | p150 |
| zimage | 1 | 0.0 | p150 |
Model coverage
Info: Full list of supported models is available in the assets section.
| Model task | Model architecture | Model variant | Model framework | Inference | Training | n150 | n300 | p150 | Single device | Data parallel | Tensor parallel | Model source |
|---|