Skip to content

1.3.0.dev20260629003203

Choose a tag to compare

@github-actions github-actions released this 29 Jun 01:32
· 3 commits to main since this release
2d0a04d

Installation

Via PyPI

pip install pjrt-plugin-tt==1.3.0.dev20260629003203 --extra-index-url https://pypi.eng.aws.tenstorrent.com/
pip install vllm-tt==1.3.0.dev20260629003203 --extra-index-url https://pypi.eng.aws.tenstorrent.com/

Via Docker

docker pull ghcr.io/tenstorrent/tt-xla-slim:1.3.0.dev20260629003203

What's Changed

  • Uplift third_party/tt-mlir to ba8aa2aca750fa530773edf5c903bf3964915bd1 2026-06-28 by @vmilosevic in #5399
  • Extend n150 LLM checks and enable perf regression checks for uplift OnPR by @vkovacevicTT in #5393

Full Changelog: 1.3.0.dev20260628003030...1.3.0.dev20260629003203


LLM Performance

Model Token/sec/user Batch Token/sec ttft (ms) Hardware
Qwen/Qwen2.5-0.5B-Instruct 228.0 32 7296.0 825.94 n150
Qwen/Qwen2.5-1.5B-Instruct 194.0 32 6208.0 1450.84 n150
Qwen/Qwen2.5-3B-Instruct 185.0 32 5920.0 2389.74 n150
Qwen/Qwen2.5-7B-Instruct 157.0 32 5024.0 4871.82 n150
Qwen/Qwen3-0.6B 218.0 32 6976.0 1384.65 n150
Qwen/Qwen3-1.7B 199.0 32 6368.0 1716.64 n150
Qwen/Qwen3-4B 175.0 32 5600.0 3183.85 n150
Qwen/Qwen3-8B 140.0 32 4480.0 4136.02 n150
meta-llama/Llama-3.1-8B-Instruct 164.0 32 5248.0 3902.18 n150
meta-llama/Llama-3.2-1B-Instruct 249.0 32 7968.0 978.8 n150
meta-llama/Llama-3.2-3B-Instruct 187.0 32 5984.0 2245.63 n150
microsoft/phi-1 402.0 32 12864.0 1877.18 n150
microsoft/phi-1_5 390.0 32 12480.0 1879.85 n150
microsoft/phi-2 283.0 32 9056.0 4111.14 n150
mistralai/Ministral-8B-Instruct-2410 155.0 32 4960.0 4107.6 n150
mistralai/Mistral-7B-Instruct-v0.3 285.0 32 9120.0 3662.98 n150
pytorch_DeepSeek-V3.2_deepseek_v3_2_exp_modified_nlp_causal_lm_custom 2.0 128 256.0 7536.72 n150
pytorch_Falcon_3_10B_Base_nlp_causal_lm_huggingface 5.0 32 160.0 1969.5 n150
pytorch_Falcon_3_10B_Base_nlp_causal_lm_huggingface 11.0 32 352.0 947.39 p150
pytorch_Falcon_3_1B_Base_nlp_causal_lm_huggingface 7.0 32 224.0 832.78 n150
pytorch_Falcon_3_3B_Base_nlp_causal_lm_huggingface 7.0 32 224.0 1036.36 n150
pytorch_Falcon_3_3B_Base_nlp_causal_lm_huggingface 13.0 32 416.0 494.55 p150
pytorch_Falcon_3_7B_Base_nlp_causal_lm_huggingface 6.0 32 192.0 1344.4 n150
pytorch_Falcon_3_7B_Base_nlp_causal_lm_huggingface 11.0 32 352.0 630.94 p150
pytorch_GPT-OSS_20B_nlp_causal_lm_huggingface 14.0 1 14.0 623.15 n150
pytorch_Gemma_1.1_2B_IT_nlp_causal_lm_huggingface 5.0 32 160.0 1235.69 n150
pytorch_Gemma_1.1_2B_IT_nlp_causal_lm_huggingface 8.0 32 256.0 651.36 p150
pytorch_Llama_3.1_70B_Instruct_nlp_causal_lm_huggingface 2.0 32 64.0 6451.96 n150
pytorch_Llama_3.1_8B_Instruct_nlp_causal_lm_huggingface 4.0 32 128.0 2366.26 n150
pytorch_Llama_3.1_8B_Instruct_nlp_causal_lm_huggingface 10.0 32 320.0 703.02 p150
pytorch_Llama_3.2_1B_Instruct_nlp_causal_lm_huggingface 8.0 32 256.0 743.83 n150
pytorch_Llama_3.2_1B_Instruct_nlp_causal_lm_huggingface 13.0 32 416.0 405.44 p150
pytorch_Llama_3.2_3B_Instruct_nlp_causal_lm_huggingface 7.0 32 224.0 793.19 n150
pytorch_Llama_3.2_3B_Instruct_nlp_causal_lm_huggingface 11.0 32 352.0 423.63 p150
pytorch_Mistral_7B_INSTRUCT_v03_nlp_causal_lm_huggingface 22.0 32 704.0 587.9 p150
pytorch_Mistral_Ministral_8B_Instruct_nlp_causal_lm_huggingface 11.0 32 352.0 884.0 p150
pytorch_Mistral_Nemo_INSTRUCT_2407_nlp_causal_lm_huggingface 11.0 32 352.0 996.09 p150
pytorch_Mistral_Small_24B_INSTRUCT_2501_nlp_causal_lm_huggingface 4.0 32 128.0 1956.94 n150
pytorch_Phi-1.5_Phi_1_5_nlp_causal_lm_huggingface 12.0 32 384.0 682.59 n150
pytorch_Phi-1.5_Phi_1_5_nlp_causal_lm_huggingface 20.0 32 640.0 343.25 p150
pytorch_Phi-1_Phi_1_nlp_causal_lm_huggingface 11.0 32 352.0 706.73 n150
pytorch_Phi-1_Phi_1_nlp_causal_lm_huggingface 20.0 32 640.0 349.89 p150
pytorch_Phi-2_Phi_2_nlp_causal_lm_huggingface 15.0 32 480.0 599.26 p150
pytorch_Qwen 2.5 Coder_32B_Instruct_nlp_causal_lm_huggingface 3.0 32 96.0 2890.96 n150
pytorch_Qwen 2.5 Coder_32B_Instruct_nlp_causal_lm_huggingface 5.0 32 160.0 1609.67 p150
pytorch_Qwen 2.5_0.5B_Instruct_nlp_causal_lm_huggingface 7.0 32 224.0 724.82 n150
pytorch_Qwen 2.5_0.5B_Instruct_nlp_causal_lm_huggingface 11.0 32 352.0 433.57 p150
pytorch_Qwen 2.5_1.5B_Instruct_nlp_causal_lm_huggingface 7.0 32 224.0 782.83 n150
pytorch_Qwen 2.5_1.5B_Instruct_nlp_causal_lm_huggingface 10.0 32 320.0 451.3 p150
pytorch_Qwen 2.5_14B_Instruct_nlp_causal_lm_huggingface 5.0 32 160.0 1158.75 p150
pytorch_Qwen 2.5_3B_Instruct_nlp_causal_lm_huggingface 6.0 32 192.0 908.31 n150
pytorch_Qwen 2.5_3B_Instruct_nlp_causal_lm_huggingface 10.0 32 320.0 495.43 p150
pytorch_Qwen 2.5_7B_Instruct_nlp_causal_lm_huggingface 5.0 32 160.0 1125.92 n150
pytorch_Qwen 2.5_7B_Instruct_nlp_causal_lm_huggingface 8.0 32 256.0 565.58 p150
pytorch_Qwen 3_0_6B_nlp_causal_lm_huggingface 6.0 32 192.0 1374.88 n150
pytorch_Qwen 3_0_6B_nlp_causal_lm_huggingface 10.0 32 320.0 722.14 p150
pytorch_Qwen 3_14B_nlp_causal_lm_huggingface 5.0 32 160.0 1273.3 p150
pytorch_Qwen 3_1_7B_nlp_causal_lm_huggingface 6.0 32 192.0 996.84 n150
pytorch_Qwen 3_1_7B_nlp_causal_lm_huggingface 10.0 32 320.0 504.75 p150
pytorch_Qwen 3_32B_nlp_causal_lm_huggingface 2.0 32 64.0 5023.13 n150
pytorch_Qwen 3_32B_nlp_causal_lm_huggingface 4.0 32 128.0 1973.44 p150
pytorch_Qwen 3_4B_nlp_causal_lm_huggingface 5.0 32 160.0 1208.54 n150
pytorch_Qwen 3_4B_nlp_causal_lm_huggingface 9.0 32 288.0 645.27 p150
pytorch_Qwen 3_8B_nlp_causal_lm_huggingface 5.0 32 160.0 1857.06 n150
pytorch_Qwen 3_8B_nlp_causal_lm_huggingface 5.0 32 160.0 1119.88 p150
tiiuae/Falcon3-1B-Base 249.0 32 7968.0 1246.39 n150
tiiuae/Falcon3-3B-Base 199.0 32 6368.0 1976.73 n150
tiiuae/Falcon3-7B-Base 269.0 32 8608.0 1689.15 n150

Non-LLM Performance

Model Batch Sample/sec Hardware
BAAI/bge-m3 1 42.0 n150
Qwen/Qwen3-Embedding-4B 1 9.0 n150
Wan2.2-I2V-A14B-DiT 1 0.0 p150
Wan2.2-I2V-A14B-UMT5-Text-Encoder 1 17.0 p150
Wan2.2-I2V-A14B-VAE-Decoder 1 0.0 p150
Wan2.2-I2V-A14B-VAE-Encoder 1 1.0 p150
playground-v2.5 1 0.0 n150
playground-v2.5 1 0.0 p150
pytorch_BERT_emrecan/bert-base-turkish-cased-mean-nli-stsb-tr_nlp_embed_gen_huggingface 8 44.0 n150
pytorch_BERT_emrecan/bert-base-turkish-cased-mean-nli-stsb-tr_nlp_embed_gen_huggingface 8 114.0 p150
pytorch_BGE-M3_Base_nlp_embed_gen_custom 4 9.0 n150
pytorch_BGE-M3_Base_nlp_embed_gen_custom 4 19.0 p150
pytorch_EfficientNet_Timm_B0_cv_image_cls_timm 8 348.0 n150
pytorch_EfficientNet_Timm_B0_cv_image_cls_timm 8 796.0 p150
pytorch_MNIST_Cnn_Dropout_cv_image_cls_custom 32 14745.0 n150
pytorch_MNIST_Cnn_Dropout_cv_image_cls_custom 32 32936.0 p150
pytorch_MobileNetV2_Mobilenet_v2_cv_image_cls_torch_hub 12 1234.0 n150
pytorch_MobileNetV2_Mobilenet_v2_cv_image_cls_torch_hub 12 3030.0 p150
pytorch_Qwen 3_Embedding_4B_nlp_embed_gen_huggingface 32 46.0 n150
pytorch_Qwen 3_Embedding_4B_nlp_embed_gen_huggingface 32 105.0 p150
pytorch_ResNet_ResNet50_HuggingFace_cv_image_cls_huggingface 8 1342.0 n150
pytorch_ResNet_ResNet50_HuggingFace_cv_image_cls_huggingface 8 2832.0 p150
pytorch_SegFormer_B0_Finetuned_Ade_512_512_cv_image_seg_huggingface 1 38.0 n150
pytorch_SegFormer_B0_Finetuned_Ade_512_512_cv_image_seg_huggingface 1 85.0 p150
pytorch_Swin_S_cv_image_cls_torchvision 1 10.0 n150
pytorch_Swin_S_cv_image_cls_torchvision 1 23.0 p150
pytorch_U-Net for Conditional Generation_Base_conditional_generation_huggingface 1 5.0 n150
pytorch_U-Net for Conditional Generation_Base_conditional_generation_huggingface 1 9.0 p150
pytorch_Ultra-Fast Lane Detection v2_TuSimple_ResNet34_Backbone_cv_image_seg_github 1 136.0 n150
pytorch_Ultra-Fast Lane Detection v2_TuSimple_ResNet34_Backbone_cv_image_seg_github 1 253.0 p150
pytorch_VGG19-UNet_base_cv_image_seg_custom 1 140.0 n150
pytorch_VGG19-UNet_base_cv_image_seg_custom 1 308.0 p150
pytorch_ViT_Base_cv_image_cls_huggingface 8 226.0 n150
pytorch_ViT_Base_cv_image_cls_huggingface 8 548.0 p150
pytorch_VoVNet_Ese_Vovnet19b_Dw.ra_In1k_cv_image_cls_timm 8 666.0 n150
pytorch_VoVNet_Ese_Vovnet19b_Dw.ra_In1k_cv_image_cls_timm 8 1518.0 p150
sdxl-lightning 1 0.0 n150
sdxl-lightning 1 0.0 p150

Model coverage

Info: Full list of supported models is available in the assets section.

Model task Model architecture Model variant Model framework Inference Training n150 n300 p150 Single device Data parallel Tensor parallel Model source
conditional generation U-Net for Conditional Generation Base pytorch View Source
cv image cls AlexNet Custom 1x2 jax View Source
cv image cls DINOv2 Small pytorch View Source
cv image cls EfficientNet B0 pytorch View Source
cv image cls MNIST Cnn Batchnorm jax View Source
cv image cls MNIST Cnn Dropout jax View Source
cv image cls MNIST Cnn Dropout pytorch View Source
cv image cls MNIST Cnn Nodropout pytorch View Source
cv image cls MNIST Mlp Custom jax View Source
cv image cls MNIST Mlp Custom jax View Source
cv image cls MNIST Mlp Custom 1x2 jax View Source
cv image cls MobileNetV1 Mobilenet v1 pytorch View Source
cv image cls MobileNetV2 Mobilenet v2 pytorch View Source
cv image cls ResNet ResNet50 HuggingFace High Resolution pytorch View Source
cv image cls SegFormer Mit B0 pytorch View Source
cv image cls Swin S pytorch View Source
cv image cls VGG HF Vgg19 pytorch View Source
cv image cls ViT Base pytorch View Source
cv image cls VoVNet Ese Vovnet19b Dw.ra In1k pytorch View Source
cv image seg Ultra-Fast Lane Detection TuSimple ResNet18 Backbone pytorch View Source
cv image seg VGG19-UNet base pytorch View Source
cv img to img Autoencoder linear pytorch View Source
cv object det Attention DenseUNet Base pytorch View Source
cv object det DETR ResNet50 Backbone pytorch View Source
cv object det OWL-ViT Base Patch32 pytorch View Source
cv object det PointPillars pointpillars pytorch View Source
cv object det YOLOP Default pytorch View Source
cv object det YOLOS Small Small pytorch View Source
cv object det YOLOv4 Base pytorch View Source
cv object det YOLOv7 Default pytorch View Source
cv object det YOLOv9 T pytorch View Source
cv object det ssd512 ssd512 pytorch View Source
mm action prediction OpenVLA-OFT Finetuned Libero 10 pytorch View Source
mm action prediction pi_0 pi0 base pytorch View Source
mm image text similarity CLIP Base Patch16 pytorch View Source
mm image text similarity SigLIP Base Patch16 224 pytorch View Source
mm visual qa Llama 3.2 11B Vision Instruct pytorch View Source
mm visual qa Mistral base pytorch View Source
nlp causal lm ALLaM 7B Instruct pytorch View Source
nlp causal lm Command_A_Reasoning command-a-reasoning-08-2025 pytorch View Source
nlp causal lm Falcon 3 10B Base pytorch View Source
nlp causal lm Falcon 3 1B Base pytorch View Source
nlp causal lm Falcon 3 3B Base pytorch View Source
nlp causal lm Falcon 3 7B Base pytorch View Source
nlp causal lm GPT-2 Base jax View Source
nlp causal lm GPT-2 Xl jax View Source
nlp causal lm GPT-OSS 20B pytorch View Source
nlp causal lm Gemma 1.1 2B IT pytorch View Source
nlp causal lm Gemma 1.1 7B IT pytorch View Source
nlp causal lm Gemma 2 27B IT pytorch View Source
nlp causal lm Gemma 2 2B IT pytorch View Source
nlp causal lm Gemma 2 9B IT pytorch View Source
nlp causal lm Llama 3.1 70B pytorch View Source
nlp causal lm Llama 3.1 70B Instruct pytorch View Source
nlp causal lm Llama 3.1 8B Instruct pytorch View Source
nlp causal lm Llama 3.2 1B pytorch View Source
nlp causal lm Llama 3.2 3B pytorch View Source
nlp causal lm Llama 3.3 70B Instruct pytorch View Source
nlp causal lm Mistral 7B INSTRUCT v03 pytorch View Source
nlp causal lm Mistral Devstral Small 2505 pytorch View Source
nlp causal lm Mistral Large INSTRUCT 2411 pytorch View Source
nlp causal lm Mistral Magistral Small 2506 pytorch View Source
nlp causal lm Mistral Ministral 8B Instruct pytorch View Source
nlp causal lm Mistral Nemo INSTRUCT 2407 pytorch View Source
nlp causal lm Mistral Small 24B INSTRUCT 2501 pytorch View Source
nlp causal lm Phi-1 Phi 1 jax View Source
nlp causal lm Phi-1 Phi 1 pytorch View Source
nlp causal lm Phi-1 Phi 1 pytorch View Source
nlp causal lm Phi-1 LoRA Phi 1 pytorch View Source
nlp causal lm Phi-1.5 Phi 1 5 jax View Source
nlp causal lm Phi-1.5 Phi 1 5 pytorch View Source
nlp causal lm Phi-2 Phi 2 jax View Source
nlp causal lm Phi-2 Phi 2 pytorch View Source
nlp causal lm Phi-3 Mini 128K Instruct pytorch View Source
nlp causal lm Phi-3 Mini 4K Instruct pytorch View Source
nlp causal lm Phi-3 Mini Instruct pytorch View Source
nlp causal lm Phi-4 Phi 4 pytorch View Source
nlp causal lm Qwen 2 Qwq 32B pytorch View Source
nlp causal lm Qwen 2.5 0.5B jax View Source
nlp causal lm Qwen 2.5 0.5B Instruct jax View Source