v0.10.0 #172
nvluxiaoz
announced in
Announcements
v0.10.0
#172
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
TensorRT Edge-LLM 0.10.0 Release 2026-08-12
We are excited to announce the release 0.10.0 of TensorRT Edge-LLM!
TensorRT Edge-LLM 0.10.0 provides Day-0 support for NVIDIA Nemotron-3.5 Lightning with MTP and DFlash. It also adds support for Cosmos3-Edge, DiffusionGemma, Nemotron-3.5-ASR, and DSpark speculative decoding. This release also introduces an experimental direct TensorRT engine builder without ONNX export, multi-turn KV-cache reuse, and video input for the experimental OpenAI-compatible server.
Key Features
Other Important Features
Runtime and Performance
Export and Quantization
Server and API
/v1/audio/speechfor Qwen3-TTS, and/v1/audio/transcriptionsfor Qwen3-ASR.Examples and Documentation
NVIDIA Contributors
@nvluxiaoz @nvamberl @xiangg-nv @willg-nv @mahu888 @nv-samcheng @duofant @Caohanwen0 @JCalafato @duanyaqi @zhazhang-nv @zhaoyuanh-nvidia @zhijial-nvidia @ever-wong @Jasper-NV @jhalabi-nv @levichen-nvidia @sunghyunp-nvidia @nvyocox @ruocheng-nv @yuanyao-nv @qikail-ctrl
This discussion was created from the release v0.10.0.
All reactions