[Discussion] Performance & Architectural Differences: ONNX Runtime (TensorRT EP) vs. Native TensorRT Backend #8868
Unanswered
RodrickSia
asked this question in
Q&A
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Title Suggestion:
[Discussion] Performance Comparison: ONNX Runtime (TensorRT Accelerator) vs. Native TensorRT
Hey everyone,
I am looking into deploying an AI model on an NVIDIA GPU and want to understand the exact performance differences between two setups:
trtexec).Since both setups use the same TensorRT engine under the hood to speed up the model, I'd love to know what the real-world performance gap looks like.
If anyone has run direct speed tests or benchmarks comparing these two approaches, please share your numbers and experiences below!
All reactions