Repository navigation
v2.1.0 - Speed up profiling 2.4x and remove the broken fx profiler (#141)
🌟 Summary
THOP v2.1.0 makes PyTorch model profiling substantially faster and simpler, delivering up to 2.4× lower profiling overhead while preserving bit-exact results. ⚡
📊 Key Changes
-
Faster profiling through optimized bookkeeping
- Replaced temporary float64 tensor buffers with lightweight Python integer attributes for operation and parameter totals.
- Converted counting functions to return native Python numbers instead of allocating tensors.
- Combined operation counting and parameter counting into a single forward hook.
- Calculates each module’s parameter count once during hook registration rather than during every forward pass.
-
Significant performance improvements
- Reduced isolated profiling overhead from 13.4 ms to 5.6 ms, a 2.4× speedup.
- Across 66 Ultralytics model configurations,
get_flopsbecame approximately 1.31× faster overall. - Real-world gains vary because model inference time still dominates total runtime.
-
Removed the broken
torch.fxprofiler- Deleted the unused and unreliable
fx_profile.pyimplementation. - This reduces maintenance burden and avoids exposing users to profiler behavior that was incomplete or inconsistent.
- Deleted the unused and unreliable
-
Improved compatibility for custom operation rules
- Custom rules that accumulate tensor values continue to work.
- Profiling outputs are normalized to standard Python
floatvalues. - Added tests covering custom operations, built-in counting rules, and expected MAC and parameter totals. ✅
-
Cleaner and more efficient counting internals
- Simplified RNN, GRU, LSTM, pooling, normalization, activation, and upsampling counters.
- Removed deprecated calculation helpers and unnecessary NumPy, warning, and tensor dependencies.
🎯 Purpose & Impact
- Faster model analysis: Developers can measure MACs, FLOPs, and parameter counts with less profiling overhead, especially for models containing many modules or hooks. 🚀
- More responsive tooling: Ultralytics utilities that rely on profiling, such as
get_flops, should complete faster without changing reported results. - Lower runtime overhead: Fewer hooks and fewer temporary tensor allocations reduce unnecessary PyTorch bookkeeping.
- Safer results: Bit-exact validation means the speed improvements do not alter profiling totals.
- Easier custom integrations: Third-party modules can continue using custom counting rules while receiving predictable Python numeric outputs.
- Upgrade consideration: Code that directly imports or depends on the removed
thop.fx_profilemodule will need to migrate to the standardprofile()API.
What's Changed
- Speed up profiling 2.4x and remove the broken fx profiler by @glenn-jocher in #141
Full Changelog: v2.0.22...v2.1.0