Skip to content

v2.1.2 - Correct adaptive pooling and zero-op counting rules (#148)

Choose a tag to compare

@github-actions github-actions released this 28 Jul 18:26
· 26 commits to main since this release
487cb79

🌟 Summary

THOP 2.1.2 improves PyTorch model profiling accuracy, expands module-rule coverage, fixes parameter and formatting edge cases, and adds automated testing across supported environments. 🎯

📊 Key Changes

  • Corrected adaptive average pooling counts 🧮
    Counts now follow the integer pooling windows PyTorch actually uses, producing more accurate and consistent MAC estimates for output sizes that do not evenly divide the input.

  • Fixed upsampling operation rules 🚀
    Nearest and nearest-exact upsampling are now correctly treated as requiring zero arithmetic operations. Costs for other interpolation modes remain assigned to their owning hook.

  • Expanded normalization and zero-operation coverage 🧩

    • GroupNorm and RMSNorm now use the existing normalization rule.
    • Constant-padding and dropout families are handled through shared base classes.
    • ReLU and LeakyReLU are explicitly treated as zero-MAC operations.
    • Lazy normalization modules receive more reliable rule resolution.
  • Improved counting-rule inheritance 🧬
    Rules are now resolved through a module type’s inheritance hierarchy. Custom subclasses, parametrized layers, and many lazy modules can inherit the nearest supported rule instead of being incorrectly reported as unsupported.

  • Improved parameter counting 📊
    Parameter totals now come directly from the model’s parameter tree rather than profiling hooks. This correctly handles shared weights, unused branches, unsupported module types, and parameters owned by parent modules.

  • Preserved model state during profiling 🛡️
    Profiling now restores each module’s original training or evaluation mode, avoiding unintended changes in models with mixed training states.

  • Fixed clever_format() edge cases ✨
    Exact powers of 1,000 and negative values now receive the correct units—for example, 1000 becomes 1.00K and -1e9 becomes -1.00G.

  • Added continuous integration testing ✅
    A new GitHub Actions workflow runs the test suite on pull requests, pushes, and nightly schedules across the supported Python and PyTorch range, including Python 3.8 with PyTorch 1.8.0 and Python 3.14 with the latest PyTorch.

  • Updated documentation and benchmarks 📚
    README badges, custom-rule guidance, Chinese documentation, and model benchmark rows—including YOLO11 and YOLO26 entries—now reflect the corrected counts.

  • Version updated to 2.1.2 📦
    This patch release packages the profiling, reliability, documentation, and testing improvements together.

🎯 Purpose & Impact

  • More trustworthy MAC estimates: Pooling and interpolation results better match actual PyTorch behavior, which improves model-complexity comparisons and deployment planning.
  • Fewer false warnings and missing counts: Inherited rules cover more real-world architectures, including custom and lazy layers.
  • More reliable parameter totals: Shared, unused, and previously uncounted parameters are represented correctly without double-counting.
  • Safer profiling: Users can profile models without permanently altering their training modes or leaving temporary hooks and attributes behind.
  • Stronger release confidence: Automated tests help catch compatibility regressions across old and new supported PyTorch environments.
  • No new MultiheadAttention rule shipped: The proposed rule was intentionally removed because PyTorch 1.8 hooks cannot observe keyword-only arguments reliably; avoiding partial support prevents valid models from crashing. ⚠️

What's Changed

  • Run the test suite in CI by @raimbekovm in #147
  • Format an exact power of 1000 and a negative value in the right unit by @raimbekovm in #146
  • Count parameters from the module tree instead of the profiling hooks by @raimbekovm in #143
  • Resolve counting rules through the module type's mro by @raimbekovm in #144
  • Leave the profiled model as it was found and correct the counting rules by @raimbekovm in #145
  • Correct adaptive pooling and zero-op counting rules by @raimbekovm in #148

Full Changelog: v2.1.1...v2.1.2