Skip to content

2.3.13

Choose a tag to compare

@PINTO0309 PINTO0309 released this 19 Mar 23:04
· 1932 commits to main since this release
e3e54f5

2.3.13

Important

Starting with onnx2tf v2.4.0, tf_converter will be deprecated and the default backend will be switched to flatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated to flatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them into flatbuffer_direct. I'll incorporate ai-edge-quantizer when I feel like it, but that will probably be about 10 years from now.

Summary

This PR improves flatbuffer_direct conversion stability, native PyTorch package generation robustness, and constant-folding quality for several previously failing or regressing models.

The main goals of this branch were:

  • reduce or eliminate write pytorch regressions in the fast path
  • repair layout/shape corruption in generated native PyTorch packages
  • improve constant folding for ALIKE-style exported graphs
  • keep previously successful models from regressing while fixing new failures

What Changed

1. Improved native PyTorch export repair coverage

This branch expands the fast precanonicalization repair logic in pytorch_exporter.py to handle several classes of broken generated code more reliably.

Key improvements include:

  • propagation of channel-first aliases so downstream softmax, concat, and related layout-sensitive ops are repaired consistently
  • repair of malformed DepthToSpace-adjacent gather patterns in NHWC/CF bridge code
  • restoration of broken or missing target shapes in resize, pool, scalar-binary, and alignment chains
  • preservation and restoration of missing internal/output shape annotations for generated Dynamo ONNX artifacts
  • repair of malformed pool/LRN/conv chains and stage-local layout bridges that previously caused accuracy loss or runtime failure
  • targeted repair for nanodet-plus-m_416 stage-0 max-pool layout corruption, which also removes a severe write pytorch slowdown

These changes were validated against the regressions reported during branch development, including:

  • alike_t_opset11_192x320
  • efficientformer_l1
  • nanodet-plus-m_416
  • FastestDet
  • age_googlenet

2. Better constant folding for ALIKE and similar exported graphs

This branch adds new constant-folding coverage in both the TFLite-side lowering flow and the generated Dynamo ONNX sanitization flow.

Highlights:

  • constant ScatterND evaluation/folding
  • constant Reshape folding
  • constant binary op folding for Add/Sub/Mul/Div
  • TFLite-side folding of constant ScatterND and follow-up binary chains on cloned export IR only, so the original IR remains safe for PyTorch package generation

These changes simplify generated artifacts and remove unnecessary constant computation chains in ALIKE-derived models.

3. Safer handling of integer-sensitive division lowering

lower_from_onnx2tf.py now avoids over-aggressive reciprocal-multiply lowering for precision-sensitive paths that later feed integer Cast consumers.

This preserves exact division semantics where needed, which was necessary to fix descriptor/indexing accuracy regressions in ALIKE without weakening the general constant-division optimization path.

4. AveragePool exclude-pad correction behavior refined

This branch also tightens AveragePool(count_include_pad=0) handling so that:

  • TFLite export avoids incorrect extra correction in SAME cases that should already behave correctly
  • native PyTorch package generation still retains the correction behavior required to preserve ONNX semantics for models such as efficientformer_l1

This separation was important to fix TFLite accuracy while avoiding PyTorch-package regressions.

5. Regression coverage expanded substantially

The test suite has been extended with focused regression tests for:

  • ALIKE dynamic score sampling and constant-folding cases
  • missing Dynamo ONNX output/internal shape restoration
  • constant ScatterND / constant Reshape / constant binary folding
  • EfficientFormer attention scalar-mul target repair
  • dynamic and malformed pool target-shape repair
  • NanoDet fast-path detection and stage-0 pool bridge repair
  • TFLite-side integer-division preservation and constant scatter folding
  • AveragePool count_include_pad=0 handling across SAME, explicit pad, and ceil-mode cases

Validation

In addition to the new unit tests, the branch was checked with real model conversions on the flatbuffer_direct path.

Confirmed as passing on this branch:

  • alike_t_opset11_192x320: ONNX/TFLite pass=True, ONNX/PyTorch pass=True
  • efficientformer_l1: ONNX/TFLite pass=True, ONNX/PyTorch pass=True
  • FastestDet: ONNX/TFLite pass=True, ONNX/PyTorch pass=True
  • age_googlenet: ONNX/TFLite pass=True, ONNX/PyTorch pass=True
  • nanodet-plus-m_416: ONNX/PyTorch pass=True, and the severe write pytorch slowdown was removed

For nanodet-plus-m_416, the native PyTorch package regression was fixed and the generation time returned to a practical range, but ONNX/TFLite still reports a remaining numerical mismatch on this branch. That issue is not introduced by this PR; the work here focuses on removing the package-generation/runtime regression and preserving previously validated models.

Version

  • bumped package version from 2.3.12 to 2.3.13
  • updated README image/tag examples accordingly

What's Changed

  • Improve flatbuffer_direct export robustness and repair coverage by @PINTO0309 in #915

Full Changelog: 2.3.12...2.3.13