Repository navigation
2.3.9
2.3.9
Important
Starting with onnx2tf v2.4.0, tf_converter will be deprecated and the default backend will be switched to flatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated to flatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them into flatbuffer_direct. I'll incorporate ai-edge-quantizer when I feel like it, but that will probably be about 10 years from now.
Summary
This PR improves the feat-torch5 line of work before merging back into main, with a focus on flatbuffer_direct lowering quality and generated PyTorch package reliability.
The branch combines structural cleanup in the PyTorch native codegen path with a broad set of behavior-preserving export fixes and graph optimizations that were validated against real-world models such as PIDNet, YOLOX, MobileBERT, SwinIR, ts_ad_model, human_segmentation_pphumanseg, iat_llie, and bread.
What Changed
1. Native PyTorch codegen refactor
- Extracted the giant native codegen implementation behind
pytorch_exporter.pyinto private pipeline/common modules. - Kept
onnx2tf/tflite_builder/pytorch_exporter.pyas the stable entrypoint while moving the heavy orchestration logic out of the main file. - Added a dedicated native codegen pipeline and state containers to reduce editor/Pylance complexity in the public module.
- Preserved generated package behavior and added structural regression coverage.
2. flatbuffer_direct lowering improvements
- Resolved static
RESHAPE-1dimensions when the input shape is fully known. - Added handling for
0copy-dim semantics whenallowZero=False, while keepingallowZero=Truebehavior unchanged. - Folded redundant consecutive
RESHAPEchains more aggressively after late shape reconciliation. - Rewrote
Div(variable, constant)patterns intoMul(variable, reciprocal_constant)when safe. - Folded consecutive constant
Mulchains into a singleMul. - Reduced redundant
Transpose/Reshapechains in multiple real models.
3. Generated PyTorch package/runtime fixes
- Fixed TorchScript export blockers in generated runtime helpers.
- Fixed Dynamo ONNX and ExportedProgram failures caused by layout-bridge handling and symbolic export incompatibilities.
- Corrected mixed NHWC/NCHW codegen around binary ops, concat/split, resize bridges, gather_nd, and broadcast constants.
- Improved channel-first shape inference for tensors whose names suggest channel-last layout but whose runtime layout is already channel-first.
- Added a safe ExportedProgram fallback path when post-export graph optimizations violate exported-program signature validation.
- Fixed generated
load_state_dicttyping so it matches the PyTorch base class contract and avoids Pylance override warnings.
4. Exported graph optimization work
- Removed redundant layout bridges from generated Dynamo ONNX and ExportedProgram artifacts using pattern-based rewrites instead of model-specific one-off handling.
- Simplified several NHWC<->NCHW permute chains around concat, add/mul trees, softmax, pooling, and segmentation head logic.
- Preserved output correctness while reducing graph noise in generated artifacts.
5. Test coverage
- Expanded regression tests for:
- static reshape concretization
- reshape chain folding
- direct builder optimization behavior
- TorchScript/Dynamo ONNX/ExportedProgram export regressions
- layout-bridge folding and mixed-layout binary paths
- native PyTorch package generation for representative models
Why This Matters
These changes make flatbuffer_direct outputs cleaner and more predictable, while also improving the reliability of the optional generated PyTorch artifacts (jit.pt, dynamo.onnx, ep.pt2).
In practice, this branch reduces redundant graph structure, resolves several export-time crashes, and improves maintainability of the PyTorch native codegen path without changing the public entrypoint.
Validation
Representative validation on this branch included:
pytest -q tests/test_tflite_builder_direct.py -k flatbuffer_directpytest -q tests/test_pytorch_exporter.py -k 'not when_model_is_available and not convert_flatbuffer_direct'- targeted regression tests in
tests/test_pytorch_exporter.py - targeted regression tests in
tests/test_tflite_builder_direct.py - repeated end-to-end CLI verification for models including:
pidnet_S_cityscapes_192x320.onnxyolov7_tiny_head_0.768_post_480x640.onnxyolox_s.onnxlite_model_mobilebert_1_metadata_1.onnxhuman_segmentation_pphumanseg_2021oct.onnxswinir-m_64x64_12.onnxts_ad_model.onnxiat_llie_180x320.onnxbread_180x320.onnx
Notes
- This branch includes both functional fixes and non-trivial internal refactoring, so review is easiest if split by area: native codegen pipeline,
flatbuffer_directlowering, and PyTorch artifact/export cleanup. - The public
pytorch_exporter.pyentrypoint remains intact even though the native codegen internals were reorganized.
What's Changed
- Improve flatbuffer_direct lowering and PyTorch export reliability by @PINTO0309 in #911
Full Changelog: 2.3.8...2.3.9