Repository navigation
2.3.14
2.3.14
Important
Starting with onnx2tf v2.4.0, tf_converter will be deprecated and the default backend will be switched to flatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated to flatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them into flatbuffer_direct. I'll incorporate ai-edge-quantizer when I feel like it, but that will probably be about 10 years from now.
Summary
This PR significantly improves the feat-torch10 native PyTorch export path and the flatbuffer_direct workflow, with an emphasis on replacing model-specific fixes with generalized structural logic.
The branch expands native package generation so that difficult ONNX graphs can stay on the native PyTorch backend more reliably, finish export faster, and preserve accuracy across a broad set of real-world models.
What changed
1. Generalized native PyTorch layout and structural repairs
A large portion of the work in this branch moves the exporter away from brittle model-name- or op-name-based routing and toward shape-, layout-, and dataflow-based repair rules.
Representative improvements include:
- generalized mixed-layout repair logic for
Add,Concat,Resize,Reshape,Slice, pooling, channel-shuffle patterns, and public output bridges - generalized structural repair rules for ALIKE, PIDNet, ShadowFormer, HumanSeg, NanoDet+, FastestDet, detection heads, and feature-last / channel-first bridge patterns
- generalized rank-3 and rank-4 canonicalization fixes in native package generation
- generalized depth-to-space and public-head bridge rewrites
- generalized skip / avoid-model-ir predicates so fast native canonicalization can be selected based on graph structure instead of exact generated code shapes
The practical effect is that many previously fragile native PyTorch conversions now succeed through common structural rules instead of accumulating one-off exceptions.
2. Faster native export routing and reduced pathological raw canonicalization cost
The branch adds broader fast-path routing and avoid-model-ir detection for native package export. This is especially important for models that are small but were previously taking an unexpectedly long time in write pytorch because they fell into expensive raw canonicalization.
Examples of the improvements in this area include:
- generalized fast-path detection for multiscale detection tails
- generalized avoid-model-ir predicates for structurally safe native exports
- native export timeout handling for child export steps
- native layout inference and native codegen shape inference speedups
A representative case is version-RFB-640.onnx, where the write pytorch phase now returns to a practical runtime while preserving native execution and accuracy.
3. Better evaluation robustness and tooling
This branch also improves supporting tooling around export validation:
- a new
flatbuffer_direct_bulk_runnerutility and its test coverage - updates to native package runtime helpers
- improved accuracy evaluation behavior for low-energy outputs, preventing false negatives caused by cosine-only gating in cases where absolute error is already negligible
4. Expanded regression coverage
The test suite was significantly expanded, especially around the native PyTorch exporter and structural repair logic. This includes:
- targeted unit tests for newly generalized repair rules
- native package generation checks for representative model families
- tests for the new bulk runner and accuracy evaluator behavior
Why this matters
The main goal of this branch is not just to fix isolated regressions, but to improve the maintainability and reproducibility of the native PyTorch export path.
Instead of continuing to accumulate exact-name patches, the branch pushes the exporter toward a more defensible strategy:
- infer intent from graph structure
- preserve layout semantics explicitly
- canonicalize only when evidence is strong
- prefer reusable repair rules that cover families of models
This should reduce future regression risk and make new model support less dependent on ad hoc exceptions.
Validation
The branch was validated with native PyTorch output comparison on the following regression set, all passing:
age_googlenet.onnxalike_t_opset11_192x320.onnxAtan_11.onnxbaseline_simplified.onnxbread_180x320.onnxbread_nonfm_180x320.onnxdeeplabv3_mobilenet_v3_large.onnxdetpth_to_space_17.onnxdetr_demo.onnxdigits.onnxefficientformer_l1.onnxFastestDet.onnxhuman_segmentation_pphumanseg_2021oct.onnxiat_llie_180x320.onnxmobilenetv2-10.onnxnanodet-plus-m_416.onnxpidnet_S_cityscapes_192x320.onnxreducemax_softmax_workaround.onnxresnet18-v1-7.onnxrfdn_64x64.onnxshadowformer_istd_160x240.onnxsinet_320_op.onnxswinir-m_64x64_12.onnxts_ad_model.onnxversion-RFB-640.onnxyolox_s.onnx
In addition, targeted exporter tests covering the restored version-RFB-640 and yolox_s fixes were re-run and passed.
Notes for reviewers
Because this branch contains a long sequence of generalization work, the most useful review lens is by subsystem rather than by individual commit:
onnx2tf/tflite_builder/pytorch_exporter.pyonnx2tf/tflite_builder/_pytorch_exporter_native_codegen_pipeline.pyonnx2tf/tflite_builder/accuracy_evaluator.pyonnx2tf/utils/flatbuffer_direct_bulk_runner.pytests/test_pytorch_exporter.pytests/test_flatbuffer_direct_bulk_runner.pytests/test_accuracy_evaluator_seeded_input.py
Recommended focus areas:
- native export routing and canonicalization boundaries
- generalized structural repair predicates
- postprocess rewrites in native code generation
- regression coverage for mixed-layout and detection-tail cases
What's Changed
- Improve native PyTorch export robustness and flatbuffer_direct tooling by @PINTO0309 in #917
Full Changelog: 2.3.13...2.3.14