Skip to content

2.3.14

Choose a tag to compare

@PINTO0309 PINTO0309 released this 23 Mar 09:04
· 1772 commits to main since this release
85a6284

2.3.14

Important

Starting with onnx2tf v2.4.0, tf_converter will be deprecated and the default backend will be switched to flatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated to flatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them into flatbuffer_direct. I'll incorporate ai-edge-quantizer when I feel like it, but that will probably be about 10 years from now.

Summary

This PR significantly improves the feat-torch10 native PyTorch export path and the flatbuffer_direct workflow, with an emphasis on replacing model-specific fixes with generalized structural logic.

The branch expands native package generation so that difficult ONNX graphs can stay on the native PyTorch backend more reliably, finish export faster, and preserve accuracy across a broad set of real-world models.

What changed

1. Generalized native PyTorch layout and structural repairs

A large portion of the work in this branch moves the exporter away from brittle model-name- or op-name-based routing and toward shape-, layout-, and dataflow-based repair rules.

Representative improvements include:

  • generalized mixed-layout repair logic for Add, Concat, Resize, Reshape, Slice, pooling, channel-shuffle patterns, and public output bridges
  • generalized structural repair rules for ALIKE, PIDNet, ShadowFormer, HumanSeg, NanoDet+, FastestDet, detection heads, and feature-last / channel-first bridge patterns
  • generalized rank-3 and rank-4 canonicalization fixes in native package generation
  • generalized depth-to-space and public-head bridge rewrites
  • generalized skip / avoid-model-ir predicates so fast native canonicalization can be selected based on graph structure instead of exact generated code shapes

The practical effect is that many previously fragile native PyTorch conversions now succeed through common structural rules instead of accumulating one-off exceptions.

2. Faster native export routing and reduced pathological raw canonicalization cost

The branch adds broader fast-path routing and avoid-model-ir detection for native package export. This is especially important for models that are small but were previously taking an unexpectedly long time in write pytorch because they fell into expensive raw canonicalization.

Examples of the improvements in this area include:

  • generalized fast-path detection for multiscale detection tails
  • generalized avoid-model-ir predicates for structurally safe native exports
  • native export timeout handling for child export steps
  • native layout inference and native codegen shape inference speedups

A representative case is version-RFB-640.onnx, where the write pytorch phase now returns to a practical runtime while preserving native execution and accuracy.

3. Better evaluation robustness and tooling

This branch also improves supporting tooling around export validation:

  • a new flatbuffer_direct_bulk_runner utility and its test coverage
  • updates to native package runtime helpers
  • improved accuracy evaluation behavior for low-energy outputs, preventing false negatives caused by cosine-only gating in cases where absolute error is already negligible

4. Expanded regression coverage

The test suite was significantly expanded, especially around the native PyTorch exporter and structural repair logic. This includes:

  • targeted unit tests for newly generalized repair rules
  • native package generation checks for representative model families
  • tests for the new bulk runner and accuracy evaluator behavior

Why this matters

The main goal of this branch is not just to fix isolated regressions, but to improve the maintainability and reproducibility of the native PyTorch export path.

Instead of continuing to accumulate exact-name patches, the branch pushes the exporter toward a more defensible strategy:

  • infer intent from graph structure
  • preserve layout semantics explicitly
  • canonicalize only when evidence is strong
  • prefer reusable repair rules that cover families of models

This should reduce future regression risk and make new model support less dependent on ad hoc exceptions.

Validation

The branch was validated with native PyTorch output comparison on the following regression set, all passing:

  • age_googlenet.onnx
  • alike_t_opset11_192x320.onnx
  • Atan_11.onnx
  • baseline_simplified.onnx
  • bread_180x320.onnx
  • bread_nonfm_180x320.onnx
  • deeplabv3_mobilenet_v3_large.onnx
  • detpth_to_space_17.onnx
  • detr_demo.onnx
  • digits.onnx
  • efficientformer_l1.onnx
  • FastestDet.onnx
  • human_segmentation_pphumanseg_2021oct.onnx
  • iat_llie_180x320.onnx
  • mobilenetv2-10.onnx
  • nanodet-plus-m_416.onnx
  • pidnet_S_cityscapes_192x320.onnx
  • reducemax_softmax_workaround.onnx
  • resnet18-v1-7.onnx
  • rfdn_64x64.onnx
  • shadowformer_istd_160x240.onnx
  • sinet_320_op.onnx
  • swinir-m_64x64_12.onnx
  • ts_ad_model.onnx
  • version-RFB-640.onnx
  • yolox_s.onnx

In addition, targeted exporter tests covering the restored version-RFB-640 and yolox_s fixes were re-run and passed.

Notes for reviewers

Because this branch contains a long sequence of generalization work, the most useful review lens is by subsystem rather than by individual commit:

  • onnx2tf/tflite_builder/pytorch_exporter.py
  • onnx2tf/tflite_builder/_pytorch_exporter_native_codegen_pipeline.py
  • onnx2tf/tflite_builder/accuracy_evaluator.py
  • onnx2tf/utils/flatbuffer_direct_bulk_runner.py
  • tests/test_pytorch_exporter.py
  • tests/test_flatbuffer_direct_bulk_runner.py
  • tests/test_accuracy_evaluator_seeded_input.py

Recommended focus areas:

  • native export routing and canonicalization boundaries
  • generalized structural repair predicates
  • postprocess rewrites in native code generation
  • regression coverage for mixed-layout and detection-tail cases

What's Changed

  • Improve native PyTorch export robustness and flatbuffer_direct tooling by @PINTO0309 in #917

Full Changelog: 2.3.13...2.3.14