Skip to content

2.3.2

Choose a tag to compare

@PINTO0309 PINTO0309 released this 08 Mar 05:18
· 2036 commits to main since this release
209e3b2

2.3.2

Starting with onnx2tf v2.4.0, tf_converter will be deprecated and the default backend will be switched to flatbuffer_direct.

Summary

This PR expands and hardens the flatbuffer_direct backend so that more conversion workflows stay on the direct ModelIR/FlatBuffer path instead of falling back to the legacy tf_converter path.

The branch focuses on improving functional parity, making interrupt/rewrite features work consistently in direct mode, and tightening the documentation/tests around the new behavior.

What changed

1. Keep more workflows on the direct path

flatbuffer_direct now keeps the direct export flow for cases that previously forced a fallback or were rejected:

  • --output_h5
  • --output_keras_v3
  • --output_tfv1_pb
  • --disable_model_save
  • -it/--input_tflite_file_path direct-import workflows

Instead of dropping to tf_converter, these flows now use a ModelIR-derived SavedModel bridge when a non-TFLite artifact is required.

This means direct mode can now:

  • produce .h5, .keras, and TFv1 .pb artifacts without leaving the direct backend
  • run with --disable_model_save while still performing internal staging/validation and leaving no final artifacts in the requested output directory
  • apply the same behavior to both ONNX input and -it input when appropriate

Validation was also tightened so incompatible combinations are rejected explicitly rather than silently changing execution mode.

2. ModelIR-based interrupt handling for -inimc / -onimc

flatbuffer_direct no longer depends on ONNX graph extraction for interrupt-based cropping.

Instead, it now crops the already lowered/imported top-level ModelIR using the requested boundary tensor names. This applies to:

  • ONNX input
  • -it/--input_tflite_file_path input

Behavioral impact:

  • direct mode no longer relies on sne4onnx for this path
  • the cropped model is the basis for downstream direct export, SavedModel bridging, split planning, and evaluation
  • boundary names are validated against top-level ModelIR tensors and invalid requests fail explicitly

3. Direct-mode support for -dgc, -ebu, and -eru

Direct mode now supports these options through shared ModelIR rewrites instead of forcing the TensorFlow conversion flow:

  • -dgc / --disable_group_convolution
  • -ebu / --enable_batchmatmul_unfold
  • -eru / --enable_rnn_unroll

The implementation unifies ONNX input and imported TFLite ModelIR handling by applying rewrites after lowering/import and before split planning / SavedModel bridging / final export.

In practice this adds:

  • grouped CONV_2D decomposition in direct mode
  • BATCH_MATMUL unfolding in direct mode
  • direct recurrent unrolling support for the supported RNN/LSTM cases

If a requested rewrite cannot be applied safely, conversion now fails explicitly instead of behaving like a no-op.

4. Direct lowering for MeanVarianceNormalization with -me / --mvn_epsilon

This branch adds direct lowering support for ONNX MeanVarianceNormalization in flatbuffer_direct.

The new lowering expands MVN into builtin-friendly primitive ops:

  • MEAN
  • SUB
  • MUL
  • MEAN
  • ADD
  • SQRT
  • DIV

mvn_epsilon is now threaded through the direct backend so -me / --mvn_epsilon affects both:

  • tf_converter
  • flatbuffer_direct

for ONNX input.

This removes one more feature gap between the two backends while keeping direct export fully self-contained.

5. CLI / README / API cleanup

This branch also updates the public surface and documentation so the new direct-path behavior is visible and consistent:

  • add short option -esm for --eval_split_models
  • fix broken formatting in the README CLI parameter section
  • document the new flatbuffer_direct semantics for:
    • SavedModel-bridge based artifact generation
    • interrupt handling
    • ModelIR rewrites
    • --disable_model_save
    • direct MVN support
  • update package/dependency metadata and lockfile sync in line with the branch changes

Why this matters

This work reduces the number of cases where users must understand or work around backend switches.

Before this branch, enabling certain artifact outputs or graph-manipulation options could unexpectedly move execution away from flatbuffer_direct, or reject flows that were conceptually still compatible with direct export.

After this branch, the direct backend is much more coherent:

  • direct mode stays direct more often
  • ONNX input and -it input behave more consistently
  • interrupt/rewrite features are ModelIR-native
  • more CLI options have explicit, documented semantics instead of fallback-driven behavior

Testing

I validated the branch with focused and backend-wide tests, including:

  • targeted MVN direct-lowering tests
  • direct-path interrupt / crop tests
  • direct rewrite tests for grouped convolution, batch matmul unfolding, and recurrent unroll
  • SavedModel-bridge / artifact-generation direct-path tests
  • backend-wide flatbuffer_direct test sweep

Most recent backend-wide run:

pytest -q tests -k flatbuffer_direct

Result:

  • 599 passed
  • 141 deselected
  • 2 warnings

The warnings were non-failing and already known:

  • HDF5 legacy save warning from tf_keras
  • float16 cast overflow warning in an existing negative-infinity broadcast test

What's Changed

  • Improve flatbuffer_direct direct-path coverage and ModelIR rewrites by @PINTO0309 in #902

Full Changelog: 2.3.1...2.3.2