Skip to content

2.3.3

Choose a tag to compare

@PINTO0309 PINTO0309 released this 08 Mar 07:03
· 2031 commits to main since this release
30fa2d7

2.3.3

Important

Starting with onnx2tf v2.4.0, tf_converter will be deprecated and the default backend will be switched to flatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated to flatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them into flatbuffer_direct.

Summary

This PR improves the flatbuffer_direct backend by adding ONNX-input support for several legacy conversion options that were previously only effective on the tf_converter path, while intentionally keeping input_tflite_file_path direct-import validation unchanged.

The main goal is to close backend behavior gaps without weakening the direct backend's architecture or introducing hidden fallback behavior.

What This PR Adds

1. Legacy option support for flatbuffer_direct ONNX input

This PR wires the following options through the flatbuffer_direct export and lowering path:

  • disable_suppression_flextranspose
  • disable_suppression_flexstridedslice
  • optimization_for_gpu_delegate
  • replace_argmax_to_reducemax_and_indices_is_int64
  • replace_argmax_to_reducemax_and_indices_is_float32
  • replace_argmax_to_fused_argmax_and_indices_is_int64
  • replace_argmax_to_fused_argmax_and_indices_is_float32
  • fused_argmax_scale_ratio

not_use_onnxsim was already effective for ONNX-input conversion before backend branching, so this PR treats it as an existing capability and adds documentation/tests to make that explicit.

2. Direct lowering support for Flex suppression control

flatbuffer_direct now honors:

  • disable_suppression_flextranspose
  • disable_suppression_flexstridedslice

Behavioral effect:

  • When transpose suppression is disabled, the direct path emits a single builtin TRANSPOSE with the original rank instead of forcing rank compression/decomposition.
  • When strided-slice suppression is disabled, the direct path emits a single builtin SLICE/STRIDED_SLICE instead of rank-compressing the operation.
  • Existing identity/inverse-transpose fast paths remain intact.
  • Existing compression knobs still remain the default path when the disable flags are not set.

3. GPU delegate-oriented rewrites in the direct backend

This PR adds direct-path handling for optimization_for_gpu_delegate in places where the TensorFlow path already had special-case behavior.

Key additions:

  • Explicit broadcast materialization helper for direct binary arithmetic lowering.
  • Reuse of that helper in broadcast-heavy direct arithmetic paths.
  • GPU-delegate-specific normalization path for runtime negative Gather indices.
  • GPU-delegate-specific Gemm bias handling for the dynamic BATCH_MATMUL fallback path.
  • Direct BatchNormalization broadcast materialization improvements.

This keeps the implementation inside the ModelIR/direct-lowering pipeline rather than introducing any backend fallback.

4. ArgMax replacement modes in flatbuffer_direct

This PR adds direct backend support for the four legacy ArgMax replacement modes.

ReduceMax-based replacements

For:

  • replace_argmax_to_reducemax_and_indices_is_int64
  • replace_argmax_to_reducemax_and_indices_is_float32

The direct path now lowers ArgMax into a reduction-based subgraph that matches the legacy TensorFlow-path semantics, including:

  • reverse-index tie handling
  • keepdims support
  • final output dtype conversion to int64 or float32

Fused ArgMax replacements

For:

  • replace_argmax_to_fused_argmax_and_indices_is_int64
  • replace_argmax_to_fused_argmax_and_indices_is_float32

The direct path now supports fused handling for ONNX Resize -> ArgMax 4D patterns.

The implementation:

  • identifies the ONNX producer/consumer relationship directly in lowering context
  • performs a reduced-resolution ArgMax path for the fused mode
  • restores the requested spatial output size with nearest-neighbor resizing
  • preserves output shape behavior and requested final dtype semantics

This intentionally does not attempt to add unrelated ScaleAndTranslate support.

Validation / Behavior Boundaries

This PR keeps existing validation philosophy intact.

Specifically:

  • input_tflite_file_path still rejects these ONNX-lowering options because there is no ONNX graph stage in direct-import mode.
  • Existing mutual-exclusion rules for the four ArgMax replacement options are preserved.
  • Existing fused_argmax_scale_ratio validation is preserved.
  • No hidden fallback to tf_converter is introduced.

Internal Changes

The implementation includes:

  • option plumbing from convert() into export_tflite_model_flatbuffer_direct()
  • lowering-context extensions for new direct-backend flags
  • ONNX producer/consumer tracking for pattern-aware direct rewrites
  • direct builder updates across transpose, slice, gather, gemm, batch normalization, elementwise broadcast, and ArgMax lowering
  • README/API documentation updates to clarify ONNX-input vs direct-import support boundaries

Tests

This PR adds and updates tests covering:

  • not_use_onnxsim=True on ONNX-input flatbuffer_direct
  • transpose rank-compression bypass when suppression is disabled
  • slice rank-compression bypass when suppression is disabled
  • GPU delegate-specific Gather lowering
  • GPU delegate-specific Gemm lowering
  • GPU delegate broadcast materialization
  • ReduceMax ArgMax replacement paths
  • Fused ArgMax replacement paths
  • direct-import validation for unsupported ONNX-lowering options

Full flatbuffer_direct-focused test execution was also run:

  • tests/test_tflite_builder_direct.py
  • tests/test_tflite2sm_phase1.py
  • tests/test_flatbuffer_direct_op_error_report.py
  • tests/test_model_convert.py

Result:

  • 677 passed

Versioning

This branch also includes the package version bump to 2.3.3 and the matching uv.lock update.

Why This Matters

Before this change, flatbuffer_direct had a meaningful feature gap versus the legacy TensorFlow conversion path for several CLI options that users still rely on when debugging backend-specific issues, tuning delegate compatibility, or preserving legacy conversion behavior.

This PR narrows that gap while preserving the direct backend's core principles:

  • explicit lowering
  • no implicit converter fallback
  • clear validation boundaries
  • backend-specific behavior made visible in docs and tests

What's Changed

  • Add flatbuffer_direct support for legacy conversion options by @PINTO0309 in #903

Full Changelog: 2.3.2...2.3.3