Repository navigation
2.3.3
2.3.3
Important
Starting with onnx2tf v2.4.0, tf_converter will be deprecated and the default backend will be switched to flatbuffer_direct. With the v2.3.3 update, all backward compatible conversion options have been migrated to flatbuffer_direct, so I will only be doing minor bug fixes until April. If you provide us with ONNX sample models, I will consider incorporating them into flatbuffer_direct.
Summary
This PR improves the flatbuffer_direct backend by adding ONNX-input support for several legacy conversion options that were previously only effective on the tf_converter path, while intentionally keeping input_tflite_file_path direct-import validation unchanged.
The main goal is to close backend behavior gaps without weakening the direct backend's architecture or introducing hidden fallback behavior.
What This PR Adds
1. Legacy option support for flatbuffer_direct ONNX input
This PR wires the following options through the flatbuffer_direct export and lowering path:
disable_suppression_flextransposedisable_suppression_flexstridedsliceoptimization_for_gpu_delegatereplace_argmax_to_reducemax_and_indices_is_int64replace_argmax_to_reducemax_and_indices_is_float32replace_argmax_to_fused_argmax_and_indices_is_int64replace_argmax_to_fused_argmax_and_indices_is_float32fused_argmax_scale_ratio
not_use_onnxsim was already effective for ONNX-input conversion before backend branching, so this PR treats it as an existing capability and adds documentation/tests to make that explicit.
2. Direct lowering support for Flex suppression control
flatbuffer_direct now honors:
disable_suppression_flextransposedisable_suppression_flexstridedslice
Behavioral effect:
- When transpose suppression is disabled, the direct path emits a single builtin
TRANSPOSEwith the original rank instead of forcing rank compression/decomposition. - When strided-slice suppression is disabled, the direct path emits a single builtin
SLICE/STRIDED_SLICEinstead of rank-compressing the operation. - Existing identity/inverse-transpose fast paths remain intact.
- Existing compression knobs still remain the default path when the disable flags are not set.
3. GPU delegate-oriented rewrites in the direct backend
This PR adds direct-path handling for optimization_for_gpu_delegate in places where the TensorFlow path already had special-case behavior.
Key additions:
- Explicit broadcast materialization helper for direct binary arithmetic lowering.
- Reuse of that helper in broadcast-heavy direct arithmetic paths.
- GPU-delegate-specific normalization path for runtime negative Gather indices.
- GPU-delegate-specific Gemm bias handling for the dynamic
BATCH_MATMULfallback path. - Direct BatchNormalization broadcast materialization improvements.
This keeps the implementation inside the ModelIR/direct-lowering pipeline rather than introducing any backend fallback.
4. ArgMax replacement modes in flatbuffer_direct
This PR adds direct backend support for the four legacy ArgMax replacement modes.
ReduceMax-based replacements
For:
replace_argmax_to_reducemax_and_indices_is_int64replace_argmax_to_reducemax_and_indices_is_float32
The direct path now lowers ArgMax into a reduction-based subgraph that matches the legacy TensorFlow-path semantics, including:
- reverse-index tie handling
keepdimssupport- final output dtype conversion to
int64orfloat32
Fused ArgMax replacements
For:
replace_argmax_to_fused_argmax_and_indices_is_int64replace_argmax_to_fused_argmax_and_indices_is_float32
The direct path now supports fused handling for ONNX Resize -> ArgMax 4D patterns.
The implementation:
- identifies the ONNX producer/consumer relationship directly in lowering context
- performs a reduced-resolution ArgMax path for the fused mode
- restores the requested spatial output size with nearest-neighbor resizing
- preserves output shape behavior and requested final dtype semantics
This intentionally does not attempt to add unrelated ScaleAndTranslate support.
Validation / Behavior Boundaries
This PR keeps existing validation philosophy intact.
Specifically:
input_tflite_file_pathstill rejects these ONNX-lowering options because there is no ONNX graph stage in direct-import mode.- Existing mutual-exclusion rules for the four ArgMax replacement options are preserved.
- Existing
fused_argmax_scale_ratiovalidation is preserved. - No hidden fallback to
tf_converteris introduced.
Internal Changes
The implementation includes:
- option plumbing from
convert()intoexport_tflite_model_flatbuffer_direct() - lowering-context extensions for new direct-backend flags
- ONNX producer/consumer tracking for pattern-aware direct rewrites
- direct builder updates across transpose, slice, gather, gemm, batch normalization, elementwise broadcast, and ArgMax lowering
- README/API documentation updates to clarify ONNX-input vs direct-import support boundaries
Tests
This PR adds and updates tests covering:
not_use_onnxsim=Trueon ONNX-inputflatbuffer_direct- transpose rank-compression bypass when suppression is disabled
- slice rank-compression bypass when suppression is disabled
- GPU delegate-specific Gather lowering
- GPU delegate-specific Gemm lowering
- GPU delegate broadcast materialization
- ReduceMax ArgMax replacement paths
- Fused ArgMax replacement paths
- direct-import validation for unsupported ONNX-lowering options
Full flatbuffer_direct-focused test execution was also run:
tests/test_tflite_builder_direct.pytests/test_tflite2sm_phase1.pytests/test_flatbuffer_direct_op_error_report.pytests/test_model_convert.py
Result:
677 passed
Versioning
This branch also includes the package version bump to 2.3.3 and the matching uv.lock update.
Why This Matters
Before this change, flatbuffer_direct had a meaningful feature gap versus the legacy TensorFlow conversion path for several CLI options that users still rely on when debugging backend-specific issues, tuning delegate compatibility, or preserving legacy conversion behavior.
This PR narrows that gap while preserving the direct backend's core principles:
- explicit lowering
- no implicit converter fallback
- clear validation boundaries
- backend-specific behavior made visible in docs and tests
What's Changed
- Add flatbuffer_direct support for legacy conversion options by @PINTO0309 in #903
Full Changelog: 2.3.2...2.3.3