Repository navigation
2.3.1
2.3.1
Starting with onnx2tf v2.4.0, tf_converter will be deprecated and the default backend will be switched to flatbuffer_direct.
Summary
This PR substantially expands the flatbuffer_direct split workflow and makes the direct path more coherent across export, validation, TFLite input import, and SavedModel output.
The main goal of this branch is to turn split handling into a first-class ModelIR-based feature for flatbuffer_direct, rather than relying on the legacy ONNX/GraphSurgeon split flow.
Main improvements
1. Unified split behavior under enable_auto_split_model
flatbuffer_direct now uses a single ModelIR-based split planner for direct export.
Key behavior changes:
enable_auto_split_model=Trueis the public split entry point forflatbuffer_direct.- Split planning now runs directly against
ModelIRand produces the manifest-based artifact set. - The legacy ONNX recursive split path is no longer used for
flatbuffer_direct. - Small models still succeed and emit a valid single-partition manifest instead of failing or falling back.
This makes split behavior deterministic and consistent with the direct backend’s internal representation.
onnx2tf \
-i deim_hgnetv2_s_wholebody28_ft_1250query_fixed.onnx \
-cotof \
-tb flatbuffer_direct \
-easm \
-asms 5MB # KB or MB or GB
2. Cleaner split artifacts and partition boundaries
The partition builder was tightened so split artifacts are more faithful and easier to inspect.
Improvements include:
- embedded weights/constants are no longer exposed as partition runtime inputs
- dead branch operators are pruned from partition subgraphs
- split manifests now reflect only real runtime / cross-partition inputs
- partition TFLite graphs avoid the misleading “isolated constant/input” appearance that previously showed up in viewers such as Netron
This specifically addresses the issue where split partitions appeared to have many disconnected constants despite being otherwise executable.
3. Split SavedModel export and validation
The direct split pipeline now supports partition-level SavedModel output.
New capabilities:
flatbuffer_direct + enable_auto_split_model + flatbuffer_direct_output_saved_modelexports partition SavedModels- split SavedModel validation runs partition-by-partition in manifest order and compares against the appropriate TFLite reference
- when split SavedModel output is requested, the workflow is aligned across ONNX-input and TFLite-input direct paths
This closes the gap between split TFLite output and SavedModel output for the direct backend.
4. -it / TFLite-input direct split support
TFLite-import (-it) can now participate in the direct split flow.
That means:
- imported TFLite can be lowered into
ModelIR - the same split planner can be applied to imported TFLite models
- split artifacts and optional split SavedModels can be generated from TFLite input as well
This makes the direct backend’s split functionality available beyond ONNX-originated conversions.
5. Evaluation and validation cleanup
Split evaluation and direct validation behavior were simplified and clarified.
Changes include:
eval_split_modelsis now the single split evaluation interface and directly encodes the reference mode (onnxorunsplit_tflite)- the separate
eval_split_referenceoption was removed - split evaluation now adapts NHWC partition inputs correctly, fixing false failures caused by input layout mismatches
-cotofno longer emits a SavedModel inference warning when the workflow did not actually produce a SavedModel- logs and reports now distinguish unsplit base accuracy evaluation from split-manifest evaluation more clearly
These changes reduce ambiguity in both CLI usage and generated reports.
6. External-data safety for ONNX shape inference
The branch also hardens ONNX handling by enforcing a repository-wide runtime policy:
onnx.shape_inference.infer_shapesis not run when an ONNX model usesexternal_data- the same policy is applied to the shared lowering helper and relevant test helpers
- symbolic fallback inference is also skipped in that case
This avoids unsafe shape-inference calls on external-data models while preserving the existing behavior for regular in-memory ONNX models.
Public-facing behavior changes
Notable interface changes:
auto_split_tflite_by_sizewas removed as a publicflatbuffer_directsplit entry pointenable_auto_split_modelis the public split trigger forflatbuffer_directauto_split_max_sizeremains the split target-size controleval_split_referencewas removedeval_split_modelsnow takes the reference mode directly (onnxorunsplit_tflite)
These changes intentionally reduce redundant split/evaluation options and make the direct backend easier to operate.
Testing
I ran the flatbuffer-direct-focused regression suite:
pytest -q \
tests/test_tflite_builder_direct.py \
tests/test_tflite_builder_op_coverage.py \
tests/test_flatbuffer_direct_op_error_report.py \
tests/test_accuracy_evaluator_input_layout.py \
tests/test_accuracy_evaluator_name_map.py \
tests/test_accuracy_evaluator_seeded_input.py \
tests/test_accuracy_evaluator_subprocess.py \
tests/test_tflite_split_planner.py \
tests/test_tflite_builder_gridsample_validation.pyResult:
625 passed, 1 warning
The remaining warning is an existing float16 cast overflow warning during one direct-path test and does not fail the suite.
Why this PR matters
Before this branch, flatbuffer_direct split support was fragmented across multiple partially overlapping flags and code paths. This branch consolidates the split workflow around the direct backend’s own ModelIR, aligns TFLite and SavedModel outputs, improves artifact correctness, and removes several confusing validation/reporting edge cases.
As a result, feat-split turns split export from an experimental side path into a much more coherent and reviewable feature set for flatbuffer_direct.
What's Changed
- feat: expand flatbuffer_direct split export and validation workflows by @PINTO0309 in #901
Full Changelog: 2.3.0...2.3.1