Repository navigation
2.6.2
2.6.2
Summary
This PR substantially refactors the TensorFlow-free flatbuffer_direct conversion path around deterministic, indexed, transactional ModelIR passes while preserving the existing CLI/Python API and compatibility entry points.
The main goal is to make converter behavior easier to reason about and safer to extend. Previously, many layout and cleanup rules lived in the central lowerer, repeatedly scanned the complete graph, and depended on implicit call order. This branch gives those rules explicit ownership, stable ordering, shared graph/layout state, transactional validation, focused characterization tests, and observable pass diagnostics.
The branch contains 202 commits and changes 104 files. Much of the apparent diff size comes from mechanically extracting existing rule implementations and their large inline tests from the central lowerer into dedicated pass and test modules.
What changed
Deterministic pass infrastructure
- Introduced and expanded the ordered pass manager with stable pass IDs, phases, priorities, preconditions, explicit change results, maximum iteration counts, and graph fingerprints.
- Added shared
ModelIRPassStateso repeated pass groups reuse oneModelIRGraphIndexand oneLayoutStateinstead of rebuilding graph maps for every rule. - Added differential ModelIR mutation APIs for input/output replacement and operator insertion/removal. Producer, consumer, duplicate-producer, and operator-position indices stay synchronized as rewrites are applied.
- Added cheap graph preflight predicates so irrelevant pass groups return before creating snapshots, fingerprints, or full state.
- Added deterministic pass diagnostics and bulk-runner metrics for invocation count, candidate visits, snapshots, fingerprints, changes, rollback, and timing.
- Transactional validation now compares invariant problems before and after a rewrite. A pass is rolled back only when it introduces a new problem; inherited temporary issues are not incorrectly attributed to a later pass. Final ModelIR validation remains strict.
Layout state and rewrite safety
- Centralized logical/physical layout tracking in
LayoutStateand synchronized it through indexed graph mutations. - Protected graph inputs/outputs, public intermediate tensors, fan-out boundaries, shared constants, quantization metadata, and tensor lineage during rewrites.
- Added copy-on-write handling where layout conversion would otherwise mutate a constant used by an unrelated branch.
- Replaced model-name-oriented coverage with semantic graph-pattern guards. Production passes do not add model-name exceptions.
- Preserved compatibility wrappers in
lower_from_onnx2tf.pywhile moving implementation ownership to focused pass-family modules.
Pass-family extraction and migration
The branch extracts and/or migrates the following families from the central lowerer into dedicated modules under onnx2tf/tflite_builder/passes/:
- boundary input, passthrough, Pad, high-rank MatMul, constant fold, Cast, and graph cleanup;
- quantized PReLU, quantized Reshape, Q/DQ and quantization cleanup;
- singleton Reshape/MaxPool/channel/spatial cleanup and consecutive/flatten-Concat-Reshape cleanup;
- Transpose cleanup, unary/fan-out bridges, Gather-axis/channel fan-out, and trailing-output adapters;
- NCHW/NHWC channel shuffle, Mean, LayerNorm statistics, and terminal Mean layout;
- attention/QKV, SE, elementwise gate, multi-branch gate, complementary gate, NDHWC/Conv3D gate, and post-add/post-convolution gate families;
- cost-volume/ScatterND, dual-Mul/Concat, Add/Concat constant suffix, axis-3 constant Concat bridge, Dequantize/Concat/Quantize, Concat/unary/Conv, generic SPP, and NDHWC pre-Concat families.
Each migrated family follows a staged process:
- characterize current success and rejection boundaries with compact generated ModelIR/ONNX fixtures;
- mechanically extract the existing implementation without changing behavior;
- migrate traversal and mutation to the shared index/layout state;
- register a stable transactional pass and replace raw production calls;
- run focused, architecture, efficiency, and representative model regression tests.
Reporting and numeric evaluation
- Extracted coverage and tensor-correspondence reporting into
tflite_builder/reporting.py, retaining thin compatibility wrappers in the lowerer. - Kept reporting within the TensorFlow-free dependency boundary.
- Added sequential accuracy-evaluation fallback for fully static models:
- evaluate with LiteRT default delegates first;
- retry without default delegates only if the primary result fails either the existing
evaluation_passrule or the hardmax_abs_error <= 1e-1contract; - never overlap inference processes;
- deterministically select the better completed report;
- retain the primary report if fallback evaluation fails.
- The public accuracy JSON schema and command behavior are unchanged.
Compatibility and packaging
- Existing CLI/Python API entry points and legacy lowerer helper signatures remain available through adapters/wrappers.
- No new dependency package was introduced.
- Normal
-tb flatbuffer_directconversion and-cotofremain TensorFlow-free. Optional TensorFlow exporters remain behind their existing optional boundary. - Updated the package and documented container version from
2.6.1to2.6.2;uv.lockremains consistent.
Validation
All local commands were run in the uv environment. Model conversion and inference were executed sequentially with at most one inference subprocess at a time.
Automated tests
- Latest complete sequential direct regression selection before the final recovery fixes:
1170 passed, 5 deselected, 2 warnings in 162.39s
- Focused and related suites after the final pass-validation and evaluator fixes:
- 92 tests passed
- 65 tests passed
- 157 distinct tests passed in total
- These selections cover transaction rollback, graph/index consistency, layout state, graph cleanup, pass ordering and efficiency, evaluator input layout/name mapping/seed behavior, subprocess isolation, architecture boundaries, and optional-TensorFlow import blocking.
- Python bytecode compilation of the final changed Python files passed.
uv lock --checkpassed.- Importing the package reports
onnx2tf.__version__ == "2.6.2". git diff --checkpassed.
Model regression coverage
- A representative eight-model sentinel set passed.
- A 74-model blast-radius cohort covering the affected duplicate-Transpose, Reshape, and terminal-Mean passes passed conversion and the strict numeric contract.
- The four regressions attributed to the refactored pass-manager behavior were recovered:
birdnet.onnx:max_abs_error=6.4849853515625e-05imageclassifier.onnx:6.67572021484375e-06model_resnet15x224_swish-072.onnx:7.43865966796875e-05resnet18-v1-7.onnx:7.152557373046875e-06
- The three delegate-evaluation regressions were recovered:
keras_rnn.onnx:6.258487701416016e-07pose_estimation_mediapipe_2023mar.onnx:0.05517578125pose_estimation_mediapipe_2023mar_int8bq.onnx:0.0443115234375
- All seven recovered models returned exit status zero, reported
evaluation_pass=true, and satisfiedmax_abs_error <= 1e-1.
Performance-oriented changes
- Reuse graph indices and constant buffers instead of rebuilding equivalent maps for each pass.
- Skip irrelevant pass groups before allocating snapshots or computing fingerprints.
- Use differential updates instead of repeated whole-ModelIR scans after each mutation.
- Collect deterministic per-pass metrics so future changes can identify unexpected repeated scans, snapshots, or cycles.
- Execute only requested evaluation/export work; evaluator fallback is not run when the primary result already passes the strict contract.
Formal tier-wide warm-up/three-run median timing and peak-RSS gates are not claimed as complete in this PR.
Known remaining work
- The full fixed-order Tier 0 through Tier 4 corpus has not been rerun after the final two recovery commits. The final evidence currently consists of the seven recovered models, eight sentinels, and the 74-model blast-radius cohort.
- Expected-timeout models remain excluded from the active regression corpus, and DEIM is intentionally classified as successful.
- Fifteen previously investigated non-timeout candidates remain unresolved:
- two DPT producer-rank cases (known first-bad commit
881b714); double_gru.onnx(known first-bad commit94af81b);- numeric mismatches in
model1.onnxandfast_acvnet_generalization_opset16_192x320.onnx; - ten inherited runtime/report-generation failures documented in the handoff.
- two DPT producer-rank cases (known first-bad commit
- Tier 5 models (2,000 or more ONNX operations) remain intentionally deferred until the Tier 0 through Tier 4 contract is fully revalidated.
- GitHub Actions had no run records for this branch at the final local checkpoint, so this PR does not claim external CI completion.
Detailed implementation and validation history is recorded in:
docs/flatbuffer_direct_architecture.mddocs/flatbuffer_direct_handoff_2026-07-12.mddocs/flatbuffer_direct_handoff_2026-07-13.md
What's Changed
- Refactor flatbuffer_direct around indexed transactional passes by @PINTO0309 in #948
Full Changelog: 2.6.1...2.6.2