2.1.2
2.1.2
1. Content and background
In the flatbuffer_direct path, we had recurring issues in practical use:
- Some ops such as
TopKandMultiHeadAttentionwere not lowered as built-ins and ended up as custom ops, which made compatibility and accuracy validation unstable. - In some models, broken
shape_signaturepropagation and side effects from layout optimizations causedRESHAPEprepare errors and intermediate tensor comparison failures. - During
-cotofvalidation, memory usage could spike heavily, leading to SWAP thrashing and near-freeze behavior. - Progress of the OP error helper was hard to observe, so long runs could look stalled.
- Some ORT contrib ops that should be treated as
com.microsoftdomain arrived as default domain, reducing runtime compatibility.
This PR addresses these issues together to improve built-in coverage, stability, and memory behavior in flatbuffer_direct.
2. Summary of corrections
-
Expanded built-in lowering coverage
- Extended built-in lowering for
TopK. - Implemented both
largest=1(descending) andlargest=0(ascending-equivalent) usingTOPK_V2plus auxiliary ops. - Added an explicit warning for
sorted=0: TFLite always returns descending-sorted output, so index order may not exactly match ONNX reference. - Added built-in lowering for
MultiHeadAttentionwith strict validation (3-input Q/K/V form, rank-3, FLOAT16/FLOAT32,num_headsconstraints, etc.). - Extended lowering for MatMul vector-rhs and dynamic
Slice. - Fixed
DepthToSpaceCRDmode lowering in built-in form (channel reordering +DEPTH_TO_SPACE).
- Extended built-in lowering for
-
Stabilized shape_signature and layout optimization behavior
- Reinforced shape/signature propagation in
index.py,shape.py,shared.py, andmodel_writer.py. - Reduced
shape_signatureinconsistency risk aroundTopK,Slice, andReshape. - Hardened transpose optimizations in
lower_from_onnx2tf.py. - Prevented unsafe
Transpose -> Reshaperewrites in axis-semantic cases (e.g., rank-3), avoiding meaning drift inLogSoftmax/Transpose/Addpaths.
- Reinforced shape/signature propagation in
-
Strengthened low-memory validation flow (memmap)
- Enabled ONNX dummy-inference output memmap path by default (can be disabled via
--disable_onnxruntime_output_memmap). - Added
--onnxruntime_output_memmap_dirto control memmap storage location. - Updated
accuracy_evaluatorto pass worker inputs/outputs via memmap, reducing large IPC tensor copies. - Added fallback to in-memory outputs when ONNX output memmap is not possible due to dynamic output shapes.
- Enabled ONNX dummy-inference output memmap path by default (can be disabled via
-
Temp directory management and cleanup on abnormal termination
- Added
onnx2tf/utils/tempdir_cleanup.py. - Added managed tempdir cleanup with
atexitand signal handlers (SIGTERM/SIGINT/SIGHUP/SIGQUIT). - Added stale tempdir cleanup using owner PID markers.
- Added
-
Improved OP error helper operability
- Added a single-line spinner (
| / - \) during helper execution. - Replaced timeout-centric behavior with clearer classification/logging for crash, prepare failure, and generic failure.
- Unified seeded input generation in helper path for better reproducibility.
- Added a single-line spinner (
-
Improved ORT compatibility (Microsoft domain)
- Added utility to supplement
com.microsoftdomain for selected contrib ops (includingMultiHeadAttention) during conversion. - Applied temporary rewrites for runtime-check paths and added informative logs when applied.
- Added utility to supplement
-
Added standalone low-memory inference script
- Added
lowmem_test_infer.py. - Supports ONNX/TFLite low-memory random inference in isolated worker processes (single backend or both).
- Intended for reproducible, low-memory test-inference workflows.
- Added
-
Test expansion
- Added/updated regression coverage mainly in
tests/test_tflite_builder_direct.py. - Representative local tests that passed:
test_flatbuffer_direct_min_topk_dynamic_k_loweringtest_flatbuffer_direct_matmul_vector_rhs_loweringtest_flatbuffer_direct_slice_dynamic_end_prefix_rank2_loweringtest_flatbuffer_direct_slice_dynamic_start_end_single_axis_loweringtest_flatbuffer_direct_logsoftmax_lowering
- Added/updated regression coverage mainly in
3. Before/After (If there is an operating log that can be used as a reference)
-
Before (examples)
Custom ops lowered: ... op_types=[ONNX_TOPK] ...Custom ops lowered: ... op_types=[ONNX_MULTIHEADATTENTION] ...tflite/kernels/reshape.cc:94 num_input_elements != num_output_elements ...OP error report generation was skipped ... helper process timed out
-
After (behavior with this PR)
TopKandMultiHeadAttentionnow have built-in lowering paths.TopK(sorted=0)now emits an explicit warning for potential index-order mismatch.- OP error helper shows live spinner progress.
- Dynamic-shape cases that cannot use memmap safely fall back to in-memory path.
- Tempdir lifecycle and cleanup behavior are strengthened for both normal and signal-exit paths.
4. Issue number (only if there is a related issue)
N/A
What's Changed
- flatbuffer_direct: expand builtin op coverage and stabilize low-memory evaluation paths by @PINTO0309 in #893
Full Changelog: 2.1.1...2.1.2