Skip to content

2.1.2

Choose a tag to compare

@PINTO0309 PINTO0309 released this 03 Mar 15:08
· 2089 commits to main since this release
a119448

2.1.2

1. Content and background

In the flatbuffer_direct path, we had recurring issues in practical use:

  • Some ops such as TopK and MultiHeadAttention were not lowered as built-ins and ended up as custom ops, which made compatibility and accuracy validation unstable.
  • In some models, broken shape_signature propagation and side effects from layout optimizations caused RESHAPE prepare errors and intermediate tensor comparison failures.
  • During -cotof validation, memory usage could spike heavily, leading to SWAP thrashing and near-freeze behavior.
  • Progress of the OP error helper was hard to observe, so long runs could look stalled.
  • Some ORT contrib ops that should be treated as com.microsoft domain arrived as default domain, reducing runtime compatibility.

This PR addresses these issues together to improve built-in coverage, stability, and memory behavior in flatbuffer_direct.

2. Summary of corrections

  1. Expanded built-in lowering coverage

    • Extended built-in lowering for TopK.
    • Implemented both largest=1 (descending) and largest=0 (ascending-equivalent) using TOPK_V2 plus auxiliary ops.
    • Added an explicit warning for sorted=0: TFLite always returns descending-sorted output, so index order may not exactly match ONNX reference.
    • Added built-in lowering for MultiHeadAttention with strict validation (3-input Q/K/V form, rank-3, FLOAT16/FLOAT32, num_heads constraints, etc.).
    • Extended lowering for MatMul vector-rhs and dynamic Slice.
    • Fixed DepthToSpace CRD mode lowering in built-in form (channel reordering + DEPTH_TO_SPACE).
  2. Stabilized shape_signature and layout optimization behavior

    • Reinforced shape/signature propagation in index.py, shape.py, shared.py, and model_writer.py.
    • Reduced shape_signature inconsistency risk around TopK, Slice, and Reshape.
    • Hardened transpose optimizations in lower_from_onnx2tf.py.
    • Prevented unsafe Transpose -> Reshape rewrites in axis-semantic cases (e.g., rank-3), avoiding meaning drift in LogSoftmax/Transpose/Add paths.
  3. Strengthened low-memory validation flow (memmap)

    • Enabled ONNX dummy-inference output memmap path by default (can be disabled via --disable_onnxruntime_output_memmap).
    • Added --onnxruntime_output_memmap_dir to control memmap storage location.
    • Updated accuracy_evaluator to pass worker inputs/outputs via memmap, reducing large IPC tensor copies.
    • Added fallback to in-memory outputs when ONNX output memmap is not possible due to dynamic output shapes.
  4. Temp directory management and cleanup on abnormal termination

    • Added onnx2tf/utils/tempdir_cleanup.py.
    • Added managed tempdir cleanup with atexit and signal handlers (SIGTERM/SIGINT/SIGHUP/SIGQUIT).
    • Added stale tempdir cleanup using owner PID markers.
  5. Improved OP error helper operability

    • Added a single-line spinner (| / - \) during helper execution.
    • Replaced timeout-centric behavior with clearer classification/logging for crash, prepare failure, and generic failure.
    • Unified seeded input generation in helper path for better reproducibility.
  6. Improved ORT compatibility (Microsoft domain)

    • Added utility to supplement com.microsoft domain for selected contrib ops (including MultiHeadAttention) during conversion.
    • Applied temporary rewrites for runtime-check paths and added informative logs when applied.
  7. Added standalone low-memory inference script

    • Added lowmem_test_infer.py.
    • Supports ONNX/TFLite low-memory random inference in isolated worker processes (single backend or both).
    • Intended for reproducible, low-memory test-inference workflows.
  8. Test expansion

    • Added/updated regression coverage mainly in tests/test_tflite_builder_direct.py.
    • Representative local tests that passed:
    • test_flatbuffer_direct_min_topk_dynamic_k_lowering
    • test_flatbuffer_direct_matmul_vector_rhs_lowering
    • test_flatbuffer_direct_slice_dynamic_end_prefix_rank2_lowering
    • test_flatbuffer_direct_slice_dynamic_start_end_single_axis_lowering
    • test_flatbuffer_direct_logsoftmax_lowering

3. Before/After (If there is an operating log that can be used as a reference)

  • Before (examples)

    • Custom ops lowered: ... op_types=[ONNX_TOPK] ...
    • Custom ops lowered: ... op_types=[ONNX_MULTIHEADATTENTION] ...
    • tflite/kernels/reshape.cc:94 num_input_elements != num_output_elements ...
    • OP error report generation was skipped ... helper process timed out
  • After (behavior with this PR)

    • TopK and MultiHeadAttention now have built-in lowering paths.
    • TopK(sorted=0) now emits an explicit warning for potential index-order mismatch.
    • OP error helper shows live spinner progress.
    • Dynamic-shape cases that cannot use memmap safely fall back to in-memory path.
    • Tempdir lifecycle and cleanup behavior are strengthened for both normal and signal-exit paths.

4. Issue number (only if there is a related issue)

N/A

What's Changed

  • flatbuffer_direct: expand builtin op coverage and stabilize low-memory evaluation paths by @PINTO0309 in #893

Full Changelog: 2.1.1...2.1.2