Skip to content

2.2.2

Choose a tag to compare

@PINTO0309 PINTO0309 released this 06 Mar 11:21
· 2061 commits to main since this release
5201514

2.2.2

Summary

This PR strengthens flatbuffer_direct conversion quality for DAMO-style graphs and broadens built-in operator coverage while keeping model contracts stable at graph boundaries.

Functional Improvements

  • Added robust Float32 export path from FP16-heavy IR:
    • Introduced clone_model_ir_with_float32 to promote FP16 tensors/options into FP32.
    • Added prune_identity_cast_operators and optimize_redundant_transpose_operators to simplify generated graphs before writing TFLite.
    • Applied the same transpose cleanup to FP16 export and preserved reduced-precision metadata handling.
  • Improved late-stage dynamic shape correctness:
    • _resolve_dynamic_reshape_shapes now supports prefer_runtime_inferable_from_onnx_raw=True to preserve ONNX runtime-inferable -1 templates instead of stale concretized shapes.
    • Extended dynamic-lineage preservation for RANGE outputs with runtime-dependent lengths.
    • Prevented aggressive reshape passthrough folding when preserveDynamicShape=True is set.
  • Added LiteRT compatibility fallback for unsupported SPLIT input dtypes:
    • Rewrites unsupported SPLIT into SLICE chains (with optional cast in/out) so conversion succeeds for dtype combinations not accepted by LiteRT SPLIT.
  • Added a new graph optimizer for DAMO-like normalization/padding layouts:
    • _optimize_transpose_flatten_globalnorm_pad_prepost_nhwc_chains removes redundant pre/post transpose wrappers and rewrites affected shape/pad metadata into NHWC-consistent form.

Operator Coverage Extensions

  • Inverse:
    • Expanded built-in lowering from only 2x2/3x3 to generic square NxN (up to 16x16) by Gauss-Jordan-style tensor-op decomposition.
  • GatherND:
    • Added lowering support for batch_dims > 0 by flattening batch prefix, synthesizing batch indices (RANGE + TILE + CONCAT), gathering, then reshaping back.
  • Gemm/MatMul -> FULLY_CONNECTED path:
    • Added explicit I/O cast handling so FP16 interfaces can safely execute with internal FP32 compute.
  • Bidirectional LSTM path:
    • Added FP16<->FP32 cast bridges to keep built-in kernel use while preserving external dtype contracts.
  • Slice shape inference:
    • Improved static metadata resolution for strided slicing with known bounds; avoids incorrect oversized static lengths when runtime clipping may occur.
  • GridSample:
    • Improved handling of partially-unknown shape metadata by merging with ONNX raw shape hints and inferring expected output dimensions when valid.

Validation/Registry Alignment

  • Updated validators (Inverse, GatherND, GridSample) to match new lowering capabilities and constraints.
  • Expanded op registry declarations for GatherND and shape/index helper ops (RANGE, SHAPE) required by new lowerings.

Tests

  • Expanded tests/test_tflite_builder_direct.py with targeted regression tests covering:
    • FP16->FP32 IR promotion and cast/transpose cleanup
    • Inverse static 8x8 lowering
    • GatherND(batch_dims>0) and GridSample unknown-shape scenarios
    • FP16 FC cast bridges
    • unsupported-SPLIT fallback rewriting
    • reshape runtime -1 preference behavior
    • transpose/flatten/globalnorm/pad chain optimization
  • Local test result:
    • pytest -q tests/test_tflite_builder_direct.py
    • 568 passed, 1 warning

Release Metadata

  • Bumped package version to 2.2.2 (pyproject.toml, onnx2tf/__init__.py).
  • Updated Docker tag examples in README to 2.2.2.

What's Changed

  • Improve flatbuffer_direct coverage and DAMO conversion robustness by @PINTO0309 in #899

Full Changelog: 2.2.1...2.2.2