Skip to content

2.5.2

Choose a tag to compare

@PINTO0309 PINTO0309 released this 09 Jul 13:20
· 1704 commits to main since this release
b48af48

2.5.2

1. Content and background

This PR improves flatbuffer_direct strict integer quantization so more mixed-dtype and detection-postprocess-adjacent graphs can be converted without producing invalid full-integer artifacts. It also keeps the int8 and full-int8 outputs available when optional int16-activation variants are not supported by LiteRT kernels.

2. Summary of corrections

  • Added strict full-integer quantization support for additional ops including GATHER_ND, SCATTER_ND, TOPK_V2, ARG_MAX, EXP, CAST, LESS, LOGICAL_NOT, WHERE, and SHAPE.
  • Extended same-qparams policy handling so mixed input/output ops only quantize activation/value tensors and leave index, axis, shape, condition, and integer outputs unquantized.
  • Added fallback qparam derivation from existing qparams, calibration ranges, or float constants when calibration cannot observe an output tensor.
  • Added explicit skip reporting for int16-activation variants when a model contains SCATTER_ND, because LiteRT does not support INT16 updates for that kernel. The int8 and full-int8 variants still validate and are emitted.
  • Added unit/integration coverage for the newly supported mixed-dtype ops and int16-activation skip reporting.
  • Bumped package/docker references from 2.5.1 to 2.5.2.

3. Before/After (If there is an operating log that can be used as a reference)

Before:

  • Pre-NMS YOLOX conversion could write int8 artifacts, then fail the overall -oiqt command when *_integer_quant_with_int16_act.tflite validation reached SCATTER_ND:
    Updates of type 'INT16' are not supported by scatter_nd.

After:

  • python -m onnx2tf -i /tmp/onnx2tf_yolox_quant/yolox_nano_ti_lite_26p1_41p8_pre_nms.onnx -o /tmp/onnx2tf_yolox_quant/pre_nms_quant_out_after_skip -oiqt -cind images calib_data_416x416_n200.npy "[[[[0.0,0.0,0.0]]]]" "[[[[1.0,1.0,1.0]]]]" completed successfully.
  • Generated and validated:
    • yolox_nano_ti_lite_26p1_41p8_pre_nms_integer_quant.tflite
    • yolox_nano_ti_lite_26p1_41p8_pre_nms_full_integer_quant.tflite
  • The int16-activation variants are now explicitly skipped with the reason:
    SCATTER_ND: LiteRT SCATTER_ND does not support INT16 updates.

Validation run locally:

  • pytest -q tests/test_strict_integer_quantization.py tests/test_tensor_buffer_builder.py tests/test_tflite_builder_direct.py::test_flatbuffer_direct_integer_quantized_smoke
  • ruff check onnx2tf/tflite_builder/quantization.py onnx2tf/tflite_builder/__init__.py tests/test_strict_integer_quantization.py tests/test_tflite_builder_direct.py
  • npx --yes pyright onnx2tf/tflite_builder/quantization.py onnx2tf/tflite_builder/__init__.py tests/test_strict_integer_quantization.py tests/test_tflite_builder_direct.py
  • python -m py_compile onnx2tf/onnx2tf.py onnx2tf/tflite_builder/__init__.py onnx2tf/tflite_builder/quantization.py tests/test_strict_integer_quantization.py

4. Issue number (only if there is a related issue)

Related issue: #929

onnx2tf \
-i yolox_nano_ti_lite_26p1_41p8_pre_nms.onnx \
-coion \
-oiqt \
-cind "images" "./calib_data_416x416_n200.npy" "[[[[0.485,0.456,0.406]]]]" "[[[[0.229,0.224,0.225]]]]"

onnx2tf \
-i yolox_nano_ti_lite_26p1_41p8_pre_nms.onnx \
-coion \
-oiqt \
-qt per-tensor \
-cind "images" "./calib_data_416x416_n200.npy" "[[[[0.485,0.456,0.406]]]]" "[[[[0.229,0.224,0.225]]]]"

What's Changed

  • Improve flatbuffer_direct strict int8 quantization by @PINTO0309 in #944

Full Changelog: 2.5.1...2.5.2