2.5.2
2.5.2
1. Content and background
This PR improves flatbuffer_direct strict integer quantization so more mixed-dtype and detection-postprocess-adjacent graphs can be converted without producing invalid full-integer artifacts. It also keeps the int8 and full-int8 outputs available when optional int16-activation variants are not supported by LiteRT kernels.
2. Summary of corrections
- Added strict full-integer quantization support for additional ops including
GATHER_ND,SCATTER_ND,TOPK_V2,ARG_MAX,EXP,CAST,LESS,LOGICAL_NOT,WHERE, andSHAPE. - Extended same-qparams policy handling so mixed input/output ops only quantize activation/value tensors and leave index, axis, shape, condition, and integer outputs unquantized.
- Added fallback qparam derivation from existing qparams, calibration ranges, or float constants when calibration cannot observe an output tensor.
- Added explicit skip reporting for int16-activation variants when a model contains
SCATTER_ND, because LiteRT does not support INT16 updates for that kernel. The int8 and full-int8 variants still validate and are emitted. - Added unit/integration coverage for the newly supported mixed-dtype ops and int16-activation skip reporting.
- Bumped package/docker references from 2.5.1 to 2.5.2.
3. Before/After (If there is an operating log that can be used as a reference)
Before:
- Pre-NMS YOLOX conversion could write int8 artifacts, then fail the overall
-oiqtcommand when*_integer_quant_with_int16_act.tflitevalidation reachedSCATTER_ND:
Updates of type 'INT16' are not supported by scatter_nd.
After:
python -m onnx2tf -i /tmp/onnx2tf_yolox_quant/yolox_nano_ti_lite_26p1_41p8_pre_nms.onnx -o /tmp/onnx2tf_yolox_quant/pre_nms_quant_out_after_skip -oiqt -cind images calib_data_416x416_n200.npy "[[[[0.0,0.0,0.0]]]]" "[[[[1.0,1.0,1.0]]]]"completed successfully.- Generated and validated:
yolox_nano_ti_lite_26p1_41p8_pre_nms_integer_quant.tfliteyolox_nano_ti_lite_26p1_41p8_pre_nms_full_integer_quant.tflite
- The int16-activation variants are now explicitly skipped with the reason:
SCATTER_ND: LiteRT SCATTER_ND does not support INT16 updates.
Validation run locally:
pytest -q tests/test_strict_integer_quantization.py tests/test_tensor_buffer_builder.py tests/test_tflite_builder_direct.py::test_flatbuffer_direct_integer_quantized_smokeruff check onnx2tf/tflite_builder/quantization.py onnx2tf/tflite_builder/__init__.py tests/test_strict_integer_quantization.py tests/test_tflite_builder_direct.pynpx --yes pyright onnx2tf/tflite_builder/quantization.py onnx2tf/tflite_builder/__init__.py tests/test_strict_integer_quantization.py tests/test_tflite_builder_direct.pypython -m py_compile onnx2tf/onnx2tf.py onnx2tf/tflite_builder/__init__.py onnx2tf/tflite_builder/quantization.py tests/test_strict_integer_quantization.py
4. Issue number (only if there is a related issue)
Related issue: #929
onnx2tf \
-i yolox_nano_ti_lite_26p1_41p8_pre_nms.onnx \
-coion \
-oiqt \
-cind "images" "./calib_data_416x416_n200.npy" "[[[[0.485,0.456,0.406]]]]" "[[[[0.229,0.224,0.225]]]]"
onnx2tf \
-i yolox_nano_ti_lite_26p1_41p8_pre_nms.onnx \
-coion \
-oiqt \
-qt per-tensor \
-cind "images" "./calib_data_416x416_n200.npy" "[[[[0.485,0.456,0.406]]]]" "[[[[0.229,0.224,0.225]]]]"- yolox_nano_ti_lite_26p1_41p8_pre_nms.onnx.zip
- yolox_nano_ti_lite_26p1_41p8_pre_nms_integer_quant.tflite.zip
What's Changed
- Improve flatbuffer_direct strict int8 quantization by @PINTO0309 in #944
Full Changelog: 2.5.1...2.5.2