Skip to content

2.6.3

Choose a tag to compare

@PINTO0309 PINTO0309 released this 14 Jul 10:13
· 1041 commits to main since this release
a864015

2.6.3

Summary

This pull request is the accumulated fb-refactor4 checkpoint for the flatbuffer_direct backend. It restructures the converter around explicit internal contracts, indexed graph state, ordered and validated passes, request-driven artifact generation, and a substantially decomposed PyTorch export stack.

The public CLI and Python API remain compatible. The default backend remains flatbuffer_direct, normal direct TFLite conversion and -cotof do not import TensorFlow or tf-keras, no new third-party dependency is introduced, and TensorFlow remains available only behind explicitly requested TensorFlow-family artifacts.

The package version is updated from 2.6.2 to 2.6.3 in the package metadata, lock file, and documented container examples.

Why this refactor was needed

The previous direct backend had accumulated a very large set of interacting graph rules and repeated whole-graph scans. This made small changes expensive to reason about, allowed ordering assumptions to remain implicit, duplicated producer/consumer and layout knowledge, and made optional artifact generation perform work that had not been requested.

The PyTorch exporter had the same problem at a different layer: native code generation, source rewriting, state-dict handling, runtime wrappers, fallback selection, TorchScript/Dynamo ONNX/ExportedProgram export, naming, shape policy, and layout policy were concentrated in a monolithic module with many implicit compatibility bindings.

This branch turns those implicit relationships into bounded internal interfaces while preserving the established conversion behavior.

Main architectural changes

1. Explicit conversion pipeline contracts

The direct path is organized around this stage order:

  1. normalize public arguments into ConversionRequest and ArtifactPlan;
  2. preprocess ONNX;
  3. create a single ConversionSession, GraphIndex, and LayoutState;
  4. execute registered passes in deterministic PassPhase order;
  5. lower operators into ModelIR;
  6. validate ModelIR invariants;
  7. invoke only exporters selected by ArtifactPlan;
  8. adapt ConversionResult back to the legacy return dictionary.

Raw option dictionaries no longer need to propagate through the new internals. Artifact controls, split thresholds, reporting options, and quantization calibration controls are resolved only when the associated artifact is requested.

2. Indexed and transactional graph mutation

GraphIndex and ModelIRGraphIndex provide shared producer, consumer, duplicate-producer, operator-position, and operator-type views. Rewriters update those indexes differentially instead of rebuilding maps or repeatedly scanning the entire graph.

Passes have stable IDs, phases, priorities, maximum iteration counts, explicit change results, and graph fingerprints for deterministic cycle termination. Risky rewrites can run transactionally and roll back when shape, dtype, layout, unresolved-tensor, duplicate-producer, or public-output invariants fail.

LayoutState is carried through the session and reconciled with graph changes so layout knowledge is no longer reconstructed independently by every rule.

3. Semantic pass extraction and indexed lowering support

Large rule families were moved out of the central lowerer into focused pass modules. This includes dynamic reshape, graph cleanup, high-rank binary and MatMul handling, rank-4/rank-5 Concat families, quantized layout families, Pad, split fallback, channel shuffle, multi-branch gates, NDHWC gates, cost-volume/ScatterND patterns, PyTorch compatibility, recurrent/control-flow preparation, and layout validation.

The extracted passes retain compatibility wrappers where existing imports require them. Production call sites use registered runners, indexed root enumeration, explicit guards, shared pruning utilities, and transactional layout reconciliation. Model-name checks are not introduced; repairs are expressed as semantic graph patterns.

4. Request-driven artifact and memory behavior

The artifact pipeline now avoids entering unrequested exporters, quantization paths, split planning, and report writers. Constant buffers and graph indexes are shared where possible, precision clones have bounded lifetimes, and redundant ModelIR copying and repeated validation-index construction are reduced.

The direct backend releases legacy graph objects before ModelIR lowering when they are no longer needed. Sequential ONNX/TFLite accuracy checking materializes large prepared ONNX graphs once in a managed temporary file rather than transferring serialized model bytes through a multiprocessing pipe.

5. PyTorch exporter decomposition

The PyTorch stack is split into dedicated, mostly Torch-free policy and implementation owners, including:

  • artifact exporters and example-input/export metadata support;
  • native codegen context, stages, emitters, indexing, values, and naming;
  • capability and fallback-package selection;
  • binary, concat, reshape, reduction, constant, expression, layout-bridge, fusion, NMS, recurrent, and shape policies;
  • generated source parsing and graph/source rewrites;
  • state-dict and package-source support;
  • runtime-wrapper and StringNormalizer exporters;
  • ONNX artifact metadata/layout handling;
  • ExportedProgram child execution and archive cleanup.

Native codegen reuses ModelIRGraphIndex, and compatibility bindings required by the dynamically compiled legacy body are explicit and covered by architecture tests. Torch is imported only after a PyTorch-family artifact has actually been requested.

TorchScript, Dynamo ONNX, and ExportedProgram generation share the generated native package and common artifact support. TFLite-backed and SavedModel-backed fallbacks reuse the same package scaffolding without pulling TensorFlow into ordinary direct conversion.

6. Source canonicalization hardening

The latest checkpoints repair several inherited PyTorch exporter failures without broad layout-policy changes:

  • structured subtraction parsing and rank-3 softmax-mask handling;
  • lazy generated-model class resolution;
  • restored native-codegen helper bindings after module extraction;
  • TensorFlow-free TFLite-input PyTorch artifact paths;
  • scalar materialization for CONCAT;
  • signed and dynamic CONCAT target-shape parsing;
  • rank-4 evidence requirements for NCHW CONCAT rewriting;
  • preservation of scalar tensors when they are the direct input to torch.reshape.

The last group removes all five inherited NoneType.group() CONCAT canonicalization crashes. Three of those tests now pass completely; the two remaining Yolov7 cases reach older structural assertions rather than crashing.

7. Managed sequential regression workflow

The repository now contains a managed Tier 0-4 corpus workflow, immutable quick-run manifests/results, normalized failure classifications, timeout exclusions, and SWAP detection for the active converter process tree.

Validation is always sequential. The runner does not use a process pool or parallel inference workers. A model that causes SWAP is classified and added to the managed exclusion policy before a later run.

The 2,000 threshold refers to ONNX graph node count for Tier classification. It is not a source-file line limit.

Compatibility and dependency boundaries

  • Existing CLI and Python entry points are retained.
  • The default backend and artifact naming remain unchanged.
  • Direct TFLite conversion and -cotof do not import TensorFlow or tf-keras.
  • SavedModel, H5, Keras, and TFv1 remain optional TensorFlow-family outputs.
  • PyTorch remains optional and is imported only for requested PyTorch-family artifacts.
  • No dependency was added; all development and validation use uv.
  • Generated FlatBuffer schema code is unchanged by the architectural policy.
  • Inference validation is sequential, including isolated subprocess execution.

Validation performed

Tier 0-4 quick corpus

A fixed short-runtime selection of 54 Tier 0-4 models was run sequentially:

Result Count
Pass 46
Expected accuracy failure 4
Missing report 2
60-second quick-run timeout 2
Conversion error 0
SWAP detected 0

No regression specific to fb-refactor4 was confirmed. Focused comparison showed that silero_vad.onnx fails identically on fb-refactor3, while d3net_dnn_double_44.onnx reaches the same quick ceiling on both branches. nchw.onnx passes direct execution on both branches with byte-identical artifacts and metrics; only the instrumented bulk run consumes the timeout headroom.

Apart from the explicitly accepted DEIM result, the largest maximum absolute error among passing models was 0.0443115234375, below the required 1e-1 ceiling. DEIM remains accepted according to its recorded near-tied TopK policy.

PyTorch exporter regression investigation

The complete central PyTorch exporter suite contains 1,120 tests. At checkpoint 3127e39d, with seven previously classified long-running cases deselected, the exact result was:

Classification Count
Passed 1,019
Non-timeout failures 94
Long-running exclusions 7

All 94 non-timeout failures were reproduced against fb-refactor3 with the same exception type and normalized first message. Six belong to the explicitly excluded bread model family. Therefore, no non-timeout regression specific to fb-refactor4 was confirmed in that comparison.

The latest focused CONCAT/NMS work additionally passed 18 relevant tests, and the architecture suite passed 97 tests. The two remaining Yolov7 assertions are documented rather than hidden: one is a local cv64_in versus cv64_in_cf spelling expectation for the same channel-first tensor, and the other compares an unsanitized direct Dynamo ONNX graph with the helper path that intentionally applies the common ONNX sanitizer/optimizer.

PyTorch artifact smoke

A sequential rfdn_64x64.onnx integration gate generated and validated float32/float16 TFLite, a native PyTorch package and state dict, TorchScript, Dynamo ONNX, and ExportedProgram artifacts.

  • TFLite maximum absolute error: 0.0000371933
  • PyTorch maximum absolute error: 0.0000433922
  • SWAP detected: no

Focused and architecture tests

The branch has also run the modular PyTorch policy/exporter tests, architecture/ownership checks, bulk-runner tests, TensorFlow import blockers, public-contract characterization, pass-level fixtures, quantization/split/report tests, and targeted compatibility-binding checks. The exact commands and checkpoint-specific results are retained in:

  • docs/flatbuffer_direct_architecture.md
  • docs/flatbuffer_direct_quick_regression_2026-07-14.md
  • docs/flatbuffer_direct_pytorch_regression_2026-07-14.md
  • docs/flatbuffer_direct_handoff_2026-07-14.md

The final version update was checked with uv lock --check and an import assertion confirming onnx2tf.__version__ == "2.6.3".

Known limitations and remaining work

This PR is a substantial checkpoint, not the end of the complete refactor plan.

  • Tier 5 models (2,000 or more ONNX nodes) were intentionally not used as an early development gate.
  • The full artifact matrix and complete Tier 0-4 corpus still require final end-to-end reruns after later Goal phases.
  • Seven inherited PyTorch exporter tests remain classified as long-running.
  • The remaining inherited source/layout assertions require semantic, model-independent fixes and should not be addressed with broad source-text rewrites.
  • d3net_dnn_double_44.onnx and instrumentation-sensitive nchw.onnx are unsuitable for the 60-second quick profile.
  • Optional TensorFlow-family output compatibility still needs its final dedicated matrix run.
  • Warm-run conversion-time and peak-RSS measurements across the active tiers remain future work.

What's Changed

  • Refactor the flatbuffer_direct pipeline and PyTorch export stack by @PINTO0309 in #949

Full Changelog: 2.6.2...2.6.3