Skip to content

mps-pointops v0.7.0

Choose a tag to compare

@gamzerA gamzerA released this 01 Oct 23:14
· 158 commits to main since this release
8d145ea

mps-pointops v0.7.0

This release measures compact voxel downsampling and adds an opt-in experimental Metal CSR pooling backend for Apple Silicon MPS. It also makes the current kNN width limit explicit across the dense, flat, and PyG MPS paths. The published performance results are tied to the named devices, versions, input distributions, and benchmark scripts; no universal PyG compatibility or fused speedup is claimed.

Added and changed

  • voxel_downsample(..., pool_backend="fused_csr") reduces positions and mean/sum features in one Metal dispatch after the existing integer cell/inverse/CSR maps have been constructed. Backward gathers through inverse. The default remains PyTorch index_add_.
  • MPS dense, flat, and PyG kNN entry points now raise ValueError when the effective requested width exceeds MAX_K=256; the flat path no longer silently starts a slow PyTorch fallback. CPU reference inputs retain their wider-k behavior.
  • Public flat API and compact voxel benchmark reports include source-pinned raw JSON and synchronized MPS measurements.

Validation and limits

  • The M5 Pro compact voxel matrix covers 20k, 100k, and 500k points; uniform and ragged batches; and dense and sparse cells, in separate Safe/Fast Math processes. At 500k uniform dense, the default path's full-step CPU/MPS medians are 78.42/21.73 ms Safe and 83.79/23.09 ms Fast. Small 20k fixtures favored CPU. See the report.
  • The optional fused backend matches the bounded integer-map, float-output, and first-order-gradient fixtures. One severe-cancellation case is a documented expected failure against MPS index_add_; no general speedup or lower peak-memory claim is made. See the contract and full-call ablation.
  • The integrated M5 Pro suite with MPS fallback disabled passed 383 tests / 38 skips / 1 expected failure in Safe Math and 382 tests / 39 skips / 1 expected failure in Fast Math.
  • The physical M1 Safe/Fast compact voxel matrix completed all 24 cases per mode with matching input hashes. At 500k uniform dense, default-path full-step CPU/MPS medians were 134.30/69.22 ms (Safe) and 133.66/65.96 ms (Fast). CPU led all 20k fixtures, while 100k crossed over with occupancy and batch shape. M1 uses PyTorch 2.12.0 and a different source snapshot from M5 Pro, so the comparison is descriptive. Physical M1 focused fused CSR runs passed 12 tests plus one expected failure in each mode; see the report and contract.
  • Phase 3 coverage is bounded to the pinned PyG 2.8 and legacy shim surfaces in the compatibility matrix. Custom scatter reductions, universal PyG compatibility, and M2–M4 physical validation remain open.

Install and cite

python -m pip install mps-pointops==0.7.0

Version DOI: 10.5281/zenodo.23087369 · All-version concept DOI: 10.5281/zenodo.23076057 · PyPI distribution

Release commit: 8d145ea7752f4a3d5a1efabbf086b25d68b65f03 · Archived source ZIP SHA-256: fd93cf53a25b939aa73103136d00281a31024d2ddbfd46c70e3ffe85b41213f4

Author: YeYoung Lee (ORCID 0009-0001-8245-1803). The repository is Apache-2.0; its Ball Query component retains the MIT notice in LICENSES/MIT-ball-query.txt.