Repository navigation
mps-pointops v0.7.0
mps-pointops v0.7.0
This release measures compact voxel downsampling and adds an opt-in experimental Metal CSR pooling backend for Apple Silicon MPS. It also makes the current kNN width limit explicit across the dense, flat, and PyG MPS paths. The published performance results are tied to the named devices, versions, input distributions, and benchmark scripts; no universal PyG compatibility or fused speedup is claimed.
Added and changed
voxel_downsample(..., pool_backend="fused_csr")reduces positions and mean/sum features in one Metal dispatch after the existing integer cell/inverse/CSR maps have been constructed. Backward gathers throughinverse. The default remains PyTorchindex_add_.- MPS dense, flat, and PyG kNN entry points now raise
ValueErrorwhen the effective requested width exceedsMAX_K=256; the flat path no longer silently starts a slow PyTorch fallback. CPU reference inputs retain their wider-k behavior. - Public flat API and compact voxel benchmark reports include source-pinned raw JSON and synchronized MPS measurements.
Validation and limits
- The M5 Pro compact voxel matrix covers 20k, 100k, and 500k points; uniform and ragged batches; and dense and sparse cells, in separate Safe/Fast Math processes. At 500k uniform dense, the default path's full-step CPU/MPS medians are 78.42/21.73 ms Safe and 83.79/23.09 ms Fast. Small 20k fixtures favored CPU. See the report.
- The optional fused backend matches the bounded integer-map, float-output, and first-order-gradient fixtures. One severe-cancellation case is a documented expected failure against MPS
index_add_; no general speedup or lower peak-memory claim is made. See the contract and full-call ablation. - The integrated M5 Pro suite with MPS fallback disabled passed 383 tests / 38 skips / 1 expected failure in Safe Math and 382 tests / 39 skips / 1 expected failure in Fast Math.
- The physical M1 Safe/Fast compact voxel matrix completed all 24 cases per mode with matching input hashes. At 500k uniform dense, default-path full-step CPU/MPS medians were 134.30/69.22 ms (Safe) and 133.66/65.96 ms (Fast). CPU led all 20k fixtures, while 100k crossed over with occupancy and batch shape. M1 uses PyTorch 2.12.0 and a different source snapshot from M5 Pro, so the comparison is descriptive. Physical M1 focused fused CSR runs passed 12 tests plus one expected failure in each mode; see the report and contract.
- Phase 3 coverage is bounded to the pinned PyG 2.8 and legacy shim surfaces in the compatibility matrix. Custom scatter reductions, universal PyG compatibility, and M2–M4 physical validation remain open.
Install and cite
python -m pip install mps-pointops==0.7.0Version DOI: 10.5281/zenodo.23087369 · All-version concept DOI: 10.5281/zenodo.23076057 · PyPI distribution
Release commit: 8d145ea7752f4a3d5a1efabbf086b25d68b65f03 · Archived source ZIP SHA-256: fd93cf53a25b939aa73103136d00281a31024d2ddbfd46c70e3ffe85b41213f4
Author: YeYoung Lee (ORCID 0009-0001-8245-1803). The repository is Apache-2.0; its Ball Query component retains the MIT notice in LICENSES/MIT-ball-query.txt.