Skip to content

mps-pointops v0.8.0

Choose a tag to compare

@gamzerA gamzerA released this 01 Oct 23:44
· 121 commits to main since this release
fddac26

mps-pointops v0.8.0

This release adds a bounded, opt-in Pointcept v1.2.1 Point Transformer V1 Seg26 pointops compatibility subset and elevates the supported squared-L2 Chamfer contract to pinned upstream PyTorch3D CPU-oracle CI. It does not claim full Pointcept, full PyTorch3D Chamfer, or universal 3D-model compatibility.

Added

  • mps_pointops.compat.install(pointcept=True) registers five PTv1 Seg26 pointops calls: farthest-point sampling, kNN query, grouping, query-and-group, and interpolation. Flat float32 3D coordinates use cumulative batch offsets and global int32 indices. The shim preserves missing-neighbor slots across batches, including nonfinite reference coordinates. The fixed synthetic Seg26 forward/backward probe passed on M5 Pro Safe/Fast after one documented temporary CUDA-constructor substitution in the pinned upstream checkout; the official repository was not changed.
  • A dedicated hosted MPS Chamfer workflow builds the official PyTorch3D 0.7.9 CPU extension at commit 88e182f989c80836f4bd744e0d9cb1852762ce01, changing only the two C++ standard build flags required by PyTorch 2.14.1. Safe and Fast Math run in separate processes with MPS fallback disabled. Run 36899829775 passed 160 cases and 1,080 output/gradient checks per mode with zero failed elements; the case-by-case JSON is committed in the CI report.
  • A pinned M5 Pro and physical M1 large bidirectional Chamfer study covers 32,768 and 65,536 points with exact nearest-index and analytic-gradient controls in Safe/Fast Math. At 65,536 points, concentrated selection raised M1 full-call time by 4.19× Safe / 4.23× Fast relative to uniform selection, while M5 Pro full-call ratios were 1.00× / 1.03×. The current default remains native PyTorch scatter; a dedicated M1 reduction is a candidate for a measured same-input ablation.

Limits

  • The Pointcept shim covers one pinned PTv1 Seg26 model path; other model families, unchanged upstream CUDA-only constructors, CUDA binary parity, and training convergence remain open.
  • The Chamfer gate covers finite float32 squared-L2 inputs, supported lengths, weights, point/batch reduction modes, and first derivatives. L1, normals, Pointclouds, second derivatives, near ties, and nonfinite valid coordinates are outside this direct gate.
  • A dedicated Metal Chamfer backward reduction is not in this release. The M1 synthetic fan-in result prioritizes a candidate ablation including grouping, buffer creation, gradients, and full-call time. No blanket MPS performance advantage is claimed.

Install and cite

python -m pip install mps-pointops==0.8.0

Version DOI: 10.5281/zenodo.23092167 · All-version concept DOI: 10.5281/zenodo.23076057 · PyPI distribution

Release commit: fddac26332d705d7acdbfed631e8ba60cb56f665 · Archived source ZIP SHA-256: a5742c2671de1ed5a3afdb20c91a930ddb91e4ba1d90824e008c93767ff4912a

Author: YeYoung Lee (ORCID 0009-0001-8245-1803). The repository is Apache-2.0; its Ball Query component retains the MIT notice in LICENSES/MIT-ball-query.txt.