Repository navigation
Releases: gamzerA/mps-pointops
Release list
mps-pointops v1.0.0
mps-pointops v1.0.0
Native Metal point-cloud operators for PyTorch on Apple Silicon. This release freezes the documented, tested public API subset while keeping research backends and private sparse modules explicitly scoped.
python -m pip install mps-pointops==1.0.0Included in this release
- An opt-in
SpatialIndexfacade with full scan and a bounded two-level Morton BVH, reusable across queries. Existing dense and flat APIs retain their default kernels. Automatic kNN selection uses the measured M5 Pro Safe Math window and a sampled density guard; automatic radius selection remains scan. - Order-preserving BVH Ball Query, with the established first-K selection, strict radius comparison, padding, and first-order coordinate-gradient contract.
- Experimental Chamfer extensions for L1 distance, normal-vector loss,
Pointcloudsinput, and variable-dimensional tensors within the documented supported domains. The required squared-L2 upstream MPS CI gate remains in place. - Private sparse SubM, strided, and saved-key inverse research modules with first-order gradients, a bounded local OpenPCDet adapter, and independently recorded M1/M5 Pro and CUDA toy evidence.
- Source-pinned benchmark records, sanitized Instruments summaries, the interactive explainer, and API/support/migration documentation.
Validation and measurements
- Physical M1 sparse validation: 92 tests passed with no skips in each Safe/Fast mode at
6959fd590deecbefc98a163d46c23debaaa67c9e; raw logs, manifests, source hashes, and synchronized SubM benchmarks are archived in the repository. - Direct
spconv2.3.8 comparison on RTX 2080: three fixed toy fixtures matched the private CPU reference in outputs and first gradients. This is a limited oracle comparison, not a claim of general MPS/CUDA or trained-model parity. - M1 Chamfer direct comparisons cover 732 finite float32 cases per Safe/Fast mode. See the case-level evidence for supported arguments and limits.
- Instruments distinguishes sampled Metal allocations, process footprint, allocator peaks, and query-window GPU Active intervals. These are not exact total physical GPU-memory peaks, named shader durations, or measured hardware occupancy.
- The release metadata PR #80 passed all seven required CI checks before tagging; those checks include three macOS configurations, Linux, packaging, PyG 2.8 MPS, and the pinned PyTorch3D Chamfer MPS comparison.
Scope and migration
The earlier v0.9.0/v0.10.0 milestones were planning targets, not published versions. v1.0.0 follows v0.8.0 directly. Unfinished items remain open: uneven-batch multigroup FPS, broader k > 256 support, complete upstream-library compatibility, sparse transpose convolution, broader hardware coverage, and upstream acceptance. Private sparse modules do not provide a public spconv replacement. Strided/inverse coordinate rulebooks still build on CPU.
See CHANGELOG.md, migration and API policy, and the support matrix.
Citation and archive
YeYoung Lee · ORCID 0009-0001-8245-1803
- Version DOI: 10.5281/zenodo.23107348
- Concept DOI: 10.5281/zenodo.23076057
- Licensing: Apache-2.0, with the separate MIT Ball Query component documented in
LICENSES/MIT-ball-query.txt.
Exact release commit: ac3789ceaa341f9801658409203a70214327f8bf.
Attached source ZIP: mps-pointops-v1.0.0-tag.zip (10,245,217 bytes). All 740 tracked files were compared byte-for-byte with the tagged Git tree: zero missing, extra, or mismatched files.
SHA-256: e5fc5a708391768a887b5bb6e8c4381aec24ad310d714540c9e47f9dc599fca1.
The same ZIP is publicly archived at Zenodo. A fresh public download matched the SHA-256 above, and the version DOI resolves to that record under the existing concept DOI. The public PyPI wheel and sdist also match the publishing workflow hashes; the README FPS/kNN/Ball Query example passed on MPS with CPU fallback disabled after an isolated installation of the published wheel.
mps-pointops v0.8.0
mps-pointops v0.8.0
This release adds a bounded, opt-in Pointcept v1.2.1 Point Transformer V1 Seg26 pointops compatibility subset and elevates the supported squared-L2 Chamfer contract to pinned upstream PyTorch3D CPU-oracle CI. It does not claim full Pointcept, full PyTorch3D Chamfer, or universal 3D-model compatibility.
Added
mps_pointops.compat.install(pointcept=True)registers five PTv1 Seg26pointopscalls: farthest-point sampling, kNN query, grouping, query-and-group, and interpolation. Flat float32 3D coordinates use cumulative batch offsets and global int32 indices. The shim preserves missing-neighbor slots across batches, including nonfinite reference coordinates. The fixed synthetic Seg26 forward/backward probe passed on M5 Pro Safe/Fast after one documented temporary CUDA-constructor substitution in the pinned upstream checkout; the official repository was not changed.- A dedicated hosted MPS Chamfer workflow builds the official PyTorch3D 0.7.9 CPU extension at commit 88e182f989c80836f4bd744e0d9cb1852762ce01, changing only the two C++ standard build flags required by PyTorch 2.14.1. Safe and Fast Math run in separate processes with MPS fallback disabled. Run 36899829775 passed 160 cases and 1,080 output/gradient checks per mode with zero failed elements; the case-by-case JSON is committed in the CI report.
- A pinned M5 Pro and physical M1 large bidirectional Chamfer study covers 32,768 and 65,536 points with exact nearest-index and analytic-gradient controls in Safe/Fast Math. At 65,536 points, concentrated selection raised M1 full-call time by 4.19× Safe / 4.23× Fast relative to uniform selection, while M5 Pro full-call ratios were 1.00× / 1.03×. The current default remains native PyTorch scatter; a dedicated M1 reduction is a candidate for a measured same-input ablation.
Limits
- The Pointcept shim covers one pinned PTv1 Seg26 model path; other model families, unchanged upstream CUDA-only constructors, CUDA binary parity, and training convergence remain open.
- The Chamfer gate covers finite float32 squared-L2 inputs, supported
lengths, weights, point/batch reduction modes, and first derivatives. L1, normals,Pointclouds, second derivatives, near ties, and nonfinite valid coordinates are outside this direct gate. - A dedicated Metal Chamfer backward reduction is not in this release. The M1 synthetic fan-in result prioritizes a candidate ablation including grouping, buffer creation, gradients, and full-call time. No blanket MPS performance advantage is claimed.
Install and cite
python -m pip install mps-pointops==0.8.0Version DOI: 10.5281/zenodo.23092167 · All-version concept DOI: 10.5281/zenodo.23076057 · PyPI distribution
Release commit: fddac26332d705d7acdbfed631e8ba60cb56f665 · Archived source ZIP SHA-256: a5742c2671de1ed5a3afdb20c91a930ddb91e4ba1d90824e008c93767ff4912a
Author: YeYoung Lee (ORCID 0009-0001-8245-1803). The repository is Apache-2.0; its Ball Query component retains the MIT notice in LICENSES/MIT-ball-query.txt.
mps-pointops v0.7.0
mps-pointops v0.7.0
This release measures compact voxel downsampling and adds an opt-in experimental Metal CSR pooling backend for Apple Silicon MPS. It also makes the current kNN width limit explicit across the dense, flat, and PyG MPS paths. The published performance results are tied to the named devices, versions, input distributions, and benchmark scripts; no universal PyG compatibility or fused speedup is claimed.
Added and changed
voxel_downsample(..., pool_backend="fused_csr")reduces positions and mean/sum features in one Metal dispatch after the existing integer cell/inverse/CSR maps have been constructed. Backward gathers throughinverse. The default remains PyTorchindex_add_.- MPS dense, flat, and PyG kNN entry points now raise
ValueErrorwhen the effective requested width exceedsMAX_K=256; the flat path no longer silently starts a slow PyTorch fallback. CPU reference inputs retain their wider-k behavior. - Public flat API and compact voxel benchmark reports include source-pinned raw JSON and synchronized MPS measurements.
Validation and limits
- The M5 Pro compact voxel matrix covers 20k, 100k, and 500k points; uniform and ragged batches; and dense and sparse cells, in separate Safe/Fast Math processes. At 500k uniform dense, the default path's full-step CPU/MPS medians are 78.42/21.73 ms Safe and 83.79/23.09 ms Fast. Small 20k fixtures favored CPU. See the report.
- The optional fused backend matches the bounded integer-map, float-output, and first-order-gradient fixtures. One severe-cancellation case is a documented expected failure against MPS
index_add_; no general speedup or lower peak-memory claim is made. See the contract and full-call ablation. - The integrated M5 Pro suite with MPS fallback disabled passed 383 tests / 38 skips / 1 expected failure in Safe Math and 382 tests / 39 skips / 1 expected failure in Fast Math.
- The physical M1 Safe/Fast compact voxel matrix completed all 24 cases per mode with matching input hashes. At 500k uniform dense, default-path full-step CPU/MPS medians were 134.30/69.22 ms (Safe) and 133.66/65.96 ms (Fast). CPU led all 20k fixtures, while 100k crossed over with occupancy and batch shape. M1 uses PyTorch 2.12.0 and a different source snapshot from M5 Pro, so the comparison is descriptive. Physical M1 focused fused CSR runs passed 12 tests plus one expected failure in each mode; see the report and contract.
- Phase 3 coverage is bounded to the pinned PyG 2.8 and legacy shim surfaces in the compatibility matrix. Custom scatter reductions, universal PyG compatibility, and M2–M4 physical validation remain open.
Install and cite
python -m pip install mps-pointops==0.7.0Version DOI: 10.5281/zenodo.23087369 · All-version concept DOI: 10.5281/zenodo.23076057 · PyPI distribution
Release commit: 8d145ea7752f4a3d5a1efabbf086b25d68b65f03 · Archived source ZIP SHA-256: fd93cf53a25b939aa73103136d00281a31024d2ddbfd46c70e3ffe85b41213f4
Author: YeYoung Lee (ORCID 0009-0001-8245-1803). The repository is Apache-2.0; its Ball Query component retains the MIT notice in LICENSES/MIT-ball-query.txt.
mps-pointops v0.6.0
mps-pointops v0.6.0
This release extends the experimental Apple Silicon MPS point-cloud stack with feature-space kNN and bounded graph/grid integration. It does not establish complete PyG compatibility or a speedup for every workload and Apple GPU.
Added
- Native Metal feature-space kNN for float32 dimensions beyond 3, including D=64 and D=128, through the dense, flat, and PyG
pyg::knnMPS paths. The feature path has an explicitk <= 256limit; see the numerical contract. - Legacy
torch_clustercompatibility paths fornearest,grid_cluster,graclus_cluster, andrandom_walkon their documented CPU/MPS input subsets.random_walkuses PyTorch tensor operations; it is not a native Metal kernel or a registration of PyG 2.8'storch.ops.pyg.random_walk. - MPS dispatch for PyG 2.8
pyg::grid_clusterand a separate experimental compact voxel API with floor-based cells and mean/sum feature reduction. Raw grid IDs differ between the two APIs. Fused Metal voxel pooling is not included.
Validation
- Synthetic DGCNN classification and PointNet++ SSG segmentation forward/backward fixtures were compared with their pinned upstream implementations. The PointNet++ fixture met
atol=rtol=1e-4for reported outputs and gradients; six of 12 raw local-index arrays differed after an FPS tie, as explained in the parity report. Dataset accuracy and training convergence were not measured. - Physical M1 Safe/Fast full-suite runs reported 394 passed / 11 skipped and 393 passed / 12 skipped with MPS fallback disabled. The M1 report records synchronized Chamfer and feature-kNN timings and a concentrated-destination
scatter_add_slowdown. M2–M4 physical validation remains open. - Bounded PyG 2.8 synthetic graph fixtures checked
avg_poolandvoxel_grid→avg_pool_x, including topology and first-order gradients. These pooling paths use PyG/PyTorch reductions rather than project Metal kernels. See the Phase 3 results.
Install and cite
python -m pip install mps-pointops==0.6.0Version DOI: 10.5281/zenodo.23086417 · All-version concept DOI: 10.5281/zenodo.23076057 · PyPI distribution
Release commit: 55ebf58b22ef2212ad7ba6226746f9fd23620cf8 · Archived source ZIP SHA-256: c3ba548d1292a03145baf27562e913620868e6faddc3861c4492e85fe36c3ced
Author: YeYoung Lee (ORCID 0009-0001-8245-1803). The repository is Apache-2.0; its Ball Query component retains the MIT notice in LICENSES/MIT-ball-query.txt.
mps-pointops v0.5.0
mps-pointops v0.5.0
This release adds experimental PointNet++ feature propagation and squared-L2 Chamfer operators on Apple Silicon MPS. Dense SIMD Ball Query, the PyTorch3D Ball Query adapter, large single-cloud FPS, and the PyG 2.8 bridge remain available.
New experimental APIs
three_nn(unknown, known): three nearest points, Euclidean distances and int32 indices.three_interpolate(features, indices, weights): three-point feature interpolation with first-order feature gradient. Its requested weight gradient is an explicit zero tensor to match the original PointNet++ wrapper; see the contract.chamfer_distance(x, y, ...): bidirectional squared-L2 distance with supported lengths, weights, point/batch reductions and first-order gradients; see the contract.
These APIs remain experimental. High-dimensional kNN, complete PointNet++ segmentation validation, and broader Chamfer options are still on the roadmap.
Validation and measurements
- Original PointNet++ CUDA comparison: two fixed fixtures passed on MPS Safe and Fast against the executed official CUDA extension on an RTX 2080. Indices matched exactly; the largest Fast Math distance difference was
2.384185791e-7. - PyTorch3D Chamfer comparison: 160 finite-input cases and 1,080 checks per port device passed against a compiled official CPU extension; the largest MPS loss difference was
9.5367e-7. - M5 Pro full regression: 260 passed, 12 skipped in separate Safe and Fast processes with MPS fallback disabled. The Safe and Fast logs record the skips.
- Synchronized M5 Pro public-API timings, all raw samples, environment, and source SHA-256: Safe and Fast. The README shows the six measured cases. These measurements do not claim a CPU or CUDA speedup.
Install with python -m pip install mps-pointops==0.5.0.
Version DOI: 10.5281/zenodo.23080506. The archive uses the existing concept DOI lineage. The repository is Apache-2.0 with the documented MIT Ball Query component notice.
mps-pointops v0.4.0
v0.4.0
This release makes three already-tested additions available through PyPI:
- Dense Ball Query uses an order-preserving SIMD prefix scan. It retains the first-K input-order contract and improves sorted-input performance.
mps_pointops.pytorch3d.ball_queryadds PyTorch3D-style arguments, lengths,return_nn, andskip_points_outside_cube(accepted as a result-preserving hint).- Large single-cloud FPS can use a multi-threadgroup Metal path. On the measured M5 Pro,
strategy="auto"selects it from 500,000 points when at least two samples are requested. Uneven multi-cloud batches and automatic selection on other Apple GPUs remain follow-up work.
The M5 Pro paired 100,000-point Safe Math Ball Query ablation measured 21.43 to 2.91 ms on x-sorted input and 7.66 to 1.40 ms on random input. At 1,024 FPS samples, paired public-API measurements gave 193.08 to 29.37 ms for 500,000 points and 421.69 to 49.39 ms for 1,000,000 points. These are device- and workload-specific results; see the raw benchmark files in this release.
M5 Pro local tests: 201 passed and 12 expected skips in each of separate Safe and Fast Math processes. All six required GitHub CI checks passed on the release PR.
Install: python -m pip install mps-pointops==0.4.0
Source commit: 848926f715298542d5a4d579e2d4345aaac8c643.
Archived version DOI: 10.5281/zenodo.23078860. The version remains under concept DOI 10.5281/zenodo.23076057.
The repository combines Apache-2.0 material with the MIT-licensed Ball Query component; see LICENSE and LICENSES/MIT-ball-query.txt.
mps-pointops 0.3.0
First PyPI release
Install with python -m pip install mps-pointops on Apple Silicon macOS with Python 3.10+ and PyTorch 2.7+.
PyG 2.8 MPS operators
mps_pointops.pyg.register_mps() registers Metal-backed MPS implementations of pyg-lib's pyg::fps, pyg::knn, and pyg::radius operators. PyG 2.8's fps, knn, radius, knn_graph, and radius_graph can then use MPS tensors. Install a pyg-lib wheel that matches your PyTorch version first; the README gives a tested PyTorch 2.12 / PyG 2.8 / pyg-lib 0.7 example.
The MPS path supports three-dimensional coordinates: float32 for FPS and kNN, float32 or float16 for radius. Float16 radius uses float32 intermediate arithmetic and can differ from pyg-lib's half-precision boundary behavior. CPU and CUDA dispatch remain with pyg-lib.
The release also bridges PyG 2.8's batch-to-pointer conversion on MPS and adds a pinned integration test on macOS. See the changelog for details.
Archive and citation
The exact v0.3.0 source tree at commit 339b1c7b5f5dc70a234efde3e150a726583c02ba is preserved on Zenodo. Cite this version with DOI 10.5281/zenodo.23076058. The archived source ZIP has SHA-256 546d7989d04e16bd6e6811d4ebd3f6f7bdc6f8c002bf2a66467d4af646ff4aa2.
mps-pointops 0.2.0
Added
- Flat, variable-length 3D
fps,knn, andradiusAPIs with sorted batch vectors, separate reference and query point sets, global indices, and compact[query, reference]edge tensors. - Native Metal kernels for the flat MPS paths. Radius search uses SIMD prefix ranks to retain the first matching reference points in input order.
- Optional
torch_cluster1.6.3-style imports andknn_graph/radius_graphwrappers. The point-cloud entry points were exercised through PyG 2.7.0 on MPS.
Numerical compatibility
- FPS count calculation now follows the original CPU and GPU degree-conversion rules and preserves scalar versus length-one tensor ratio behavior.
- Flat float32 radius search uses the torch-cluster threshold obtained by computing
r * rin double precision and rounding to float32. Dense Ball Query keeps its PyTorch3D threshold contract.
Verification and scope
- Apple M5 Pro, PyTorch 2.7.0: 147 passed, 7 expected skips in separate Safe and Fast Math runs. The raw logs identify the tested code commit.
- GitHub CI covers macOS with Python 3.10 and 3.12, the oldest supported PyTorch 2.7.0 combination, Linux CPU tests, and wheel contents.
- This release supports 3D point coordinates. PyG 2.8.0 calls separate
torch.ops.pygoperators that the shim does not register. The float16 radius path and very small radii have the numerical limitations described in the README.
Install
python -m pip install "git+https://github.com/gamzerA/mps-pointops.git@v0.2.0"The attached wheel can also be installed directly. SHA256SUMS contains hashes for the wheel and source distribution.
Full changes: v0.1.1...v0.2.0
mps-pointops 0.1.1
Patch release. Use this instead of 0.1.0.
Fixed
- Minimum PyTorch version is 2.7. 0.1.0 declared
torch>=2.6, but the kernels are compiled withtorch.mps.compile_shader, which PyTorch 2.6 does not have (71 of 85 tests fail there). 0.1.1 requirestorch>=2.7, and a PyTorch build withoutcompile_shadernow gets a clear error on first use instead of anAttributeError. ball_querychecksradiusandKthe same way on every device. The CPU path used to accept negative, infinite and NaN radii (a negative radius behaved like its absolute value) and positive radii below the supported minimum, all of which the MPS path rejects. It also checks input shapes now.- The source distribution includes
docs/,bench/,examples/andtools/, which the README links to.
Added
- CI on GitHub Actions. On GitHub's Apple Silicon macOS runners (Apple M1, virtual) the MPS tests run on the runner GPU, in Metal safe and fast math, for Python 3.10 and 3.12 with the latest PyTorch and for the oldest supported combination, Python 3.10 with PyTorch 2.7.0: 105 tests passed in each macOS configuration. Linux runs the CPU tests, and a packaging job checks that the wheel carries the Metal kernels and license files.
CHANGELOG.md.
Install
pip install "git+https://github.com/gamzerA/mps-pointops.git@v0.1.1"Or install the wheel attached below; SHA256SUMS lists the checksums.
Full changes: v0.1.0...v0.1.1
mps-pointops 0.1.0
Correction: this release declares
torch>=2.6, but it needs PyTorch 2.7 or later (torch.mps.compile_shaderis new in 2.7). Use 0.1.1, which fixes the requirement.
First public release: point cloud ops for PyTorch on Apple Silicon GPUs (MPS), written as Metal kernels, plus drop-in stand-ins for the CUDA-only pointnet2_ops and knn_cuda packages.
Highlights
- Metal kernels for farthest point sampling, k nearest neighbors and ball query (PyTorch3D-style first-K contract, with coordinate gradients).
- Existing CUDA-only code runs unmodified:
mps_pointops.compat.install()servespointnet2_ops(furthest_point_sample,gather_operation,grouping_operation,ball_query) andknn_cuda(KNN). - Same results as on CUDA: MulSen-AD's Point-MAE 3D anomaly detector, re-fit on the Mac GPU through the stand-ins, gave the same object and 3D-label AUROC/AP as the original CUDA runs in all 45 runs (15 categories x 3 seeds). Per-sample scores differ by at most 9.85e-5 relative, with the same ranking in every run.
Performance
Apple M5 Pro, PyTorch 2.14.1, 100,000 points, batch 1, median of 5 runs. Timings move by a few ms between runs.
| op | mps-pointops (Metal) | best CPU library | plain PyTorch on MPS |
|---|---|---|---|
| FPS, 1024 samples | 33.3 ms | 174.0 ms (fpsample) | 247.5 ms |
| kNN, 1024 queries, k=128 | 6.1 ms | 17.6 ms (scipy cKDTree) | 120.9 ms |
| Ball query, 1024 queries, K=64, r=0.1 | 13.2 ms | 21.3 ms (scipy cKDTree) | 286.4 ms |
On 30 real MulSen-AD point clouds (21k to 117k points), MulSen-AD's own grouping step (FPS + kNN) took a median of 36.0 ms, against 163.4 ms for fpsample + scipy and 330.6 ms for plain PyTorch on MPS.
The Metal kernels returned the same indices as the references on every benchmark input. Full tables and the correctness checks behind them are in bench/results/.
Install
Needs an Apple Silicon Mac and PyTorch with MPS (torch>=2.6). Tested with PyTorch 2.14.1 on macOS 26.5, M5 Pro.
pip install "git+https://github.com/gamzerA/mps-pointops.git@v0.1.0"Or install the wheel attached below. SHA256SUMS lists the checksums of both files.
Known limitations
- Tested on one machine (M5 Pro). Results from other Apple GPUs are welcome.
- The kernels assume 32-wide simdgroups; the first call checks this and raises an error otherwise.
- On MPS: float32 input for FPS and kNN, float32 or float16 for ball query, and
k <= 256for kNN. - Ball query on spatially sorted input is slower than scipy's KD-tree (32.5 ms against 23.2 ms at 100k points), because the first-K-in-input-order contract rules out reordering the scan.
- The ball query API is not full PyTorch3D compatibility yet: no
return_nn, 3D coordinates only, and nolengthsarguments in the public API. - The CUDA stand-ins can resolve near ties differently from the CUDA packages. As in
pointnet2_ops,pointnet2_utils.ball_queryreturns index 0 in every slot for a query with no neighbor. - No
torch_cluster/ PyG style API (flat inputs with batch vectors) yet.
License
Apache-2.0. The ball query kernel and its Python wrapper are MIT; see LICENSES/MIT-ball-query.txt.
Developed with the help of AI coding assistants.