Skip to content

mps-pointops v1.0.0

Latest

Choose a tag to compare

@gamzerA gamzerA released this 02 Oct 17:43
· 8 commits to main since this release
ac3789c

mps-pointops v1.0.0

Native Metal point-cloud operators for PyTorch on Apple Silicon. This release freezes the documented, tested public API subset while keeping research backends and private sparse modules explicitly scoped.

python -m pip install mps-pointops==1.0.0

Included in this release

  • An opt-in SpatialIndex facade with full scan and a bounded two-level Morton BVH, reusable across queries. Existing dense and flat APIs retain their default kernels. Automatic kNN selection uses the measured M5 Pro Safe Math window and a sampled density guard; automatic radius selection remains scan.
  • Order-preserving BVH Ball Query, with the established first-K selection, strict radius comparison, padding, and first-order coordinate-gradient contract.
  • Experimental Chamfer extensions for L1 distance, normal-vector loss, Pointclouds input, and variable-dimensional tensors within the documented supported domains. The required squared-L2 upstream MPS CI gate remains in place.
  • Private sparse SubM, strided, and saved-key inverse research modules with first-order gradients, a bounded local OpenPCDet adapter, and independently recorded M1/M5 Pro and CUDA toy evidence.
  • Source-pinned benchmark records, sanitized Instruments summaries, the interactive explainer, and API/support/migration documentation.

Validation and measurements

  • Physical M1 sparse validation: 92 tests passed with no skips in each Safe/Fast mode at 6959fd590deecbefc98a163d46c23debaaa67c9e; raw logs, manifests, source hashes, and synchronized SubM benchmarks are archived in the repository.
  • Direct spconv 2.3.8 comparison on RTX 2080: three fixed toy fixtures matched the private CPU reference in outputs and first gradients. This is a limited oracle comparison, not a claim of general MPS/CUDA or trained-model parity.
  • M1 Chamfer direct comparisons cover 732 finite float32 cases per Safe/Fast mode. See the case-level evidence for supported arguments and limits.
  • Instruments distinguishes sampled Metal allocations, process footprint, allocator peaks, and query-window GPU Active intervals. These are not exact total physical GPU-memory peaks, named shader durations, or measured hardware occupancy.
  • The release metadata PR #80 passed all seven required CI checks before tagging; those checks include three macOS configurations, Linux, packaging, PyG 2.8 MPS, and the pinned PyTorch3D Chamfer MPS comparison.

Scope and migration

The earlier v0.9.0/v0.10.0 milestones were planning targets, not published versions. v1.0.0 follows v0.8.0 directly. Unfinished items remain open: uneven-batch multigroup FPS, broader k > 256 support, complete upstream-library compatibility, sparse transpose convolution, broader hardware coverage, and upstream acceptance. Private sparse modules do not provide a public spconv replacement. Strided/inverse coordinate rulebooks still build on CPU.

See CHANGELOG.md, migration and API policy, and the support matrix.

Citation and archive

YeYoung Lee · ORCID 0009-0001-8245-1803

Exact release commit: ac3789ceaa341f9801658409203a70214327f8bf.

Attached source ZIP: mps-pointops-v1.0.0-tag.zip (10,245,217 bytes). All 740 tracked files were compared byte-for-byte with the tagged Git tree: zero missing, extra, or mismatched files.

SHA-256: e5fc5a708391768a887b5bb6e8c4381aec24ad310d714540c9e47f9dc599fca1.

The same ZIP is publicly archived at Zenodo. A fresh public download matched the SHA-256 above, and the version DOI resolves to that record under the existing concept DOI. The public PyPI wheel and sdist also match the publishing workflow hashes; the README FPS/kNN/Ball Query example passed on MPS with CPU fallback disabled after an isolated installation of the published wheel.