Skip to content

mps-pointops v0.4.0

Choose a tag to compare

@gamzerA gamzerA released this 01 Oct 09:10
· 209 commits to main since this release
848926f

v0.4.0

This release makes three already-tested additions available through PyPI:

  • Dense Ball Query uses an order-preserving SIMD prefix scan. It retains the first-K input-order contract and improves sorted-input performance.
  • mps_pointops.pytorch3d.ball_query adds PyTorch3D-style arguments, lengths, return_nn, and skip_points_outside_cube (accepted as a result-preserving hint).
  • Large single-cloud FPS can use a multi-threadgroup Metal path. On the measured M5 Pro, strategy="auto" selects it from 500,000 points when at least two samples are requested. Uneven multi-cloud batches and automatic selection on other Apple GPUs remain follow-up work.

The M5 Pro paired 100,000-point Safe Math Ball Query ablation measured 21.43 to 2.91 ms on x-sorted input and 7.66 to 1.40 ms on random input. At 1,024 FPS samples, paired public-API measurements gave 193.08 to 29.37 ms for 500,000 points and 421.69 to 49.39 ms for 1,000,000 points. These are device- and workload-specific results; see the raw benchmark files in this release.

M5 Pro local tests: 201 passed and 12 expected skips in each of separate Safe and Fast Math processes. All six required GitHub CI checks passed on the release PR.

Install: python -m pip install mps-pointops==0.4.0

Source commit: 848926f715298542d5a4d579e2d4345aaac8c643.

Archived version DOI: 10.5281/zenodo.23078860. The version remains under concept DOI 10.5281/zenodo.23076057.

The repository combines Apache-2.0 material with the MIT-licensed Ball Query component; see LICENSE and LICENSES/MIT-ball-query.txt.