Skip to content

v0.3.0

Choose a tag to compare

@miniufo miniufo released this 20 Sep 01:13
· 17 commits to master since this release

xinvert v0.3.0

Complete CUDA GPU back-end, a rename pass on the module and public API, and a
thread-safe progress reporter. This release also folds in everything that had
accumulated since v0.2.0 — v0.2.1 and v0.2.2 were developed but never
published, so 0.3.0 is the next version on PyPI after 0.2.0. (Version numbers
have drifted from PyPI before: v0.1.4–v0.1.6 were released here on GitHub but
never uploaded to PyPI.)

Highlights

GPU back-end for every equation type

architect='gpu' now works for all seven solvers, not just the standard 2-D
Poisson case:

equation type CPU kernel GPU kernel
standard 1D invert_standard_1D invert_standard_1D_gpu
standard 2D invert_standard_2D invert_standard_2D_gpu
standard 2D full (flux form) invert_standard_2D_full invert_standard_2D_full_gpu
standard 3D invert_standard_3D invert_standard_3D_gpu
general 2D invert_general_2D invert_general_2D_gpu
general 3D invert_general_3D invert_general_3D_gpu
general 2D biharmonic invert_general_bih_2D invert_general_bih_2D_gpu

Red-Black SOR kernels with a shared loop skeleton (_run_sor_2d_loop,
_run_sor_3d_loop), adaptive convergence-check intervals, and an optional
per-call thread-block shape via iParams['gpu_block2d']. GPU solves are
bounded to one at a time (gpus._GPU_MAX_CONCURRENT) so dask can pipeline
disk I/O against them, and CUDA context creation is done on the main thread.

Measured on an RTX 3090:

  • single timestep, 300 x 720: 17.3x (1.73 s GPU vs 29.99 s on one core)
  • 12 timesteps, 300 x 720: ~1.6x over 12-core dask CPU (21.3 s vs 33.3 s)
    — at this size the solve is launch-bound, not execution-bound
  • large grids, fixed 1000 iterations: up to 17.1x

See the benchmark docs
for the full sweep and the launch-overhead analysis.

Module rename: numbas.py → cpus.py

The CPU kernel module is now named for what it contains. xarray's dask
backend remains the multi-core path.

API renames (breaking)

Four public names lose a _test suffix that misdescribed them as test-only —
they implement real discretizations used by shipped applications:

old new
invert_standard_2D_test invert_standard_2D_full
inv_standard2D_test inv_standard2D_full
invert_GillMatsuno_test invert_GillMatsunoFlux
invert_Stommel_test invert_StommelFlux

Users calling the old names must update; there are no aliases. The private
__coeffs_*_test helpers became __coeffs_*Flux.

Thread-safe live progress output

printInfo output from dask worker threads and distributed clusters no longer
interleaves. Each message is emitted as a single atomic write with the right
parent header, so Jupyter and worker logs stay readable.

Tests

  • Seven CPU/GPU consistency suites (tests/test_Gpu*.py plus
    tests/test_CpuGpuConsistency.py, ~1700 lines) asserting the analytic
    solution where one exists and CPU/GPU agreement otherwise; GPU cases skip
    cleanly when no CUDA device is present.
  • Benchmark suite under tests/, with machine-readable results in
    tests/results/*.json and figure regeneration via tests/plot_benchmarks.py.
  • CI matrix: Python 3.9 – 3.12, 38 passed, 37 skipped.

Housekeeping

  • .gitattributes line-ending policy plus a one-off renormalization — the
    repository had a CRLF/LF mix that inflated diffs by thousands of lines.
  • deriv() with the central scheme no longer converts numpy-backed input to
    dask; the rechunk is applied only to genuinely chunked data.
  • Documentation: GPU status table, benchmark page and the parallel-inversions
    notebook brought in line with the implementation.

Upgrade notes

  • Breaking: the four API renames above.
  • CUDA support is optional — import xinvert still works on machines without a
    GPU, and architect='cpu' remains the default.
  • If you pin gpu_block2d, the default is still (16, 16).

Full diff: v0.2.0...v0.3.0