Repository navigation
v0.3.0
xinvert v0.3.0
Complete CUDA GPU back-end, a rename pass on the module and public API, and a
thread-safe progress reporter. This release also folds in everything that had
accumulated since v0.2.0 — v0.2.1 and v0.2.2 were developed but never
published, so 0.3.0 is the next version on PyPI after 0.2.0. (Version numbers
have drifted from PyPI before: v0.1.4–v0.1.6 were released here on GitHub but
never uploaded to PyPI.)
Highlights
GPU back-end for every equation type
architect='gpu' now works for all seven solvers, not just the standard 2-D
Poisson case:
| equation type | CPU kernel | GPU kernel |
|---|---|---|
| standard 1D | invert_standard_1D |
invert_standard_1D_gpu |
| standard 2D | invert_standard_2D |
invert_standard_2D_gpu |
| standard 2D full (flux form) | invert_standard_2D_full |
invert_standard_2D_full_gpu |
| standard 3D | invert_standard_3D |
invert_standard_3D_gpu |
| general 2D | invert_general_2D |
invert_general_2D_gpu |
| general 3D | invert_general_3D |
invert_general_3D_gpu |
| general 2D biharmonic | invert_general_bih_2D |
invert_general_bih_2D_gpu |
Red-Black SOR kernels with a shared loop skeleton (_run_sor_2d_loop,
_run_sor_3d_loop), adaptive convergence-check intervals, and an optional
per-call thread-block shape via iParams['gpu_block2d']. GPU solves are
bounded to one at a time (gpus._GPU_MAX_CONCURRENT) so dask can pipeline
disk I/O against them, and CUDA context creation is done on the main thread.
Measured on an RTX 3090:
- single timestep, 300 x 720: 17.3x (1.73 s GPU vs 29.99 s on one core)
- 12 timesteps, 300 x 720: ~1.6x over 12-core dask CPU (21.3 s vs 33.3 s)
— at this size the solve is launch-bound, not execution-bound - large grids, fixed 1000 iterations: up to 17.1x
See the benchmark docs
for the full sweep and the launch-overhead analysis.
Module rename: numbas.py → cpus.py
The CPU kernel module is now named for what it contains. xarray's dask
backend remains the multi-core path.
API renames (breaking)
Four public names lose a _test suffix that misdescribed them as test-only —
they implement real discretizations used by shipped applications:
| old | new |
|---|---|
invert_standard_2D_test |
invert_standard_2D_full |
inv_standard2D_test |
inv_standard2D_full |
invert_GillMatsuno_test |
invert_GillMatsunoFlux |
invert_Stommel_test |
invert_StommelFlux |
Users calling the old names must update; there are no aliases. The private
__coeffs_*_test helpers became __coeffs_*Flux.
Thread-safe live progress output
printInfo output from dask worker threads and distributed clusters no longer
interleaves. Each message is emitted as a single atomic write with the right
parent header, so Jupyter and worker logs stay readable.
Tests
- Seven CPU/GPU consistency suites (
tests/test_Gpu*.pyplus
tests/test_CpuGpuConsistency.py, ~1700 lines) asserting the analytic
solution where one exists and CPU/GPU agreement otherwise; GPU cases skip
cleanly when no CUDA device is present. - Benchmark suite under
tests/, with machine-readable results in
tests/results/*.jsonand figure regeneration viatests/plot_benchmarks.py. - CI matrix: Python 3.9 – 3.12,
38 passed, 37 skipped.
Housekeeping
.gitattributesline-ending policy plus a one-off renormalization — the
repository had a CRLF/LF mix that inflated diffs by thousands of lines.deriv()with the central scheme no longer converts numpy-backed input to
dask; the rechunk is applied only to genuinely chunked data.- Documentation: GPU status table, benchmark page and the parallel-inversions
notebook brought in line with the implementation.
Upgrade notes
- Breaking: the four API renames above.
- CUDA support is optional —
import xinvertstill works on machines without a
GPU, andarchitect='cpu'remains the default. - If you pin
gpu_block2d, the default is still(16, 16).
Full diff: v0.2.0...v0.3.0