pyTurb is live
The "atmosphere" release: pyturb goes from a phase-screen library to a complete,
benchmarked, GPU-native AO atmosphere.
Added
Atmosphere— layered atmosphere summed to pupil OPD, with per-layer
wind, airmass/zenith scaling, off-axisdirections=, field-of-view
oversizing, and integratedr0/seeing/theta0/tau0/
greenwood_frequency. Built from named profiles viafrom_profile.- Two frozen-flow engines.
engine="spectral"(default): exact sub-pixel
shift-theorem translation, all layers in one FFT, but periodic. Boiling
(tau_boil) via a spectral AR(1).engine="extrude": Assémat–Wilson row
extrusion in a wind-aligned ring buffer with rotated sub-pixel sampling —
unbounded, non-periodic, any wind direction. InfinitePhaseScreenrewritten with a ring buffer and sub-pixel
advance()(Catmull-Rom / linear), memory bounded over arbitrarily long runs.- Named profiles:
paranal-median,mauna-kea,keck,las-campanas,
cerro-pachon,armazones,hv57,single-layer,two-layer;
discretize_cn2(method=...)with moment-conserving"equivalent"
(conservestheta0andtau0),"centroid", and"optimal_grouping"—
the last chooses bin edges by dynamic programming to minimise the
Cn²-weighted within-group spread ofh^{5/3}, for MCAO/tomography layer
compression (Saxenhuber et al. 2017). Atmosphere.evolve(dt)— single-step, in-seconds frozen-flow stepper
(mirrors HCIPy'sevolve_until); repeated calls reproduceframes(dt).interp="lanczos"(6-tap Lanczos-3) sub-pixel readout on the extruder and
InfinitePhaseScreen— a flatter sub-Nyquist kernel that cuts the extruder's
finest-scale travel-phase flicker (~10% → ~3.5%) and structure-function
deficit versus the default cubic, with no change to the extrusion statistics.pyturb.analysis: Zernike basis/decomposition, Noll (1976) mode
variances, temporal PSD + power-law fit, angular decorrelation.- I/O:
pyturb.save/pyturb.loadfor.npzand FITS (optional astropy)
with metadata;Atmosphere.metadata. - Chromatic OPD:
dispersion="edlen"(dry air) anddispersion="ciddor"
with awet_fractionwater-vapour term for the mid-IR/interferometric
"wet–dry" problem;pyturb.air_refractivityand
pyturb.water_vapour_refractivity. - LGS cone effect:
Atmosphere(lgs_altitude=...)on both engines — the
extruder samples its ring buffer on a magnified grid, the spectral engine
zoom-resamples each layer's screen about the pupil centre by the same factor.
On the spectral engine the cone now composes withtau_boilboiling,
closing the previous cone/boiling mutual exclusivity. - Non-Kolmogorov spectra:
PhaseScreen(power_law=..., inner_scale=...). - Threaded CPU FFT:
pyturb.set_fft_workers(). - GPU test path: GPU tests marked
@pytest.mark.gpu, run with
pytest --run-gpu(adevicefixture parameterises statistics tests over
CPU/GPU);.github/workflows/gpu.ymlruns them on a self-hosted GPU runner. pyturb.benchmark()convenience;benchmarks/bench_suite.py
(per-use-case throughput sweep across CPU/GPU) and
benchmarks/bench_compare.pyhead-to-head vs aotools/soapy/HCIPy;
validation/validate.pygallery.- Docs:
docs/comparison.md,docs/interop.md,docs/validation.md; examples
gallery (examples/01–05). py.typedmarker; version single-sourced from package metadata.
Changed
Atmosphereoutput is OPD in metres (achromatic); passwavelength=for
phase.PhaseScreen/InfinitePhaseScreenstill return radians.discretize_cn2default method is now"equivalent"(moment-conserving).
Performance
- Spectral engine: collapse the layer axis before the transform. The
inverse FFT and subharmonic outer product are linear and shared across
layers, soAtmosphere._integratenow sums the shifted spectra to one
(n, n)array and inverse-FFTs once instead of once per layer (and sums
each subharmonic level's3x3coefficients before the shared basis product).
Identical output; measured on an RTX 5090, 9-layer paranal-median: CPU
25 → 87 fps at 512² (3.4×), GPU 865 → 1232 fps (1.4×). - Batch all subharmonic levels into one matmul. The low-frequency
subharmonic correction shares one(3, n)sinusoid basis across levels and
layers, soAtmosphere._integrate,PhaseScreen.generate,
FourierFlowScreen.translateand the boiling update now evaluate every level
in a couple of batched matmuls instead of a Python loop over levels (which was
launch-latency bound on the GPU — ~78% of a frame). Identical output.
Measured on an RTX 5090, 9-layer paranal-median frozen flow: 1,232 → 3,004
fps at 512² GPU, 3,234 fps at 256²; single-layer Monte-Carlo generation
14,000 → 31,000 screens/s at 512² (55,000 → 108,000 at 256²); CPU
spectral ~87 → ~130 fps at 512² before the accel extra below. - Fused GPU/CPU extruder readout kernel. Every layer's ring buffer is a slab
of one contiguous(L, cap, W)array, and the per-frame rotated, sub-pixel,
per-layer-wind-shifted pupil gather runs in a single pass for the"cubic"
and"lanczos"interpolators: a hand-written CUDA kernel on the GPU and a
fusedprangeNumba kernel on the CPU (see the accel extra below), bit-exact
with the previous tap-broadcast gather. Measured on an RTX 5090, 9-layer
paranal-median: 121 → 4,484 fps at 256² GPU (37×), 118 → 1,730 fps at
512² (15×), 50 → 602 fps at 1024²; the"lanczos"readout is now a
fused kernel too (~334 fps at 512² GPU, from ~120 fps). - Optional Numba CPU acceleration (
pip install pyturb[accel]). The CPU
frozen-flow hot paths — the spectral engine's fused layer sum and the
extruder's fused bicubic/Lanczos readout — run through Numba when it is
importable, with a NumPy fallback otherwise (identical results to float
round-off). Measured on a 32-core CPU, 9-layer paranal-median: spectral
~130 → 270 fps at 512², extruder 6 → 164 fps at 512² (27×) and
28 → 966 fps at 256² (34×). - Geometry-derived extruder buffer sizing. The shared ring buffer is now
sized to the largest along-wind/off-axis requirement actually present among
the layers (each layer's own wind direction and altitude), not a blanket
every-layer-at-45-degrees-and-max-altitude assumption. Measured (n=512): a
ground-layer-only atmosphere withfield_of_view=30"uses ~68% less buffer
memory; an axis-aligned atmosphere uses ~41% less even at
field_of_view=0. Never worse than before.
Full Changelog: https://github.com/jacotay7/pyturb/commits/0.2.0