Skip to content

Releases: jacotay7/makewfs

makewfs 2.1.1

Choose a tag to compare

@jacotay7 jacotay7 released this 07 Oct 19:15
2d0ca25

Fixed

  • CPU and GPU pyramids used different propagation grids for some
    geometries
    (#15). The FFT
    size came from scipy.fft.next_fast_len on the CPU and
    cupyx.scipy.fft.next_fast_len on the GPU, and SciPy also admits a factor
    of 11. Whenever it picked one (requests such as 33, 66 or 99 samples), the
    CPU padded to 33/66/99 while the GPU padded to 35/70/100, and the photon
    rates differed by up to 46% (24-pixel pupil, 9-pixel separation, 12-point
    modulation at 3 lambda/D). Both backends now use one rule: the smallest
    length with no prime factor above 7, which equals CuPy's choice, so GPU
    results are unchanged
    and CPU and GPU now agree to rounding.
    • Behaviour change on the CPU for every pyramid whose old SciPy size had a
      factor of 11: the grid grows by at most 9.1% and the rate changes. At the
      default fft_oversampling = 2 this affects detector widths
      (pixels_across_pupil + pupil_separation_pixels + 2 * detector_margin_pixels)
      of 11, 22, 33, 38, 43, 44, 55, 65, 66, 76, 77, 82, 88, 99, 109, 110,
      113-115, 121, 129-132, 136, 137, 148, 151-154, 163-165, 176, 181, 197, 198,
      217-220, 226-231, 241, 242, 246, 247, 263-269, 271-275 and 295-297 pixels
      (up to 300). Overall, 1,342 of the requests 1-4096 change size.
    • Unaffected: the shipped example and benchmark pyramids (grids of 288,
      108, 160 and 216 samples) and every Shack-Hartmann configuration, whose FFT
      sizes follow from the configured sampling.
    • Frames now record the grid as wfs_pyramid_fft_size_px.
  • The 2.1.0 GPU pyramid's CUDA graph could replay freed cuFFT memory. It
    used plans from CuPy's global plan cache, which evicts them once other
    transforms fill it (for example many sensors or user FFTs in one process),
    and replay then read freed memory (cudaErrorIllegalAddress). The pyramid
    now owns its plans, built exactly as CuPy builds them, so results are
    unchanged.

makewfs 2.1.0

Choose a tag to compare

@jacotay7 jacotay7 released this 07 Oct 18:52
69927fc

Performance

  • Pruned FFTs on CPU and GPU. Shack-Hartmann spots and the pyramid
    transform only the FFT lines that hold data and keep only the cropped
    output lines; the pyramid applies its mask on the unshifted grid instead of
    shifting four full grids per state. CPU results are unchanged (bit-identical
    in 451 of 468 arrays of a before/after matrix, otherwise within 2.5e-7
    relative in float32).
  • CUDA-graph replay of the pyramid. On a GPU the pyramid's fixed-shape
    propagation is captured once and replayed with one launch.
  • Compiled CUDA executor for integer-FFT Shack-Hartmann grids, which
    previously ran the array path. Results agree to float rounding.
  • End-to-end frames/s on an Ampere Neoverse-N1 (12 cores) and an RTX 4060,
    2.0.0 -> now: SH 20x20 float32 CPU 38.5 -> 126.1 (3.3x), GPU 496 -> 813
    (1.6x); SH 60x60 float64 CPU 10.4 -> 29.8 (2.9x), GPU 199 -> 640 (3.2x);
    nine-sample SH CPU 52.0 -> 96.3 (1.9x), GPU 189 -> 764 (4.1x); pyramid 40
    CPU 923 -> 1,101 (1.2x), GPU 397 -> 759 (1.9x); pyramid 60 mod-8 CPU 173 ->
    272 (1.6x), GPU 408 -> 762 (1.9x); pyramid 80 mod-32 float64 CPU 11.4 ->
    20.9 (1.8x), GPU 204 -> 277 (1.4x). See docs/performance.md.

Fixed

  • Compiled CUDA Shack-Hartmann executor misread float32 rotated or offset
    lenslet grids.
    Such grids resample the OPD in float32, but the kernel read
    it as float64, so the GPU spot pattern was wrong (82-96% error). The OPD is
    now widened (exactly) before the launch.

  • Compiled CUDA Shack-Hartmann executor failed to launch for large
    lenslet sampling
    (e.g. 32x32 samples, 1024 threads) when register use left
    fewer threads per block. Such geometries now fall back to the array path.

  • Benchmark artifacts named the wrong GPU on multi-GPU hosts.
    benchmarks/run.py recorded the first line of nvidia-smi, which ignores
    CUDA_VISIBLE_DEVICES. It now asks CuPy for the device the run used. On
    Arm hosts, whose /proc/cpuinfo has no model name, it reads the CPU model
    from lscpu (e.g. Neoverse-N1) instead of reporting aarch64.

Added

  • Arm benchmark data point (benchmarks/device-results-neoverse-n1.*): the
    device table on an Ampere Neoverse-N1 host (16 pinned cores) with an RTX
    4060 and an RTX A400.

makewfs 2.0.0

Choose a tag to compare

@jacotay7 jacotay7 released this 07 Oct 05:21
74b5369

Breaking: the frame metadata key wfs_input_opd_rms_m now has a different meaning. It is the pupil-weighted, piston-removed RMS defined in aocore CONVENTIONS 4.1. The old value, the unweighted whole-grid RMS with piston included, is still recorded under wfs_input_opd_rms_unweighted_m. To get the 1.x number, read that key. See "Migrating to 2.0" in the stability guide: https://jacotay7.github.io/makewfs/stability/

Breaking

  • wfs_input_opd_rms_m is now the pupil-weighted, piston-removed RMS.
    It follows aocore CONVENTIONS.md 4.1: the RMS of the input OPD in metres,
    weighted by the intensity of the pupil the optics use and with the
    intensity-weighted mean (piston) removed. Before 2.0 it was the quadratic
    mean over the whole input grid, with every pixel counted equally, pixels
    outside the pupil included and piston kept. For the same wavefront the new
    value is usually smaller; a pure piston now reports 0. The weights are the
    configured pupil's intensity (the amplitude the engines use, squared) on the
    input grid: an analytic pupil is evaluated on input.shape with the
    configured numerics.pupil_supersampling, the same map
    WavefrontSensor.pupil_illumination() returns; a custom mask, which exists
    only on the engine's pupil grid, is area-averaged onto the input grid.
    Phase input is converted to OPD first, and expose_integrated reports the
    RMS of the mean OPD, as before. The value is still reduced on the device and
    crosses to the host in the same single batched transfer as the captured
    photon rate. A pupil with no transmission on the input grid is now rejected
    when the sensor is built.
  • makewfs.provenance.metadata, an internal helper, takes the two
    reduced RMS values as required opd_rms_m and opd_rms_unweighted_m
    arguments and no longer accepts opd_m.

See "Migrating to 2.0" in the stability guide.

Added

  • wfs_input_opd_rms_unweighted_m frame metadata keeps the 1.x quantity,
    the unweighted RMS over the whole input grid with piston included, so no
    information is lost. Read it wherever the old number is still wanted.
  • Both sensor engines expose configured_pupil, the pupil amplitude on their
    own pupil grid, and makewfs.sampling.area_rebin area-averages a map onto
    another grid of the same extent, for any shape ratio.

Changed

  • The closed_loop_injection.py and showcase.py examples take the residual
    RMS they plot from each frame's wfs_input_opd_rms_m instead of a
    whole-grid np.std, so the numbers they print are pupil-weighted. The
    checked-in makewfs_showcase.webp was not regenerated.
  • Requires aocore>=0.1.3,<0.2. makewfs.sampling.block_sum drops its
    own factor-two shortcut and delegates every factor to aocore.block_sum,
    whose strided CPU adds and single CuPy kernel measured at least as fast on
    the spot stacks the Shack-Hartmann actually bins (for example
    (400, 16, 16) float32: CPU 112 vs 173 us, Quadro P620 65 vs 94 us;
    (3600, 12, 12) float64: CPU 0.90 vs 1.27 ms, GPU 98 vs 173 us). The
    additions happen in a different order, so Shack-Hartmann float32 images
    differ from 1.2.0 by float32 rounding (at most about 1e-7 relative) and
    float64 ones by about 2e-16; pyramid images are bit-for-bit unchanged.
    Pixel-centre coordinates for the input, pupil and DFT detector grids are
    now built on the selected device by aocore.centered_coordinates
    (ArrayBackend.centered_coordinates), with identical values.
  • The input-RMS tests check the unweighted key against
    aocore.rms_unweighted. The per-frame path keeps its own device reduction,
    because the aocore functions return host floats and would each add a
    synchronization.

Full changelog: https://github.com/jacotay7/makewfs/blob/main/CHANGELOG.md

makewfs 1.2.0

Choose a tag to compare

@jacotay7 jacotay7 released this 07 Oct 03:35
de2ab3d
  • Fixed: phase input was reported as metres. With
    input.quantity = "phase", OpticalResult.opd_m and the
    wfs_input_opd_rms_m frame metadata held the raw phase in radians, so a
    13 nm wavefront was recorded as 0.083 m. Both now carry the input converted
    to OPD metres at input.reference_wavelength_m. Images were never affected,
    and OPD input is unchanged.

  • Changed: generic optics helpers now come from
    aocore.
    aocore>=0.1.2,<0.2 moves
    from the dev extra to a core dependency. Under CONVENTIONS.md §9, makewfs
    now imports these primitives instead of keeping its own copies:
    centered_coordinates for the pupil, field-stop and DFT detector grids;
    ARCSEC_TO_RAD/RAD_TO_ARCSEC for source angles and the Shack-Hartmann
    plate scale; opd_to_phase/phase_to_opd for phase input and the sensor
    phasors; and block_sum for pixel integration. No public name changes.
    makewfs.sampling.block_sum keeps its signature, error messages and
    factor-two fast path, and delegates other factors to aocore. The behaviour
    already followed the conventions. Float32 renders of the example
    configurations are bit-for-bit unchanged, and float64 renders differ only by
    floating-point rounding (at most about 6e-16 relative), from the changed
    operation order in the angle and phase conversions and, for oversampling
    factors above two, in pixel binning. The pupil masks, the
    ArrayBackend and its fftshift-convention centred FFTs stay local
    because aocore has no equivalent for them; the module docstrings say why.

  • Tests: conformance with the AO stack conventions. tests/test_conformance.py
    runs aocore's checks against both
    sensors. Shack-Hartmann spots centre between pixels for a flat wavefront and
    move towards +x for a +x OPD ramp, and pyramid images are centred
    (CONVENTIONS.md 1.3 and 3.1). aocore joins the dev extra.

  • Fixed: magnitude-normalized photon rates ignored spiders, segment gaps and
    custom masks.
    The rate came from the analytic annulus area
    pi / 4 D^2 (1 - eps^2) and was then distributed over the sampled pupil's own
    flux, so obstructions inside the annulus never removed photons. Both sensors
    now scale a magnitude-normalized rate by the sampled pupil's clear fraction
    of the annulus, exposed as engine.clear_aperture_fraction: about 0.96 for
    four 2 % wedge spiders, and exactly 1 for a plain annulus, so such
    configurations are unchanged. Direct detector_photon_rate sources are never
    rescaled; the Keck HAKA example already computes its rate from the masked
    area and is unaffected.

  • Fixed: Shack-Hartmann flux creation and ghost spots for wide subaperture
    windows
    (#4). A lenslet
    field sampled at s points per lenslet has a far field that repeats every
    s lenslet lambda/d; when the detector window
    pixels_per_subaperture / spot_sampling_pixels_per_lambda_over_d exceeded
    s, the sampled DFT summed the replicas as light. A 20x20 sensor with 4
    pixels at 0.25 pixel per lambda/d and 6 pupil samples per lenslet returned
    8.6 times the configured photon rate (79 times for 16 pixels at 0.32), and
    tilts beyond +-s/2 lambda/d aliased to the wrong side. Each lenslet field is
    now propagated on the smallest integer refinement of the configured pupil
    grid that is at least as fine as the widest window, evaluated at the shortest
    configured wavelength and bounded by any field stop. The configured grid,
    custom masks, and lenslet illumination are unchanged; the OPD is interpolated
    linearly onto the refined grid. Configurations that already satisfied the
    rule are bit-for-bit unchanged, and the CPU reference path and the compiled
    CUDA executor share the fix. The engine reports pupil_samples_per_lenslet,
    field_upsampling, samples_per_lenslet, and
    detector_window_lambda_over_d; see the new "Lenslet-field sampling"
    section of the Shack-Hartmann guide.

    The Keck HAKA example changes. Its 4 samples per lenslet were below the
    4.4 lambda/d window at 673 nm and the 7.2 lambda/d window at its 411 nm
    quadrature node, which captured 1.56 times the light at that node and 1.06
    times the configured rate overall at zero OPD. It now propagates at 8 samples
    per lenslet and captures 0.93 (about 12% less signal; warm optical render
    1.29 to 1.66 ms on GPU, 671 to 751 ms on CPU). The checked-in HAKA
    real-versus-simulation and LUT artifacts were produced before this fix and
    have not been regenerated.

1.1.0

Choose a tag to compare

@jacotay7 jacotay7 released this 01 Oct 05:57
a188a82

[1.1.0] - 2026-08-24

  • Declared the license as a PEP 639 SPDX expression (license = "MIT" plus
    license-files) instead of the deprecated license = { text = "MIT" } table,
    and dropped the now-redundant License :: classifier. The built distribution
    carries License-Expression: MIT and License-File: LICENSE. No change to the
    license itself.

  • Requires getframes>=2.2.0. The detector.readout_mode = "cds" path calls
    getframes.Camera.correlated_double_sample[_spectral], which is released in
    getframes 2.2.0; the previous >=2.1.1 floor would have installed a getframes
    without it and failed at first CDS readout rather than at resolve time.

  • detector.background_photon_rate_per_s adds a uniform incident sky or
    thermal background in photons/s/pixel, passed to the getframes background
    term on every readout path including correlated double sampling and the
    spectral variants. It is light, so it collects charge and carries shot noise;
    detector dark current remains the camera preset's own. makewfs does not
    compute the rate, because converting a sky surface brightness to a rate per
    pixel needs the field stop and plate scale, which belong to the instrument.

  • The Keck example's photon budget covers arbitrary passbands, and no longer
    extrapolates its extinction curve silently.
    broadband_budget gained
    band_min_nm, band_max_nm, quadrature_order, and extinction_paths, so a
    near-infrared sensing arm can use the same budget as the visible one. Multiple
    extinction tables are concatenated in wavelength, keeping the Gemini optical
    measurement and the near-infrared continuum separately attributed, and a band
    reaching past the last tabulated point now raises instead of silently taking
    numpy.interp's flat continuation. That check immediately caught the existing
    HAKA configuration: its 400-950 nm band runs 50 nm past the optical curve, so
    it had been clamping extinction at the 900 nm value. mauna_kea_extinction_nir.csv
    supplies the measured J/H/K coefficients from Leggett et al. (2006, MNRAS
    373, 781; UFTI on UKIRT over 21 photometric nights), held flat across each
    MKO passband so integrating the table over that passband returns the
    published number. The entries between the passbands sit in the 1.4 and 1.9 um
    telluric water bands and are indicative only. KECK_ALUMINUM_MIRROR_REFLECTIVITY_NIR
    records that
    aluminium is about 0.97 per surface in the near infrared against 0.88 in the
    visible -- a 30% flux difference over three reflections.

  • Correlated double sampling as a selectable detector readout mode.
    detector.readout_mode = "cds" routes the adapter to the released
    getframes.Camera.correlated_double_sample[_spectral] instead of
    expose[_spectral], returning the signed int32 difference of the two reads
    of one global-reset ramp. This is how nondestructive-readout IR arrays such as
    the C-RED One are actually operated, and it is the natural readout for a
    pyramid sensor on one. Ownership is unchanged: makewfs still supplies only a
    photon-rate map and every noise term stays in getframes.
    detector.cds_pedestal_interval_s models a finite reset-to-pedestal-read
    delay. Note that exposure_s is the read-to-read integration, not the frame
    period: a C-RED One at its 1750 Hz maximum CDS rate integrates for 1/3500 s,
    because the other half of the frame period is the reset and pedestal read.
    CDS rejects binning > 1 and caller-owned out storage rather than silently
    ignoring them. See examples/cds_readout.py.

  • examples/showcase.py renders an animated WebP of four sensor
    configurations -- 20x20 SH, 60x60 SH, a modulated pyramid, and a
    range-elongated sodium LGS SH -- watching one wind-blown pyturb atmosphere, each panel
    overlaid with the end-to-end throughput that configuration sustained on the
    running machine. The clip is the README header image. The atmosphere uses
    engine="extrude" so a long clip never replays turbulence a periodic spectral
    screen would have wrapped.

  • The README follows the documentation-first structure used across the
    sibling projects: docs link, showcase clip, install, quickstart, benchmarks,
    then the feature list.

  • Renamed benchmarks/configs/shack_hartmann_broadband_lgs.toml to
    shack_hartmann_quadrature_9sample.toml.
    "Broadband LGS" is a contradiction:
    a sodium beacon returns the 589 nm D2 line broadened by roughly 0.003 nm,
    while the config sweeps 585--595 nm, about three orders of magnitude wider.
    Its spectral axis is a quadrature load for the polychromatic path, not
    beacon physics, and the new name says so; the config's optical content and its
    benchmark numbers are unchanged, so results remain comparable across the
    rename. The README, docs/performance.md, and a new header comment in the
    config itself all state the distinction. Benchmark artifacts recorded before
    this change (benchmarks/reference-results.json,
    benchmarks/reference-table.md) still refer to the old filename.
    examples/showcase.py now runs a monochromatic 589 nm beacon sampled at five
    altitudes through the sodium layer, which is where spot elongation actually
    comes from; examples/realistic_broadband.py remains the genuine broadband
    demonstration, on a natural guide star.

  • docs/performance.md carries the refreshed CPU/GPU table and its matching
    environment, which had drifted from benchmarks/device-results.md.

  • Refreshed benchmarks/device-results.{json,md} on the RTX 5090 / Ryzen 9
    9950X3D reference machine against makewfs 1.0.0, getframes 2.1.1, pyturb 1.0.0,
    NumPy 2.2.6, and CuPy 14.1.1.

  • Benchmark provenance now records the CuPy version when CuPy is installed
    under a CUDA-specific wheel name (cupy-cuda12x/cupy-cuda11x); the metadata
    block previously reported cupy: null on exactly the machines that had
    produced the GPU column.

  • mkdocs.yml declares site_url, so the published documentation emits
    canonical links and a sitemap.

  • Compatible GPU Shack--Hartmann optics now use a first-use-JIT compiled
    executor.
    CuPy specializes one CUDA kernel for the fixed lenslet, temporal,
    spectral, precision, field-stop, and detector-sampling geometry, then reuses
    its process and disk caches. It fuses sampled-DFT propagation through photon
    mosaics and composes detector-owned focal charge diffusion with native-pixel
    integration once at startup. The array implementation remains the exact
    reference/fallback for CPU, FFT, continuous/native optical blur, and oversized
    CUDA geometries. A cold isolated-cache, alternating 20-frame physical HAKA
    benchmark on a Quadro P620 measured 24.85 ms versus 569.67 ms p50 (22.93x),
    after a 3.176 s first-use compile, with photon/spectral/captured-rate relative
    disagreements below 1e-7.

  • WavefrontSensor.expose() and expose_integrated() now accept an optional
    caller-owned detector out array. The detector adapter pairs it with
    getframes.DetectorWorkspace, including wavelength-resolved exposures, so a
    high-rate owner can keep one stable contiguous ADU destination. Ordinary calls
    retain independent frame lifetimes; only explicit out calls alias the
    caller's storage.

  • Persistent Shack--Hartmann sensors now cache propagation geometry. Half-sample
    FFT ramps, sampled-DFT kernels, field-stop masks, and backend blur kernels are
    built once per compatible source geometry instead of once per frame. The ordinary
    large FFT path is intentionally unchanged; matched local timings improved the
    geometry-heavy field-stop/DFT paths by roughly 6--20% with optical parity tests.

  • A temporally integrated exposure now renders its samples in one pass.
    ShackHartmannEngine.render_integrated builds the fields for every
    (temporal sample, source state) pair as a single batch, and
    expose_integrated uses it when the engine provides it. The motivation is
    not transform size: on a HAKA-scale configuration the transforms are about
    four percent of a render, and the cost is dominated by the fixed dispatch
    overhead of the many small elementwise operations around them, which is paid
    per call however much data the call carries. Presenting the whole exposure at
    once amortises that: the reference HAKA exposure drops from 13.7 ms to
    11.2 ms.

    Averaging the spot intensities before the mosaic is legitimate because
    everything downstream of them -- mosaic assembly, flux scaling, and the
    captured-rate accounting -- is linear in the spots, and there is a test
    asserting the batched and sequential paths agree rather than leaving that as
    an argument. Agreement is to float round-off from the changed summation
    order, about 1e-7 relative in single precision, not bit-for-bit.

  • Detector charge diffusion now reaches the Shack--Hartmann spots. The
    measured OCAM2K value was previously carried as
    shack_hartmann.optical_blur_fwhm_pixels and applied after pixel
    integration, where a 0.37-pixel FWHM Gaussian is a numerical no-op, so a
    measured detector property changed nothing. Charge diffusion is detector
    physics, so getframes now owns both the value
    (CameraConfig.charge_diffusion_fwhm_px, declared by the andor_ocam2k
    preset) and the kernel model; the Shack--Hartmann engine asks getframes for
    the operator at its own focal-plane oversampling and applies it to the
    oversampled irradiance ahead of the pixel-area integration that collects the
    diffused charge. Configuration that cannot represent the width now fails with
    the required fft_oversampling instead of applying nothing.
    This changes delivered spot profiles and slope gains for any detector
    declaring a nonzero width, so recorded HAKA evidence must be regenerated.
    makewfs.charge_diffusion_fwhm_px and makewfs.resolve_camera_config are new
    public helpers for consumers needing a sensor property before a frame exists.

  • Added `Wavefront...

Read more

1.0.0

Choose a tag to compare

@jacotay7 jacotay7 released this 26 Jul 20:08

[1.0.0] - 2026-07-26

  • Prepared the first stable public release with versioned package metadata,
    PyPI/CI badges, citation and release documentation, and a trusted-publishing
    GitHub Actions workflow using the pypi environment.

  • Fixed the benchmark runner on Python 3.10 by using the portable
    datetime.timezone.utc API.

  • Raised the detector dependency to released getframes>=2.1.1, made
    wavelength-resolved detector QE and full spectral truth part of the supported
    contract, and removed the pre-release integrated-signal compatibility path.

  • Added an R-band HAKA camera-LUT analysis using representative A0 V through M3
    V continua and generated open-loop Maunakea states. It reports mean active
    4x4-lenslet intensity SNR with OCAM2K photon, EM-excess, dark, CIC, read, and
    quantization noise, fits a smooth ceiling-aware broken-power-law cadence floor
    only to the R>=10 fine-adjustment tail, asymptotes to the true 2067 Hz OCAM2K
    limit, and emits a smooth saturation-constrained policy that never slows below
    that empirical model merely to recover per-frame SNR.

  • Generalized the HAKA broadband photon-budget helper from fixed Johnson V
    normalization to an explicit Johnson normalization band, retaining V as the
    showcase and eng519 default.

  • Fixed temporal integration to average and forward wavelength-resolved photon
    cubes to the detector, preserving configured spectral QE instead of silently
    falling back to scalar QE.

  • Fixed even-sized Shack-Hartmann focal-plane registration: zero slope now lies
    at the intersection of the central four detector pixels, using half-integer
    Fourier samples rather than an asymmetric integer-grid crop.

  • Added a Keck II HAKA open-loop worked example with a generated 36-segment
    Keck pupil including a live-data-fitted circle-plus-hexagon secondary shadow
    and six 26 mm support arms, exact
    57x57-by-4x4 (228x228) Shack-Hartmann/OCAM2K geometry,
    magnitude-dependent EM gain and frame rate, temporally integrated pyturb
    Maunakea OPD, exposure-matched master-dark subtraction, GIF/MP4 output, and a
    reproducibility manifest with per-frame photon/electron/count flux auditing.
    The supplied real eng519 V=10.16 RTC cube constrains the roughly 54-lenslet pupil
    diameter, compact quadcell sampling, and the eight-output 4x2 OCAM geometry,
    outside-pupil dark/bias levels, and relative conversion gains. With no matched
    dark cube, the RTC comparison subtracts a per-output/repeated-4x4 template and
    per-frame output drift inferred outside the pupil, then reports real and
    simulated lenslet signal/morphology without global rescaling. The
    eng519 comparison now simulates V=10.16 at 750 fps, retains every tenth
    generated phase-screen exposure like the telemetry, and writes a side-by-side
    GIF. The magnitude showcase advances frozen flow by a visible minimum cadence
    at every magnitude without changing the physical detector exposure. HAKA NGS
    photon formation now integrates a V-normalized 6600 K spectrum over the full
    400--950 nm band, applies measured Mauna Kea extinction at the observed
    airmass, applies 0.88 reflectivity to the aluminum primary, secondary, and
    tertiary, uses the sampled clear-pupil collecting area, and passes the
    resolved spectral cube through OCAM2K's wavelength-dependent QE. The reference
    renders now use the Keck-characterized approximately 28 output e-/ADU OCAM2K
    conversion. Team-confirmed independent bench measurements now establish the
    downstream HAKA throughput as 28.7%; it is applied as physical radiometry after
    the telescope mirrors rather than inferred or fitted by the RTC comparison.
    The regenerated eng519 comparison has a real/simulation signal ratio of
    1.00029.

  • Added a reproducible warm HAKA CPU/GPU benchmark. It times non-periodic Mauna
    Kea atmosphere evolution, the full eight-wavelength 57x57 Shack--Hartmann
    propagation, and noisy OCAM2K exposure while excluding static setup and
    synchronizing CUDA batches. Its local GIF shows CPU/GPU detector streams over
    equal wall-clock playback with measured FPS, real-time factor, frame counter,
    and atmosphere time overlays.

  • Removed generated PNG/GIF artifacts from version control and ignore them
    globally; example scripts continue to create them locally on demand.

  • Added public end-to-end GPU execution through numerics.device = "gpu".
    CuPy OPD, SH/PWFS optics, wavelength-resolved photon maps, the getframes
    detector chain, truth, and ADU remain device-resident. The runtime reports an
    actionable error when the installed getframes lacks its GPU camera contract.

  • Added direct pyturb GPU OPD → Shack–Hartmann → GPU ADU integration coverage,
    updated SH/PWFS CUDA parity tests, synchronized GPU benchmark mode, and measured
    detector-only timing.

  • Added a paired CPU/GPU bulk-throughput artifact and rendered comparison for
    representative SH, broadband LGS, and modulated pyramid workflows, with
    README and performance-guide results plus exact reproduction commands.

  • Optimized persistent SH/PWFS execution by caching source/range geometry,
    modulation phasors, resampling grids, flux normalization, and monochromatic
    spectral views; using native orthonormal FFT scaling and an intensity-only SH
    transform; removing redundant validations/resampling; and batching GPU
    metadata scalar transfers. On the RTX 5090 reference matrix this improves CPU
    throughput by 1.19x–1.73x and GPU throughput by 1.42x–2.58x over the initial
    end-to-end implementation while retaining the physics/parity gates.

  • Expanded optical verification with a direct-DFT pyramid reference,
    multi-amplitude HCIPy SH response curves, HCIPy low-order pyramid response
    maps, supplementary local OOPAO comparisons, and quantitative SH/pyramid
    metrics in the deterministic validation report.

  • Fixed the pyramid propagation grid to honor numerics.fft_oversampling, so
    the diffraction halo no longer wraps onto the pupil rims; cropped flux is
    reported as captured rate and independent HCIPy parity improved.

  • Reconfigured the shipped example TOMLs to be representative demonstrations:
    pyramid pupils are now separated (pupil_separation_pixels larger than
    pixels_across_pupil) and source photon rates correspond to a bright guide
    star so detector frames show spots above read noise.

  • Added deterministic broadband/finite-source quadrature, measured SED and
    transmission curves, physical SH sampling, field stops, optical blur,
    detector margins, and sodium-range SH elongation examples.

  • Added strict configuration-reference documentation for every v1 table and
    key, plus validation and benchmark smoke reports in CI.

  • Added headless worked-example CI smoke tests, deterministic plotting backend
    selection, and a 90% enforced branch-coverage gate.

  • Added configuration-relative three-column angular source kernels for measured
    or resolved guide-star morphologies, with normalized state provenance.

  • Added rotated analytic segment-gap pupils and a cached physical-coordinate
    lenslet-grid rotation/offset path with aligned-grid parity tests.

  • Added an optional HCIPy ideal-pyramid cross-check and a dedicated validation
    CI job; HCIPy remains outside runtime dependencies.

  • Added configuration-relative measured SH optical blur kernels with unit-sum
    validation, cached convolution, and provenance hashes.

  • Added a public API/configuration stability audit and same-run benchmark
    regression envelopes for representative CPU kernels.

  • Added versioned benchmark snapshots and isolated non-editable-wheel
    interoperability verification for pyturb 1.0 and getframes.

  • Added a versioned labelled SVG capability gallery with units, color bars,
    seeds, configuration digests, and modeling notes.

  • Added wavelength-resolved detector QE through the public getframes spectral
    cube contract, with truth preservation and a shipped comparison example.

  • Formalized the private optical ArrayBackend boundary and added static
    leakage/parity checks so a future device backend does not require sensor
    mathematics to be rewritten.

  • Added the original private CUDA 12 CuPy optical path with SH/pyramid parity
    tests; it is retained as a compatibility hook underneath the public
    configuration-driven GPU path.

  • Added the monochromatic CPU four-face pyramid engine, modulation support, a
    complete pyramid example configuration, and symmetry/flux/detector tests.

  • Added the implementation roadmap and agent guide.

  • Added the initial configuration and numerical implementation foundation.