Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 4 additions & 3 deletions dashboard.html
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@
<body>
<h1>📋 PyAutoMind Dashboard</h1>
<p>Every task the Mind is holding. Tap a task's 📋 and its <code>/start_dev</code> command is on your clipboard — paste it into a Claude Code chat to route Claude straight to that task.</p>
<p class="muted">In flight 4 · Parked 3 · Planned 6 · Backlog 145 · <a href="https://github.com/PyAutoLabs/PyAutoMind/blob/main/dashboard.md">markdown version</a></p>
<p class="muted">In flight 4 · Parked 3 · Planned 6 · Backlog 146 · <a href="https://github.com/PyAutoLabs/PyAutoMind/blob/main/dashboard.md">markdown version</a></p>
<h2>Start here</h2>
<h3>Highest priority <span class="facets">(filed as high) — showing 12 of 33</span></h3>
<div class="task"><button class="copy" data-cmd="/start_dev draft/triage/jax_zero_contour.md" aria-label="Copy the Claude command">📋</button><p><a href="https://github.com/PyAutoLabs/PyAutoMind/blob/main/draft/triage/jax_zero_contour.md">TRIAGE: needs manual review before routing</a> — <span class="facets">medium · safe · high</span></p></div>
Expand Down Expand Up @@ -84,9 +84,9 @@ <h2>Planned <a class="mdsrc" href="https://github.com/PyAutoLabs/PyAutoMind/blob
<div class="task"><button class="copy" data-cmd="/route start the planned PyAutoMind task latent-nan-guard-honest-run — its record is in planned.md" aria-label="Copy the Claude command">📋</button><p><b>latent-nan-guard-honest-run</b></p></div>
</details>
<h2>Backlog <a class="mdsrc" href="https://github.com/PyAutoLabs/PyAutoMind/tree/main/draft">markdown version</a></h2>
<p class="muted">145 filed prompts, not started — sorted most-pickable first (priority, then size).</p>
<p class="muted">146 filed prompts, not started — sorted most-pickable first (priority, then size).</p>
<details>
<summary>bug — 36</summary>
<summary>bug — 37</summary>
<div class="task"><button class="copy" data-cmd="/start_dev draft/bug/autoarray/numba_first_call_garbage_psf_weighted_data.md" aria-label="Copy the Claude command">📋</button><p><a href="https://github.com/PyAutoLabs/PyAutoMind/blob/main/draft/bug/autoarray/numba_first_call_garbage_psf_weighted_data.md">Numba sparse-operator likelihood: first-call garbage / intermittent worker corruption</a> — <span class="facets">autoarray · medium · supervised · high</span></p></div>
<div class="task"><button class="copy" data-cmd="/start_dev draft/bug/autofit/ep_scale_collapse_basin_cure_or_caveat.md" aria-label="Copy the Claude command">📋</button><p><a href="https://github.com/PyAutoLabs/PyAutoMind/blob/main/draft/bug/autofit/ep_scale_collapse_basin_cure_or_caveat.md">EP hierarchical parent-scale collapse: cure the basin, or document the</a> — <span class="facets">autofit · too-large · human-required · high</span></p></div>
<div class="task"><button class="copy" data-cmd="/start_dev draft/bug/autolens/pixelization_eager_vs_jit_divergence.md" aria-label="Copy the Claude command">📋</button><p><a href="https://github.com/PyAutoLabs/PyAutoMind/blob/main/draft/bug/autolens/pixelization_eager_vs_jit_divergence.md">Investigate eager <code>FitImaging.figure_of_merit</code> vs JIT/step-by-step divergence in rectangular pixelization</a> — <span class="facets">autolens · too-large · supervised · high</span></p></div>
Expand All @@ -100,6 +100,7 @@ <h2>Backlog <a class="mdsrc" href="https://github.com/PyAutoLabs/PyAutoMind/tree
<div class="task"><button class="copy" data-cmd="/start_dev draft/bug/autofit/jax_011_message_log_partition_tuple_shape.md" aria-label="Copy the Claude command">📋</button><p><a href="https://github.com/PyAutoLabs/PyAutoMind/blob/main/draft/bug/autofit/jax_011_message_log_partition_tuple_shape.md">jax 0.11 breaks beta/gamma message log_partition under jit ('tuple' object</a> — <span class="facets">autofit · small · supervised · medium</span></p></div>
<div class="task"><button class="copy" data-cmd="/start_dev draft/bug/pyautoheart/script_timing_baselines_orphaned_and_window_filled.md" aria-label="Copy the Claude command">📋</button><p><a href="https://github.com/PyAutoLabs/PyAutoMind/blob/main/draft/bug/pyautoheart/script_timing_baselines_orphaned_and_window_filled.md">Heart script_timing baselines are orphaned by path moves and filled</a> — <span class="facets">pyautoheart · small · supervised · medium</span></p></div>
<div class="task"><button class="copy" data-cmd="/start_dev draft/bug/autolens_workspace_test/jax_grad_local_assertions_fail_but_pass_in_ci.md" aria-label="Copy the Claude command">📋</button><p><a href="https://github.com/PyAutoLabs/PyAutoMind/blob/main/draft/bug/autolens_workspace_test/jax_grad_local_assertions_fail_but_pass_in_ci.md">jax_grad scripts fail assertions locally that PASS in CI</a> — <span class="facets">autolens_workspace_test · medium · supervised · medium</span></p></div>
<div class="task"><button class="copy" data-cmd="/start_dev draft/bug/autoarray/numba_kernel_shift_axes_swapped.md" aria-label="Copy the Claude command">📋</button><p><a href="https://github.com/PyAutoLabs/PyAutoMind/blob/main/draft/bug/autoarray/numba_kernel_shift_axes_swapped.md">Numba PSF gathers derive the y/x kernel shifts from the</a> — <span class="facets">autoarray · low · supervised · medium</span></p></div>
<div class="task"><button class="copy" data-cmd="/start_dev draft/bug/autoarray/pynufft_scipy_pinv2_dev_extra.md" aria-label="Copy the Claude command">📋</button><p><a href="https://github.com/PyAutoLabs/PyAutoMind/blob/main/draft/bug/autoarray/pynufft_scipy_pinv2_dev_extra.md">PyNUFFT dev extra is incompatible with current SciPy on Python</a> — <span class="facets">autoarray · small · supervised · normal</span></p></div>
<div class="task"><button class="copy" data-cmd="/start_dev draft/bug/autofit/loggaussian_prior_declares_own_support.md" aria-label="Copy the Claude command">📋</button><p><a href="https://github.com/PyAutoLabs/PyAutoMind/blob/main/draft/bug/autofit/loggaussian_prior_declares_own_support.md"><code>LogGaussianPrior</code> misreports its own support as <code>(-inf, inf)</code></a> — <span class="facets">autofit · small · supervised · normal</span></p></div>
<div class="task"><button class="copy" data-cmd="/start_dev draft/bug/autofit/plot_functions_discard_kwargs.md" aria-label="Copy the Claude command">📋</button><p><a href="https://github.com/PyAutoLabs/PyAutoMind/blob/main/draft/bug/autofit/plot_functions_discard_kwargs.md"><code>autofit.plot</code> functions accept <code>**kwargs</code> and silently discard them</a> — <span class="facets">autofit · small · supervised · normal</span></p></div>
Expand Down
14 changes: 11 additions & 3 deletions dashboard.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ Every task the Mind is holding, on one page: what is in flight, what is parked,
| [In flight](#in-flight) (`active/`) | 4 |
| [Parked](#parked) (`parked.md`) | 3 |
| [Planned](#planned) (`planned.md`) | 6 |
| [Backlog](#backlog) (`draft/`) | 145 |
| [Backlog](#backlog) (`draft/`) | 146 |

## Start here

Expand Down Expand Up @@ -273,10 +273,10 @@ Scoped but not started; some are not yet prompt files. Full detail in [`planned.

## Backlog

**145** filed prompts, not started. Each section is sorted most-pickable first (priority, then size).
**146** filed prompts, not started. Each section is sorted most-pickable first (priority, then size).

<details>
<summary><b>bug</b> — 36</summary>
<summary><b>bug</b> — 37</summary>

<details><summary>📋 <a href="draft/bug/autoarray/numba_first_call_garbage_psf_weighted_data.md">Numba sparse-operator likelihood: first-call garbage / intermittent worker corruption</a> — autoarray · medium · supervised · high</summary>

Expand Down Expand Up @@ -382,6 +382,14 @@ Scoped but not started; some are not yet prompt files. Full detail in [`planned.

</details>

<details><summary>📋 <a href="draft/bug/autoarray/numba_kernel_shift_axes_swapped.md">Numba PSF gathers derive the y/x kernel shifts from the</a> — autoarray · low · supervised · medium</summary>

```
/start_dev draft/bug/autoarray/numba_kernel_shift_axes_swapped.md
```

</details>

<details><summary>📋 <a href="draft/bug/autoarray/pynufft_scipy_pinv2_dev_extra.md">PyNUFFT dev extra is incompatible with current SciPy on Python</a> — autoarray · small · supervised · normal</summary>

```
Expand Down
86 changes: 86 additions & 0 deletions draft/bug/autoarray/numba_first_call_garbage_psf_weighted_data.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,3 +55,89 @@ as the regression probe.
(`_materialize_all`). Worth testing: `cache=False`, numba version pin, and
whether symptom 2 reproduces with symptom 1 fixed (they may share a cause).
- Keep the profiling-harness corruption counters as the acceptance test.

## Root cause — confirmed 2026-08-21

Not a numba codegen or caching bug. `psf_weighted_data_from` gathers the
weight map at `[ip0_y + k0_y + kernel_shift_y, ip0_x + k0_x + kernel_shift_x]`
with **no bounds check**. numba `@jit()` does not bounds-check array reads, so
for any unmasked pixel within `kernel_shape // 2` of the array edge the gather
reads uninitialized heap memory instead of raising `IndexError`. Negative
indices are unsafe in the same way — they wrap to the opposite edge.

Proof: compiling the shipped source unchanged under `boundscheck=True` raises
`IndexError: index is out of bounds` for a mask reaching the array edge, and is
clean for an interior-only mask. With the guard added, the numba output matches
the zero-padded numpy twin exactly across array sizes and kernel sizes.

This explains **both** symptoms, and explains why the inputs were verified
identical between call 1 and call 2 — they were; the function reads memory
*outside* its inputs:

- **Symptom 1** — a cold-cache first call runs right after numba's compilation
has churned the heap, so the memory next to the freshly allocated weight map
holds compiler garbage (~1e299). Warm cache: no compile, benign neighbour.
- **Symptom 2** — each forked worker has a different heap layout, so whether
the neighbouring memory is poisonous varies per worker and per run. Hence
2/8 corrupted in one map and 0/24 in the next.

The `np.isnan` guard was doing real work (masked border is `0/0 = NaN`) but
never protected the array bounds. The sibling `psf_precision_value_from` was
already hardened against exactly this — `psf_weighted_data_from` was missed.

The inf suspect (finite image / zero noise passing the `isnan` guard) is **not
reachable** via the caller: `.native` zeroes both data and noise outside the
mask, giving `0/0 = NaN`, never `inf`. An inf would require a zero noise value
*inside* the mask, which is a data-validation error and should stay loud. The
`isnan` guard is therefore left as-is.

Fix: @PyAutoArray PR #456 (branch `claude/autoarray-numba-psf-garbage-hfxnjv`) — bounds
guard mirroring the sibling, plus a numba-vs-numpy equivalence regression test
on an edge-touching mask (fails without the fix, passes with it). Full
`test_autoarray` suite: 1034 passed, 3 pre-existing pynufft failures unrelated
to this change.

## Reproduced on the euclid dataset — 2026-08-21

Symptom 1 reproduced exactly, on the real profiling dataset
(`autolens_profiling`, `dataset/imaging/euclid`, mask radius 3.5", PSF 21x21),
by calling `psf_weighted_data_from` directly on the masked dataset:

| | `max abs(psf_weighted_data)` | sum |
|---|---|---|
| pre-fix (`1c33850`) | **4.901e300** | 1.333e301 |
| post-fix | **298.312** | 559433.417 |

The bug report's own numbers were "max abs = 4.8e299 on fit #1, 2.98e02 on fit
#2". The post-fix value **298.31 = 2.98e02** matches the report's *correct*
value exactly, and the pre-fix value reproduces the uninitialized-memory scale.
1244 of the 3841 unmasked pixels (32%) drive the gather off the array.

Note on the mask padding — it does **not** protect this path. `apply_mask`
emits no padding warning and leaves `data.native` and `data.mask` at (71, 71);
only `derive_mask.blurring_from(allow_padding=True)` pads, to (89, 89), and
that padded blurring mask is used by the dense convolver, not by the sparse
numba path. `psf_weighted_data_from` reads the unpadded (71, 71) array via
`data.mask.derive_indexes.native_for_slim`, so the mask sits flush against the
array edge and the 21x21 kernel reads past it.

## The acceptance probe is a weak detector — use the direct check instead

`parallel_scaling/pixelization_numba.py` was run at P=2, 24 evals, 2 map
repeats, cold `NUMBA_CACHE_DIR`, both pre-fix and post-fix. **Both runs
reported `corrupt_evals_first_map = 0` and `corrupt_evals_steady_maps = [0, 0]`,
and both had a finite warm-up likelihood.** That is not evidence of no bug: the
values an out-of-bounds read returns are whatever the allocator left next to
the weight map, so in that process they happened to be benign. The pre-fix
warm-up likelihood still drifted from the post-fix one in the 6th decimal
(5860.175003698866 vs 5860.175922117387) — the same reads, landing on small
values instead of huge ones. This is exactly the reporter's own intermittency
(2/8 corrupted in one map, 0/24 in the next).

So the counters can sit at zero on a run where the bug is fully present. Prefer
the direct `max abs(psf_weighted_data)` check above as the regression probe —
it is deterministic within a process and reproduces the reported magnitudes.

Split out while fixing this: `draft/bug/autoarray/numba_kernel_shift_axes_swapped.md`
— both numba gathers derive the y/x kernel shifts from the transposed kernel
axes, harmless for square kernels but wrong for non-square odd PSFs.
67 changes: 67 additions & 0 deletions draft/bug/autoarray/numba_kernel_shift_axes_swapped.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,67 @@
# Numba PSF gathers derive the y/x kernel shifts from the wrong kernel axes

Type: bug
Target: autoarray
Repos:
- @PyAutoArray
Difficulty: low
Autonomy: supervised
Priority: medium
Status: formalised

Found 2026-08-21 while fixing
`draft/bug/autoarray/numba_first_call_garbage_psf_weighted_data.md` (the
out-of-bounds gather in `psf_weighted_data_from`). Split out under the
one-prompt-one-task rule: separate defect, separate blast radius.

## Symptom

Both numba PSF gathers in
`autoarray/inversion/inversion/imaging_numba/inversion_imaging_numba_util.py`
compute their kernel half-widths from the **transposed** kernel axes:

```python
kernel_shift_y = -(kernel_native.shape[1] // 2) # shape[1] is x
kernel_shift_x = -(kernel_native.shape[0] // 2) # shape[0] is y
```

at `psf_weighted_data_from` (line ~48) and `psf_precision_value_from`
(line ~294). The y shift must come from `shape[0]` and the x shift from
`shape[1]`.

The zero-padded numpy twin
(`imaging/inversion_imaging_util.py:psf_weighted_data_from`) gets it right and
is the reference:

```python
Ky, Kx = kernel_native.shape
ph, pw = Ky // 2, Kx // 2
```

## Reachability

Harmless for square kernels (`shape[0] == shape[1]`), which is the common
case and why no test catches it. It is **not** unreachable: kernels are
validated as *odd* in each axis, not square — `exc.KernelException("Convolver
Convolver must be odd")` in `operators/convolver.py:268` and
`structures/grids/uniform_2d.py:1153` check parity only. A non-square odd PSF
(e.g. 3x5) therefore mis-centres the gather, sampling the weight map / noise
map off-centre along both axes.

With the bounds guard now in place the mis-centred reads are clipped rather
than reading uninitialized memory, so this is a silent wrong-answer bug, not
a crash or a garbage-value bug.

## Fix

Swap the two right-hand sides in both functions. Fix them **together** — they
must agree on kernel orientation, and correcting only one would make the
`psf_weighted_data` and `psf_precision_operator` paths disagree.

## Acceptance

Extend the numba-vs-numpy equivalence test added by the OOB fix
(`test_autoarray/inversion/inversion/imaging/test_inversion_imaging_util.py::
test__psf_weighted_data_from__unmasked_pixels_on_array_edge`) to a non-square
odd kernel (e.g. 3x5). It passes today only because that test uses a square
kernel; with a non-square kernel the two implementations diverge.
Loading