Skip to content

2026.30.1

Choose a tag to compare

@Kastier1 Kastier1 released this 24 Jul 23:46
63c0697
Stop sorting a whole column to size error-bar caps (#264)

`test_first_payload_errorbar_large` is the most expensive benchmark in the
suite (~557 ms on CodSpeed). Two thirds of it is work neither answer needs.

`_auto_cap_size` wants the median adjacent gap between distinct positions,
and reached for `np.unique` — which sorts unconditionally, an O(N log N)
pass over the full column. Error-bar positions are usually an ordered
independent variable, so one O(N) diff both proves the column is already
non-decreasing and yields those gaps directly; only an out-of-order column
still pays for the sort. Same distinct values in the same order, so the
median is identical.

`_zero_baseline_anchor` compacted `base`/`value` down to their finite rows
before asking three all-or-nothing questions about them, allocating and
copying two full columns per axis per build. Each question is really "does
any row violate this?", which the finite mask answers in place. NaN rows
are excluded by the mask exactly as the compaction excluded them.

Measured on the benchmark's shape (1M points, yerr=1.0, decimated tier),
min-of-3:

  first payload   48.07 ms -> 14.09 ms   -70.7%
  _auto_cap_size  27.48 ms ->  2.92 ms   -89.4%  (1M sorted positions)
  _zero_baseline   4.44 ms ->  1.99 ms   -55.2%  (3M rows)

Both sit on real build paths, not just the benchmark: cap sizing runs for
every auto-cap error bar, and the zero-baseline probe runs for every bar
and histogram build.

`_auto_cap_size` returns identically over sorted, unsorted, descending,
duplicate-heavy, all-equal, single, empty, NaN-bearing and -0.0-bearing
columns; `_zero_baseline_anchor` over all-positive, all-negative, mixed,
nonzero-base, NaN-bearing, all-NaN and all-zero columns.