Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 7 additions & 1 deletion spec/benchmarks/results.md
Original file line number Diff line number Diff line change
Expand Up @@ -739,7 +739,13 @@ comparison—not a same-render-target speedup claim.

> The kernel microbench numbers above predate the kernel-parallelization pass:
> `bin_2d`, `histogram`, `m4`, `range_indices`, and `normalize` now fan out
> across cores above 512k rows. Zone maps have no merge traffic and use an
> across cores at per-cost-class gates: compute-bound scans (`m4`,
> `histogram`) from 128k rows, general scans (`bin_2d`, `range_indices`,
> `normalize`, min/max, sortedness) from 512k. Where each worker owns a
> private accumulator those row counts are only a floor: `bin_2d` also caps
> workers at points per cell and `histogram` at points per bin, so an
> accumulator as large as the input stays serial. Zone maps have no merge
> traffic and use an
> earlier, chunk-aware crossover: two complete 65,536-row chunks, with workers
> capped by the number of chunks. All paths are bitwise-deterministic; see
> `src/kernels.rs`.
Expand Down
17 changes: 12 additions & 5 deletions spec/design/rust-engine.md
Original file line number Diff line number Diff line change
Expand Up @@ -228,11 +228,18 @@ formats only; the SVG and PDF export paths have their own text contracts.
(`0`, `usize::MAX`, `-1`/`-2`) rather than one documented negative error
enum; unifying them is a work item for the next ABI bump.
- **E5 — threading stays inside**: parallel kernels use `std::thread::scope`
inside the call; the ABI stays synchronous. General row scans cross over at
512k values and scale to at most 18 workers. Zone maps cross earlier, at two
complete 65,536-row chunks, because chunks are independent and require no
merge; worker count is also capped by actual chunks. CodSpeed stays serial
because its simulator sums thread instructions rather than wall time.
inside the call; the ABI stays synchronous. Fan-out gates are per cost
class, each a named constant in `kernels.rs` with its measured serial
ns/row and break-even: compute-bound scans (M4, uniform histogram,
~2.5–3 ns/row) cross at 128k rows; general scans (min/max, sortedness,
range/validity, f32 encode/normalize) keep the 512k gate. Where each worker
owns a private accumulator the row gate is only a floor: `bin_2d` caps
workers at points per cell and `histogram_uniform` at points per bin, so an
accumulator as large as the input stays serial. Zone maps cross
earlier, at two complete 65,536-row chunks, because chunks are independent
and require no merge; worker count is also capped by actual chunks. All
gates scale to at most 18 workers. CodSpeed stays serial because its
simulator sums thread instructions rather than wall time.
Incremental build = handle + `xy_pyramid_append`, which mutates a live
pyramid in place under the registry lock. A polling entry point
(`xy_pyramid_poll`) for genuinely async builds is not built; if one is ever
Expand Down
95 changes: 85 additions & 10 deletions src/kernels.rs
Original file line number Diff line number Diff line change
Expand Up @@ -2502,7 +2502,7 @@ pub fn m4_indices(x: &[f64], y: &[f64], x0: f64, x1: f64, n_buckets: usize) -> V
n_buckets,
start,
end,
par_threads(end - start),
par_threads_above(end - start, PAR_THRESHOLD_COMPUTE).min(n_buckets),

@cubic-dev-ai cubic-dev-ai Bot Aug 1, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: Clustered or duplicate-x M4 windows in the new 128k–512k fan-out band can regress badly: every proposed split scans through the same giant bucket, then only one segment is executed. Consider finding bucket boundaries with a single pass/binary search, or falling back before paying these repeated scans.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At src/kernels.rs, line 2505:

<comment>Clustered or duplicate-x M4 windows in the new 128k–512k fan-out band can regress badly: every proposed split scans through the same giant bucket, then only one segment is executed. Consider finding bucket boundaries with a single pass/binary search, or falling back before paying these repeated scans.</comment>

<file context>
@@ -2502,7 +2502,7 @@ pub fn m4_indices(x: &[f64], y: &[f64], x0: f64, x1: f64, n_buckets: usize) -> V
         start,
         end,
-        par_threads(end - start),
+        par_threads_above(end - start, PAR_THRESHOLD_COMPUTE).min(n_buckets),
     )
 }
</file context>
Fix with cubic

@cubic-dev-ai cubic-dev-ai Bot Aug 1, 2026

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: The per-cost-class routing docs updated in this PR enumerate the worker caps for bin_2d (points per cell) and histogram (points per bin), but the newly added m4 cap at n_buckets is not documented. Since m4's parallel path only ever splits at bucket boundaries (at most n_buckets distinct segments are possible), capping workers at n_buckets is a benign optimization — but it is freshly introduced routing behavior. Per the repo's spec-tracking contract (AGENTS.md: "Every decimation/tier decision is recorded in the spec, never silent"), it would be worth one sentence alongside the existing accumulator-cap note in spec/design/rust-engine.md (E5) and spec/benchmarks/results.md so the documented fan-out contract matches the implementation.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At src/kernels.rs, line 2505:

<comment>The per-cost-class routing docs updated in this PR enumerate the worker caps for `bin_2d` (points per cell) and `histogram` (points per bin), but the newly added `m4` cap at `n_buckets` is not documented. Since m4's parallel path only ever splits at bucket boundaries (at most `n_buckets` distinct segments are possible), capping workers at `n_buckets` is a benign optimization — but it is freshly introduced routing behavior. Per the repo's spec-tracking contract (AGENTS.md: "Every decimation/tier decision is recorded in the spec, never silent"), it would be worth one sentence alongside the existing accumulator-cap note in `spec/design/rust-engine.md` (E5) and `spec/benchmarks/results.md` so the documented fan-out contract matches the implementation.</comment>

<file context>
@@ -2502,7 +2502,7 @@ pub fn m4_indices(x: &[f64], y: &[f64], x0: f64, x1: f64, n_buckets: usize) -> V
         start,
         end,
-        par_threads(end - start),
+        par_threads_above(end - start, PAR_THRESHOLD_COMPUTE).min(n_buckets),
     )
 }
</file context>
Fix with cubic

)
}

Expand Down Expand Up @@ -3174,22 +3174,29 @@ fn bin_2d_count_f32_scalar(
}
}

/// Row-scan kernels fan out across cores only past this size, where thread
/// Row-scan kernels fan out across cores only past these sizes, where thread
/// spawn + merge cost is well amortized. Threading stays inside the call —
/// the ABI remains synchronous (engine doc E5).
/// the ABI remains synchronous (engine doc E5). Compute-bound scans cross over
/// earlier: at ~2.5–3 ns/row serial, M4 and uniform histogram amortize a
/// fan-out from ~128k rows, the same crossover zone maps use.
const PAR_THRESHOLD: usize = 1 << 19;
const PAR_THRESHOLD_COMPUTE: usize = 1 << 17;

fn par_threads(n: usize) -> usize {
par_threads_above(n, PAR_THRESHOLD)
}

fn par_threads_above(n: usize, threshold: usize) -> usize {
static CODSPEED: std::sync::OnceLock<bool> = std::sync::OnceLock::new();
if *CODSPEED.get_or_init(|| std::env::var_os("CODSPEED_ENV").is_some()) {
return 1;
}
let cores = std::thread::available_parallelism().map_or(1, |p| p.get());
par_threads_for(n, cores)
par_threads_above_for(n, threshold, cores)
}

fn par_threads_for(n: usize, cores: usize) -> usize {
if n >= PAR_THRESHOLD {
fn par_threads_above_for(n: usize, threshold: usize, cores: usize) -> usize {
if n >= threshold {
cores.clamp(1, MAX_ROW_THREADS)
} else {
1
Expand Down Expand Up @@ -4596,7 +4603,14 @@ pub fn stratified_sample_mask<T: Copy + Sync + Into<u64>>(
pub fn histogram_uniform(data: &[f64], lo: f64, hi: f64, out: &mut [f64]) -> u64 {
assert!(x1_gt_x0(lo, hi));
assert!(!out.is_empty());
histogram_uniform_impl(data, lo, hi, par_threads(data.len()), out)
histogram_uniform_impl(data, lo, hi, histogram_threads(data.len(), out.len()), out)
}

/// Workers for a uniform histogram. Each owns a full bin vector and the merge
/// is `threads * n_bins`, so points per bin caps the row gate — same rule as
/// `bin_2d_threads`, whose per-thread grids are this kernel's per-thread bins.
fn histogram_threads(n: usize, n_bins: usize) -> usize {
(n / n_bins.max(1)).clamp(1, par_threads_above(n, PAR_THRESHOLD_COMPUTE))
}

/// Per-bin u64 counting shared by the serial and parallel paths. Stays scalar
Expand Down Expand Up @@ -6854,9 +6868,70 @@ mod tests {
// shape): per-thread grids + merge dwarf the scan — stay serial.
assert_eq!(bin_2d_threads(1 << 20, 1 << 20), 1);
assert_eq!(bin_2d_threads(2_100_000, 2048 * 2048), 1);
assert_eq!(par_threads_for(PAR_THRESHOLD - 1, 64), 1);
assert_eq!(par_threads_for(PAR_THRESHOLD, 64), MAX_ROW_THREADS);
assert_eq!(par_threads_for(PAR_THRESHOLD, 1), 1);
assert_eq!(
par_threads_above_for(PAR_THRESHOLD - 1, PAR_THRESHOLD, 64),
1
);
assert_eq!(
par_threads_above_for(PAR_THRESHOLD, PAR_THRESHOLD, 64),
MAX_ROW_THREADS
);
assert_eq!(par_threads_above_for(PAR_THRESHOLD, PAR_THRESHOLD, 1), 1);
}

#[test]
fn histogram_threads_bin_aware() {
let n = PAR_THRESHOLD_COMPUTE;
let cap = par_threads_above(n, PAR_THRESHOLD_COMPUTE);
// Below the compute gate: serial regardless of bin count.
assert_eq!(histogram_threads(PAR_THRESHOLD_COMPUTE - 1, 8), 1);
// Chart-sized bin counts (512 is the benchmark config): pure row gate.
assert_eq!(histogram_threads(n, 512), cap);
assert_eq!(histogram_threads(n, 8), cap);
// Bins as numerous as the rows: merge dwarfs the scan — stay serial.
assert_eq!(histogram_threads(n, n), 1);
assert_eq!(histogram_threads(200_000, 1_000_000), 1);
// In between, workers track points per bin.
assert_eq!(histogram_threads(1 << 18, 1 << 15), 8.min(cap));
}

#[test]
fn histogram_high_bin_count_matches_serial() {
// 128k-512k band: high bin counts route serial, and the public path
// stays bit-identical to the serial impl either way.
let n = PAR_THRESHOLD_COMPUTE + 1_000;
let bins = 300_000;
let data: Vec<f64> = (0..n).map(|i| (i % 977) as f64 * 0.25).collect();
assert_eq!(histogram_threads(n, bins), 1);
let mut serial = vec![0f64; bins];
let ts = histogram_uniform_impl(&data, 0.0, 244.0, 1, &mut serial);
let mut routed = vec![0f64; bins];
let tr = histogram_uniform(&data, 0.0, 244.0, &mut routed);
assert_eq!(ts, tr);
assert_eq!(serial, routed);
// Same rows, ordinary bin count: still fans out.
let cap = par_threads_above(n, PAR_THRESHOLD_COMPUTE);
assert_eq!(histogram_threads(n, 512), cap);
let mut small_serial = vec![0f64; 512];
let ss = histogram_uniform_impl(&data, 0.0, 244.0, 1, &mut small_serial);
let mut small_routed = vec![0f64; 512];
let sr = histogram_uniform(&data, 0.0, 244.0, &mut small_routed);
assert_eq!(ss, sr);
assert_eq!(small_serial, small_routed);
}

#[test]
fn par_thread_gates_route_by_cost_class() {
const { assert!(PAR_THRESHOLD_COMPUTE < PAR_THRESHOLD) }
let t_c = PAR_THRESHOLD_COMPUTE;
for cores in [1, 4, 20, 64] {
let cap = cores.clamp(1, MAX_ROW_THREADS);
// Compute-bound kernels fan out from 128k rows.
assert_eq!(par_threads_above_for(t_c - 1, t_c, cores), 1);
assert_eq!(par_threads_above_for(t_c, t_c, cores), cap);
// ...and stay serial in the band the general gate still reserves.
assert_eq!(par_threads_above_for(t_c, PAR_THRESHOLD, cores), 1);
}
}

#[test]
Expand Down
Loading