[PHYSICS] Consolidate current pipeline to main (GPD-verified sound; cluster closure pending)#34
Merged
Merged
Conversation
…nalization Physics-change protocol requirement (assert OLD numerical results before the change so the [PHYSICS] diffs are verifiable): - NEW test_bayesian_statistics_host_z_kernel.py: exact pins of single_host_likelihood (volume_deconv + local_ratio, with/without BH mass, low-z window clamp) via stubbed worker globals — the first tests to numerically execute the production host-z kernel. - test_constants.py: pin HOST_DRAW_Z_MAX == 0.5 and GALAXY_CATALOG_REDSHIFT_UPPER_LIMIT == 0.55 (pre-#20 values) + the host-draw <= population-depth ordering constraint. - physical_relations_test.py: d_L anchors at z = 0.5 / 1.5 incl. the closure-truth case dist(1.5, h=0.67) > dist(1.5, h=0.73) that the current fiducial-h parameter-space cap violates. - test_pp_coverage.py: exact-float pins of the tiny-config harness output for both kernels (the determinism test does not pin across code changes). Prep for issues #20 (depth 1.5) and #16 (sigma_v marginalization). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
User decision on issue #20 (2026-07-03): 'first go for z=1.5 and then see the results and HPC performance'. The pre-dt2 justification (horizon z ~ 0.18, truncation exact) is retired — post-dt2 the EMRI horizon reaches z ~ 1.5+, so the depth is a deliberate population-model choice matching Model1CrossCheck.max_redshift = 1.5. - constants.py: HOST_DRAW_Z_MAX 0.5 -> 1.5; GALAXY_CATALOG_REDSHIFT_UPPER_LIMIT 0.55 -> 1.55 (documented as UNWIRED: the reduced CSV is full-depth, load-time depth is max_redshift via _get_pruned_galaxy_catalog). - main.py injection_campaign: z_cut is now HOST_DRAW_Z_MAX instead of a hardcoded 0.5 — the P_det injection grid must span the host-draw volume or the selection function is blind above the cut. - cosmological_model.py: parameter-space d_L cap computed at the LOWEST campaign h (H_MIN/100 = 0.60 -> 13.0 Gpc) instead of the fiducial h = 0.73 (10.686 Gpc), which silently dropped z >~ 1.35 events in closure runs at h_true = 0.67 (dist(1.5, 0.67) = 11.643 Gpc) via ParameterOutOfBoundsError; plus an explicit ordering guard HOST_DRAW_Z_MAX <= max_redshift (the d_L pre-screen derivation relies on it). - handler.py: rewrite the stale 'truncation is exact' draw docstrings. - main.py fig23: sky-averaged completeness figure now rendered to the campaign depth. - tests: constants pins flipped 0.5->1.5 / 0.55->1.55 in the same diff (physics-change protocol); NEW campaign-depth machinery validation (Schechter/gammaincc finite, bounded, monotone over z in [0.5, 1.5] on synthetic + committed frozen m_th map; frozen-map f_bar(1.0) ~ 0); pixelated dark-draw W_k test pinned to an explicit z_max = 0.5 window (sampler is depth-agnostic; W_k contrast lives at low z — at 1.5 the near-uniform weights drown the Pearson check in multinomial noise). Refs: population reach anchors dist(1.5, h=0.73) = 10.686 Gpc, dist(1.5, h=0.60) = 13.002 Gpc (physical_relations_test.py pins). Full fast suite: 735 passed. Closes #20. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… kernel (issue #16) User decision on issue #16 (2026-07-03): 'for the next campaign we marginalize it and test the isolated effect in parallel'. Inference-side only (re-evaluate tier) — no re-simulation, no catalogue rebuild. Old kernel (bayesian_statistics.single_host_likelihood): N(z; z_g, sigma_z_cat) [+/- 4 sigma_z_cat window] New kernel: N(z; z_g, sigma_z_eff), sigma_z_eff^2 = sigma_z_cat^2 + sigma_z_pv^2 sigma_z_pv = (1 + z_g) * SIGMA_V_PEC_KM_S / c, SIGMA_V_PEC_KM_S = 200 References: - (1+z) factor: Davis et al. (2011), arXiv:1012.2912, Eqs. (1)/(A1) (z_obs = z_cos + (1 + z_cos) v_pec/c); Davis & Scrimgeour (2014), arXiv:1405.0105, Eq. (9). - Quadrature with catalogue sigma_z: Mastrogiovanni et al. (2023), arXiv:2305.10488, Sec. IV (icarogw/GLADE+ practice). - sigma_v = 200 km/s: Fishbach et al. (2019), arXiv:1807.05667, Sec. 2.2; Chen et al. (2018), arXiv:1712.06531. LISA-EMRI precedent (Laghi et al. 2021, arXiv:2102.01708) uses 500 km/s with the (1+z) factor — kept as a systematics-budget row. Dimensional analysis: [km/s] / [km/s] = dimensionless redshift error; quadrature of two dimensionless sigmas. Limiting cases: sigma_v -> 0 recovers the old kernel exactly; z -> 0 gives sigma_z_pv = sigma_v/c = 6.67e-4 (the low-z LVK convention); at z = 1.5 the factor 2.5 gives 1.67e-3 ~ the catalogue's parse-time PV floor. Applied ONCE at function entry — all eight downstream consumption points (window bounds, Z_g renorm, prior pdf, per-host D_g, with-BH numerator/ denominator, MC proposal + sampling_pdf) flow through the single norm() object, so no double counting inside the likelihood. The catalogue z_error's parse-time PV-CORRECTION floor (0.0015) is a DISTINCT term (correction uncertainty vs residual dispersion) — documented in constants.py. Ball-tree candidate window / pruning intentionally keep the bare catalogue z_error (second-order candidate-list effect). Also: integration-testing twin mirrors the quadrature; pp_coverage gains an inert sigma_z_pv knob (default 0.0 — committed anchor runs stay bit-identical, proven by the unchanged exact-value pins). Kernel regression pins updated in this diff (protocol): at z_g = 0.1, sigma_z_cat = 0.0015 the kernel widens ~11%, integrals shift 0.2-0.5% (e.g. volume_deconv numerator 1622.007 -> 1629.370); pre-change values in the parent commit. Full fast suite: 735 passed. #16 stays OPEN pending the parallel isolated PV value-correction test. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…nance in DATA_INVENTORY - CHANGELOG [Unreleased]: campaign depth z=1.5 (incl. z_cut wiring, low-h d_L cap fix, ordering guard) and residual-PV marginalization (sigma_v = 200 km/s, (1+z) factor, references). - DATA_INVENTORY: new 'Galaxy Catalogue (reduced GLADE+)' section — the z_cmb 8-col full-depth CSV was previously untracked here; records schema, frame provenance (18e9608), rebuild recipe + append-mode gotcha, retired backups, and the frozen m_th map C1 coupling (unchanged by the depth constants since no CSV rebuild occurred). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…, injection provenance Readiness-sweep fixes (adversarially verified, 2026-07-03) closing the window between catalogue deepening (#20) and new-injection arrival: Stale-pool gates (A2-STALE-POOL-GATE, campaign-blocking): - SimulationDetectionProbability gains expected_z_max/allow_shallow_pool: rejects pools with max z < 0.9 x expected depth and pools MIXING provenance (different z_cut stamps, or stamped + legacy files — the partial-rsync signature). Default None keeps synthetic/test pools ungated. Production constructors (evaluate + combine paths) pass expected_z_max=HOST_DRAW_Z_MAX; the combine path previously had NO check at all. - BayesianStatistics.evaluate: validate_coverage is now a HARD gate (raise below 95% coverage; --allow_low_pdet_coverage escape) instead of a warning buried in 1 of 41 task logs. - Injection writer stamps z_cut + code_rev into every row: h_inj alone cannot discriminate pre-/post-dt2 or shallow/deep pools (0.73 in every era, identical filenames). - Local 80-file pre-dt2 z_cut=0.5 pool archived to simulations/injections_RETIRED_predt2_zcut0p5_20260703/ (filename collision with the regenerated pool); DATA_INVENTORY retirement note. test_sky_selection data-gated tests now self-skip until the depth-1.5 pool lands. [PHYSICS] M_z out-of-grid policy (A2-EXTRAP): p_det with-BH-mass queries outside the injected M_z support were silently LINEARLY extrapolated (scipy fill_value=None semantics) while four comments claimed 'nearest'. Old: linear extrapolation, clipped to [0,1]. New: clamp M_z to the grid edge — true nearest, the documented Phase 44 boundary intent. Affects only out-of-support queries (rare once the pool covers the campaign range); regression test pins edge equality. Injection-loop consistency (A1/A3): - _TIMEOUT_S 30 -> 90 s, aligned with the simulation loop: timed-out events are DROPPED from the pool, so a timeout-rate correlation with (d_L, M_z) at depth 1.5 would bias the p_det grid; counter logged for smoke-test binning. - Symmetric M_z <= M.upper_limit truncation (the CRB path structurally excludes that corner via the Fisher bounds guard; the pool must match). - Pre-dt2 comments retired (93.5% z-rejection, 24/69500 horizon claim). Ops/instrumentation: - --prescreen_audit CLI flag: bypass the quick-SNR skip while logging PRESCREEN_AUDIT (quick, full, params) lines — enables the smoke-run false-negative measurement for PRE_SCREEN_SNR_FACTOR + issue #19. - --combine fast path now writes run_metadata_combine.json (TC-10: previously no git_commit/args recorded for the combine stage). - ParameterSpace d_L default 7 -> 13.1 Gpc (dist(1.5, h=0.60) = 13.0015; the old default silently rejected z >~ 0.9 events in bare constructions); dead luminostity_detection_threshold removed; completeness-plot default z_max tracks HOST_DRAW_Z_MAX. - Integration fixture pool deepened to z <= 1.5 + provenance-stamped (mirrors the production writer; exercises the gates end-to-end). mypy clean (144 files); full fast suite 734 passed / 15 data-gated skips. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… safe resubmit, provenance probes Readiness-sweep fixes TC-01..TC-15 (adversarially verified where campaign-blocking; measured anchors 2026-07-03, jobs 5732036/5735964/5735965): - evaluate.sbatch: 128 cpu/15 min -> 16 cpu/6 h placeholder (anchor: 56-76 min per h-value @ 3355 events / 16 cpus; the old pair guaranteed 38/38 TIMEOUTs + a dead chained combine); partition cpu,cpu_il; 41-value hybrid h-grid (0.655/0.665/0.675 added so the 0.005-dense window covers the 0.67 closure truth); per-task --seed (EVAL_SEED + task id) + --simulation_index metadata; per-h skip-if-output idempotency guard. - combine.sbatch: 45 -> 90 min; cluster-side figures OFF by default (RENDER_FIGURES=1 to opt in) — the figures phase is what killed job 5735965; posteriors-only anchor ~20 min. - ALL array templates (simulate/evaluate/combine/inject): private per-run CWD (/cwd with simulations + master_thesis_code symlinks) replaces the shared /simulations symlink dance — concurrent multi-seed pipelines could silently cross-contaminate CRB/ posterior writes (TC-03); merge.sbatch already --workdir-safe. - submit_pipeline.sh: --array derived from the evaluate.sbatch H_VALUES line (single grid source, TC-14); --h_true flag (threads H_VALUE, suffixes RUN_DIR) and REQUIRED --injection_pool flag (stages the pool into RUN_DIR/simulations/injections; loud failure on empty dir); submit-time posterior archiving (replaces the racy task-0 in-job archive, TC-06). - resubmit_failed.sh: TIMEOUT removed from the default state filter (gpu_h100_short tasks TIMEOUT by design — the old filter would delete ALL outputs of a healthy run); h_value recovered from run_metadata before cleanup with conflict-abort (closure-truth contamination fix); abort if the merged CRB already exists (append-duplication guard). - preflight.sh + cluster.env: O(1) catalogue fingerprint (row-1 z field: 0.001733 = z_cmb, 0.00099... = STALE z_helio — both are 8-col, so the column count cannot discriminate) + shallow-injection-pool depth probe. - datasets.yaml: seed43000_Mz/seed700 RETIRED (pre-dt2, z_cut=0.5); depth15_campaign placeholder pending regeneration. - submit_resimulate_phase50.sh: stale INJECTION_SOURCE default removed (must be set explicitly). - LAUNCHING_JOBS.md / README.md: private-CWD mental model, 2026-07-03 walltime anchors, new flags + resubmit signature. bash -n clean on all 10 shell files; 41-label convention verified against bayesian_statistics posterior filenames. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…alidity boundary measured
Referee blocker REF-P001/S006 (Paper A): sigma_z in {0.10, 0.15, 0.25} x
{bare, volume} kernels, 250 realizations x 250 events, paired seed
20260701, truths {0.62, 0.72, 0.84}.
Verdict: bare kernel rails 100% at every sigma_z >= 0.10 (unusable);
volume kernel near-nominal to sigma_z/z ~ 0.5-0.8 (cov68 0.62-0.70,
|bias| <= 0.011) and demonstrably degrades at sigma_z/z >~ 1
(sigma_z = 0.25: cov68 0.33-0.44, bias +0.04-0.06). The paper's
calibration claim gets a MEASURED validity boundary instead of the
asserted decisiveness the referee dinged. Details in SUMMARY.md.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The smoke run's quick-SNR false-negative measurement (--prescreen_audit, issue #19 / PRE_SCREEN_SNR_FACTOR) had no path from sbatch to the CLI. PRESCREEN_AUDIT=1 in the submitting shell now appends the flag (set -u-safe empty-array expansion). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…mission Fixed BEFORE any campaign data exists (referee asserted-decisiveness critique): multi-seed MAP accuracy (2-SEM over 4 seeds), closure-truth recovery inside per-run 68% HPD, per-seed pp_coverage near-nominal (bounded by the 2026-07-03 sigma_z/z validity scan), and the verdict language for either outcome. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… resolved from package root The 46900 smoke pool came back WITHOUT z_cut/code_rev: _INJECTION_COLUMNS (the pd.DataFrame columns whitelist) silently dropped the new keys — textbook W-PRE-12 'check EVERY output writer'. Columns added + warning comment. Also _get_git_commit() ran git in the process CWD, which under the new private $RUN_DIR/cwd is not a repo (-> 'unknown' provenance); now anchored at the package's realpath so the cwd symlink resolves back to the checkout. Smoke anchors recorded (job 5739442, depth 1.5, A100): 50 events in 2:29-5:29 per task (~3-6.6 s/event SNR-only); timeouts 4/~230 (@90s); M_z > 1e6 population-bound skips ~10% (expected at (1+z) lift against Model1CrossCheck's M cap — symmetric with the CRB path by design). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tion validated, variant built Isolated PV value-correction impact test (handoff §7b), phase 1: - GLADE+'s helio->CMB convention identified and VALIDATED on the 21,925,647 flag2==0 control rows: multiplicative (1+z_cmb) = (1+z_helio)(1+(v_sun/c) cos theta), Planck dipole (369.82 km/s toward l=264.021, b=48.253); residual median 5.28e-5, p99 1.98e-4. Additive and sign-flipped conventions ruled out 10-30x. Small z-dependent effective-amplitude deficit (~350-363 km/s) characterized; biases the flagged-row reconstruction by <= 2.1e-5, ~30x below the median PV correction. - Variant reduced_galaxy_catalogue_noPVcorr.csv (1.7 GB, NOT committed; git-excluded): identical to the live z_cmb catalogue except the 709,117 flag2==1 rows (3.13%, z <= 0.109) carry the frame-only transform of z_helio — i.e. the BORG PV value-correction REMOVED. Median removed |dz| = 6.30e-4 (p99 3.11e-3, max 5.19e-2). Exact row parity with the live CSV; all unflagged rows byte-identical. - Phase 2 (pending): production --evaluate on frozen seed600 events against live vs variant catalogue (identical events, same pool, --allow_low_pdet_coverage for the archived shallow baseline); posterior MAP/mean shift = the isolated PV value-correction impact, an UPPER bound for the campaign (seed600 events are all low-z). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…5% false negatives (fix #19 margin re-measurement) Old: skip events with quick(1-yr) SNR < SNR_THRESHOLD * 0.3 before the full 5-yr computation (sqrt(T) bound 0.447 + chirp margin, calibrated pre-dt2 at z <= 0.5). New: PRE_SCREEN_SNR_FACTOR = 0.0 — gate disabled; the 1-yr quick waveform is skipped entirely unless --prescreen_audit is set. Reference (measurement, not literature): depth-1.5 smoke audit, cluster job 5740080 (2026-07-03), 543 (quick, full) SNR pairs recorded via --prescreen_audit (archived at results/campaign_phase2_smoke_20260703/prescreen_audit_pairs.txt): - 3/543 FALSE NEGATIVES with full SNR >= 20 at quick SNR 0.25/1.48/5.34 (full 44.1/29.0/30.2) — sources plunging in years 2-5 accumulate SNR the 1-yr check generator cannot see; quick == full for 362/543 (plunge-within-1yr), so NO positive factor separates the populations. - A lossy gate is a selection-function inconsistency against the gate-free injection pool (the p_det grid counts those events as detectable). ~30% throughput saving not worth a 0.55% biased loss concentrated in the loud late-plunging corner. Limiting case: factor -> 0 recovers the injection-consistent selection exactly (every candidate gets the same full 5-yr SNR the pool got). Also closes the #19 remainder: the smoke measured ZERO d_L pre-screen hits at depth 1.5 (bound 11.220 Gpc = 1.05 x dist(1.5, h=0.73) vs max drawable d_L 10.686 Gpc) — margin 1.05 confirmed adequate and non-cutting on post-dt2 data. Full fast suite: 734 passed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ion state - datasets.yaml: depth15_campaign is CURRENT (500 files / 50k events, two-batch provenance incl. the benign mixed-code_rev note). - DATA_INVENTORY: Phase-2 campaign section — pool, design (4+2 seeds, 100x40, 41-grid, per-task eval seeds), smoke anchors, seed1000 submission (jobs 5743694-97), pre-registered criterion pointer. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ependent of dev-box SSH
Submits the remaining Phase-2 pipelines (seeds 2000/3000/4000 @ 0.73 +
closures 5000 @ 0.67, 6000 @ 0.77) one at a time whenever the expanded
queue depth drops below 250 (submit cap measured between 294 and 544 on
2026-07-03). Idempotent via run-dir existence checks, so a login-node
purge or restart cannot double-submit; one automatic retry per seed;
logs to $WS/campaign_orchestrator{,_submissions}.log.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ce-staged fallback) Lets the orchestrator run from a copy on the workspace filesystem when $HOME git sync is unavailable (Lustre EIO on pack writes, 2026-07-03). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Never advance past a failed seed: retry same seed with 5-min backoff (up to 500 attempts ~ 41 h), $HOME read probe (cat of every file the submission path needs) before each attempt. - MAX_PENDING 250 -> 150 (cap measured in (294, 544]; stay clear at the low end). - Defensive cleanup of unsubmitted run dirs (no logs/) after a failed attempt so restart idempotency stays truthful. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… away-section - results/campaign_phase2_runs/watch_and_retrieve.sh: detached dev-box loop, 30-min cadence — cluster status snapshots + rsync-back of every campaign run dir and the orchestrator logs (per-poll SSH connections; nothing persistent to drop). - Runbook §4c: the three self-driving layers, Monday checklist, and the $HOME Lustre EIO situation (cluster repo pinned at b233375, git pulls blocked — code-equivalent to the campaign tag). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
combine_posteriors gained allow_shallow_pool (from the existing --allow_low_pdet_coverage flag): the combine path got the stale-pool depth gate in 3273fa5 but not the deliberate-shallow escape the evaluate path has — blocking legitimate re-evaluations of archived shallow baselines (first hit: the frozen-seed600 PV test combine). mypy clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ow plan + campaign-safety rules Fresh-session entry point while the cluster is 2FA-blocked (until Monday): first finalize the completed §7b PV analysis (post to #16), then a Workflow-orchestrated whole-codebase review — 12 dimensions with informed leads from today's recon, adversarial verification of majors, neutral-only fix wave (sim-/inference-semantic findings become issues, never landed against the running campaign), report artifact + CLAUDE.md known-bugs reconciliation. Includes the campaign-safety classification hard rule and the one running-campaign risk to check first (eval-node memory at the 4-5x larger in-RAM catalogue). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
CLAUDE.md + hpc-gpu.md drifted from the code: - Known Bug #1 (unconditional cupy import) struck — guarded since 4894648 in all four GPU modules. - Known Bug #7 (bayesian_inference.py 10% distance error) marked moot — Pipeline A was removed in c1571a2; production uses per-source Cramer-Rao bounds. - Known Bug #8 (WMAP cosmology) rewritten as the documented G11 design choice (OMEGA_M=0.2726 matches Barausse-2012 M1; Planck mismatch is a tracked systematic). - Architecture bullets fixed to name bayesian_statistics.py / simulation_detection_probability.py / posterior_combination.py instead of the deleted Pipeline-A modules; bayesian_inference.py removed from the physics-trigger list. - Parameter-estimation bullet updated (5-point stencil is default since Phase 10). Docs-only; no code touched (campaign-safe). Part of the 2026-07-04 code review. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Neutral cluster-script fixes (do NOT restart live jobs — deploy on Monday's git pull): - CLU-03: '|| true' on the grep-miss branches in resubmit_failed.sh and submit_pipeline.sh so the friendly empty-result / diagnostic branches actually run instead of set -euo pipefail aborting first. - CLU-06: orchestrator stop-doc uses the pkill bracket idiom so it can't signal its own ssh wrapper shell. - CLU-07: watcher run-dir glob widened to run_202607*_seed* so seeds whose run dirs are datestamped 07-07+ (orchestrator can delay a seed ~41 h under queue-cap pressure) are still mirrored. Structural cluster items (CLU-01 double-submit, CLU-02 blind drift monitor, CLU-05 --mem, CLU-09 combine afterok) tracked in issue #27 for Monday. Part of the 2026-07-04 code review. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ty (issue #16) Isolated the GLADE+ peculiar-velocity VALUE correction on the frozen seed600 event set (live PV-corrected catalogue vs a frame-only variant with the correction removed on the 709k flag2==1 low-z rows). Production --evaluate, 17-value H0 grid, shared 3342-event intersection: 1D (redshift-distance): Δmean(live-noPV) = -0.0142 (-3.3σ), ΔMAP -0.010 2D (with BH mass): Δmean = +0.0012 (+0.2σ) — PV-insensitive, un-railed seed600 is the designed worst case (all-low-z, all 709k corrections). The σ_v marginalization (8568d9f) covers this; no campaign re-scope warranted. Closed #16. Commits the provenance (ANALYSIS.md, README.md, builder, queue, stats), not the 1.7 GB variant CSV or the per-h posteriors. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Neutral robustness/provenance fixes from the 2026-07-04 code review (data_simulation and injection_campaign are on the simulation path, not --evaluate; the campaign runs pinned code on the cluster, so these take effect on the next git pull): - SIM-01: cancel the 90s SIGALRM in a try/finally on EVERY path (success, each exception-continue, the quick-gate continue) so no stale alarm fires in inter-iteration code and kills an unattended task. - SIM-02: scope warnings-as-errors with warnings.catch_warnings()/simplefilter so it is restored on every exit path and cannot leak into the host-refill / population sampling between iterations (and stops the global filter list growing unbounded). - SIM-03: per-(stage, exception-class) skip tally, logged at end of data_simulation, so a parameter-correlated CRB-stage drop rate (post-SNR-cut) is measurable. - SIM-04: register a SIGTERM flush handler in injection_campaign so a wall-time cap kill loses at most the events since the last periodic flush, not up to 1999 full-5yr-generator SNRs (mirrors data_simulation's existing handler). - SIM-07: bare 'raise' for the unmatched ValueError instead of raise ValueError(e). - REP-02/03: run_metadata cli_args now serialises the FULL parsed namespace via Arguments.to_dict() (captures normalization_mode, pdet_*, catalog_only, ...), and Arguments.seed is cached so it can't return a fresh random value on repeated access. Full fast suite green (739 passed), mypy clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Covers the previously-untested parse_to_reduced_catalog -> read_reduced_galaxy_catalog
round trip on a synthetic raw file: flag {1,3} filter, PV-error NaN->0.0015 floor +
quadrature fold into the redshift error, trailing integer redshift flag, the 8-column
positional reorder, and the PHI_S/THETA_S rename on read. The 8-headerless-column
contract is load-bearing for every downstream reader and drifted once before
(the .stale6col_mar28 / .zhelio_20260702 variants). 5 tests, all green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
dist()/cached_dist()/dist_vectorized() accept w_0/w_a but evaluate the ΛCDM-only hypergeometric form, so a wCDM call silently returned the ΛCDM result. Add a shared _reject_unsupported_wcdm() guard that raises NotImplementedError on w_0 != -1 or w_a != 0. Every production call uses the defaults (verified by grep), so this changes no computed value — it only turns a silent wrong answer into a loud one. A real wCDM distance (numerical 1/E quadrature) stays deferred to /physics-change. Also PHY-02: dist_derivative now forwards Omega_m/Omega_de/w_0/w_a to hubble_function (which implements the full CPL E(z)) instead of silently using module defaults; and the module/dist docstrings now say ΛCDM-only for the analytic distance. Full fast suite green (739 passed), mypy clean. GitHub #4. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Pipeline A was deleted in c1571a2; these are its last unreachable remnants (verified zero live callers by grep across source, tests, scripts, cluster): - datamodels/galaxy.py (synthetic GalaxyCatalog) + its only consumer test_benchmarks.py (the whole file was just that one slow benchmark). - constants TRUE_HUBBLE_CONSTANT, GALAXY_REDSHIFT_ERROR_COEFFICIENT, FRACTIONAL_LUMINOSITY_ERROR, FRACTIONAL_BLACK_HOLE_MASS_CATALOG_ERROR, LUMINOSITY_DISTANCE_THRESHOLD_GPC — all were galaxy.py-only after c1571a2. - handler.parse_to_reduced_catalog_with_reduced_errors — a no-op (built a local DataFrame and discarded it) with zero callers; drops the now-orphaned dist_to_redshift_error_proagation import + a docstring mention. - bayesian_statistics.single_host_likelihood_grid — a print-and-return-[] debug stub with zero callers. Production single_host_likelihood is untouched. (The dev-only single_host_likelihood_integration_testing cross-check twin is kept.) - scripts/quick_snr_calibration.py — orphaned Phase-12 one-off (SNR_THRESHOLD long settled), referenced by nothing. ruff + mypy clean, full fast suite green (739 passed). Closes the code side of #7. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
All changes are value-neutral for production (dead code or unreachable branches or comments-only): - HPC-06: power_spectral_density now raises ValueError on an unknown TDI channel instead of silently returning an all-zero PSD (which would make inner products inf/nan). Unreachable in production (callers iterate ESA_TDI_CHANNELS = 'AE'). - PHY-07: delete DarkEnergyScenario.de_equation — dead (zero callers) and wrong (divides by w_a instead of the CPL multiply; ZeroDivisionError at the fiducial w_a=0). - cosmological_model.py: fix a stale comment pointing at the removed detection_probability module (now simulation_detection_probability). - PHY-09: correct the M derivative_epsilon comment — the log-uniform midpoint is 10^5.5 ≈ 3e5 M_sun (not the stated ~3e3), and the real rationale is a phase-coherence bound for the oscillatory waveform, not Vallisneri's eps_mach^(1/4)|x|. The epsilon value (1.0) is UNCHANGED. Full fast suite green (739 passed), mypy clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…sweep Neutral plotting/comment fixes from the code review (figures only; no likelihood or simulation semantics): - PLT-03/04: figure truth-lines and the CRB-derived redshift now use constants.H (TRUTH_H) instead of a hardcoded 0.73 literal, so every 'Injected' marker agrees if the fiducial changes (main.py fig01/fig02, paper_figures axvlines). - PLT-05: the interactive tension-explorer x-axis widened to [0.55, 0.88] so the top of the production 0.60-0.86 h-grid (h=0.86) is no longer rendered off-screen. - PLT-10: catalog_plots plot_glade_completeness docstring says Gpc to match its axis label (was Mpc); dropped the stale 'GalaxyCatalog plotting methods' reference. - Stale-comment sweep: TQ-04 test docstrings say HOST_DRAW_Z_MAX is 1.5 (was 0.5); the bias_investigation test_27 comment no longer points at the deleted datamodels/galaxy.py:64. Full fast suite green (739 passed), ruff + mypy clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Verified-findings table (landed vs deferred-to-issue vs verified-correct), the 9 commits on this branch, the campaign risk assessment (eval-node OOM refuted), coverage gaps, and the §7b PV bound summary. Companion to issues #23-#27. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
CHANGELOG [Unreleased]: Fixed/Removed/Added entries for the review's neutral changes. TODO: PHYS-6 (wCDM guard) marked done; PHYS-2/PHYS-7 marked moot (datamodels/galaxy.py deleted). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…low, debrief + model discipline Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…model probe) Add sigma_dl_model_in_likelihood (config + --sigma-model-in-likelihood) to pp_coverage.py: the GW-likelihood factor uses the z-dependent model/true-distance width sigma_f*A(z)/h (carrying its own 1/sigma(z) normalization via _norm_pdf) instead of the constant observed-distance sigma_f*dL_obs. Applies to the host kernel numerator (every mixture_mode) and the completion B_num; the p_det selection integrals (D(h), gray D_g_i) are unchanged (they integrate p_det, not the GW likelihood). Combined with --pdet-in-numerator it is the fully-consistent exact conditional for the latent-thresholded generative model. Probes the sharpened floor candidate from results/pp_coverage_pdetnum_20260711 SUMMARY.md sec.2: the constant observed-distance inference sigma is an O(sigma_f^2) mismatch vs the sigma_f*dL_true generative noise, the scale of the +0.002..+0.005 sigma_z-independent residual floor. Default off is bit-identical (regression guard: exact zs=0.3 sz=0.035 results block byte-identical to the committed exactmode JSON). ruff + ruff-format + mypy clean; 33 pp_coverage tests pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…he dominant floor component
2x2 {const-sigma, model-sigma} x {p_det-inside off, on} on exact deep cells
(zs 0.2/0.3) + inert controls (zs 0.5/1.0), plus n_events scaling (250/1000/4000)
and a fine-grid confirm (h_step 0.001). Pre-registered predictions in RUNBOOK.md
(written before the runs).
- model-sigma + p_det-inside (the fully-consistent exact conditional for this
latent-thresholded model) removes ~85-90% of the +0.002..+0.005 floor: MAP bias
-> <= +0.0008 on the deep cells AND nulls the -0.002..-0.004 control offset, with
cov68 nominal at campaign-scale n. Neither half alone works (model-sigma alone
over-corrects negative; p_det alone was the 27m refutation) — they are the two
halves of one conditional.
- n-scaling: the const-sigma floor is FLAT in n with cov68 COLLAPSING (h=0.72
0.63->0.38->0.12) => a real asymptotic model bias, not a finite-sample MAP-skew.
A tiny second-order residual (~+0.0005, ~15x below campaign sigma_boot) survives
the corrected estimator, visible only at n=4000.
- fine-grid confirm: coarse(0.004) == fine(0.001) biases to +-0.0001 — not a
grid-quantization artifact.
Floor decomposition COMPLETE: (i) dominant membership-support kernel leak
(260711-117) + (ii) inference-noise-model floor (this task) + (iii) negligible
second-order residual. Practically subdominant for Paper B; a required design input
to the user-gated soft-f(z)-kernel production correction (do NOT add p_det alone).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…edger (decomposition COMPLETE) H_sigma CONFIRMED: the sigma_z-independent floor is the inference noise-model approximation (sigma(dL_obs)-vs-sigma(dL_true) width + p_det-inside, the two halves of the latent-threshold exact conditional). Both together remove ~85-90%; const-sigma floor is a real asymptotic bias (flat in n, cov68 collapses), not finite-sample skew; tiny 2nd-order residual ~15x below campaign sigma_boot. Floor decomposition COMPLETE. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…N-4 shallow-venue probe Add d50_gpc/w_pdet_gpc (config + --d50-gpc/--w-pdet-gpc) to pp_coverage.py, threaded through detection_probability and every call site (population sampler, D(h), beta_G, host/completion p_det factors). Lets the harness model a shallower detection horizon: d50=0.23, w=0.037 -> z_median 0.044 (the seed600 shallow venue, vs the default 1.85 -> z_median 0.28). For N-4(a): test whether the calibrated volume-kernel/no-truncation estimator develops the seed600 +0.0132 offset when the venue is made shallow. Default (1.85/0.30) is bit-identical (regression guard: exact zs=0.3 sz=0.035 results byte-identical to committed exactmode). ruff/format/mypy clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…stimator-intrinsic (large sigma_z/z at low z) (a) Depth ladder (calibrated volume kernel, no truncation, sigma_z=0.035): calibrated at the commission depth (z_med 0.28, bias -0.002) but develops a strong POSITIVE bias as the venue shallows -> +0.011 at z_med 0.056, +0.030 at z_med 0.044 (seed600 depth), cov68 collapsing. P-B confirmed (P-A refuted). (b) sigma_z sweep at the shallow rung: bias VANISHES at sigma_z<=0.015 (-0.002, calibrated) and appears only at sigma_z=0.035 -> it is a sigma_z/z effect (~0.8 at z_med 0.044): the host-z kernel truncates at z>=0 and the volume/Eddington correction (derived for an untruncated kernel) stops cancelling. (b) jackknife on the on-disk seed600 run_live JSONs (no re-eval): reproduces the raw +0.0132; the residual is broad/systematic (62% of events tilt high, Gini 0.65, trimming the top-|tilt| events GROWS the residual) -> not outlier-driven, matching the per-event depth mechanism. Load-bearing caveat: full attribution to seed600 needs its low-z redshift-error model (sigma_z/z ~ O(1)?); cross-seed systematic-vs-scatter needs the campaign. A z>=0-truncation-aware volume kernel would address BOTH the deep membership leak and this shallow effect (user-gated /physics-change, not this task). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…edger Shallow-venue 1D +0.0132 EXPLAINED: estimator-intrinsic sigma_z/z-at-low-z truncated-volume-kernel Eddington effect (venue depth ladder + sigma_z sweep; seed600 jackknife shows it is broad/systematic). Full seed600 attribution pending its low-z sigma_z model; cross-seed needs the campaign. A z>=0-truncation-aware volume kernel would address BOTH the deep membership leak and this shallow effect. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…hallow explained), N-5 optional, production-kernel convergence Both ranked probes done: deep-incompleteness bias fully decomposed (kernel leak + noise-model floor, hx1) and shallow +0.0132 explained (sigma_z/z-at-low-z volume-kernel truncation, iic). Remaining local: N-5 (optional 2D) + the seed600 low-z sigma_z model check that closes N-4 attribution. Both regimes converge on one user-gated production fix (z>=0-truncation-aware volume kernel). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
… are photo-z (σ_z/z~O(1)) Closes the load-bearing caveat from [L8]: seed600's effective redshift uncertainty at z≈0.046 is large-fractional photo-z, so the σ_z/z-at-low-z truncated-volume-kernel Eddington effect applies and the shallow +0.0132 residual is attributed to it (not an unexplained offset). Measured on the reduced GLADE+ catalogue seed600 evaluated (inline, no re-eval), z-shell 0.03-0.06 (z_med 0.046, n=767,552): 89.7% photometric, σ_z median 0.0344, σ_z/z median 0.65 — near-exact match to the harness σ_z=0.035 rung that gave +0.030; spec-z minority (10.3%) at σ_z/z≈0.033 is the calibrated counterweight the jackknife saw. Code trace (airtight): the likelihood host-z kernel width IS the catalogue σ_z (bayesian_statistics.py:2243, host_z_error_eff=sqrt(σ_z²+σ_z_pv²)) and the [PHYSICS] z>=0 clamp is active for low-z photo-z hosts z_g<4·σ_z (:2234-2239); at z_g=0.046, 4σ_z=0.14>z_g. Un-truncated-derived volume/Eddington correction stops cancelling. Docs only (no code/physics change). Cross-seed systematic-vs-scatter still needs the campaign. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…presentation gate (user-gated) Scoping ONLY — no production code. Assembles the physics-change presentation the hard gate requires so the user can decide whether/how to proceed with the z>=0-truncation-aware / photo-z-marginalized volume host-z kernel that both bias regimes converge on ([L7] deep membership leak + [L8] shallow σ_z/z Eddington). Contents: OLD formula (bayesian_statistics.py:2243-2317 volume_deconv kernel); the identified σ_z/z~O(1) limitation (commission-d2 calibrated at z_med 0.28 → -0.002 vs [L8] shallow z_med 0.044 → +0.030); candidate directions (A truncated-normal- consistent, B photo-z-marginalized soft membership); the [L7] distance-error coupling constraint (do NOT fix the z-kernel alone); references (Gray 2020, MFG 2019, CFH 2018, Mastrogiovanni/ICAROGW 2305.10488, Alfradique/Bom 2023-2026 full-photo-z-PDF practice, Wang-Chen 2408.10382 tolerance); dimensional analysis; six binding regression gates (must reproduce commission-d2 deep calibration AND remove the shallow bias); open user decisions D1/candidate/coupling/validation. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…n-in-M impact now -0.0022 Optional handoff item N-5 (2D-channel subsample dependence). Re-ran the G7row9 494-event seed600 driver at HEAD (713fbd1 D_g fix present) on the archived shallow venue. Under current code the 494-event 2D subsample is well-behaved: edge_mass 0.216->0.003 (pre-fix railing toward 0.86 GONE), mean 0.790->0.768; it sits +0.0135 above the full-venue 0.7546 = subsample-selection offset, not a code defect. 1D subsample 0.745 reproduces the venue +0.013 residual. No additional 2D subsample/grid pathology; venue-level +0.025 2D residual stays campaign-gated (D4). Caveat: the pre-fix artifact is NOT a clean D_g-only baseline (its 1D 0.730 predates #29/z-clamp changes); the clean D_g attribution stays in the L-B full-venue A/B. Bonus: post-D_g-fix Eddington-in-M impact on the 2D mean is -0.0022 (was -0.020 pre-fix) => the bayesian_statistics.py:2400-2401 comment + quoted value are now stale (flagged, not edited - physics-trigger file). Driver: threaded allow_shallow_pool=True through BOTH evaluate() and combine_posteriors() (both build a SimulationDetectionProbability guarding the campaign-depth pool) so the archived shallow seed600 venue re-runs, same precedent as the seed600 A/B. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…d, N-5 done, production fix scoped) N-4 shallow attribution CLOSED (seed600 low-z σ_z/z~O(1) photo-z), N-5 2D subsample well-behaved post-D_g-fix, production kernel fix scoped (user-gated /physics-change). Remaining is cluster-gated (EXP-40, campaign) + user decisions D1-D5 + the production /physics-change awaiting approval. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…nt all) D1 evidence-driven (defer framing to EXP-40/campaign); D2 one combined deployment (#22->#31->#32 + #27 CLU); D3 Paper A on hold until pipeline+results satisfy, then upgrade; D4 defer 2D +0.025 to campaign (no more local bias work); PROD implement all (full truncated-normal x volume-prior + soft photo-z membership + distance-error coupling kernel, new normalization_mode, volume_deconv stays golden). Execution sequence (physics hard gate) recorded. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…dy to execute) Nails down the code-level plan after reading the production kernels: volume_trunc is SHALLOW-only (z>=0 floor + unified numerator support; no-op on deep venue by construction; deep leak is a separate z_support-edge change). Substantive change = integrate N_g over the per-host galaxy window (not the shared GW window), restructuring the batched hot path; volume_deconv stays byte-identical. Decisive gate is the seed600 494-event A/B (genuine empirical uncertainty whether the numerator-window is the lever). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
Part 1 formula approved via /physics-change gate; execution deferred to a fresh session per user (hot-path batched-kernel restructuring wants full context). Pasteable kickoff with code locations, golden-guard, and the decisive seed600 494-event empirical gate. STATE marked NEXT SESSION START HERE. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…t (2026-07-12) Advisory pre-Phase-2 convention trace from the orbiter coverage model (read-only on code; nothing submitted). No live divergence found; HOST_DRAW_Z_MAX already fixed 0.5->1.5; all 4 historical divergence classes mutually consistent. - CONVENTIONS-MANIFEST.md: 8-row skeleton of convention-bearing quantities (incompleteness declared) - COVERAGE.md: coverage-map/v1; finding C-MTC-20260712-001 (first missing-coverage finding), C-003 pp_coverage depth (WATCH — under active exploration, not a blocker), C-004 M_z injection invariant test (APPROVED FOR IMPLEMENTATION — a future session should add a CI assertion that the injection catalog "M" column == M_source*(1+z)) - tracer-verdict-2026-07-12.md: per-quantity verdict + refuter pass; ADVISORY, ratified 2026-07-12 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016thJ2geJxGSRFmQ3FRPyyu
…lly FALSIFIED Implements the approved Part 1 shallow-venue correction behind a new, isolated normalization_mode="volume_trunc": the calibrated volume kernel with the in-catalogue NUMERATOR integrated over the per-host galaxy window [z_g-4sigma, z_g+4sigma] (shared with Z_g/D_g) and the lower z-limit floored at 0 instead of 1e-6. Scalar and batched kernels; volume_deconv/local_ratio stay BYTE-IDENTICAL (golden regen additions-only), batch==scalar bit-identical on all volume_trunc cases, full CPU suite 889 passed. The decisive seed600 494-event shallow-venue A/B (scripts/volume_trunc_ab.py) REJECTS it: volume_trunc worsens the shallow bias — 1D mean 0.745 -> 0.800, MAP 0.73 -> 0.80, posterior collapses onto h=0.80 (moved AWAY from truth 0.73 by ~4x the +0.013 residual it was meant to remove). The volume_deconv arm reproduces the established reference exactly (1D 0.745, 2D 0.768), so the driver/data are sound. Mechanism (results/volume_trunc_ab_20260712/): (1) the shared fixed_quad(n=50) is numerically invalid over the wide host window — the sparse Gauss-Legendre nodes alias the narrow GW peak (n=50 -> 0.0 vs exact 0.24-0.65), h-dependently; (2) even the exact host-window numerator tilts high in the shallow regime. Both push H0 high. Verdict: the numerator-window unification is NOT the +0.013 shallow lever and, as specified, is numerically broken. volume_trunc is retained as an EXPERIMENTAL / FALSIFIED mode (not CLI-wired, not for production) so the finding is reproducible; volume_deconv remains the golden default. Tests: volume_trunc pins + sigma_z->0 limiting case + prior-shape h-independence. Ref: Gray 2020 A.10; G2b §1.4; scoping §7b (.planning/PRODUCTION-KERNEL-FIX-SCOPING-20260712.md). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…(2026-07-12) STATE.md + DECISIONS §PROD-Part-1-OUTCOME + new HANDOFF-VOLUME-TRUNC-FALSIFIED: the approved Part 1 volume_trunc kernel worsens the shallow venue (1D mean 0.745->0.800, posterior collapses to h=0.80); numerator-window is not the +0.013 lever and is numerically broken (fixed_quad n=50 aliases the GW peak over the wide host window). Redirect to Candidate B soft membership + [L7] coupling with a peak-aware quadrature (user-gated). Code/finding committed in c4a1c7d. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
Records the volume_trunc host-z kernel (Part 1) under Unreleased: implemented behind an isolated normalization_mode, volume_deconv byte-identical, but FALSIFIED at the seed600 shallow A/B (worsens the bias) and retained only as an experimental/falsified diagnostic. Not CLI-wired, not for production. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…2 verified) User insight: the catalogue BH-mass error is ~60% (Reines-Volonteri 0.24 dex intrinsic scatter, linear sigma_M/M floor 0.58), not 10% — so the LINEAR-Gaussian host-mass kernel hits the same untruncated-vs-truncated inconsistency as the low-z photo-z redshift kernel, at the [M_min,M_max]=[1e4,1e7] EMRI bounds. Mass is 2D-only => candidate for the 2D +0.025 residual + the info-monotonicity violation. Two probes (results/mass_kernel_truncation_20260713/): - mass_trunc_probe.py: P(M<0)=4.8% under the linear kernel; G2d effective mass is <1% off in the interior but 15% low near M_min (wrong sign) / 22% high near M_max; 65% of R_eff-weighted EMRI hosts sit in the M_max boundary zone. - mass_kernel_h0_toy.py: controlled 2D mass<->z<->H0 toy, production (linear+G2d) vs correct (lognormal x R_eff truncated). Mass-kernel H0 shift = HIGH, seed-robust: +0.0027+/-0.0002 (sigma_z/z=0.15), +0.0089+/-0.0008 (sigma_z/z=0.30); grows with photo-z leverage; control (correct arm) ~unbiased at near-spec-z. Sign puzzle RESOLVED (marginalisation sees the kernel shape -> HIGH, not the naive LOW). Unified: 1D = redshift effect (+0.013); 2D adds the mass effect => 2D bias > 1D bias. Indicated fix (user-gated /physics-change): lognormal x R_eff truncated mass kernel. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…rrection Adds a behavior-preserving config.clamp_zgal knob to the pp_coverage harness (default True = bit-identical; 29 harness tests green) and a clamp-isolation diagnostic. At the shallow seed600-matched venue (d50=0.23, z_med~0.044, sigma_z=0.035), volume kernel, multi-seed: clamp ON : map_bias +0.0240 +/- 0.0022, cov68 ~0.61 (the [L8] +0.030) clamp OFF : map_bias -0.0056 +/- 0.0020, cov68 ~0.68 (vanishes, cov recovers) Deep venue is clamp-independent (control). => the shallow high bias is the boundary clamp on the OBSERVED photo-z + a naive Gaussian kernel that does not model the clamp; the volume/Eddington correction is a red herring (present both ways). Refines [L8]: the fix is a censored-measurement likelihood, not a re-derived volume kernel. Production relevance is OPEN: the reduced catalogue redshift is NOT hard-clamped (min -0.0003, 17/500k <=0), so production may not fully share the harness artifact; needs a low-z photo-z pileup/censoring inspection to settle whether the seed600 +0.013 1D residual is this effect. results/h1_zclamp_20260713/FINDINGS.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…+0.030 is largely an artifact Direct check of the reduced catalogue z<0.10 shell (871k gal): no pileup/spike at 0 (n(z==0)=0, min -0.0003, smooth rising histogram), only 3.7% of low-z hosts have the kernel crossing 0. => production does NOT reproduce the harness generative clamp, so the +0.030 shallow bias is substantially a harness artifact. Weakens [L8]'s seed600 +0.013 attribution and suggests the redshift-half production fix is largely unnecessary; mass-channel (H2) bias is independent and stands. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…egime Wide h-grid (no railing) mass-kernel H0 differential vs photo-z leverage (3 seeds): sigma_z/z=0.15 -> +0.0025, 0.30 -> +0.0081, 0.50 -> +0.0165, 0.75 -> +0.0214. At the real shallow-shell leverage (sigma_z/z~0.5-0.65) the mass-kernel bias is +0.016 to +0.02 HIGH -- a large fraction of the +0.025 2D residual. Combined with the H1 clamp finding (redshift half largely a harness artifact), the MASS kernel is the primary production-relevant driver of the 2D residual; the lognormal x R_eff truncated mass kernel is the load-bearing production fix. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…ted as 2D bias driver New isolated normalization_mode="mass_trunc": the 2D (with-BH-mass) channel's host-mass prior replaced from the linear-Gaussian G2d moment match (eddington_shifted_host_mass) to the truncated lognormal × R_eff prior on [M_MIN, M_MAX] — Reines & Volonteri (2015) lognormal error × Babak et al. (2017) R_eff population weight, renormalised on the physical EMRI mass window. Numerator: Gauss-Hermite on the narrow GW M_z peak (peak-aware — the fix for the fixed_quad aliasing that falsified volume_trunc). Denominator: Gauss-Legendre in ln M over a per-host peak-aware window (the erf-sum closed form is Gaussian-prior-only). volume_deconv / local_ratio / volume_trunc stay byte-identical (kernel-parity golden regenerated additions-only: 9 new *_mt_* pins, zero existing pins changed; mt_3d == vd_3d exactly; single_host_likelihood_batch bit-identical to scalar; limiting cases in test_mass_trunc_kernel.py). Full CPU suite 940 passed; mypy/ruff clean. Decisive seed600 494-event shallow-venue A/B (scripts/mass_trunc_ab.py): Δ2D mean = +0.0029 (small, WRONG sign), Δ1D = 0 exact → the mass-kernel truncation is NOT the 2D +0.025 residual driver. The isolated single-host toy over-stated it by omitting the selection denominator, which cancels the numerator shift in the full ratio-of-sums pipeline; the linear-Gaussian G2d approximation is thereby empirically validated as adequate for the 2D channel. Retained as an experimental/exonerated mode (not CLI-wired); volume_deconv stays the golden default. Refs: Reines & Volonteri (2015) arXiv:1508.06274 §4.1; Babak et al. (2017) arXiv:1703.09722; Abramowitz & Stegun 25.4.46. Finding: results/mass_trunc_ab_20260713/. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
…ingle venue Both "X is NOT the 2D-bias driver" conclusions rest on the SAME seed600 494-event shallow subsample — no cross-venue confirmation, so a shared venue idiosyncrasy would fool both. Also separate the CLEAN A/B delta (same events both arms) from the cross-venue extrapolation: the A/Bs run on the SUBSAMPLE (2D mean 0.768 / +0.038), not the full venue (0.7546 / +0.025) the "+0.025 residual" refers to. Add a binding rule to the plan of record (§5): negative conclusions are venue-scoped and provisional pending the campaign (D4/§4b), not universal facts. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis
This was referenced Jul 16, 2026
Closed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Consolidates the cumulative, GPD-verified physics onto
mainso the default branch reflects thecurrent best pipeline. Supersedes/subsumes #22, #31, #32 (all their content is contained here).
Physics landing (10
[PHYSICS]commits)B_num/D(no longer silently dropped)max_redshift(production no-op; horizonz_max ≤ 1.5)--h_valuesgrid (3.8×)volume_trunc(FALSIFIED) andmass_trunc(exonerated) — not in the CLI--normalization_modechoices.Verification (GPD soundness gate — PASS)
All 7 live changes verified physically sound (HIGH confidence, internal correctness):
hosts-present bit-equal for #29; #30 num/den move together; #16
σ²_eff=σ²_cat+σ²_pvrecovers old kernelas σ_v→0; exact-vs-quad denominator agree to 5.1e-13; spline d_L 2.5e-11 rel; perf parity bit-
==.Gate clean: ruff + mypy clean, 940 passed / 6 skipped, 149 physics/parity tests pass.
Honest caveat (not a soundness blocker)
The empirical H₀ MAP/bias/coverage closure on the deep z=1.5 pool is data-dependent and pending the
cluster run (the deep injection pool lives only on the cluster). If the cluster run surfaces a problem,
it will be addressed in a follow-up investigation/branch — the pipeline code here is verified sound.
Known limitations remain documented honestly in
docs/source/limitations.rst/H0_BIAS_RESOLUTION.md.🤖 Generated with Claude Code
https://claude.ai/code/session_01QDfQjz3Bv6oKNUQycTSDis