Skip to content

goldenmatch v3.7.0

Choose a tag to compare

@github-actions github-actions released this 21 Jul 23:03
Immutable release. Only release title and notes can be modified.
3ba5956

Fixed

  • Zero-config Fellegi-Sunter recall no longer collapses at scale (candidate-pair
    projection fix + recall-safe compounding + memory-aware budget).
    Zero-config
    FS recall collapsed 1.0 at ≤2.4M → 0.82 at 4.8M → ~0.02 at 30M (F1 0.030,
    the 30M single-box proof) — entirely scale-dependent and invisible below ~5M.
    Three fixes, the first being the actual scale-dependent root cause:

    • Candidate-pair projection (the root cause). _project_pass_pairs
      extrapolated each blocking block's SIZE by the full row ratio
      (cnt * n_full / sample_n). That is only right for a SATURATED low-cardinality
      key; a NEAR-UNIQUE key keeps producing new values as N grows (blocks stay
      ~constant size, the COUNT grows), so growing its size invents ~C(ratio, 2)
      PHANTOM pairs per sample singleton. At 30M (ratio ~150) a near-unique
      (zip, email) compound was projected at ~2.2B pairs and DROPPED by the
      pair-gate, collapsing blocking to a single first_name pass (dups have typo'd
      first names → recall ~0.02). Fixed by growing block size only by the key's
      sample collision headroom (1 - distinct/sample_n): saturated keys grow by the
      full ratio (byte-identical), near-unique keys barely grow (singletons stay
      singletons → 0 phantom pairs). Small data (n_full == sample_n, the whole
      bench-probabilistic panel) is unaffected — no extrapolation runs.
    • Recall-safe compounding. When the pair-gate DOES bound an over-budget
      coarse pass, it now compounds with an exact-agreement identity field
      (email / identifier / phone) at full value
      , before the corruption-prone
      name/geo initials. Duplicates share those exactly, so the compound keeps every
      true pair together while collapsing the block to near-singletons — recall-safe
      AND a stronger reducer (30M person shape: zip + first-initial = recall 0.82;
      zip + email = recall 1.0 at 0.9M pairs). The old "most selective" reducer (a
      name initial) split true pairs on any typo'd name.
    • Memory-aware budget. _fs_total_pair_budget is now
      max(300M floor, available_ram_gb * ~40M) (anchored to the 25M-on-64GB proof
      where ~2.1B bounded pairs peaked at ~28 GB), so a big box does less
      compounding to begin with and keeps coarse passes pure. Byte-identical below
      the trigger (small boxes keep the 300M floor + all #1803 tuning).
      Measured F1 0.9005 → 1.0000 at 4.8M (P=R=1.0); the identity reducer holds
      recall 1.0 at 4.8M even at a tight 300M budget (3.3M candidate pairs). Not EM:
      the EM within-block-pair sample cap is irrelevant to this (100K/200K/400K give
      identical F1); blocking recall exactly tracks pipeline recall. The
      gate-fs-zeroconfig nightly now runs at 10M (was 1M — below where this class is
      visible) with set -o pipefail so the F1-floor failure is no longer masked by
      | tee.
  • Weak-positive blocking-pass pruning now runs on the FS arrow lane.
    GOLDENMATCH_BLOCKING_PRUNE_PASSES=1 invoked select_passes (polars-native:
    with_row_index / group_by) directly on df, but the FS routed / arrow
    lane passes a pyarrow Table -- so it threw AttributeError, was swallowed
    into "keep all passes", and pruning was a silent no-op for every arrow-lane FS
    caller. _maybe_prune_blocking_passes now coerces a Table / LazyFrame to
    polars first. With the pruner actually running on a representative sample it
    cuts a redundant 6-pass zero-config FS scheme (3 first-name + 2 last-name
    transform variants + zip) to 3 passes (one name axis per field + zip),
    measured 71s -> 32.6s at 1M, F1 unchanged 1.000 (P=R=1.0) on realistic
    person data -- recall-safe because it drops only redundant transform variants
    of the same field, keeping each blocking axis. Still opt-in; a default-on
    flip for the FS path is gated on the bench-probabilistic panel.

  • **Learned blocking no longer clobbers the #1207 strong-identifier union at

    =50k rows (#1316).** Auto-config forced strategy="learned" unconditionally
    at total_rows >= 50_000, discarding the per-identifier blocking union that
    null-sparse multi-source strong-id data depends on. Measured on that shape at
    50k, learned blocking under-blocks catastrophically -- candidate-pair recall
    collapses from 1.0 (union) to 0.0 (the learner trains on a <=5k sample, finds
    no pairs above its recall threshold, falls back to one column, and
    skip_oversized drops every resulting oversized block). The >=50k gate now
    keeps a strong-identifier union it detects and only upgrades non-union large
    shapes to learned blocking.

  • Zero-config Fellegi-Sunter admits shared identity identifiers (email/phone)
    at cardinality 1.0.
    FS auto-config dropped every card == 1.0 exact field as
    a "perfect surrogate", but that also discarded identity-bearing identifiers
    (email/phone) that duplicates carry verbatim -- FS's single strongest signal.
    Because cardinality is measured on a config sample that can under-represent
    duplicates, this silently collapsed the EM model to zero matches at scale
    (measured: zero-config FS F1 0.0 at 1M on realistic person data; recovers to
    1.0 with the identifier admitted). build_probabilistic_matchkeys now admits
    email/phone at card >= 1.0 (an FS comparison field self-regulates -- a
    true PK to neutral, a shared identifier to a large weight) while still
    excluding the ambiguous bare identifier type (row PKs) for config hygiene.
    FS path only; the weighted/exact matchkey path is unchanged.

  • Zero-config FS blocking pair-budget now prunes at scale (~6x wall). The FS
    blocking pair-budget (_bound_probabilistic_blocking_pairs) is documented to
    extrapolate each pass's candidate pairs to the full population, but
    auto_configure_probabilistic_df never passed n_rows_full -- so on the
    auto-config sample the bound measured pairs at sample scale (a 66M-at-1.2M pass
    reads as ~1.8M at a 200K sample), stayed under budget, and never pruned, leaving
    redundant giant-block soundex passes. Threading the full row count lets the
    bound bound the oversized name passes at true scale: measured zero-config FS
    wall 410s -> 71s at 1.2M (F1 unchanged at 1.000). The FS routing call site and
    the bench/gate helper now pass n_rows_full.

  • FS missing="unobserved": a partial-observation pair no longer normalizes to
    1.0 (#1854).
    The min-max score range accumulated only over the OBSERVED
    fields, so a pair agreeing on its single observed field had total == pair_max
    and rescaled to 1.0 — maximal confidence from minimal evidence. The range now
    spans EVERY matchkey field (the sum stays over observed only), so a
    one-of-many-observed agreement is correctly uncertain (e.g. 0.75 on a
    two-field key). Cross-surface: fs-core::score_fs_pair (the native/unobserved
    runtime + fs-wasm), the four Python reference paths in core/probabilistic.py
    (vectorized ×2, scalar, score_pair_probabilistic). Identical when every field
    is observed (missing="disagree" and fully-populated pairs are byte-unchanged;
    auto-config routes null-heavy data to disagree, so the default path is
    unaffected). Measured under forced unobserved: historical_50k
    f1_probabilistic recovers to ~0.63 from the collapsed ~0.33; febrl3 −0.002
    (within the quality-gate tolerance).

Changed

  • Out-of-core FS scorer batches blocks into the native kernel (opt-in path,
    parity-exact) + GOLDENMATCH_FS_OOC_DEBUG progress.
    score_fs_out_of_core
    scored one block per score_probabilistic_bucket_native call; on person data
    (tens of thousands of tiny blocks per pass) that made the FFI fan-out the wall
    (~60s for a single 200K pass, hours at 25M). It now hands a whole
    block-contiguous wave to the kernel in one call per worker-chunk, with the
    per-block size_list isolating blocks — mirroring the in-memory
    _score_one_bucket batched call, so the emitted pair set + scores are
    byte-identical (the existing parity tests against the per-block reference
    gate it). The numpy vectorized path likewise batches via
    score_probabilistic_vectorized_batch; unsupported scorers keep the per-block
    fallback. Measured ~8× on the per-pass scoring at 200K. GOLDENMATCH_FS_OOC_DEBUG=1
    prints a per-phase / per-pass timing line (load, block-map, scan+score, block
    count) so a long >=25M streaming leg shows live progress instead of a blank
    spinner. NOTE: this fixes the per-block fan-out only; the OOC path still runs
    one full-dataset scan+score per blocking pass, so low-cardinality passes and
    multi-pass depth remain a separate scale lever.
  • Out-of-core streaming FS refinements (review follow-ups, opt-in path only).
    _prep_all_ids returns a range instead of a 25–50M-element Python list when
    __row_id__ is contiguous (the pipeline-generated common case), avoiding a
    multi-GB transient before the pyarrow int64 array on the ≥40M streaming path.
    stream_fs_dedupe_output and dedupe_to_parquet's in-memory fallback now
    remove a stale golden.parquet left by a prior run into the same out_dir
    when a run produces no golden rows, so the on-disk file set matches the
    returned golden_path=None.

Added

  • Bounded bucket streaming for the in-RAM FS route
    (GOLDENMATCH_FS_BLOCK_SOURCE=frame, default OFF) — cuts the ≥1M
    frame-residency peak.
    The scale branch of the FS (probabilistic) bucket
    scorer (score_buckets._score_single_pass, height ≥ n_buckets) used to
    partition_by the keyed frame into all n_buckets eager frames up front — a
    ~2× transient at partition time whose freed pages jemalloc retains straight
    through cluster/golden, the dominant remaining single-node FS peak once the EM
    build_blocks fixes landed. With the flag on, the scale branch keeps the
    single bucketed frame resident and slices each bucket out on demand
    (filter_eq inside the worker), so peak holds the bucketed frame plus at most
    max_workers in-flight slices instead of all N partitions. Byte-identical to
    the eager path: filter_eq preserves within-bucket row order ==
    partition_by(maintain_order), so each bucket's scorer output is unchanged,
    and cross-bucket append order is order-invariant downstream (pairs
    canonicalized). Measured (synthetic person 1M, local 4c/15GB, jemalloc-decay
    env): whole-pipeline peak 3244 → 2875 MB (−11.4%), byte-identical output
    (850,714 clusters both).
    Default OFF keeps the eager path until a CI ≥1M
    peak gate validates the flip; scoped to the FS route (the weighted path is
    untouched); the DuckDB (above-RAM) source is the separate
    GOLDENMATCH_FS_OUT_OF_CORE path below. Spec:
    docs/superpowers/specs/2026-07-20-fs-frame-residency-bucket-streaming-design.md.

  • Out-of-core single-box streaming Fellegi-Sunter dedupe — the ≥40M scale
    path (GOLDENMATCH_FS_OUT_OF_CORE=1, default OFF).
    The probabilistic (FS)
    route had no out-of-core or distributed path: _fs_use_bucket_route hands
    backend=duckdb/ray to a single-node scorer, so the whole prepared frame
    stayed resident and the single-box FS wall was ~40M on 64 GB
    (CI-measured: 25M @ 40.3 GB / 16 min; 50M projected to ~82 GB OOM), while F1
    stayed scale-stable. backends/fs_out_of_core.py adds three bounded
    mechanisms: score_fs_out_of_core streams block groups one at a time from a
    DuckDB-resident (file-spilled) prepared table (scoring peak = one block
    group, byte-parity with score_buckets absent oversized blocks);
    stream_fs_dedupe_output writes unique/dupes via DuckDB COPY ... TO parquet
    with no result frame; run_fs_dedupe_streaming ties prep → DuckDB file →
    free frame → score → cluster → stream. New public
    gm.dedupe_to_parquet(*files, out_dir=...) reaches it (and falls back to the
    in-memory pipeline + parquet write when the config is not FS-eligible or the
    flag is off, so it always yields the same files). The default path (no
    output_dir) is byte-unchanged. Spec:
    docs/superpowers/specs/2026-07-20-fs-frame-residency-bucket-streaming-design.md.

  • FS EM build_blocks memory-peak fixes (both default ON, byte-identical
    output).
    The FS memory peak is EM's build_blocks, not score_buckets.
    GOLDENMATCH_FS_EM_BLOCK_SLIM projects each EM block-frame to
    [__row_id__] + blocking fields before materialization (width 14→6);
    GOLDENMATCH_FS_EM_AGG_BLOCKS builds the EM-only blocks as compact int64
    row-id arrays via one group_by().agg() per pass, never materializing
    per-block frames (supersedes the slim lever). Measured whole-pipeline peak on
    person 100K: 2126 → 527 MB (−75%); regime-dependent above ~1M where the
    EM-sample cap already bounds block count. A jemalloc page-decay env
    (_RJEM_MALLOC_CONF) trims the 1M FS peak a further ~33% at ~zero wall.

Fixed

  • DedupeResult.clusters now exposes real contents to C-level consumers on
    the frames-out path (re-scoped #1961).
    The lazy cluster handle
    (LazyClusterDict, a dict subclass that builds on first Python content
    access) left its underlying storage empty until an override fired. C-level
    consumers that bypass those overrides — the goldenmatch-pg bridge's pyo3
    .extract::<HashMap>() (PyDict_Next), json.dumps (empty-dict fast path
    via PyDict_GET_SIZE) — silently observed zero clusters, so a dedupe that
    correctly formed a size-2 cluster serialized as empty (the pg p4_typed
    smoke: "expected 2 rows in a size-2 cluster, got 0"). DedupeResult.clusters
    is now a property that materializes the lazy handle to a plain dict on first
    read, so any consumer sees the real contents; a result whose .clusters is
    never read still never pays the build (the frames-out perf win is preserved).
    The pg extension's columnar SPI read (#1951 spi.rs) was correct and is
    unchanged.

  • Arrow-native auto-config no longer silently degrades blocking on
    wide/sparse frames (#1852, mode 2).
    With
    GOLDENMATCH_AUTOCONFIG_ARROW_NATIVE=1 (default since 2026-07-14), three
    build_blocking helpers still ran raw-polars idioms on the input frame:
    _id_pass_scale_safe_nonnull (the #1207 per-identifier union gate),
    _name_path_primary's geo-compound sizing, and _llm_suggest_blocking_keys.
    On a pa.Table the first two AttributeError'd into a bare except that
    returned False/continue, so the strong-identifier blocking union and
    name+geo compounding silently collapsed to name-only blocking — a recall/
    precision divergence between the arrow and polars lanes (the _llm_* path
    crashed outright on an arrow+LLM-blocking run). All three now route through
    the backend-neutral Frame seam, so arrow and polars select identical
    blocking passes. Locked by a wide/sparse config-equality parity test
    (test_build_blocking_id_union_arrow_parity). Complements the earlier
    _build_compound_blocking fix (mode 1, the crash).

  • auto_configure_df(pa.Table) no longer raises AttributeError: 'height'
    in composite blocking search (#1852 tail).
    When every exact-eligible column
    is a perfectly-unique surrogate key, auto-config goes fuzzy-only and
    build_blocking falls into composite-key search. find_composite_blocking_keys
    and estimate_avg_block_size (core/blocking_candidates.py) still ran raw
    polars idioms (df.height, df.select(...).n_unique()) on the input frame,
    which is a pa.Table by default under GOLDENMATCH_AUTOCONFIG_ARROW_NATIVE=1
    crashing on the arrow-native lane. Both are now routed through the backend-
    neutral Frame seam (to_frame + joint_n_unique), matching the earlier
    build_blocking ports. This branch is only reached on an all-unique-identifier
    (join-table / order-shaped) frame, which is why the #1852 wide/sparse gate never
    exercised it; locked by test_auto_configure_all_unique_ids_arrow_parity.