Skip to content

v0.5.27

Choose a tag to compare

@github-actions github-actions released this 31 May 04:28
· 352 commits to main since this release

v0.5.27 — TA-D3 batched dict resolution spike (-17% e2e ingest)

Phase-0 (v0.5.26) showed dict resolution is 73% of LUBM-1
ingest time, driven by 26,473 individual SPI calls. TA-D3
batches them via two new APIs:

  • dict::put_terms_batch(terms) -> Vec resolves N dict
    terms in 2 SPI calls per chunk:
    (a) bulk INSERT ... ON CONFLICT DO NOTHING
    (b) bulk JOIN-back lookup with WITH ORDINALITY for
    input-order preservation

  • loader::ingest_turtle_dict_batched + UDFs
    pgrdf.parse_turtle_dict_batched / load_turtle_dict_batched
    use a 2-pass design: collect triples + unique terms,
    pre-resolve datatype IRIs, bulk-resolve remaining terms in
    chunks (default 500), then walk triples to build s/p/o
    batches via the existing flush_batch path.

    metric baseline spike delta
    ──────────────── ──────── ────── ──────────────
    triples 100,573 100,573 identical
    elapsed_ms 1,452 1,200 -17.4% e2e
    dict_ms 1,059 739 -30.2%
    parse_ms 74 85 +16% (2-pass cost)
    insert_ms 312 376 +21% (re-walk cost)
    dict_db_calls 26,473 106 -99.6% (250x fewer)

Theory validates. Production landing (TA-7 winner) should use
one-pass with incremental dict flushes to remove the 2-pass
overhead.

tests/regression/sql/128-parse-turtle-dict-batched-parity.sql
locks the two paths produce equivalent _pgrdf_quads. 4
assertions covering: triple count parity, baseline-subset-of-
spike, spike-subset-of-baseline (excluding blank-node
subjects/objects which are parser-assigned), and the new
path = dict_batched JSONB discriminator. All evaluate to t.

  • Existing parse_turtle / load_turtle paths untouched.
    New UDFs are additive; baseline path runs by default.
  • Shmem cache integration deferred to TA-D2.
  • Production landing (TA-7) is a separate decision after
    TA-D2 + TA-11 + TA-10 + TA-D1 + TA-9 close.

Cargo.toml, pgrdf.control, compose.yml SQL mount, 00-smoke
expected.out, META.json (both fields), docs example output.
Upgrade bridge renamed sql/pgrdf--0.5.1--0.5.26.sql ->
sql/pgrdf--0.5.1--0.5.27.sql (no-op; schema unchanged).

31 / 92 -> 32 / 92 once this release advertises. TA-D3 closes.

Next: TA-D2 (shmem cache pre-warm) as v0.5.28.

Refs: _WIP/SPIKE.TRACK-A.phase0-findings.md;
tests/perf/lubm/spike-ta-d3.lubm-1.md;
SPEC.ROADMAP.TRACK.TASKS.v1.0-devel.md §6 TA-D3 (v0.5.27).