Skip to content

feat(calibration): slop-corpus replay backfill + replay-derived provider track records#8281

Merged
JSONbored merged 4 commits into
mainfrom
feat/slop-corpus-backfill
Jul 24, 2026
Merged

feat(calibration): slop-corpus replay backfill + replay-derived provider track records#8281
JSONbored merged 4 commits into
mainfrom
feat/slop-corpus-backfill

Conversation

@JSONbored

Copy link
Copy Markdown
Owner

Summary

Closes #8277. Advances #8278 (the adapter + first recorded table; the BYOK multi-provider campaign remains).

Phase-3 backfill (scripts/backfill-slop-corpus{,-core}.ts): replays the deterministic slop scorer over the archived raw-context corpus manifest and synthesizes provenance-tagged (slop_replay_backfill_v1) slop_gate_score fired/override pairs labeled by realized outcomes — the phase-1 "mapping a" framing applied to the #8224 knob. Signal-subset honesty is enforced by allowlist: only the three diff-derivable signals may contribute (caught in testing: the scorer fires empty_pr_description on an undefined description — correct live where absence is emptiness, inflating in replay where the field simply wasn't archived); risk/band recompute from exactly those weights, and every row records its computed signal codes. Idempotent upserts, dry-run default, zero GitHub traffic, wrangler/--pg dual path (with the fail-loud resolvePgConnection contract honored — its silent-D1-fallthrough warning caught a real wiring bug during rollout).

Already applied + verified on both stores (460 fired + 460 override rows each; evidence + the honest lower-bound reading posted on #8277): all replayed mass sits below the 0.60 ceiling, so these rows are neutral to ladder comparisons by construction — they feed the reliability curve and drift reads; the ladder's discriminating source remains live full-signal capture.

#8278 adapter (scripts/provider-replay-track-record.ts): cached replay verdicts → ProviderReviewSignalcomputeProviderTrackRecords (#8228, zero new math), abstentions excluded, replay-derived disclaimer baked into the output. First same-seed two-provider table posted on #8229.

Validation

  • Full local gate npm run test:ci: green except two known first-test-in-file load timeouts (satisfaction-floor-loosening-run, upstream-ruleset) while a deploy + replay ran concurrently — both pass 76/76 standalone in this worktree; neither file is touched here. npm audit --audit-level=moderate clean.
  • New suites: 8 tests on the backfill core (parse arms, allowlist honesty, skip accounting, timestamp flooring, render modes) + 2 on the adapter (vote mapping + rollup rendering), all green; typecheck clean.
  • Codecov note: scripts/** is outside the patch-coverage measurement; the cores are unit-covered per the backfill-family convention regardless.

@JSONbored JSONbored self-assigned this Jul 23, 2026
@superagent-security

Copy link
Copy Markdown
Contributor

Superagent didn't find any vulnerabilities or security issues in this PR.

@codecov

codecov Bot commented Jul 23, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 92.16%. Comparing base (5486d9e) to head (041c208).
⚠️ Report is 1 commits behind head on main.
✅ All tests successful. No failed tests found.

Additional details and impacted files
@@           Coverage Diff           @@
##             main    #8281   +/-   ##
=======================================
  Coverage   92.16%   92.16%           
=======================================
  Files         788      788           
  Lines       79175    79175           
  Branches    23933    23934    +1     
=======================================
  Hits        72974    72974           
  Misses       5062     5062           
  Partials     1139     1139           
Flag Coverage Δ
shard-1 53.80% <ø> (-4.78%) ⬇️
shard-2 50.87% <ø> (+0.89%) ⬆️
shard-3 58.02% <ø> (+3.97%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

@loopover-orb loopover-orb Bot added the gittensor:bug Gittensor-scored bug fix — scores a 0.05x multiplier. label Jul 23, 2026
@loopover-orb

loopover-orb Bot commented Jul 23, 2026

Copy link
Copy Markdown
Contributor

Warning

⏸️ LoopOver review result - manual review recommended

Review updated: 2026-07-24 01:12:40 UTC

5 files · no blockers · CI green · clean

⏸️ Suggested Action - Manual Review

  • AI review could not be completed: The dual-model AI review did not return a usable verdict for this change.

Review summary
AI review could not be completed for this PR head. LoopOver is holding this PR for manual review instead of relying on deterministic signals alone.

Nits — 2 non-blocking
  • PR author also opened the linked issue — Link an issue that was opened by a different contributor, or provide a rationale for why this self-authored issue represents genuine discovery work.
  • AI review could not be completed — The gate is held for a human reviewer rather than passed automatically; it re-evaluates on the next update.

Decision drivers

  • ✅ Code review — No blockers (No AI review summary)
  • ⚠️ Gate result — Not blocking (Advisory; not blocking this PR.)
Context & advisory signals — never blocks the verdict
Signal Result Evidence
Linked issue ✅ Linked #8277
Related work ✅ No active overlap found No same-issue or scoped active PR overlap found.
Change scope ✅ 20/20 Low review scope from cached public metadata (1 linked issue).
Validation posture ✅ 25/25 PR body includes validation/test evidence.
Contributor workload ✅ 10/10 Author activity: 14 registered-repo PR(s), 14 merged, 228 issue(s).
Contributor context ✅ Confirmed Gittensor contributor JSONbored; Gittensor profile; 14 PR(s), 228 issue(s).
Improvement ✅ Minor risk: clean · value: minor
Review context
  • Author: JSONbored
  • Role context: owner (maintainer lane)
  • Public audience mode: oss maintainer
  • Lane context: Repository is configured for direct PR review.
  • Public profile languages: Python, TypeScript, Ruby, Go, MDX, Shell, Solidity, JavaScript
  • Official Gittensor activity: 14 PR(s), 228 issue(s).
  • PR-specific overlap: none found.
Contributor next steps
  • Start here: Treat this as maintainer-lane context rather than normal contributor-lane activity.
Signal definitions
  • Related work = same linked issue, overlapping active PRs, or title/path similarity.
  • Change scope = cached public metadata such as size labels, draft state, and review-burden hints.
  • Validation posture = whether the PR provides enough public validation/test evidence for maintainer review.
  • Contributor workload = public contributor activity and cleanup pressure, not a repo-wide quality failure.
  • Contributor context = public GitHub/Gittensor identity context; non-Gittensor status is not a blocker.
🧪 Chat with LoopOver

Ask LoopOver a question about this PR directly in a comment — grounded only in the same cached, public-safe facts shown above, never a new claim.

  • @loopover ask &lt;question&gt; answers contribution-quality Q&A with source citations and freshness.
  • @loopover chat &lt;question&gt; answers in natural prose from cached decision-pack facts via local inference (maintainer/collaborator; read-only).
  • A plain-language @loopover mention with a real question is routed to the closest matching read-only command automatically — no exact syntax required.

Full command reference: https://loopover.ai/docs/loopover-commands

🧪 Experimental — new and may change.

🟩 Safe / merged · 🟦 Advisory · 🟨 Held for review · 🟥 Blocked / closed


💰 Earn for open-source contributions like this. Gittensor lets GitHub contributors earn for the work they already do — register to start earning →.

Checked by LoopOver, a quiet PR intelligence layer for OSS maintainers.

  • Re-run LoopOver review

@loopover-orb loopover-orb Bot added the manual-review Gittensor contributor context label Jul 23, 2026
…ill family (#8277)

Replays the deterministic slop scorer over the archived raw-context diffs
(the #8130/#8170 corpus manifest) and synthesizes provenance-tagged
slop_gate_score fired/override pairs labeled by realized outcomes — the
phase-1 'mapping a' counterfactual framing, applied to the #8224 knob.

Signal-subset honesty is enforced by ALLOWLIST: only the three
diff-derivable signals may contribute (the scorer fires
empty_pr_description on an UNDEFINED description — right live, inflating
here), risk/band recompute from exactly their weights, and every row
records the computed signal codes beside the provenance tag. Zero GitHub
traffic; idempotent upserts; wrangler/--pg dual path.

Applied + verified on both stores (460 fired + 460 override rows each):
distribution zero 165 / low 3 / elevated 292 / high 0 — all mass below the
0.60 ceiling, so the rows are NEUTRAL to ladder comparisons by construction
(they feed the reliability curve and drift reads without being able to
fabricate proposals; the discriminating source for the ladder band remains
live full-signal capture).
Maps the #8221 harness's cached per-fixture verdicts onto ProviderReviewSignal
(would_flag => fail, would_not_flag => pass, abstentions yield NO signal) and
aggregates with computeProviderTrackRecords — the identical function live
reviewer_vote rows feed. Offline report only: nothing persists, and the
rendered table carries the replay-derived disclaimer (#8278's segregation
rule). First same-seed two-provider table recorded on #8229.
@JSONbored
JSONbored force-pushed the feat/slop-corpus-backfill branch from 78732e6 to e564d71 Compare July 24, 2026 00:24
…ver process.exit

Both wrappers export pure helpers the unit suites import; an import-time
main().then(process.exit) fails the whole vitest run as an unhandled
rejection even with every test green (CI shard 2's exact failure — all
6568 tests passed). The audit-quality-gate-min-score.ts direct-execution
idiom fixes it; dry-run CLI behavior unchanged.
@JSONbored
JSONbored merged commit aae63ae into main Jul 24, 2026
14 checks passed
@JSONbored
JSONbored deleted the feat/slop-corpus-backfill branch July 24, 2026 01:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

gittensor:bug Gittensor-scored bug fix — scores a 0.05x multiplier. manual-review Gittensor contributor context

Projects

None yet

Development

Successfully merging this pull request may close these issues.

calibration: backfill the slop_gate_score corpus by deterministic replay over archived diffs

1 participant