docs(151): amend the TR-3 base rate — 25.3% is one sample, not the rate - #208
Conversation
An independent run over a different 40-paper sample gives 491 native table pages and TR-3 firing on 68 (13.8%), against this note's 245 pages and 62 (25.3%). Neither is wrong: this note samples every 10th paper from the first 400 by name; the other used a different list. The spread is sampling variance across two small samples of a library where table density varies sharply between papers. This matters because 25.3% had started being quoted as settled in the #204 argument -- it is the number the 'TR-3 is the better next hour' case rests on. The relative findings here are within-sample comparisons and stand, including the 31-page non-overlap with the shape gate. Separately recorded: the is_numeric_token notation gaps (#206/#207) explain only 2 of 68 firings, measured before/after on the same papers, so sampling cannot account for that result.
|
Warning Review limit reached
Next review available in: 51 minutes You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Code Review by Qodo🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0)
Great, no issues found!Qodo reviewed your code and found no material issues that require reviewTip of the day💡 Did you know, you can describe a rule in plain language on the Rules page and Qodo drafts it for you |
PR Summary by QodoDocs: clarify TR-3 base rate is sample-specific, not a settled global rate
AI Description
High-Level Assessment
Files changed (1)
|
There was a problem hiding this comment.
1 issue found across 1 file
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="docs/log/2026-08-14_gh151-b1-predicate-design.md">
<violation number="1" location="docs/log/2026-08-14_gh151-b1-predicate-design.md:126">
P3: The stated range "mid-teens to mid-twenties percent" does not include the lower of the two samples this same amendment cites: 68/491 = 13.8%, which is low teens, not mid-teens. Since the edit's whole point is honest, non-overprecise reporting, adjust the range (e.g. "low-teens to mid-twenties") so it actually spans both measured values.</violation>
</file>
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Re-trigger cubic
| > the local papers library, first 400 by name", the other used a different list. The 25.3% vs | ||
| > 13.8% spread is sampling variance across two small, differently-drawn samples of a library | ||
| > where table density varies sharply between papers. The honest statement is that TR-3 fires on | ||
| > **somewhere in the mid-teens to mid-twenties percent** of native table pages. |
There was a problem hiding this comment.
P3: The stated range "mid-teens to mid-twenties percent" does not include the lower of the two samples this same amendment cites: 68/491 = 13.8%, which is low teens, not mid-teens. Since the edit's whole point is honest, non-overprecise reporting, adjust the range (e.g. "low-teens to mid-twenties") so it actually spans both measured values.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At docs/log/2026-08-14_gh151-b1-predicate-design.md, line 126:
<comment>The stated range "mid-teens to mid-twenties percent" does not include the lower of the two samples this same amendment cites: 68/491 = 13.8%, which is low teens, not mid-teens. Since the edit's whole point is honest, non-overprecise reporting, adjust the range (e.g. "low-teens to mid-twenties") so it actually spans both measured values.</comment>
<file context>
@@ -117,6 +117,21 @@ computed on every table page and consumed by nothing — which is verbatim B1's
+> the local papers library, first 400 by name", the other used a different list. The 25.3% vs
+> 13.8% spread is sampling variance across two small, differently-drawn samples of a library
+> where table density varies sharply between papers. The honest statement is that TR-3 fires on
+> **somewhere in the mid-teens to mid-twenties percent** of native table pages.
+>
+> The *relative* findings in this section are within-sample comparisons and are unaffected —
</file context>
| > **somewhere in the mid-teens to mid-twenties percent** of native table pages. | |
| > **somewhere in the low-teens to mid-twenties percent** of native table pages. |
Summary
Amends the TR-3 base rate in
docs/log/2026-08-14_gh151-b1-predicate-design.md§2. 25.3% is one sample, not the rate.Scope: one design note, one added caveat block. No code.
Why
An independent run over a different sample gives a materially different figure:
/tmp/tr3_measure(#205)/tmp/b1probe/list.txtNeither count is wrong. Different sampling rules → different papers → different page counts, and the rate spread is sampling variance across two small, differently-drawn samples of a library where table density varies sharply between papers.
This matters because 25.3% had started being treated as settled. It is the number the "TR-3 is the better next hour" case rests on in the ranked proposals on #204. Reasoning to one decimal place from a 32-paper sample is over-reading it.
The honest statement, now in the note: TR-3 fires on somewhere in the mid-teens to mid-twenties percent of native table pages.
What is unaffected
is_numeric_tokennotation hypothesis is still refuted: bug(tables): is_numeric_token rejects .034 and U+2217 significance stars, inflating TR-3's firing rate #206/fix(206): accept leading-decimal values and Unicode significance marks #207 explain only 2 of 68 firings, measured before/after on the same papers, so sampling cannot account for that result.Process note
Both measurements recorded their corpus loosely enough that reconciling them took a re-read of the note's preamble. Any future measurement in this area should state the sampling rule, the file list and the page count up front — the amendment says so.
Test plan
Documentation only. One file, one added block, no code paths touched.
Summary by cubic
Clarifies the TR-3 base rate in
docs/log/2026-08-14_gh151-b1-predicate-design.md§2: 25.3% is from one 32-paper sample, not “the” rate; adds an amendment citing a 40-paper run at 13.8% and reframes the claim to “mid-teens to mid-twenties percent.” This matters because 25.3% was being treated as settled in #204 when prioritizing TR-3.Review notes
is_numeric_tokengaps (bug(tables): is_numeric_token rejects .034 and U+2217 significance stars, inflating TR-3's firing rate #206/fix(206): accept leading-decimal values and Unicode significance marks #207) explain only 2/68; precision remains unmeasured.Written for commit 14fe27f. Summary will update on new commits.