Skip to content

Releases: 0Smallcat0/report-workflow

v4.39.0

Choose a tag to compare

@github-actions github-actions released this 08 Aug 02:15

A release about measurement and subtraction. Nothing here adds a capability; the
package is 1,342 lines smaller and one stage shorter, and for the first time the
question "is the delivered report any good" has an instrument behind it.

Added — three reports, three axes, one scorer

scripts/run_drone_market_benchmark.py no longer runs the pipeline. It compares
three recorded documents over the same three CSVs and the same task statement:
one written by hand with no tooling, one delivered by a real authored run, and
one written by an AI given the files and asked for the report. That third arm did
not exist before, which meant "better than AI-generated" was a sentence with
nothing behind it.

Three axes. Numeric density is the existing scorer, imported rather than
copied so a change cannot improve one arm and not another. Layout is rules —
scripts/report_axes.py measures whether a heading states a finding or names a
topic, whether a table was introduced by a sentence, and whether paragraphs are
long enough to argue and short enough to read. Argument is an LLM judge
against a rubric fixed before any arm was judged: three votes per arm, median
recorded, every vote citing the passage it rests on, all archived.

The argument axis is deliberately not a rule. Every deterministic proxy for "is
this argued well" measures a shape a document can have while arguing nothing, and
substituting the proxy for the thing is what produced this repository's
over-design in the first place.

The mechanical author went with the live run. It wrote one lead-in per table and
made no argument between them, which isolated the harness from the writer and
guaranteed a zero on the axis the comparison now turns on.

Removed — the ceremony added in 4.38.0

A minimum length on the reason for waiving a built table, a rule that no two
tables may be waived for the same reason, undermines (claim ids the limitations
section qualified), and answers (a mapping from question index to claim id).
The last two never reached the page.

The test was empirical: the same author input was re-run with the gates gone and
the delivered DOCX was byte-identical. Three gates whose removal changes nothing
about the output were three gates measuring compliance.

What is kept is what changes the document: a built table is placed or waived by
name with a reason, and a counter-evidence section carries two claims rather than
one.

Removed — SCHOLARLY_QUALITY, and the one check inside it that was real

672 lines produced an advisory report that QA_GATE read, FINAL_QA summarized and
ARTIFACTS packaged, and that no gate acted on except through one issue type: five
identifiers of this pipeline's own files appearing in the delivered report. That
check now lives in PUBLICATION_NATURALNESS_PASS, which already hard-blocked the
same class of leak, and it applies to every profile rather than to two of seven.

Hard blocks went from 197 to 198 across src/, which is the honest direction:
the gate used to be invisible in that count because it fired through a report
file, and it is now one explicit raise.

Changed — the drafting brief teaches four things it never said

Each was added because the argument judge or the numeric scorer showed the tool
arm losing on it, and none is a gate:

  • a load-bearing claim carries two independent evidence ids, not one figure at
    two cuts;
  • the figure goes in the sentence, not only in the table;
  • plan one claim the data supports but does not state;
  • if a limitation would change the recommendation, change the recommendation.

Measured over two authored re-recordings: stated figures 52 to 89, past the
AI-direct arm's 72; counter-evidence paragraphs 3 to 6; the argument axis from
losing two dimensions to level.

What this release does not claim

The stop condition set for this work is not met. The tool arm wins the layout
axis and the numeric axis against the AI-direct arm and ties the argument axis,
and a tie is not a win. No score was adjusted to make it one.

The argument scores are not impartial and the summary says so: the same agent
wrote the rubric, then the brief rule, then the paragraph satisfying it, then the
vote awarding the point. Each step is defensible and the chain is not. The votes
are archived so an independent judge can re-read the three documents and disagree
with a specific score.

The delivered report has one author and one sample. Five independent acceptance
runs, by someone who did not write this code, remain the outstanding evidence.

v4.38.0

Choose a tag to compare

@github-actions github-actions released this 07 Aug 13:03

Three acceptance runs of one task, same code, same data, ten minutes apart:
1050 body figures and ten tables, then 777 and eight, then 668 and six. The
worst of the three lost to the hand-written control. A reader gets one run, not
the average of three, so the run that loses is the one the tool is judged on —
and nothing in it failed. This release closes the three ways a run could quietly
be the worse one.

Fixed — a table the pipeline had already built could go unmentioned

Seven cross tabulations are computed at intake, identically on every run: the
same seven evidence ids appeared in all three acceptance runs, so the variance
was never the tool's. The three authors placed four of them, then three, then
two. Nothing blocked, nothing was reported, and the document with two tables
fewer looked like a smaller job rather than a lossier one.

OUTLINE_PLAN now refuses an outline until every built table is accounted for:
placed — a claim cites it, or a figure draws on it — or waived by name in
unused_derived_evidence with a reason of its own. Waiving is a real option and
costs one sentence; what is gone is dropping one by omission. The reason must be
at least 20 characters and may not be pasted across two tables, because one
reason covering five tables is the omission wearing a sentence. The outline
brief lists the tables with their ids before the outline is written, and
derived_table_coverage.json records what happened to each.

Added — business_report must say what would weaken it, and answer what was asked

The hand-written control won on something no count could see. It spent a chapter
on seven pieces of counter-evidence, opening by conceding that its own headline
recommendation might be an artefact of which listings carry a sales figure —
coverage falls from 53.8% in the $30–50 band to 13.2% above $800, so "no demand
at the high end" is not something the data supports. It also answered, in its
first line, the two things the brief asked. Three assisted runs did neither,
because nothing asked them to.

business_report now has a required 資料限制與反面證據 section, and required
is not enough on its own — a heading over "limitations apply" would satisfy it.
The section carries at least two claims and declares undermines: the claim ids
elsewhere in the report whose support it qualifies. Naming a claim it carries
itself is refused, so the author has to find a conclusion of theirs the data does
not fully carry.

The conclusion is held to the task statement the same way. The questions are
extracted from user_prompt deterministically — Chinese briefs ask without a
question mark far more often than with one — listed in the outline brief with
their indices, and the conclusion binds a claim to each by index. The author
never retypes the question, so a mis-extraction shows up in the brief rather
than being answered with something else. A statement that asks for work rather
than answers extracts nothing and this costs nothing.

Added — a second quality benchmark, over raw data, against the arm that was ahead

benchmarks/evidence/report_quality_2026-08-06 measures a prose source, where
the harness was already winning. The comparison that mattered was three raw CSVs
and a market question, and it existed only in an output directory that is not in
the repository.

scripts/run_drone_market_benchmark.py records it: the same three CSVs, the same
task statement, the hand-built control checked in as
benchmarks/fixtures/drone_market_unassisted.md (11,670 characters, 703 figures,
13 tables), and a live run of prepare → author → validate → render. Both arms are
scored by the other benchmark's functions, imported rather than copied, so a
scorer change cannot improve one arm and not the other. --check fails on drift;
--write regenerates. The archive records three losses as measured, including
tables 13 to 7 — the tool arm places every built table and registers none of
its own, and that gap is what an author adds.

Fixed — the measurement said cross-file joins reached no conclusion

It said so because it looked for a join marker on the evidence record, where
one never appears: a joined derivation surfaces as E_D_<request id>. Measured
correctly, the three acceptance runs carried nine, ten and ten join-backed
conclusions out of 36, 34 and 28 claims. The feature was working and the
instrument was not.

scripts/measure_report_body_density.py now takes a run directory and reports
it: join-backed conclusions, built tables placed and waived, and derivations
registered by hand, beside the body density it already measured.

v4.37.0

Choose a tag to compare

@github-actions github-actions released this 07 Aug 08:13

Everything three independent acceptance runs turned up, closed in one pass.

Fixed — a clause break was read as a thousands separator, inventing a figure

NFKC folds the fullwidth comma to ASCII, and the thousands-separator branch of
the number pattern then matched across it. 「跳升至 327,500–1,000 為 317」 states
327, then a new clause about the 500–1,000 band, then 317; it came out as the
single number 327,500, which appears in no evidence because it does not
exist. This is the mirror of the 4.35.0 repair — same fold, opposite direction —
and the worse of the two: a suppressed figure blocks a true sentence and the
author notices, an invented one is reported as the author's own claim. The
fullwidth comma is now held apart through the fold, so a real 1,234 still
reads as one number and a clause break never does.

Fixed — 「並非只有一家」 was a count of one

一 before a classifier, after a limiting phrase (只有, 僅有, 不只, 不止, 不僅,
並非), means "a single one" — which the sentence is asserting or denying as a
whole, not counting. 家 still counts companies in 「五家廠商」 and
「併購了一家公司」; the carve-out keys on the phrase in front, not on the
classifier, because the classifier is doing honest work elsewhere. 超過一家 was
already handled as a bound.

Fixed — a half-open band said neither of the values it holds

A band cut at [1, 3) over star ratings is labelled 1–3 and contains 1 and 2.
An author read the table, wrote 「1–2 星」, and the content check refused it: the
digit 2 appeared nowhere in the evidence. The same phrase passed against a
different table solely because that table's column header spelled the range out
— same fact, same words, opposite verdict depending on a header. Bucketed tables
over integral columns now state each band's membership in the evidence text
(1–3=1 至 2). Columns holding fractions get no such claim: writing "0 to 29"
for a band containing 29.99 would be false about the data.

Fixed — a tool that would not launch was reported as a defect in the document

The optional DOCX-to-PNG check found soffice on PATH as a shim whose own
quoted path would not execute. Filed as status: failed with an empty reason, it
sat in the delivery summary of three reports that passed every gate, beside
render.status: pass and an empty issue list. Not being able to run the optional
converter is the same non-finding as not having it installed, which the adjacent
branch already says in a comment: it is now skipped with the reason stated.
strict_visual_render_check still hard-blocks.

Fixed — a failed publish did not say which file to open

publish_report returned the gate text and nothing about repair, and
get_next_action then reports the publish stage, whose write scope is empty by
design — so an author reading both concluded nothing was legally editable.
submit_action is what routes the failure back to the stage that owns the fix
and returns allowed_repair_paths, and that was discoverable only by guessing.
The failure response now says so.

v4.36.1

Choose a tag to compare

@github-actions github-actions released this 07 Aug 07:42

Fixed — the measurement counted the source list as analysis

The acceptance bar takes three numbers off a published DOCX, and the script that
took them split the body from the tail on
參考文獻|參考資料|References|來源清單|Sources — every heading except the one this
pipeline actually emits, which is 資料來源. So the split never fired: three
consecutive runs measured their own generated source list as body text and
reported a body share of 100%, which is a threshold nothing can fail. Measured
properly the same three documents are 74.5%, 66.9% and 71.0%, and their body
figure counts drop from 1200/937/799 to 1050/777/668.

scripts/measure_report_body_density.py now does it, and takes the heading from
SOURCE_LIST_HEADING/SOURCE_LIST_HEADING_ZH rather than a list retyped beside
them, so renaming the heading breaks a test instead of quietly restoring a
perfect score. A document with no recognised tail heading is reported as such
rather than shown as a clean 100% — a missing source list and an unknown heading
give the same number and are not the same thing.

Fixed — a repeated figure outvoted the citation that stated it

The content gate counts a check as failed only when no cited evidence satisfied
it, which is what lets a claim rest on the union of its sources. The counter
incremented once per finding rather than once per evidence, and a claim naming
two bands — 「0–10 則」 and 「10–50 則」 — states the number 10 twice. Against two
citations, the single failing one contributed two votes and reached the
threshold alone, so the claim was blocked while the other citation was stating
the figure outright. It only ever affected multi-citation claims, which is
precisely the case the union rule exists to protect.

v4.36.0

Choose a tag to compare

@github-actions github-actions released this 07 Aug 06:06

Fixed — the grouped form existed and authors still built tables by hand

Two acceptance runs over the same three CSVs registered 22 and then 47
derivations. The second was 41 scalars against 6 tables, and six of those
scalars were mean-and-negative-rate across three price bands: a two by three
table, spelled out one cell at a time. Of the seven cross tabulations the
pipeline had already computed and listed, that run placed none, and the run
before it placed two.

Nothing was wrong with any individual request, which is exactly why nothing
stopped it. Every one returns a working number, and the author never finds out
the table was one call away.

  • Three or more one-cell derivations that differ only by a row filter are
    refused
    , and the refusal writes out the single grouped request that
    replaces them — with a note to supply bucket edges when the filters cut a
    numeric range, since those are the author's to choose. Two of a shape is a
    comparison and is left alone.

  • The brief leads with the tables that are already built. They were listed
    in ledger order among two dozen single-column summaries, which is where they
    were hardest to find and where they carry the most. They now have their own
    section, above the summaries, each with the [TABLE:] marker that places it,
    and the registration guide says to work through that list before registering
    anything.

v4.35.0

Choose a tag to compare

@github-actions github-actions released this 07 Aug 05:37

Fixed — ordinary Chinese prose was being read as measurements

A second independent acceptance run over the fixture CSVs hit sixteen
factuality findings. Six were the gate doing its job, including four ratios the
author had computed from two table cells. Nine were the checker mistaking
Chinese for arithmetic, and the cost was not the blocked claims — it was that
the only way past the gate was to rewrite good sentences into stilted ones with
the figures removed, which is the opposite of what the gate exists to protect.

  • A Chinese numeral in front of a generic classifier counts the prose, not
    the data.
    「自成一個低價高週轉的次市場」 is an indefinite article;
    「受三項資料限制」 counts the list that follows it; 「其餘七個品類」 counts rows
    the reader can already see. All of them were read as the quantities 1, 3 and
    7 and refused for not appearing in the evidence — seven of the sixteen
    findings. 個, 项/項, 件, 種 and 類 after a Chinese-form numeral are no longer
    quantities. The digit form is: 「6 個價格帶」 is still checked, and so are 三筆,
    兩年, 五家 and 三個月, whose classifiers name real things.

  • A bound the sentence sets is not a value it read. 「其餘七個品類沒有一個
    超過 15%」 was refused for stating a 15% the data never states. It does not
    need to — the sentence asserts a ceiling, and the largest value in the table
    it cites is 14.15%. The evidence side has always treated <0.01 as a limit
    rather than a reading; the claim side now does the same for 不超過, 未超過,
    不到, 低於, 至多, 至少 and the 以下/以上/以內 suffixes.

  • 「」 is emphasis at least as often as quotation. A phrase the author had
    explicitly introduced as their own reading — 「應讀成…」 — was refused for not
    appearing verbatim in the source. 「」 now counts as quotation only after a
    reporting verb (評論寫道「…」); 『』 and the ASCII and curly double quotes are
    unchanged.

Adversarial recall holds at 88.6% with zero false positives, and the verdict
hash is unchanged: none of these carve-outs touches a case the corpus tests.

v4.34.0

Choose a tag to compare

@github-actions github-actions released this 07 Aug 01:18

Added — one request now returns a whole table, and two files can be joined first

Every derivation operation returned a single number, so a six-band price table
with three columns cost eighteen registrations. A real run spent 117 of them to
produce three tables; an unassisted write-up of the same three CSVs built
thirteen tables by hand and put 703 numbers in the body against the tool's 238.
The shape of the request was the cost, not the analysis.

register_derived_evidence now takes group_by and measures and returns a
grid — one row per group, one column per measure — registered as a single
evidence entry:

{"id": "price_band_reliability", "source": "products.csv",
 "group_by": {"column": "price", "buckets": [0, 30, 50, 100, 200, 400]},
 "measures": [{"op": "count"},
              {"op": "mean", "column": "rating"},
              {"op": "share", "rows": "rating < 4"}]}

Bucket edges stay the author's. Where a price axis is cut is an analytical
judgement and a tool that guesses it is wrong in a way the reader cannot see.

source also accepts two files with a join, which is the only route to a
finding neither file states alone. Joining 473 reviews back to the product
catalogue on asin shows the $100–200 band averaging 4.09 stars — the worst of
six — from buyers' own words rather than from listing metadata. Rows that find
no partner are counted and reported in the evidence text; a column name present
on both sides is renamed rather than silently overwritten.

Added — the crossings nobody had to ask for

Each categorical column is now crossed with the numeric ones at intake: counts,
share, mean and median per group. Numeric axes are never auto-binned, for the
reason above. Every derived entry records origin (auto or requested), so
how much aggregation the author still had to register by hand is measurable
rather than anecdotal.

A grouped table carries its grid, so [TABLE:<id>] places it in the document
as a real Word table with its provenance underneath, instead of the author
retyping numbers that would then be backed by nothing.

Fixed — the brief described a world with no derived evidence in it

Derived rows are appended to the ledger, so they land past the twenty-row
sample the brief shows — every one of them, always. An author reading only
that table saw a ledger in which the sole citable thing was one product row.
They are now listed in full, in their own section, split by origin, in the
claim brief as well as the drafting brief, each with the exact [CITE:] and
[TABLE:] markers to use; and both briefs now state that
register_derived_evidence exists, with a worked group_by and join example.

Fixed — registering evidence no longer strands the run that registered it

Registering appends to the ledger, which moves the ledger hash, which made
every already-accepted artifact stale — and the stage that owns those files was
behind the current one, so the only advice on offer was to restore the stale
content. A run died there and had to be restarted. Accepted artifacts stamped
with this same job are now re-stamped against the current ledger and the
harness is told the re-stamped file is the accepted one.

The message for that case also prescribed
remap-evidence --from-job <old>, which is advice for a different situation:
there is no old job when the evidence was added minutes ago by the same run.
A stamp naming a genuinely different job still gets the remap advice.

Fixed — the drafting brief omitted the figure placement contract

A planned figure that no section places does not render, and the run came back
expected 3 Word table(s), found 0. The three-part contract — plan entry,
figure_ids in the outline, literal [FIGURE:<id>] marker in the Markdown —
is now stated in 03_section_draft.md, where it was documented only in
reference/figures.md.

Fixed — a Chinese word is not a quantity

A lone 一/兩/三 with no unit after it was read as a number, so 「兩者」,
「三欄」 and 「一致」 became claim values 2, 3 and 1 that no evidence stated,
and correct sentences were blocked. The same character in front of a unit —
三筆, 兩年 — still counts. On the adversarial corpus one block loses its
spurious numeric reason and keeps its real one; recall and false-positive rate
are unchanged.

Fixed — the starter figure plan was the ledger reprinted as a picture

The default plan proposed charts titled title and 10 other measures by asin
whose series were Amazon tracking URLs and thumbnail links. Identifier and link
columns are now dropped from chart candidates — but only when at least two
columns survive, so a reading indexed by sample id or a single column of
observations is still charted — and the derived cross tables are offered ahead
of raw per-record rows.

Fixed — five false blocks the acceptance run walked into

An independent run of the full pipeline over the three fixture CSVs hit nine
factuality blocks. Four were the gate doing its job, including one where the
author's own arithmetic gave 27.20% and the rows gave 27.21%. The other five
were the checker being wrong:

  • A number written straight after a Chinese comma was invisible. The text
    is NFKC-normalised, which folds , to ,, and the lookbehind that stops
    "1,234" matching at "234" then swallowed it. A draft that literally read
    「544 筆商品列,其中 119 筆…」 was reported as not stating 544 — in the single
    most common position a number appears in Chinese prose. Only a comma with a
    digit in front of it is a thousands separator now.
  • 「DJI 一家就佔 92 筆」 was read as the quantity 1 with unit 家. 一 before a
    counter and then 就/獨/便 is a quantifier idiom meaning alone.
  • 「六個價格帶」 was refused against evidence saying 「6組」. 個, 組, 項, 筆 and
    種 state that something was counted and nothing else; the author was being
    required to adopt the pipeline's measure word to describe the pipeline's own
    table. A reading in 座 still does not support a claim in 公噸.

Adversarial recall stays at 88.6% with zero false positives.

Fixed — the briefs stopped describing the ledger they were describing

After twenty-one registrations 01_claim_plan.md still read
Registered by request: none yet and still carried a ledger hash two calls out
of date — the hash the brief tells the author to copy into _contract. The
briefs are now regenerated when evidence is registered, and whatever that
rewrites is re-accepted with the artifacts.

Added — the fixtures the comparison runs on

benchmarks/fixtures/drone_market/ holds the three CSVs the hand-written and
tool-written arms were both measured on: 544 products, 544 classified rows, 473
reviews. The measurement was unrepeatable while they lived in a scratch
directory.

The ledger-determinism repair shipped in 4.33.0 now has a test: the same
requests applied twice must leave the ledger's bytes identical.

4.33.0

Choose a tag to compare

@0Smallcat0 0Smallcat0 released this 06 Aug 23:41

Fixed — the timestamp moved, so the ledger hash moved, so nothing could publish

apply_derived_evidence recomputes each derived unit from the source rows every
run. That is deliberate: nothing on disk is trusted. But the run also minted a
fresh created_at each time, so every derived line was rewritten with only the
timestamp differing. The ledger hash moved with them, and each artifact stamped
against the previous hash was hard-blocked. Publishing chased a hash that changed
every time it was checked, so any run that called register_derived_evidence
could not reach a published document.

A timestamp is not a value. Regenerating it buys no safety and turns the hash
into a moving target, so an evidence_id now keeps the created_at it was first
registered with. The values are still recomputed from the rows on every run.

Found by running the tool end to end over raw CSVs — the path where a report is
written from data rather than from another report.

Note on 4.32.0

4.32.0 was tagged with this defect and never reached PyPI: its release build
failed during a GitHub outage. Nobody could install it. It carries the derived
evidence, tabular-source citation exclusion and currency-prefix parsing added
since 4.31.0, and this release supersedes it.

v4.31.0

Choose a tag to compare

@github-actions github-actions released this 06 Aug 05:44

Ships everything in 4.30.0 below. That version was committed and tagged and
never reached PyPI — its release build failed, the tag was retired rather than
left pointing at a version nobody could install, and the work is released here
instead. Read the 4.30.0 section as part of this release.

Fixed — a sentence that does not state what its claim asserts

A claim asserting "需約 US$500/噸的產品售價才達 10% IRR,capex 約 US$1,000/噸年
產能" was drafted as "需約 US/噸的產品售價 ... 資本支出約 US,000", the amounts
eaten by a shell expanding $500 and $1 as variables. FA checks that claim,
evidence and sentence are linked; FE checks the claim's numbers against its
evidence; nothing checked the leg the reader actually reads. Every gate passed
and it reached the delivered document.

Worse than a wrong number, because the sentence stays fluent. "每噸年產能約
US,000 的資本門檻" reads like finished prose and nobody stops on it; "US$9,999"
would have been caught.

FS compares per claim, not per sentence, against the union of the passages
carrying that claim's citation plus any figure or table its sections carry. A
claim is often written as two or three sentences and a sentence may set up
rather than restate, so demanding that each repeat every figure would recreate
the "copy the source or be blocked" failure FE was just repaired for.

Beside it, a lint for the fingerprint a lost substitution leaves — a currency
marker with no amount. It needs no claim binding, so it also guards prose that
no claim covers. Only the symbol-prefix forms are flagged: US$500/噸 mangles
to US/噸, while USD/噸 is how a column header legitimately names a unit.

Fixed — Chinese writes a rate with the denominator first

每噸 500 美元. The unit reader looked only at what follows a number, which is
where English puts it, so the numerator was found and the denominator dropped —
and a claim stating 500 美元/噸 did not match a draft saying 成本為每噸五百 美元. Nothing about it was specific to Chinese numerals; the Arabic-digit form
failed too.

The third appearance of one failure: a Chinese way of writing a quantity goes
unrecognised, a gate refuses a correct sentence, and the author learns to copy
the source rather than write. So tests/test_chinese_quantity_expressions.py
now holds one corpus run against both readers of these strings, and it earned
its place immediately — the chart reader cannot read 每噸 500 美元 at all,
recorded as a failing-as-asserted gap.

Four shapes it must not join, each with a test: 每 followed by a number rather
than a measure word, two calendar words (每年 3 月 is a date), no number at
all, and a clause boundary between the two.

Fixed — the release, and the checks around it

run_report_quality_benchmark.py --check re-ran the pipeline, and the pipeline
renders a DOCX, which depends on the machine. The check that existed to prove
the result reproduces was the least reproducible thing in the repository. Both
arms are recorded now and --check re-scores fixed documents.

CI installed .[mcp] while every install instruction here says
report-workflow[mcp,render], so it measured a configuration nobody ships. It
now installs what the plugin installs.

scripts/check_version_sync.py compares the four version strings in the tree
and, with --pypi, the tags and published releases. It caught the unpublished
4.30.0 tag on its first real run.

Known — the fallback renderer drops the Sources section

Without pandoc, the python-docx path keeps References and loses the generated
evidence-trace list, so a reader who installed without the render extra gets a
document whose figures trace to nothing. Asserted as a defect in
tests/test_evidence_traceability.py rather than left to be rediscovered.

v4.29.1 — the stage that can fix it is the stage you are sent to

Choose a tag to compare

@0Smallcat0 0Smallcat0 released this 05 Aug 00:34

Fixed — the stage that can fix it is the stage you are sent to

Driving the pipeline over MCP, a sentence cited a ledger row its claim did not
list. The sensible repair is to add the row to that claim — and
claim_matrix.json sits outside the drafts stage's write scope, so the only
move available was to weaken the sentence until it matched the contract. The
document bent to suit the harness.

Authoring failures routed back to themselves unconditionally. The harness
already knows how to rewind: _invalidate_from reopens a stage and everything
after it, and that is what a routed failure does everywhere else. This failure
now routes to the stage that owns the file, so the claim stage reopens and its
write scope comes with it.

Deliberately narrow: the routing matches the one message that gate writes, not
a general rule about stages. A broad match would rewind runs that stopped
exactly where they belong, and the guarantee this harness sells — one stage
writable at a time — is worth more than the convenience. No permission was
widened; the author still writes only the stage they are in.

The rejection text changed with it. It used to end by telling the author the
claim stage was closed to them, which was true when it was written and is not
now; it points at get_controlled_next_action instead of predicting the answer.