Repository navigation
Releases: vlr-code/PayneSDD
Releases · vlr-code/PayneSDD
Release list
PayneSDD 2.0.1 — the reviewer first, the full check once and alone
Fixed
- At a check point the reviewer goes first and the full check runs once, last
and alone (DIGEST.md §4, §5). In a real 2.0.0 session (Storia) the agent
started the full check in the background beside the independent reviewer; the
reviewer came back first, and its fixes and their tests ran while the full
check still ran — the load turned timing tests red for nothing once in that
check point (and once earlier that day) — and each round started a new full
check on code the next fixes then changed: on one check point three full
checks started beside a reviewer and a fourth alone, in 30 minutes. Now the reviewer reads first,
its findings are fixed with their own tests, the next round reads only the
fixes until one finds nothing, and the full check runs once on that code with
nothing else building or testing on the machine; a red that load could cause
is re-run once alone and that run counts, and a full check whose code changed under it is
stopped as stale. No-harm check on the stand's 13-turn dialog, three runs per
text, one epoch, the reading rule registered first
(benchmark/checkpoint-order-2026-10-07.json):
nothing forgotten in any run, the manager held in 3 of 3 runs against 2.0.0's
2 of 3; input tokens 10.0M against 10.7M and wall time 1864 s
against 1691 s (each summed over the three runs), reported without a verdict. The
review tightened the wording after that run (rounds repeat until one finds
nothing; the load re-run counts); that wording was not run on the stand. The gain itself needs a slow
background check, which the stand does not have.
PayneSDD 2.0.0 — one short protocol file
Changed
- The protocol is one short file.
DIGEST.mdnow holds the whole protocol
(~1.3k tokens by tiktoken o200k_base, 5,070 bytes): name the tier; write a short contract that puts
the tricky spots — contradictions, missing data, assumptions, risks — first
and asks about them; ask before code (one line on Light, a plan block naming
every irreversible act on Full); prove "done" with the machine check and its
criterion → check map; break the result (a self-review on Light, an
independent reviewer on Full); run long work through the manager (each change
checked by its own tests, the full check and the review at check points, no
«ok?» for a change the human dictated, nothing dropped); close with a verdict
and the three-header summary; questions in one block at the end; Joe in a
few lines. There is no second file to read: the full protocol (AGENT.md,
~16.7k tokens) and the rule that a Full task reads it are gone;AGENT.mdis
now a pointer toDIGEST.md. Why: the maintainer asked whether the protocol
was overdoing it — on the stand the measure of the work itself (outcome)
stayed inside each pair's acceptance cut across 0.9.8, 1.0.0 and its fix
(z +0.52, −1.0, −0.48 against cuts of about −2.2)
(benchmark/verdict-word-2026-10-06.json), while releases moved the ritual
measures.
Measured on the local stand in one epoch of 27 tasks, two runs per text
(benchmark/slim-2026-10-07.json, every
definition and the reading rules registered before the runs): against 1.0.1
the one-file text passed the acceptance rule (worst measure: checks run, 38
of 42 against 40, z −0.85, reached by about 88% of random splits of that
epoch), flagged contradictions in 8 of 8 against 8 of 8, and used 0.56 of the
input tokens (13.4M against 24.1M) in 0.63 of the mean wall time per run (100
s against 159 s) — beyond the bar registered before the run, a bar taken from
another epoch's same-text halves (this comparison's own noise not measured).
In a 13-turn dialog, three runs each, nothing was forgotten in any run; the
manager held in 2 of 3 runs against 1.0.1's 3 of 3 (the miss: 5 reviews after
the first build where the bar is 4), at 0.23 of the tokens and 0.46 of the
time. A 12-task first look preceded it (a PASS without an admissible cut;
contradictions flagged 2 of 4 against 4 of 4 there). The shipped file changes
the title line and adds a version line to the measured text. - Rewritten for one file: README (install is one
@DIGEST.mdline),
scripts/payne-check.sh(the protocol's size band is 3,000–7,000 bytes, the
version stamp now lives inDIGEST.md,AGENT.mdmust stay a pointer),
/payne-review,/payne-editand thepayne-qualityreviewer, and the
repo's ownCLAUDE.md. The rules for changing the protocol — the decision
log, the independent review, measurement on the stand, releases, dev mode and
its inbox — move toMAINTAINING.md; they are not part of the protocol an
agent runs on a project. - Upgrading from 1.x: import
DIGEST.mdand delete a host line that points
atAGENT.md(a config that loadedAGENT.mdwhole must switch); remove a
payne-speccommand copy or link; a copied Stop-hook is no longer maintained;
a talk-level block still works but its «AGENT.md, TALK LEVEL» pointer is stale;
dev mode in other projects needs@…/MAINTAINING.mdin the host config.
Removed
- The full protocol text in
AGENT.md,ROLES.md,DEPLOYMENT.md, the
enforced Stop-hook (hooks/) with its behavior suite and
settings.example.json,scripts/payne-digest-stamp.sh, the/payne-spec
command andtemplates/SPEC.template.md. Every rule not inDIGEST.mdis
gone from the protocol, among them: the analyst subagent and the depth
choice, the look-back before outward acts, the decision log in user projects
(projects no longer keep.payne/decisions.log), the ratchets, the helper
subagent for heavy checks, the no-subagent fallback, the talk level, dev mode
(now inMAINTAINING.md), the measurement-method rules and the long examples.
All of it stays in the git history (v1.0.1).
PayneSDD 1.0.1 — a task closes with its verdict word again
Fixed
- A task closes with its verdict word again (DIGEST.md, Step 6). The
digest's Step 6 line now opens with «Light and Full close with ONE verdict
word as the reply's headline ("PASS — ")», the manager's Full
sentence follows it, and the summary sits «under the verdict word». In 1.0.0
that line opened with the Full manager sentence, and a Light task — which in
1.0.0 runs on the digest alone — closed without its verdict word far more
often (counted after the run: 18 of 1.0.0's 23 misses were Light replies with
the summary headers but no verdict word). Measured on the local stand in one
epoch of 0.9.8 as installed (19bab52), 1.0.0 and this fix
(27 tasks, two runs per text, 162 runs;
benchmark/verdict-word-2026-10-06.json,
with every definition and the reading rule registered before the run): the
acceptance rule REJECTED 1.0.0 against 0.9.8 on all three seeds — the
closing verdict with its summary came in 21 of 44 runs against 37 (z −3.60;
random splits of that epoch reach it in 0.07%), and no other measure came
near the cut (next lowest: strict consent, 8 of 8 → 6 of 8, z −1.51); with
the fix it came in 39 of 44 (z +4.12 against 1.0.0, above the bar the rule
set), and the fix passed the acceptance rule against both texts (its lowest
measure: fabrication avoided, 8 of 8 → 6 of 8, z −1.51, reached by about half
of random splits; outcome 38 of 50 against 42 for 1.0.0). Its
consent measures showed no gain: the hint from the earlier nine-task run
(6 consent skips of 18 against 0) did not replicate. In the same epoch 1.0.0
used 1.34 of 0.9.8's input tokens in 1.13 of its mean wall time per run, and
the fix 1.08 and 1.12 of 1.0.0's — inside the bar registered before the run
(token noise of half-size samples of one text in another epoch — this
comparison's own noise not measured, and 0.9.8's halves reached 1.65x there;
the time's noise not measured), except the fix against 0.9.8 on tokens (1.45; a
pair the rule did not register, read after the run);
the medians per run were close (278k, 271k and 272k), the excess sitting in
a few heavy runs (counted after the run). So 1.0.0's token and time gain
over 0.9.8 is not shown: the earlier epoch's 0.78 and 0.80 did not hold.
Changed
- The digest's size band moves from 12,300 to 12,400 bytes
(scripts/payne-check.sh; a loosening of that check, approved by the
maintainer on 2026-10-07), so the fix ships exactly as measured, the pin
aside (12,376 bytes). The digest grows to 3,099 tokens always-on (o200k_base; 3,081 in
1.0.0).
PayneSDD 1.0.0 — a manager for work in progress, and Light on the digest
Added
- A manager for work in progress (Step 4), and no consent question for a
change the human dictated (Step 1.6). On Full, the first build gets the full
cycle (gate, independent pass, PASS); each later change of the same work gets
only the cheapest check that can see it — its tests — and comes back under the
headline TESTS GREEN, not PASS: no independent pass, no full gate, no
decision-log line of its own (the check point writes one [APPROVED] line for
the changes folded in), and its new tests' red proof waits for the check
point, where each is shown red by its assertion — never by an error raised
before it — against the code before the batch, or its subject broken on
purpose and reverted. The full gate and the independent pass run over every
change no pass has read at a check point: a handover (a commit or push, a pull
request, a merge, a release, a publish), the human's "done" (asked once each
time all that was asked is built, not repeated), three TESTS GREEN changes (a
chosen bound, not measured; look, timing or wording tweaks judged by eye do
not count), a change whose own diff alters a hard-floor mechanism (a condition
or computation that decides money, access, a retry, limit, expiry or
idempotency rule, a migration, concurrency, an interface published beyond this
work, a deletion, infra, a secret — not a reporting field, a text that reveals
nothing, a formatter or an accessor, even inside payment or auth code; in
doubt, it is a mechanism), before building on what is costly to unwind, or on
request; the one-line note on each change names that next check point, not
"done". At a handover or the human's "done" the tier line carries a count —
"changes 1–N; the last pass read 1–M" — and changes M+1–N are gated and passed
before the PASS; with nothing unread the close is that one line. Every request
since the work began is done or named under Remaining as the next task, none
dropped without the human's word; a separate task asked mid-work first runs
the running work's check point when a change is unread, else starts at once,
and stays a Remaining line on every handback until it starts. A PASS's
evidence names which changes its pass read. On Full a PASS alone never closes
the work — the human's "done" or a handover does — so a follow-up serving work
already built is a change of it, not a new Full task; work re-tiered to Full
shows the Full plan block once before the floor change and has what was built
under Light read at its first Full check point. A same-work change whose every
choice the human's words pin takes the directive carve-out on Full too — one
line "your words are the plan: doing X", no question; a question only when the
agent would add a fork or a default, and a costly fork, an irreversible act or
an outward act (with its look-back) still re-enters the consent gate in full.
The human's own words ("копи правки", "batch my changes") switch the
three-change bound off. Light keeps its look, timing and wording series, which
the digest states again. Besides the batching itself, approved with the plan
on 2026-10-05, three of these reduce checking on Full against the
previous text — the no-ask carve-out for a dictated change, the inline fork
sweep in place of the analyst subagent for a fully dictated request (1.5a,
below), and the red proof waiting for the check point — and are loosenings,
each approved by the maintainer by name on 2026-10-06. Why: the maintainer
kept seeing the full check and an independent review run after every change of
work that was still changing. Measured on the local stand
(benchmark/check-manager-2026-10-05.json,
with every definition): on a billing task with eight dictated changes, both
sides reading the full protocol, two runs per side, the old text asked before
every dictated change (8 of 8 in both runs; a later pair on the same task,
check-manager-2026-10-06.json e34_speed1: 8 and 1 of 8) and form 15 before
none (0 of 8 in both); the old text reviewed every change in one run and closed the other with
its last change unread by any review; form 15 reviewed at the first build, at
the retry-policy change and at the close in both runs, but its three-change
check point fired late — at the fifth and the sixth change instead of the
third (at the latest the fourth); per-change time, input tokens and model
calls, new over old, read 0.79, 1.10 and 0.87 — inside the noise (the old
text's own two runs differed 1.71x in per-change time, 1.30x in tokens and
1.20x in calls). Caveats: these figures come from forms 13 and 15; the
shipped wording was read by one small stand run only (e34-verdict1: nine
Light tasks, no manager work — in the loading-rule entry below; the second
no-harm epoch below ran the branch text
before this review, wip20); the form-15 comparison used the old
text's runs from an earlier same-day epoch, against the stand's methodology
(an earlier form compared in one epoch read 0.70, 0.96 and 0.83). The digest
stays inside its size band by shorter wording and by dropping: the NEVER lines
"self-assign Light past the hard floor" (its Step 0 hard floor) and "silently
default a costly-to-reverse choice" (its costly-to-reverse paragraph); "tier
choice has no machine enforcement" and ROLES' definition of a large task
(AGENT.md and ROLES.md keep both); and its header line ("checksum-pinned to
AGENT.md", which payne-check still enforces). The persona's closing line now
fires when the work closes, not the task. AGENT.md grows by 2,327 tokens
(o200k_base, 14,334 → 16,661), this and the two entries below together. - An inline fork sweep for a fully dictated Full request, changes in as few
steps as they allow, and no work written before the yes (Steps 1.5a, 1.6,
4). When every Behavior line of the drafted contract restates a sentence of
the request (no UI, no sibling, no default of the agent's on a
costly-to-reverse fork), the agent runs the fork sweep itself in one pass,
reports the draft's contradictions, and goes straight to the consent STOP; a
fork found, or doubt, sends the work to the analyst subagent, which reads the
request, the draft, the touched files and what a category sends it to (current
callers, the sibling) — not the whole repo — while its brief still carries the
categories, the sibling and the pinned decisions. A TESTS GREEN change takes
as few edits as it allows, its tests run once when green. Before the yes no
file of the work and no measurement script is written anywhere — a measurement
runs from the shell. The
no-ask carve-out on Full covers only a change of work already built, after its
yes. Why: the maintainer asked for wall-clock speed. Measured on the local
stand
(benchmark/check-manager-2026-10-06.json,
with every definition and each registered reading rule's text): on the long
billing task, form 17 over the old text per change — model calls 0.75, output
tokens 0.80, time 0.89 — inside the noise (a difference needs 0.6 or below;
the output-token bar's noise was not measured on committed data); turn 1 took
89 and 264 s against 191 and 458 s (descriptive); whole runs used more input
tokens, 17.5M and 23.9M against 9.9M and 19.4M (raw, no verdict); one form-17
run closed with its last change unread by any pass, the registered quality
override, which the close's count line now answers — measured only on the
dialog task, where an unread change at the close arose in one of two runs. In
a dialog with a separate task asked mid-work (form 18), neither text forgot a
request (no benefit shown on forgetting); the new text carried the close's
count line in 2 of 2 runs, the old in 0 of 2, and ran 3 and 3 reviews after
the first build against 6 and 2 — fewer in total, not fewer than each old run,
the miss its registered criterion names. On eight trap tasks, form 18 wrote
before the yes in 8 of 16 runs against the old text's 2 (REJECT by the
acceptance rule) — the cause, in the agents' words, was a measurement script
written to get honest numbers and the no-ask carve-out taken on a first
message; form 19 read
5 of 16 against 2 (inconclusive; the acceptance rule PASS), and 4 of those 5
were scratch scripts outside the project (counted after the run), which the
shipped text no longer allows. The no-harm epoch (27 tasks, 108 runs, base and
candidate in one epoch) passed the acceptance rule twice: on the text before
these speed rules, with no_premature_write as its worst measure (z −2.00), and
on the branch text before this review (wip20, b6f5160), whose worst decisive
measure was claim_sourced at z −1.03 (reached by 88% of random splits of that
epoch) with no_premature_write 40/48 against the old text's 38/48; its runs
used more input tokens than the old text's (median per run 437k against 297k,
26.2M against 23.3M in all) in about the same wall time (a mean of 144 against
151 s per run) — reported raw, no verdict. - A Light task runs on the digest; a Full task reads AGENT.md every time (the
loading rule); two guards on the manager (Steps 0, 1.5a). The digest's
loading rule changes from «every Light/Full task reads AGENT.md before the
contract» to «every Full task reads AGENT.md before the contract, no
exception; on Light this digest is the protocol», and the install line becomes
«on a Full task read it in full before the contract; on Light the digest is
the protocol». The digest now carries the Light path — the red proof of a
check you wrote, the check map's run signal and red value, CHANGED LINES NO
RUN REACHED, the mechanism search after a fixed defect, the re-tier of Light
work a floor topic turns Full, the decision log read at contract time,
«not-found is unknown», the re-read before an irreversible step and the gate
...
PayneSDD 0.9.8 — questions and warnings go in one closing block
Added
- Questions and warnings go in one closing block. A report from the human:
questions got lost in long replies. Read back in that session, a question was
followed by status text and then by several progress notes, and one was not
understood. A new section, QUESTIONS AND WARNINGS, sits in the core protocol
before the never-do list and holds for every reply, Trivial and chat included:
everything that waits for the human's answer or action goes in one block at
the very end under a «❓ Question:» line (the word in the human's language),
with nothing after it; until answered it repeats in every later reply, and a
progress note carries it as one line; when nothing runs in the background and
the host has a questioning tool, the question goes through it and the block
holds a pointer; only four things get a «⚠️ Important:» line — an irreversible
or external act, a spend past its cap, a red or unverified result handed over,
a risk of losing data. Step 1.6's consent question (its Light example now
reads «❓ Question: doing X — ok?», since agents copy the example), Step 4's
tweak handback line, Step 6's fresh-start offer, the dev-mode report,
/payne-edit §8 and DEPLOYMENT follow. The talk level keeps the block and the
warning lines whole, so its recipe changes for the first time since 0.8.0: a
host that copied it should copy it again (now ≈490 tokens, tiktoken
o200k_base, counted as AGENT.md SETTING IT says). The digest gains a four-line
paragraph; to stay inside its band it drops a maintainer note and a path hint,
and the persona's optional status moves into its heading. Evidence, on the
stand with the scorer and its rule logged before the runs (a reply counts when
its first turn ends in the ❓ block): every consent reply ended on its
question; the marker held in 3 of 6 while the Light example lacked it and in 3
of 3 once the example carried it; a plain answer carried no marker on its
first turn (2 of 2). Repetition across replies is not measured.
PayneSDD 0.9.7 — the review reads past the ticket
Changed
- What a review misses beyond the ticket becomes required output. A field
report: the same SDK feature, built for two platforms in separate sessions,
passed every gate and three independent reviews; a lead reading both platforms
side by side at release found four defects and two spec problems, because
every check compared the code only with its own ticket. The same rules written
as prose were tested first on the benchmark stand against blind-authored traps
and did not move it (decision log, 2026-09-30 and 2026-10-01). On the easy
traps the protocol without them already caught the repo rule and the sibling
divergence in every run. On two harder ones — a rule that lives only in an
AGENTS.md the CLAUDE.md merely links to, in a large repository, and an
undocumented subscription failure — they made no measured difference. Counted
after those runs, not a registered measure: on the large-repository trap 4 of
10 runs with the prose rule opened AGENTS.md (a tool call naming it) against 0
of 10 without it, and all four followed the rule; on the subscription trap the
task ran Light in 19 of 20 runs (the first tier word of the first reply, any
case) and no run named more than one edge-sweep category (category words
matched in that reply; rules and counts in the stand's evidence file). So the
fix is now output the agent must show. Every Light and Full contract opens
with a rule-files line: AGENTS.md, CLAUDE.md and CONTRIBUTING, looked for at
the root and in the touched directories with the in-repo files they link to,
read, or "none found". The edge sweep is shown as one row naming every
category with its case or "—", and gains trigger (who or what starts the
behavior, and when: at start or at a lazy first use) and setup failure. Step 5
looks for the rule files itself instead of taking the author's list, reads the
change through the lenses that apply, and a Light self-pass ends with one lens
row. On SDK or multi-platform work the human is asked, before the analyst
runs, whether the feature exists or is planned on another platform, and a
named sibling is diffed./payne-review,/payne-spec, ROLES.md and the SPEC
template carry the same; the digest gains the rule-files line, the sweep row,
the lens row and the sibling question, so its size band moves from 10,500 to
10,800 bytes (a loosening of that check, approved). Evidence: a pre-release
check on the stand (5 runs on each of two traps, its rule logged before the
run) did not pass as registered — it looked for a literal «Rule files:» label
and for the sweep row, and found neither. Read after that verdict, not a
registered measure: on the large-repository trap all 5 runs opened AGENTS.md
and kept its rule (no new lock), reporting the lookup in their own words,
against 3 of 10 without these rules and 4 of 10 with the earlier prose (an
earlier epoch); the sweep row appeared in none of the 10 runs, and the
subscription failure was still missed in 5 of 5.
PayneSDD 0.9.6 — a look-back before anything goes out
Added
- A look-back before anything goes out. Step 1.6's plan block already named
the irreversible and external acts; now each one others will see under a name
or in a form — a release, deploy or upload, a tag, a pull request, a commit to
a shared branch, a message — carries one more line: how it will look, how the
previous ones of its kind look from outside and where they were seen, and the
differences. The agent reads the published result, not the script that made
it, and only reads. A name left to a tool's default is said to be one. A
previous one it cannot see is reported as not seen, with a request for a
reference; a yes without one covers going ahead. A known difference the
approved plan did not name goes back to the human before the act. The trigger
was one live report: the plan for an SDK release named the publish command and
got its yes, nobody looked at how the last release appeared in the package
registry, and the new one went out under a label the tool generated instead of
the one the team's releases carry. The digest's Step 1.6 line now names the
irreversible acts and the look-back: 341 more bytes (wc -c)./payne-edit§7
reads the last two GitHub releases before asking to cut one. The README and
DEPLOYMENT token figures are recounted for the longer files. The benefit is
not measured: no stand task contains a publish.
PayneSDD 0.9.5 — three token-economy rules
Added
- Three token-economy rules. Every model call carries the whole
conversation, so all of it is read again at every step: in the seven largest
sessions of one maintainer's week, 96.8–99.5% of the main agent's input
tokens were cache re-reads, at 346K–560K tokens a step on average (input =
input + cache creation + cache reads of each model response, from the session
logs; figures and method:benchmark/token-economy-2026-09-17.json). Step 4
now sends a check that would bring images or a long log into the conversation
to a helper subagent, which saves the whole output with its exit status to a file and returns a
verdict per check, the state it ran against, the deciding lines verbatim —
found in that file before a verdict rests on them — and the file paths, never
the images or the whole log. A series of small look, timing or wording tweaks
the human judges on their device gets a build per tweak, with the full gate
and the Step 5 pass once at the series' end, announced in one line. Step 6
ends a task closed in a long conversation with an offer of a fresh start, the
summary as the handoff. The digest carries all three in 293 more bytes
(wc -c). No saving is measured yet — no session has run on these rules. The
stand's short two-turn tasks are not built to exercise them, and a text
search of every reply found none of their phrasings. What the stand can check
is a gross regression elsewhere, and the no-regression epoch registered before
launch (27 tasks, two runs per task per arm, claude-sonnet-5) found none: no
seed of the acceptance rule rejected; the worst gated measure's
two-proportion z was −0.62 against the rule's own-null cut of −2.46 (rule:
benchmark/README.md; figures: the JSON's stand section). Five gated measures
read slightly lower for the new text, one higher and three equal, all inside
the cut — a pass means no gross drop was seen, never "better". - Three candidate rules, tested and not shipped. A rule against re-running
an answered check unchanged, a rule that every subagent brief names what the subagent
may write, run or send, and a rule that a host-reported approval is not a yes
were each checked under a rule registered before the runs: the first two
did not show a benefit on their traps, and the third's premise — an approval made
by the host while no human answers — did not appear in the only configuration
tried (--safe-mode, two runs), so no protocol text changed. Evidence, including what the traps could
not show:benchmark/rules-2026-09-16.json.
Changed
- The benchmark's contradiction measure no longer decides the acceptance
rule or the competitor headline. The audit shipped in 0.9.4 showed its
question test credits any question mark. Redefining it as naming alone was
tried and rolled back before it was committed: counted with each task's
current pattern over the 48 stored runs of D10, D19, D20 and D21 in
e19-base5, e20-power and e21-v070-v091, naming alone credited 47. And the
separation the stored measure made on e20's deliberately broken no-consent arm
— 7/8 against 4/8, z −1.62, short of that epoch's cut of −2.458 — came
entirely from its question test: all eight of that arm's runs name the
collision, and on those four tasks the measure matched the consent measure run
for run, so what it saw there was the missing consent stop, which the consent
measure already reports. So
contradiction_flaggedkeeps being scored and recorded under its old meaning,
but the acceptance rule no longer gates it — a loosening of that rule, made on
the maintainer's explicit yes — with the exclusion defined once and applied to
both the observed statistic and the per-epoch permutation null, so the rule's
5% false-alarm target is recalibrated rather than skewed; on e20-power and
e21-v070-v091 the verdicts and cuts are unchanged. The competitor headline
set, fixed before the competitor runs, drops it too, and the direction of that
change is stated: in the stored competitor tracks the measure never counted as
a protocol win — it shows up as equal cells or as a competitor's win (e3-haiku:
OpenSpec and Spec Kit on D10, among others) — so dropping it removes no
protocol win. The older pre-committed diagnostics in
report.py, fixed for the 0.5.1 and 0.6.0 arms, still list it. Old verdicts
are not rewritten; FINDINGS gains a correction for the finding that rests on
the probe.
PayneSDD 0.9.4 — no rule changed: the contradiction measure audited, and two replacements for it rejected
Changed
- No protocol rule changed in this release. The only edits to
AGENT.md
andDIGEST.mdare the version stamp and the checksum that pins one to the
other;ROLES.md, the commands, the agent and the gate scripts are
byte-identical to 0.9.3. What ships is the evidence behind one measure of the
benchmark, and the version moves only so that evidence has a released number
to cite.
Added
- The "found the contradiction" measure, audited
(benchmark/contradiction-audit-2026-09-16.json).
Three blind label passes of one model — self-consistency, not independent
judges — labelled the sixteen runs of the four contradiction tasks. Both arms
named the collision in 8 of 8 runs; they asked which rule wins in 4 of 8 and
5 of 8, against the probe's 8 of 8 and 6 of 8. Against that label the probe
has no false misses and five false credits — the whole of the dip published
in 0.9.2's comparison, and an artifact of its question test: it counts any
question mark in the turn-1 text as the question. Nothing was re-scored; the
stored verdicts reproduce exactly. - Two replacements for that question test, registered before the run and both
rejected
(benchmark/probe-fix-attempt-2026-09-16.json).
The handback shape the consent probe already uses, and that shape plus a
question sentence about the collision, were frozen and hashed with their
acceptance bar — 14 of 16 against the blind labels — before a single number
existed. They reached 10 of 16 and 10 of 16, the second 11 of 16 under the
wider spelling its own code carried against its frozen text; all three
readings are below the bar, so the registered reading says nothing ships. The
scorer is untouched and the axis keeps the test the audit criticised until a
next cycle redefines it. The handback shape turns out to equal an existing
measure on all 16 runs and cannot tell "asked which rule wins" from "asked me
to approve my plan".
Fixed
- A claim of absence, corrected inside published evidence. The audit file
said no Python 3.10+ was runnable on the author's machine and built its tools
around that; an installed arm64 3.12 had never been looked for. The comparison
was re-run importing the real scorer, with identical numbers, and the file now
carries the correction at the sentence that made the claim — the protocol's
own "not-found is UNKNOWN" rule, failed by its author and repaired in the
record rather than quietly. The stand itself still pins an interpreter that no
longer runs on that machine; that is recorded in the evidence, not fixed here.
PayneSDD 0.9.3 — a review reads what it gets credit for, and a green binds to what it ran on
Changed
- A review describes the state it read (Steps 1.6, 5 and 6). PASS needs the
gates and the adversarial pass to have read one and the same state, and the
evidence names it once for both. After fixes the gate re-runs and the pass
reads what no pass has read yet — a fix, a later commit, a merge's resolution
— each round counting against the Step 2 budget. Before code is handed over
for acceptance (a pull request opened for review, a merge, a release, a
publish) the gate re-runs on the state handed over unless its last run read
that very state, and a pass reads anything in the change no pass has read; a
passed manual or device test names no state. Mirrored in/payne-review,
/payne-editand the digest's Step 5 and PASS lines. A pull request had
passed several review rounds and the gate, yet a code commit landed after the
last review, nobody reviewed it, and an outside reviewer found a hang in it. - Changed lines no run reached (Step 4). On code the gate map ends with the
changed lines or branches that no run executed, read from coverage or log
output; each is an unverified gap to close or escalate, and a line no output
can show as run is listed as not shown, a gap like the rest. In that pull
request the test environment always provided the input whose absence the
changed path handled, so the path never ran under a green gate. - Ratchet the code (Step 4). Fixing a defect in code that had already passed
a gate or a review, or that a person reported, starts a search of the module
for the same mechanism, with a verdict for every hit — fixed inside the
change, surfaced outside it, or fine for a stated reason; an empty search is
reported as "no hit for in ". A hang fixed at one call site had
an older sibling with the same wait in the same change, and nobody searched. - A green describes the artifact and environment it ran in (Step 4). A check
on a sibling build or in another shell than the human's proves that sibling:
run it on the exact file handed over, in the environment it will run in; the
digest's smoke-launch line carries the artifact half. A debug build had been launched
while the release file was delivered, and a PATH fix had been checked in one
shell while the human used another. - A compare that can match on nothing is decoration (Step 4). A verify step
must fail on missing or empty input first: a script hashed a release archive
that was already gone, so every local hash was the digest of empty input and
the compare reported a false mismatch; with both sides missing, the same
compare would have passed. - Copies are searched by concept, with wrapped lines joined (/payne-edit §2,
payne-quality). A word search had missed restatements in other words and a
phrase split across two lines, and each cost a later review round.
Fixed
- The gate names a broken tool. payne-check and payne-digest-stamp report a
shellcheck or sha256 tool that is present but cannot run by name, keep the
gate red and write no pin, instead of reading as a lint failure or a changed
AGENT.md; the digest size is labelled in bytes. - DEPLOYMENT's slim-core figure carries its count. The block is counted with
tiktokeno200k_base, and the digest's extra over it follows from the README
token table. - /payne-review asks for a green gate on the state a PASS names, and its
verdict sentence no longer reads as tying the closing summary to PASS alone.