Skip to content

Releases: vlr-code/PayneSDD

PayneSDD 2.0.1 — the reviewer first, the full check once and alone

Choose a tag to compare

@vlr-code vlr-code released this 07 Oct 20:10

Fixed

  • At a check point the reviewer goes first and the full check runs once, last
    and alone (DIGEST.md §4, §5).
    In a real 2.0.0 session (Storia) the agent
    started the full check in the background beside the independent reviewer; the
    reviewer came back first, and its fixes and their tests ran while the full
    check still ran — the load turned timing tests red for nothing once in that
    check point (and once earlier that day) — and each round started a new full
    check on code the next fixes then changed: on one check point three full
    checks started beside a reviewer and a fourth alone, in 30 minutes. Now the reviewer reads first,
    its findings are fixed with their own tests, the next round reads only the
    fixes until one finds nothing, and the full check runs once on that code with
    nothing else building or testing on the machine; a red that load could cause
    is re-run once alone and that run counts, and a full check whose code changed under it is
    stopped as stale. No-harm check on the stand's 13-turn dialog, three runs per
    text, one epoch, the reading rule registered first
    (benchmark/checkpoint-order-2026-10-07.json):
    nothing forgotten in any run, the manager held in 3 of 3 runs against 2.0.0's
    2 of 3; input tokens 10.0M against 10.7M and wall time 1864 s
    against 1691 s (each summed over the three runs), reported without a verdict. The
    review tightened the wording after that run (rounds repeat until one finds
    nothing; the load re-run counts); that wording was not run on the stand. The gain itself needs a slow
    background check, which the stand does not have.

PayneSDD 2.0.0 — one short protocol file

Choose a tag to compare

@vlr-code vlr-code released this 07 Oct 07:47

Changed

  • The protocol is one short file. DIGEST.md now holds the whole protocol
    (~1.3k tokens by tiktoken o200k_base, 5,070 bytes): name the tier; write a short contract that puts
    the tricky spots — contradictions, missing data, assumptions, risks — first
    and asks about them; ask before code (one line on Light, a plan block naming
    every irreversible act on Full); prove "done" with the machine check and its
    criterion → check map; break the result (a self-review on Light, an
    independent reviewer on Full); run long work through the manager (each change
    checked by its own tests, the full check and the review at check points, no
    «ok?» for a change the human dictated, nothing dropped); close with a verdict
    and the three-header summary; questions in one block at the end; Joe in a
    few lines. There is no second file to read: the full protocol (AGENT.md,
    ~16.7k tokens) and the rule that a Full task reads it are gone; AGENT.md is
    now a pointer to DIGEST.md. Why: the maintainer asked whether the protocol
    was overdoing it — on the stand the measure of the work itself (outcome)
    stayed inside each pair's acceptance cut across 0.9.8, 1.0.0 and its fix
    (z +0.52, −1.0, −0.48 against cuts of about −2.2)
    (benchmark/verdict-word-2026-10-06.json), while releases moved the ritual
    measures.
    Measured on the local stand in one epoch of 27 tasks, two runs per text
    (benchmark/slim-2026-10-07.json, every
    definition and the reading rules registered before the runs): against 1.0.1
    the one-file text passed the acceptance rule (worst measure: checks run, 38
    of 42 against 40, z −0.85, reached by about 88% of random splits of that
    epoch), flagged contradictions in 8 of 8 against 8 of 8, and used 0.56 of the
    input tokens (13.4M against 24.1M) in 0.63 of the mean wall time per run (100
    s against 159 s) — beyond the bar registered before the run, a bar taken from
    another epoch's same-text halves (this comparison's own noise not measured).
    In a 13-turn dialog, three runs each, nothing was forgotten in any run; the
    manager held in 2 of 3 runs against 1.0.1's 3 of 3 (the miss: 5 reviews after
    the first build where the bar is 4), at 0.23 of the tokens and 0.46 of the
    time. A 12-task first look preceded it (a PASS without an admissible cut;
    contradictions flagged 2 of 4 against 4 of 4 there). The shipped file changes
    the title line and adds a version line to the measured text.
  • Rewritten for one file: README (install is one @DIGEST.md line),
    scripts/payne-check.sh (the protocol's size band is 3,000–7,000 bytes, the
    version stamp now lives in DIGEST.md, AGENT.md must stay a pointer),
    /payne-review, /payne-edit and the payne-quality reviewer, and the
    repo's own CLAUDE.md. The rules for changing the protocol — the decision
    log, the independent review, measurement on the stand, releases, dev mode and
    its inbox — move to MAINTAINING.md; they are not part of the protocol an
    agent runs on a project.
  • Upgrading from 1.x: import DIGEST.md and delete a host line that points
    at AGENT.md (a config that loaded AGENT.md whole must switch); remove a
    payne-spec command copy or link; a copied Stop-hook is no longer maintained;
    a talk-level block still works but its «AGENT.md, TALK LEVEL» pointer is stale;
    dev mode in other projects needs @…/MAINTAINING.md in the host config.

Removed

  • The full protocol text in AGENT.md, ROLES.md, DEPLOYMENT.md, the
    enforced Stop-hook (hooks/) with its behavior suite and
    settings.example.json, scripts/payne-digest-stamp.sh, the /payne-spec
    command and templates/SPEC.template.md. Every rule not in DIGEST.md is
    gone from the protocol, among them: the analyst subagent and the depth
    choice, the look-back before outward acts, the decision log in user projects
    (projects no longer keep .payne/decisions.log), the ratchets, the helper
    subagent for heavy checks, the no-subagent fallback, the talk level, dev mode
    (now in MAINTAINING.md), the measurement-method rules and the long examples.
    All of it stays in the git history (v1.0.1).

PayneSDD 1.0.1 — a task closes with its verdict word again

Choose a tag to compare

@vlr-code vlr-code released this 07 Oct 06:03

Fixed

  • A task closes with its verdict word again (DIGEST.md, Step 6). The
    digest's Step 6 line now opens with «Light and Full close with ONE verdict
    word as the reply's headline ("PASS — ")», the manager's Full
    sentence follows it, and the summary sits «under the verdict word». In 1.0.0
    that line opened with the Full manager sentence, and a Light task — which in
    1.0.0 runs on the digest alone — closed without its verdict word far more
    often (counted after the run: 18 of 1.0.0's 23 misses were Light replies with
    the summary headers but no verdict word). Measured on the local stand in one
    epoch of 0.9.8 as installed (19bab52), 1.0.0 and this fix
    (27 tasks, two runs per text, 162 runs;
    benchmark/verdict-word-2026-10-06.json,
    with every definition and the reading rule registered before the run): the
    acceptance rule REJECTED 1.0.0 against 0.9.8 on all three seeds — the
    closing verdict with its summary came in 21 of 44 runs against 37 (z −3.60;
    random splits of that epoch reach it in 0.07%), and no other measure came
    near the cut (next lowest: strict consent, 8 of 8 → 6 of 8, z −1.51); with
    the fix it came in 39 of 44 (z +4.12 against 1.0.0, above the bar the rule
    set), and the fix passed the acceptance rule against both texts (its lowest
    measure: fabrication avoided, 8 of 8 → 6 of 8, z −1.51, reached by about half
    of random splits; outcome 38 of 50 against 42 for 1.0.0). Its
    consent measures showed no gain: the hint from the earlier nine-task run
    (6 consent skips of 18 against 0) did not replicate. In the same epoch 1.0.0
    used 1.34 of 0.9.8's input tokens in 1.13 of its mean wall time per run, and
    the fix 1.08 and 1.12 of 1.0.0's — inside the bar registered before the run
    (token noise of half-size samples of one text in another epoch — this
    comparison's own noise not measured, and 0.9.8's halves reached 1.65x there;
    the time's noise not measured), except the fix against 0.9.8 on tokens (1.45; a
    pair the rule did not register, read after the run);
    the medians per run were close (278k, 271k and 272k), the excess sitting in
    a few heavy runs (counted after the run). So 1.0.0's token and time gain
    over 0.9.8 is not shown: the earlier epoch's 0.78 and 0.80 did not hold.

Changed

  • The digest's size band moves from 12,300 to 12,400 bytes
    (scripts/payne-check.sh; a loosening of that check, approved by the
    maintainer on 2026-10-07), so the fix ships exactly as measured, the pin
    aside (12,376 bytes). The digest grows to 3,099 tokens always-on (o200k_base; 3,081 in
    1.0.0).

PayneSDD 1.0.0 — a manager for work in progress, and Light on the digest

Choose a tag to compare

@vlr-code vlr-code released this 06 Oct 19:06

Added

  • A manager for work in progress (Step 4), and no consent question for a
    change the human dictated (Step 1.6).
    On Full, the first build gets the full
    cycle (gate, independent pass, PASS); each later change of the same work gets
    only the cheapest check that can see it — its tests — and comes back under the
    headline TESTS GREEN, not PASS: no independent pass, no full gate, no
    decision-log line of its own (the check point writes one [APPROVED] line for
    the changes folded in), and its new tests' red proof waits for the check
    point, where each is shown red by its assertion — never by an error raised
    before it — against the code before the batch, or its subject broken on
    purpose and reverted. The full gate and the independent pass run over every
    change no pass has read at a check point: a handover (a commit or push, a pull
    request, a merge, a release, a publish), the human's "done" (asked once each
    time all that was asked is built, not repeated), three TESTS GREEN changes (a
    chosen bound, not measured; look, timing or wording tweaks judged by eye do
    not count), a change whose own diff alters a hard-floor mechanism (a condition
    or computation that decides money, access, a retry, limit, expiry or
    idempotency rule, a migration, concurrency, an interface published beyond this
    work, a deletion, infra, a secret — not a reporting field, a text that reveals
    nothing, a formatter or an accessor, even inside payment or auth code; in
    doubt, it is a mechanism), before building on what is costly to unwind, or on
    request; the one-line note on each change names that next check point, not
    "done". At a handover or the human's "done" the tier line carries a count —
    "changes 1–N; the last pass read 1–M" — and changes M+1–N are gated and passed
    before the PASS; with nothing unread the close is that one line. Every request
    since the work began is done or named under Remaining as the next task, none
    dropped without the human's word; a separate task asked mid-work first runs
    the running work's check point when a change is unread, else starts at once,
    and stays a Remaining line on every handback until it starts. A PASS's
    evidence names which changes its pass read. On Full a PASS alone never closes
    the work — the human's "done" or a handover does — so a follow-up serving work
    already built is a change of it, not a new Full task; work re-tiered to Full
    shows the Full plan block once before the floor change and has what was built
    under Light read at its first Full check point. A same-work change whose every
    choice the human's words pin takes the directive carve-out on Full too — one
    line "your words are the plan: doing X", no question; a question only when the
    agent would add a fork or a default, and a costly fork, an irreversible act or
    an outward act (with its look-back) still re-enters the consent gate in full.
    The human's own words ("копи правки", "batch my changes") switch the
    three-change bound off. Light keeps its look, timing and wording series, which
    the digest states again. Besides the batching itself, approved with the plan
    on 2026-10-05, three of these reduce checking on Full against the
    previous text — the no-ask carve-out for a dictated change, the inline fork
    sweep in place of the analyst subagent for a fully dictated request (1.5a,
    below), and the red proof waiting for the check point — and are loosenings,
    each approved by the maintainer by name on 2026-10-06. Why: the maintainer
    kept seeing the full check and an independent review run after every change of
    work that was still changing. Measured on the local stand
    (benchmark/check-manager-2026-10-05.json,
    with every definition): on a billing task with eight dictated changes, both
    sides reading the full protocol, two runs per side, the old text asked before
    every dictated change (8 of 8 in both runs; a later pair on the same task,
    check-manager-2026-10-06.json e34_speed1: 8 and 1 of 8) and form 15 before
    none (0 of 8 in both); the old text reviewed every change in one run and closed the other with
    its last change unread by any review; form 15 reviewed at the first build, at
    the retry-policy change and at the close in both runs, but its three-change
    check point fired late — at the fifth and the sixth change instead of the
    third (at the latest the fourth); per-change time, input tokens and model
    calls, new over old, read 0.79, 1.10 and 0.87 — inside the noise (the old
    text's own two runs differed 1.71x in per-change time, 1.30x in tokens and
    1.20x in calls). Caveats: these figures come from forms 13 and 15; the
    shipped wording was read by one small stand run only (e34-verdict1: nine
    Light tasks, no manager work — in the loading-rule entry below; the second
    no-harm epoch below ran the branch text
    before this review, wip20); the form-15 comparison used the old
    text's runs from an earlier same-day epoch, against the stand's methodology
    (an earlier form compared in one epoch read 0.70, 0.96 and 0.83). The digest
    stays inside its size band by shorter wording and by dropping: the NEVER lines
    "self-assign Light past the hard floor" (its Step 0 hard floor) and "silently
    default a costly-to-reverse choice" (its costly-to-reverse paragraph); "tier
    choice has no machine enforcement" and ROLES' definition of a large task
    (AGENT.md and ROLES.md keep both); and its header line ("checksum-pinned to
    AGENT.md", which payne-check still enforces). The persona's closing line now
    fires when the work closes, not the task. AGENT.md grows by 2,327 tokens
    (o200k_base, 14,334 → 16,661), this and the two entries below together.
  • An inline fork sweep for a fully dictated Full request, changes in as few
    steps as they allow, and no work written before the yes (Steps 1.5a, 1.6,
    4).
    When every Behavior line of the drafted contract restates a sentence of
    the request (no UI, no sibling, no default of the agent's on a
    costly-to-reverse fork), the agent runs the fork sweep itself in one pass,
    reports the draft's contradictions, and goes straight to the consent STOP; a
    fork found, or doubt, sends the work to the analyst subagent, which reads the
    request, the draft, the touched files and what a category sends it to (current
    callers, the sibling) — not the whole repo — while its brief still carries the
    categories, the sibling and the pinned decisions. A TESTS GREEN change takes
    as few edits as it allows, its tests run once when green. Before the yes no
    file of the work and no measurement script is written anywhere — a measurement
    runs from the shell. The
    no-ask carve-out on Full covers only a change of work already built, after its
    yes. Why: the maintainer asked for wall-clock speed. Measured on the local
    stand
    (benchmark/check-manager-2026-10-06.json,
    with every definition and each registered reading rule's text): on the long
    billing task, form 17 over the old text per change — model calls 0.75, output
    tokens 0.80, time 0.89 — inside the noise (a difference needs 0.6 or below;
    the output-token bar's noise was not measured on committed data); turn 1 took
    89 and 264 s against 191 and 458 s (descriptive); whole runs used more input
    tokens, 17.5M and 23.9M against 9.9M and 19.4M (raw, no verdict); one form-17
    run closed with its last change unread by any pass, the registered quality
    override, which the close's count line now answers — measured only on the
    dialog task, where an unread change at the close arose in one of two runs. In
    a dialog with a separate task asked mid-work (form 18), neither text forgot a
    request (no benefit shown on forgetting); the new text carried the close's
    count line in 2 of 2 runs, the old in 0 of 2, and ran 3 and 3 reviews after
    the first build against 6 and 2 — fewer in total, not fewer than each old run,
    the miss its registered criterion names. On eight trap tasks, form 18 wrote
    before the yes in 8 of 16 runs against the old text's 2 (REJECT by the
    acceptance rule) — the cause, in the agents' words, was a measurement script
    written to get honest numbers and the no-ask carve-out taken on a first
    message; form 19 read
    5 of 16 against 2 (inconclusive; the acceptance rule PASS), and 4 of those 5
    were scratch scripts outside the project (counted after the run), which the
    shipped text no longer allows. The no-harm epoch (27 tasks, 108 runs, base and
    candidate in one epoch) passed the acceptance rule twice: on the text before
    these speed rules, with no_premature_write as its worst measure (z −2.00), and
    on the branch text before this review (wip20, b6f5160), whose worst decisive
    measure was claim_sourced at z −1.03 (reached by 88% of random splits of that
    epoch) with no_premature_write 40/48 against the old text's 38/48; its runs
    used more input tokens than the old text's (median per run 437k against 297k,
    26.2M against 23.3M in all) in about the same wall time (a mean of 144 against
    151 s per run) — reported raw, no verdict.
  • A Light task runs on the digest; a Full task reads AGENT.md every time (the
    loading rule); two guards on the manager (Steps 0, 1.5a).
    The digest's
    loading rule changes from «every Light/Full task reads AGENT.md before the
    contract» to «every Full task reads AGENT.md before the contract, no
    exception; on Light this digest is the protocol», and the install line becomes
    «on a Full task read it in full before the contract; on Light the digest is
    the protocol». The digest now carries the Light path — the red proof of a
    check you wrote, the check map's run signal and red value, CHANGED LINES NO
    RUN REACHED, the mechanism search after a fixed defect, the re-tier of Light
    work a floor topic turns Full, the decision log read at contract time,
    «not-found is unknown», the re-read before an irreversible step and the gate
    ...
Read more

PayneSDD 0.9.8 — questions and warnings go in one closing block

Choose a tag to compare

@vlr-code vlr-code released this 02 Oct 05:37

Added

  • Questions and warnings go in one closing block. A report from the human:
    questions got lost in long replies. Read back in that session, a question was
    followed by status text and then by several progress notes, and one was not
    understood. A new section, QUESTIONS AND WARNINGS, sits in the core protocol
    before the never-do list and holds for every reply, Trivial and chat included:
    everything that waits for the human's answer or action goes in one block at
    the very end under a «❓ Question:» line (the word in the human's language),
    with nothing after it; until answered it repeats in every later reply, and a
    progress note carries it as one line; when nothing runs in the background and
    the host has a questioning tool, the question goes through it and the block
    holds a pointer; only four things get a «⚠️ Important:» line — an irreversible
    or external act, a spend past its cap, a red or unverified result handed over,
    a risk of losing data. Step 1.6's consent question (its Light example now
    reads «❓ Question: doing X — ok?», since agents copy the example), Step 4's
    tweak handback line, Step 6's fresh-start offer, the dev-mode report,
    /payne-edit §8 and DEPLOYMENT follow. The talk level keeps the block and the
    warning lines whole, so its recipe changes for the first time since 0.8.0: a
    host that copied it should copy it again (now ≈490 tokens, tiktoken
    o200k_base, counted as AGENT.md SETTING IT says). The digest gains a four-line
    paragraph; to stay inside its band it drops a maintainer note and a path hint,
    and the persona's optional status moves into its heading. Evidence, on the
    stand with the scorer and its rule logged before the runs (a reply counts when
    its first turn ends in the ❓ block): every consent reply ended on its
    question; the marker held in 3 of 6 while the Light example lacked it and in 3
    of 3 once the example carried it; a plain answer carried no marker on its
    first turn (2 of 2). Repetition across replies is not measured.

PayneSDD 0.9.7 — the review reads past the ticket

Choose a tag to compare

@vlr-code vlr-code released this 01 Oct 19:06

Changed

  • What a review misses beyond the ticket becomes required output. A field
    report: the same SDK feature, built for two platforms in separate sessions,
    passed every gate and three independent reviews; a lead reading both platforms
    side by side at release found four defects and two spec problems, because
    every check compared the code only with its own ticket. The same rules written
    as prose were tested first on the benchmark stand against blind-authored traps
    and did not move it (decision log, 2026-09-30 and 2026-10-01). On the easy
    traps the protocol without them already caught the repo rule and the sibling
    divergence in every run. On two harder ones — a rule that lives only in an
    AGENTS.md the CLAUDE.md merely links to, in a large repository, and an
    undocumented subscription failure — they made no measured difference. Counted
    after those runs, not a registered measure: on the large-repository trap 4 of
    10 runs with the prose rule opened AGENTS.md (a tool call naming it) against 0
    of 10 without it, and all four followed the rule; on the subscription trap the
    task ran Light in 19 of 20 runs (the first tier word of the first reply, any
    case) and no run named more than one edge-sweep category (category words
    matched in that reply; rules and counts in the stand's evidence file). So the
    fix is now output the agent must show. Every Light and Full contract opens
    with a rule-files line: AGENTS.md, CLAUDE.md and CONTRIBUTING, looked for at
    the root and in the touched directories with the in-repo files they link to,
    read, or "none found". The edge sweep is shown as one row naming every
    category with its case or "—", and gains trigger (who or what starts the
    behavior, and when: at start or at a lazy first use) and setup failure. Step 5
    looks for the rule files itself instead of taking the author's list, reads the
    change through the lenses that apply, and a Light self-pass ends with one lens
    row. On SDK or multi-platform work the human is asked, before the analyst
    runs, whether the feature exists or is planned on another platform, and a
    named sibling is diffed. /payne-review, /payne-spec, ROLES.md and the SPEC
    template carry the same; the digest gains the rule-files line, the sweep row,
    the lens row and the sibling question, so its size band moves from 10,500 to
    10,800 bytes (a loosening of that check, approved). Evidence: a pre-release
    check on the stand (5 runs on each of two traps, its rule logged before the
    run) did not pass as registered — it looked for a literal «Rule files:» label
    and for the sweep row, and found neither. Read after that verdict, not a
    registered measure: on the large-repository trap all 5 runs opened AGENTS.md
    and kept its rule (no new lock), reporting the lookup in their own words,
    against 3 of 10 without these rules and 4 of 10 with the earlier prose (an
    earlier epoch); the sweep row appeared in none of the 10 runs, and the
    subscription failure was still missed in 5 of 5.

PayneSDD 0.9.6 — a look-back before anything goes out

Choose a tag to compare

@vlr-code vlr-code released this 30 Sep 08:44

Added

  • A look-back before anything goes out. Step 1.6's plan block already named
    the irreversible and external acts; now each one others will see under a name
    or in a form — a release, deploy or upload, a tag, a pull request, a commit to
    a shared branch, a message — carries one more line: how it will look, how the
    previous ones of its kind look from outside and where they were seen, and the
    differences. The agent reads the published result, not the script that made
    it, and only reads. A name left to a tool's default is said to be one. A
    previous one it cannot see is reported as not seen, with a request for a
    reference; a yes without one covers going ahead. A known difference the
    approved plan did not name goes back to the human before the act. The trigger
    was one live report: the plan for an SDK release named the publish command and
    got its yes, nobody looked at how the last release appeared in the package
    registry, and the new one went out under a label the tool generated instead of
    the one the team's releases carry. The digest's Step 1.6 line now names the
    irreversible acts and the look-back: 341 more bytes (wc -c). /payne-edit §7
    reads the last two GitHub releases before asking to cut one. The README and
    DEPLOYMENT token figures are recounted for the longer files. The benefit is
    not measured: no stand task contains a publish.

PayneSDD 0.9.5 — three token-economy rules

Choose a tag to compare

@vlr-code vlr-code released this 17 Sep 08:57

Added

  • Three token-economy rules. Every model call carries the whole
    conversation, so all of it is read again at every step: in the seven largest
    sessions of one maintainer's week, 96.8–99.5% of the main agent's input
    tokens were cache re-reads, at 346K–560K tokens a step on average (input =
    input + cache creation + cache reads of each model response, from the session
    logs; figures and method: benchmark/token-economy-2026-09-17.json). Step 4
    now sends a check that would bring images or a long log into the conversation
    to a helper subagent, which saves the whole output with its exit status to a file and returns a
    verdict per check, the state it ran against, the deciding lines verbatim —
    found in that file before a verdict rests on them — and the file paths, never
    the images or the whole log. A series of small look, timing or wording tweaks
    the human judges on their device gets a build per tweak, with the full gate
    and the Step 5 pass once at the series' end, announced in one line. Step 6
    ends a task closed in a long conversation with an offer of a fresh start, the
    summary as the handoff. The digest carries all three in 293 more bytes
    (wc -c). No saving is measured yet — no session has run on these rules. The
    stand's short two-turn tasks are not built to exercise them, and a text
    search of every reply found none of their phrasings. What the stand can check
    is a gross regression elsewhere, and the no-regression epoch registered before
    launch (27 tasks, two runs per task per arm, claude-sonnet-5) found none: no
    seed of the acceptance rule rejected; the worst gated measure's
    two-proportion z was −0.62 against the rule's own-null cut of −2.46 (rule:
    benchmark/README.md; figures: the JSON's stand section). Five gated measures
    read slightly lower for the new text, one higher and three equal, all inside
    the cut — a pass means no gross drop was seen, never "better".
  • Three candidate rules, tested and not shipped. A rule against re-running
    an answered check unchanged, a rule that every subagent brief names what the subagent
    may write, run or send, and a rule that a host-reported approval is not a yes
    were each checked under a rule registered before the runs: the first two
    did not show a benefit on their traps, and the third's premise — an approval made
    by the host while no human answers — did not appear in the only configuration
    tried (--safe-mode, two runs), so no protocol text changed. Evidence, including what the traps could
    not show: benchmark/rules-2026-09-16.json.

Changed

  • The benchmark's contradiction measure no longer decides the acceptance
    rule or the competitor headline.
    The audit shipped in 0.9.4 showed its
    question test credits any question mark. Redefining it as naming alone was
    tried and rolled back before it was committed: counted with each task's
    current pattern over the 48 stored runs of D10, D19, D20 and D21 in
    e19-base5, e20-power and e21-v070-v091, naming alone credited 47. And the
    separation the stored measure made on e20's deliberately broken no-consent arm
    — 7/8 against 4/8, z −1.62, short of that epoch's cut of −2.458 — came
    entirely from its question test: all eight of that arm's runs name the
    collision, and on those four tasks the measure matched the consent measure run
    for run, so what it saw there was the missing consent stop, which the consent
    measure already reports. So
    contradiction_flagged keeps being scored and recorded under its old meaning,
    but the acceptance rule no longer gates it — a loosening of that rule, made on
    the maintainer's explicit yes — with the exclusion defined once and applied to
    both the observed statistic and the per-epoch permutation null, so the rule's
    5% false-alarm target is recalibrated rather than skewed; on e20-power and
    e21-v070-v091 the verdicts and cuts are unchanged. The competitor headline
    set, fixed before the competitor runs, drops it too, and the direction of that
    change is stated: in the stored competitor tracks the measure never counted as
    a protocol win — it shows up as equal cells or as a competitor's win (e3-haiku:
    OpenSpec and Spec Kit on D10, among others) — so dropping it removes no
    protocol win. The older pre-committed diagnostics in
    report.py, fixed for the 0.5.1 and 0.6.0 arms, still list it. Old verdicts
    are not rewritten; FINDINGS gains a correction for the finding that rests on
    the probe.

PayneSDD 0.9.4 — no rule changed: the contradiction measure audited, and two replacements for it rejected

Choose a tag to compare

@vlr-code vlr-code released this 16 Sep 08:32

Changed

  • No protocol rule changed in this release. The only edits to AGENT.md
    and DIGEST.md are the version stamp and the checksum that pins one to the
    other; ROLES.md, the commands, the agent and the gate scripts are
    byte-identical to 0.9.3. What ships is the evidence behind one measure of the
    benchmark, and the version moves only so that evidence has a released number
    to cite.

Added

  • The "found the contradiction" measure, audited
    (benchmark/contradiction-audit-2026-09-16.json).
    Three blind label passes of one model — self-consistency, not independent
    judges — labelled the sixteen runs of the four contradiction tasks. Both arms
    named the collision in 8 of 8 runs; they asked which rule wins in 4 of 8 and
    5 of 8, against the probe's 8 of 8 and 6 of 8. Against that label the probe
    has no false misses and five false credits — the whole of the dip published
    in 0.9.2's comparison, and an artifact of its question test: it counts any
    question mark in the turn-1 text as the question. Nothing was re-scored; the
    stored verdicts reproduce exactly.
  • Two replacements for that question test, registered before the run and both
    rejected

    (benchmark/probe-fix-attempt-2026-09-16.json).
    The handback shape the consent probe already uses, and that shape plus a
    question sentence about the collision, were frozen and hashed with their
    acceptance bar — 14 of 16 against the blind labels — before a single number
    existed. They reached 10 of 16 and 10 of 16, the second 11 of 16 under the
    wider spelling its own code carried against its frozen text; all three
    readings are below the bar, so the registered reading says nothing ships. The
    scorer is untouched and the axis keeps the test the audit criticised until a
    next cycle redefines it. The handback shape turns out to equal an existing
    measure on all 16 runs and cannot tell "asked which rule wins" from "asked me
    to approve my plan".

Fixed

  • A claim of absence, corrected inside published evidence. The audit file
    said no Python 3.10+ was runnable on the author's machine and built its tools
    around that; an installed arm64 3.12 had never been looked for. The comparison
    was re-run importing the real scorer, with identical numbers, and the file now
    carries the correction at the sentence that made the claim — the protocol's
    own "not-found is UNKNOWN" rule, failed by its author and repaired in the
    record rather than quietly. The stand itself still pins an interpreter that no
    longer runs on that machine; that is recorded in the evidence, not fixed here.

PayneSDD 0.9.3 — a review reads what it gets credit for, and a green binds to what it ran on

Choose a tag to compare

@vlr-code vlr-code released this 16 Sep 05:34

Changed

  • A review describes the state it read (Steps 1.6, 5 and 6). PASS needs the
    gates and the adversarial pass to have read one and the same state, and the
    evidence names it once for both. After fixes the gate re-runs and the pass
    reads what no pass has read yet — a fix, a later commit, a merge's resolution
    — each round counting against the Step 2 budget. Before code is handed over
    for acceptance (a pull request opened for review, a merge, a release, a
    publish) the gate re-runs on the state handed over unless its last run read
    that very state, and a pass reads anything in the change no pass has read; a
    passed manual or device test names no state. Mirrored in /payne-review,
    /payne-edit and the digest's Step 5 and PASS lines. A pull request had
    passed several review rounds and the gate, yet a code commit landed after the
    last review, nobody reviewed it, and an outside reviewer found a hang in it.
  • Changed lines no run reached (Step 4). On code the gate map ends with the
    changed lines or branches that no run executed, read from coverage or log
    output; each is an unverified gap to close or escalate, and a line no output
    can show as run is listed as not shown, a gap like the rest. In that pull
    request the test environment always provided the input whose absence the
    changed path handled, so the path never ran under a green gate.
  • Ratchet the code (Step 4). Fixing a defect in code that had already passed
    a gate or a review, or that a person reported, starts a search of the module
    for the same mechanism, with a verdict for every hit — fixed inside the
    change, surfaced outside it, or fine for a stated reason; an empty search is
    reported as "no hit for in ". A hang fixed at one call site had
    an older sibling with the same wait in the same change, and nobody searched.
  • A green describes the artifact and environment it ran in (Step 4). A check
    on a sibling build or in another shell than the human's proves that sibling:
    run it on the exact file handed over, in the environment it will run in; the
    digest's smoke-launch line carries the artifact half. A debug build had been launched
    while the release file was delivered, and a PATH fix had been checked in one
    shell while the human used another.
  • A compare that can match on nothing is decoration (Step 4). A verify step
    must fail on missing or empty input first: a script hashed a release archive
    that was already gone, so every local hash was the digest of empty input and
    the compare reported a false mismatch; with both sides missing, the same
    compare would have passed.
  • Copies are searched by concept, with wrapped lines joined (/payne-edit §2,
    payne-quality).
    A word search had missed restatements in other words and a
    phrase split across two lines, and each cost a later review round.

Fixed

  • The gate names a broken tool. payne-check and payne-digest-stamp report a
    shellcheck or sha256 tool that is present but cannot run by name, keep the
    gate red and write no pin, instead of reading as a lint failure or a changed
    AGENT.md; the digest size is labelled in bytes.
  • DEPLOYMENT's slim-core figure carries its count. The block is counted with
    tiktoken o200k_base, and the digest's extra over it follows from the README
    token table.
  • /payne-review asks for a green gate on the state a PASS names, and its
    verdict sentence no longer reads as tying the closing summary to PASS alone.