Releases: askrubberduck/skills
Release list
v2.5.0 — one name per skill
A shape and dry pass over the skill bodies, the descriptions, and the README. Fifteen skills touched; nothing added to the pipeline.
One name per skill
Nine bodies opened with a title of their own — "Decorrelated Superreview", "Obligations Critique Sweep", "Pickable Work Scan" — while the frontmatter, the README table, and the Codex interface all said Duck X. duck-proof had no title at all. Every body now opens with the name you invoke it by.
Details that did not execute as written
duck-prooftold the doer to create the defect ledger on first use and, one sentence later, never to create it silently. Its section 4 ran the baregit diffits own opener says proves nothing on a committed candidate.duck-runsaid commit-first and proof-before-SHA in one paragraph, and produced the receipts twice. The order is stated once: receipts, commit, SHA, review. Its shape and dry beats no longer carry drifted summaries of the skills they invoke.duck-sweepstep 3 deleted every branch with-Dafter step 2 forbade it for preserved ones.duck-runandduck-campaignsaid whatCUTdoes and not whatOWNER DECISIONdoes. Both exits are stated; the campaign's is inferred fromduck-frame's exit semantics.
Dried
Twelve "Measured:" incidents restated the rule beside them. Six keep their fact in present tense, scoped to the evidence; six are gone. A vault layout that had leaked into three skills as examples is generalised. The README's prerequisites have a heading, the Codex route loses an unexplained "(recommended)", and one trigger clause that invocation made redundant leaves duck-run's description.
Minor rather than patch: the name every Codex user sees changed, and two exits were added.
How it was reviewed
duck-proof ran over the candidate and found five defects in it — two rewrap bulges, three dried sentences that had turned one measurement into a law — all fixed before landing. Not gated: the owner waived the duck-review gate and the PR approval on #17 in writing, and the bump landed directly on master as on v2.4.0. CI and the validator's corruption self-test are green on the merged head.
Full changelog: v2.4.0...v2.5.0
v2.4.0 — one home per held fact
A shape pass over the skills, README, and validators, scored in concepts a reader must hold rather than lines: 106 added, 109 removed.
Skills
- One reason had four homes.
duck-run,duck-review, andduck-raceeach re-derived, in their own words, why a receipt goes to the durable records home and never the scratchpad or a candidate-branch commit. They now state the contract and citeduck-proof, which owns it. duck-race's pin proof had drifted. Its copy of the model-pin procedure omitted the roster lineduck-reviewrequires recorded, so a race run from that skill alone produced a receipt that could not prove cross-family. It cites the bar instead of carrying a copy.- Two paragraphs a reader could not stop inside.
duck-runstage 6 fused remediation lanes with the circuit breaker;duck-landstep 1 fused remote policy, history shape, and base pinning. Both split under labelled headings, text unchanged. duck-runsays what trust-touching means at first use; the term was defined only induck-review.
README
- Prerequisites had no heading and sat under "Works well with". Own section, before Install.
- Install uses
ln -sfn, so installing and relinking after a pull are the same command. Measured: a symlink install made before v2.1.0 never gainedduck-shapeuntil relinked by hand. - The skill count leaves the prose; it was a fact maintained by hand.
Validators
render-catalog.py --checkduplicated the validator's staleness check. Removed, with its CI step.- A
REQUIRED_REFERENCESentry naming a retired skill now fails the gate instead of being skipped silently. Mutation-tested. importlibreplaces theexecplus__name__trick; the dotted-entry exemption has one home.
Minor rather than patch: the validator gains a rejection and CI loses a step.
How it was reviewed
Not gated. One model family authored and self-checked it; the owner waived the duck-review gate and the PR approval on #16 in writing, and directed the version bump to land directly on master. CI and the validator's corruption self-test are green on the merged head.
Full changelog: v2.3.0...v2.4.0
v2.3.0 — gate contradictions
A review-only pass over all 19 skills found eight defects, all of them contradictions between
files that read correctly in isolation. This release closes all eight.
Blockers
duck-plan's naming rule broke duck-scan's read-only guarantee. It said "a skill named
inside another skill's step is an instruction to invoke it, not a citation" — absolutely, which
turned duck-scan's mention of duck-cut into a dispatch two lines above where duck-scan says
"the skills do the acting". Nine citations across the corpus sat under the same rule. Now scoped
to imperative steps.
A converging third review round was undispatchable. duck-review refused every N≥3 dispatch
without a committed loop-diagnosis:, but duck-run only minted one when the loop was diverging.
Both now turn on the same predicate, and a converging loop is an exit like the other four.
duck-run's Verify stage inverted its own scratchpad rule — "never the scratchpad, which is
where the gate looks". The relative clause attached to the scratchpad.
Gaps
duck-race's receipt was the only one homed in$SP, which dies before the stage that reads it.- The receipt commit path was unstated corpus-wide. It goes on a records branch, never base — a
receipt landed on base advances whatduck-landre-verifies and voids the authorization the
receipt exists to support. duck-reviewnamed the model-pin proof and declined to enforce it whileduck-racemandated the
same control, deferring to "the soul" — a reference resolving nowhere in the corpus. Now binding
in both.duck-plandeferred a co-author count toduck-review's bar without stating it. Now two.REQUIRED_REFERENCEScovered 1 of 8 load-bearing delegations, so five of the pipeline's own
links could be deleted with the validator green. Each new row is mutation-tested.
Minor rather than patch: behaviour tightened rather than only repaired.
How it was reviewed
Two rounds against a roster-proven different model family. Round 1 rejected the change on a
defect in the fix itself — the rewritten breaker claimed to fire "earlier than" its round count,
which is impossible, since remediation-born blockers need three rounds to occur twice. Deleted
rather than repaired. Round 2 approved.
Reviewed by one family rather than two: the second outaged on credits and the owner waived the
two-family requirement for round 2 only.
Full changelog: v2.2.0...v2.3.0
v2.2.0 — duck-shape
duck-shape
A new skill: shape code by the concepts a reader must hold to change it safely — never line count, never nesting depth.
The distinction that makes it different from a size metric: depth is not the defect. A five-level hierarchy where every level names what it is for is cheap; two levels called Manager and Helper are expensive, and no nesting metric tells them apart. The target is clarity, not flatness.
It cuts both ways, which a lines-of-code lens cannot: a golfed one-liner is fewer lines and more held facts. A regex replacing a named parser is one line, and the reader now holds the grammar.
Top of the measure is "does this concept need to exist at all?" — deleting a concept outranks naming it well. Deleting two hundred lines that were one concept removes one held fact; deleting a single line that was a mode flag removes a held fact from every reader of every branch downstream.
Roles
duck-shape— the doer, at change timeduck-roastangle 2 — the same lens at milestone altitude, with churn as evidence on a finding, never a rankduck-review— reports what change time missed
duck-run: eight stages to six
duck-shape and duck-dry moved from separate stages into the per-unit cycle inside Execute: failing test, minimal pass, shape, dry. They belong there because the diff is not committed yet — the one window where neither costs ceremony. Defer them and shaping becomes a restructure, drying becomes a sweep, and both then need their own commit and their own trip through the gate.
Gate
- Required cross-references must exist, not merely resolve. A load-bearing link could previously be deleted with the whole gate green.
- Skill descriptions are bounded — every host loads all of them, every session.
- A directory under
skills/with noSKILL.mdis reported rather than silently uncounted.
Also
duck-break owns "red team it"; duck-shape owns "duck simplification"; trigger descriptions normalized across the collection.
v2.1.1 — manifest homepage, and the gate green again
A metadata patch and the gate fix it exposed. No skill changed.
What changes for you
homepagein both plugin manifests, pointing at the repository. It is the link a plugin directory listing shows, and it was absent.- The root
plugin.jsonstub is gone. No host reads a root manifest: Claude reads.claude-plugin, Codex.codex-plugin, Agy.agents. - The validator no longer requires that stub. Deleting it left
MANIFESTSnaming a file that no longer existed, sogatefailed on every master push between that commit and this release.
Gate
This release did not pass its own gate. It landed through the OrganizationAdmin bypass of the master ruleset with no duck-review dispatch — the third consecutive release on a waiver. What stands behind it is the mechanical gate only: render-catalog.py --check, validate-distribution.py --self-test, and claude plugin validate --strict, all green on master.
Full log: v2.1.0...v2.1.1
v2.1.0 — the gate rebuilt from the roast's owner decisions
Ten owner decisions from a two-round adversarial roast of the whole collection (Claude Opus 5, GPT-5.6-Sol, Gemini 3.1 Pro; 72 standing findings, stopped before the loop ran dry).
What changes for you
- Receipts move out of the scratchpad.
proof-rN.mdandbreak-rN.mdnow live in the project's durable records home — the gate refuses to dispatch without them, and the session that wrote them has usually already ended. They are never committed on the candidate branch, which would advance the head past the SHA under review. duck-reviewwill not dispatch until the export is authorized. The gate works by sending your repository's contents to model vendors outside your machine. A host that refuses on those grounds has asked the owner's question, so it routes toduck-deciderather than to a retry. If you have not authorized the export for a repository, the gate fails closed and says so.- Reviewer family comes from the harness roster, not a self-report. One binary can serve several families; a harness that prints no roster establishes no family, which is the unknown identity the gate already refuses to count.
duck-runproves before it freezes the candidate SHA.duck-prooffixes what it finds, so the old order left the authorization pointing at code nobody reviewed.duck-campaignowns its continuation by booking the next session per packet, instead of handing off to a driver the collection never shipped.duck-cutrunsduck-scanas its only locator, and a CUT commit now carries the deleted item's text — deletion is the one verdict with no recovery path.duck-proofandduck-drylose the sections that duplicatedduck-runandduck-campaign.- The validator gains two checks: the release version must agree across the three files carrying it, and a skill showing a reviewer dispatch must also state the by-path rule.
Known gaps, stated rather than discovered
Nothing in the distribution says which controls are mechanically enforced and which are discipline. Pin validation and receipt existence are discipline. That disclosure was written and then removed by owner decision; the constraint it satisfied is knowingly unmet.
This release did not pass its own gate. It landed under a written owner waiver of duck-land's APPROVE precondition, merged through the OrganizationAdmin bypass of the master ruleset's last-push-approval rule. Both were recorded before the fact. It is the second consecutive release of gate files on a waiver.
Full log: v2.0.2...v2.1.0
v2.0.2 — the map stops lying
A skill collection routes on two things: the description your agent matches a request against, and
the map you read when you want to know what hands off to what. Both were wrong in v2.0.1.
Seven descriptions now say what their bodies do
A description is the routing table. When it drifts from the body the skill either never fires or
fires on the wrong request, and nothing tells you which.
duck-runsaid "plan, implement, test, and independently review". It also frames, verifies,
and lands. It says so now.duck-drysaid comments and docstrings. It has always set the bar for commit messages and PR
descriptions too.duck-cutclaimed it would "finish every viable item autonomously". It closes, cuts, merges,
and unblocks whatever does not need you — it does not build.duck-planoffered "independent critique of a draft plan". What it does is decorrelated
co-authorship: a second family writes the plan with you, not a review after you.duck-campaignsaid "prioritized" workstreams; it carves independent ones.duck-dietnow names stage routing — which model tier and agent type a stage runs on —
because it decides that too.duck-landdropstagandcut a releasefrom its triggers, because it does neither.
That is a gap, not a cleanup: "cut a release" now routes to nothing, and this release was cut by
hand.
The map covers every skill
duck-why and duck-dry were absent from it. duck-decide was drawn off to one side when it is a
REJECT exit. Campaign planning sat inside each packet, when a campaign plans every packet before
any build starts.
Two vertical diagrams now, one per line — a single change through duck-run, a backlog through
duck-campaign — with the conditional edges beside the stage they leave: the CUT exit, duck-break
before proof, the diverging-loop breaker. The map moved above the skills table, because the map is
what groups the skills and the table is for alphabetical lookup.
duck-run names a cause before it patches
On REJECT the pipeline routes through duck-why when a blocker reports a symptom and the defect
behind it is not already obvious from the diff. A reviewer tells you where it hurt, not where it
broke.
The gate
Shipped under an owner waiver. Two families reviewed this span with no outage —
gemini-3.1-pro-high (Google) and gpt-oss-120b-medium (OpenAI) — and both returned REJECT.
Two of their three blockers died on evidence: a claimed CI failure the run log shows did not happen,
and a claimed contradiction in the duck-why routing that duck-why's own text answers. One
survived adjudication — this span removes the README section stating which rules are mechanically
checked and which are only discipline, under a justification that covers a different fact. The owner
read the finding and chose to ship without restoring it.
codex was unavailable for the third gate running (ERROR: Your workspace is out of credits), so
both reviewers came from one harness under different pinned models. The pin was proved before use:
a deliberately invalid --model errors with the full roster.
v2.0.1 — the blockers the gate never got to raise
v2.0.0 shipped on a NOTE because one of its two required reviewers outaged. That reviewer was re-run afterwards and found two real blockers. This closes both.
The outage was fixable and the fix was already recorded: agy's headless failure mode is a permission-denied tool call provoked by a prompt that invites shell, and the v2.0.0 review prompt was full of git vocabulary — SHAs, a PR number, a commit range. Stripped of it, the same model reviewed the same material without incident. v2.0.0 did not have to ship unreviewed.
Blocker 1 — a numeric stage reference survived
duck-run still said "Stage 1's design" after v2.0.0 added the rule against referring to its own stages by number. The grep behind the claim it was gone was case-sensitive and returned zero, so the receipt asserted the opposite of what shipped.
Blocker 2 — a rule contradicted by the example beneath it
duck-review and duck-race both stated "never inline", then demonstrated $(cat …) dispatch examples that inline.
Fixed by narrowing the rule to the evidence it cites rather than softening it. The 28-of-41 verdict-flip measurement compared a pasted corpus against one read from disk — it never tested the instruction prompt. Both files now say the prompt is the command's argument and the material it refers to is a path the reviewer opens, which is what the measurement supports and what every dispatch in this session actually did.
The gate
Two families, both APPROVE, no outage: gemini-3.1-pro-high (Google) and gpt-oss-120b-medium (OpenAI). Model pins proved by rejection before use, inputs passed by path, one scratchpad directory per reviewer — the rules v2.0.0 added, exercised on v2.0.1.
The merge still used an admin bypass of the require_last_push_approval rule, recorded in the outcome. The review gate itself was not waived this time.
v2.0.0 — the duck asks why
BREAKING: /duck-pingpong is deleted with no alias. Its capability survives as duck-race rally mode; the command does not.
duck-race now has two modes
Race — two families attempt the problem in parallel, independently, and executed evidence picks the winner. Rally — they alternate, one writing a failing test and the other satisfying it. The eight machinery blocks the two skills each carried now have one home. Fixes a prose/code mismatch where the text said "dispatch in the background" while the block captured the diff synchronously, recording empty attempts as forfeits.
duck-why is new
Names the cause of a failure and hands the fix on, holding the same doer/judge separation as duck-break. Of eighteen skill descriptions, exactly one mentioned a bug or a failing test — so "fix the layout bug" matched nothing. It does now.
Every gate corrected
Each from a finding with executed evidence behind it:
duck-review: an outage on a required reviewer now bars APPROVE. Measured — a v1.0.0 release gate approved on a single family because the other outaged for the third time that gate and nothing stopped it.duck-review: inputs by path, never inlined. The same pinned model disagreed on 28 of 41 verdicts between an inlined corpus and one read from disk, every flip toward the finding standing. Prompt delivery is a verdict-integrity control, not a performance note.duck-review: prove the model pin. Send a deliberately invalid--modeland require the harness to reject it with a roster. A model's self-report is unfalsifiable; a harness rejection is not.duck-break: a dirty candidate does not survivegit worktree add. Measured — an uncommitted marker was present in the source and absent from the copy, so every attack passed against the base commit.duck-land+duck-sweepshare one evidence chain. Landing records the SHA; the sweep reads it instead of routing every squash-merge to a decision it can never resolve.duck-runreferences its stages by name. Numbers went stale twice in one session while renumbering.duck-landno longer spells out branch-protection bypass mechanics — the rule they taught survives without the recipe.- The README states which rules are mechanically checked and which are discipline. A rule that reads like a guarantee and isn't is worse than no rule.
Validator: 631 → 290 lines
Structure only. It checks what exists, what links to what, what shape a file has, and whether generated files match their source — never content. The skill set is discovered from disk instead of enumerated, so adding duck-why required no validator edit.
Removed with it: version agreement across the manifests, and the description YAML-safety guard. A malformed description now fails silently in a host; all eighteen were verified by hand for this release.
How this shipped
Under two written owner waivers. The review gate returned NOTE, not APPROVE — one of two required reviewers outaged, and the new bar refused to approve its own release. The merge then bypassed the ruleset's last-push-approval rule via admin. Both recorded before the push; the decorrelated review is deferred, not passed.
v1.1.1 — the README learns to talk
No skills changed behaviour this release. The README did.
What changed
The README led with installation instructions and described the duck's temperament in the third
person. It now leads with what the collection buys you, and the two registers are separated: the
skills are dry, straightforward, and skeptical; the README's narrator is the one having fun.
Skills and The soul moved above Install. The mermaid flowchart became an ASCII map.
Status collapsed to a version and a link — the count it carried was hand-typed with nothing
checking it.
Manifest descriptions now read "the duck is skeptical, dry, and straightforward". That is a change
of meaning, not a spelling fix: the skills demand evidence, they do not distrust motives.
Separately, spelling normalized to American English across the collection, matching
duck-proof's own frontmatter.
What the gate caught
Four rounds. Eleven substantiated blockers, nine of them the ASCII map asserting handoffs that no
skill body supports — a duck-campaign connector invented three different ways, and conditional
stages drawn as unconditional, then bracketed, then explained, each fix generating the next
finding.
The circuit breaker fired twice. Both times the exit was deletion rather than a better
explanation: the campaign row left the diagram, then duck-plan and duck-break did. What
remains is the spine duck-run runs unconditionally, and one sentence that names the conditional
skills without restating a threshold that lives in their bodies — so it cannot go stale when one
changes.
The doer's receipts overclaimed in three consecutive rounds: a six-word grep reported as "zero UK
spellings remain", a grep result described rather than read, a line count reported as an
occurrence count. Each was caught by a reviewer, not by the doer. The receipt now cites captured
stdout and quotes it.
Gate record
| Round | OpenAI gpt-5.6-sol |
Google gemini-3.1-pro-high |
|---|---|---|
| r1 | REJECT | REJECT |
| r2 | REJECT | REJECT |
| r3 | outage — out of credits | REJECT |
| r4 | seat waived by owner | APPROVE |
The OpenAI seat went to outage at r3. agy's gpt-oss-120b-medium was evaluated as a substitute
and rejected: it self-reports as Gemini, so it is either the seated family or of unknown identity,
and neither counts as decorrelated. The owner waived the second seat explicitly. r4 is a
single-family APPROVE under a recorded waiver, not a quorum.