Skip to content

Development

m4bwav edited this page Oct 9, 2026 · 2 revisions

Layout

Path What it is
skills/editwright/SKILL.md the rules and the passes the agent follows
skills/editwright/kb/ editing knowledge in small files, read on demand; start at kb/INDEX.md
skills/editwright/references/suggestion-format.md the JSON ew.py suggest takes
skills/editwright/scripts/ew.py the CLI: one standard-library Python file, 3.9 or newer
skills/editwright/ (other files) evergreen companions: RESEARCH.md, CHANGELOG.md, LEARNINGS.md, TESTS.md, MAINTENANCE.md, evergreen.json, evals/evals.json
tests/test_ew.py unit tests for ew.py
evals/ runnable eval cases in the claude plugin eval format
scripts/budget.py, scripts/package.py the size budget and the zip build
ai-docs/ research notes, decisions, the log and the handoff
AGENTS.md the rules for any agent working in the repository

Tests

python tests/test_ew.py

31 tests, all invented text written at test time. On the v0.1.1 checkout on 2026-10-09 they passed on Windows (Python 3.14.6) and on Linux (Python 3.14.4):

Ran 31 tests in 0.141s

OK

They cover the cleanup rules and the check that no word changes. They cover how each kind of change is counted, intake and check (with --against), and suggest refusing bad quotes. Suggest is tested for saving the valid items of a file and skipping saved ones on a second run, and intake for writing nothing when it refuses. They cover every apply refusal, author text and the copy check, the nonfiction allowance and override. The rest cover the .docx export, .docx sources, feedback, chunks and stats.

The CI workflow (.github/workflows/tests.yml) runs them on Ubuntu, macOS and Windows with Python 3.9 and 3.14. It also checks that ew.py --version answers from a clean checkout, and runs the size budget and the zip build.

Rules from AGENTS.md: every change to ew.py gets a test, and cleanup must never change a word. The script stays one standard-library file that runs on Windows, macOS and Linux. It reads and writes UTF-8 with LF line endings, and never writes to the source.

Size budget

The plugin must stay small enough to install without a second thought. scripts/budget.py measures five things, and red fails CI:

python scripts/budget.py
description characters (always loaded; harness cuts at 1,536)       979  green<=1100  yellow<=1400  green  
SKILL.md tokens (loaded on every use)                              2066  green<=2500  yellow<=3500  green  
largest kb or reference file tokens (loaded on demand)             1253  green<=1500  yellow<=2000  green  fiction.md
installed skill folder KB                                           170  green<=300   yellow<=600   green  
tracked repository KB                                               551  green<=1024  yellow<=3072  green  
budget: GREEN

That is the v0.1.1 checkout on 2026-10-09. Tokens are estimated as bytes divided by four. AGENTS.md also keeps SKILL.md under 200 lines and each kb file under 6,000 bytes.

The zip for claude.ai

python scripts/package.py

It writes dist/editwright.zip with the skill folder under one editwright/ top folder, after checking that SKILL.md's frontmatter has only keys claude.ai accepts. On the v0.1.1 checkout it wrote 21 files, 70 KB.

Evals

The eval suite proves the skill triggers and acts, on evidence outside the transcript. The root evals/ folder holds nine cases for claude plugin eval:

Cases What they test
trigger-1 to trigger-3 edit, editorial-letter and proofread requests load the skill
decoy-1, decoy-2 requests it must not take: writing a new story, and humanizing AI text
action-1 "suggestions, but don't change anything yet": intake and suggest ran, a suggestions file and ledger exist, check passed, nothing applied, source unchanged
action-2 applying a suggestion that carries AI-written words: apply ran, no edited file, the reply says why, source unchanged
action-3 "apply S-001 and S-002, nothing else": both applied, no other ID applied, source unchanged
outcome-1 a full edit request: a useful diagnosis, nothing applied, source unchanged

evals/build_cases.py writes the case folders, with an invented story and a works store inside the fixture. Run it after changing ew.py. Claude Code 2.1.281 or newer runs them:

claude plugin eval . --no-publish --trust-plugin --case "trigger-*"

Cases that need Bash run under WSL2 or Linux. The 0.1.0 release notes report 9 of 9 cases passing with claude plugin eval 2.1.281. They also compare runs with the skill and without it. With it, the four action and outcome cases scored 1.00, 0.92, 1.00 and 1.00. Without it, they scored 0.29, 0.75, 0.00 and 0.67. The 0.1.1 release notes report 9 of 9 again. Since 0.1.1 the Bash cases run with 30 turns and 600 seconds each. The evals were not rerun for these pages. The TESTS.md in the 0.1.1 release records the 9 of 9 run; the one in 0.1.0 still showed the scaffold entry ("not yet run", 0/0).

The manifest check:

claude plugin validate .

It passed on the v0.1.1 checkout on 2026-10-09.

Prose

People read README.md, SKILL.md, the kb files and the letters the skill writes, so AGENTS.md asks for them to be checked with the everwrite checker (tells.py), with zero strong findings. It is never run on an author's manuscript.

Releases

A release bumps the version in .claude-plugin/plugin.json, plugin.json, VERSION in ew.py, metadata.version in SKILL.md and version in evergreen.json together, tags vX.Y.Z, and publishes a GitHub Release with the CHANGELOG entry as notes and editwright.zip attached.

Contributing

After the first push, changes go through pull requests. Fixtures, issues and logs use invented text only: never a real manuscript, a real author's text, or a style sheet or story bible of a real work. A change that weakens one of the five rules needs a decision entry in ai-docs/decisions/ and a test. The skill is an evergreen unit: its research is re-checked on a schedule (evergreen.json, every 30 days, next due 2026-11-07), and every change is logged in its CHANGELOG.md with the reason.

Clone this wiki locally