Emails people actually answer — 8 enforceable rules, benchmarked. Most published email skills are marketing machinery: cold sequences, drip campaigns, subject-line A/B tests. The email that actually eats your day — the ask to a colleague, the tricky multi-question reply, the follow-up after silence, the polite no — has almost no skill coverage, and what exists is either voice-mimicry that needs samples of your sent mail to work, or template structures that mandate the very warm-up paragraphs readers skip. This skill packages the discipline nobody else codifies: checkable rules a draft either passes or doesn't.
Built and bench-tested by SkillProof, the tested Claude skills directory.
We surveyed 87 email-writing skills (from a 16,682-skill dataset) plus the standalone field before writing this one. Four techniques showed up in none of them as enforceable rules — they are this skill's core:
- One goal per email, ask in the first two lines. Context comes after the ask, never before. Two unrelated asks → two emails.
- Word budgets by type, numeric. Cold ≤120 words, follow-up ≤60, reply ≤150, internal update ≤150. Over budget means cut, not apologize.
- Reply discipline. Read the whole thread, then answer their questions first, in their order, before adding your own ask. No skill we found codifies this at all.
- The recipient's-perspective pass. Reread as the recipient: "what do I do after reading this?" If the answer isn't one obvious low-effort action, rewrite.
Plus the rest: subject line = the ask (≤50 chars, deadline included), a mechanically scannable banned-fluff list ("I hope this email finds you well", "I wanted to reach out", "just checking in", the AI-tell vocabulary), exactly one CTA with a real date, and a strict output contract — subject + body, nothing else, variants only on request.
Explicitly out of scope: marketing campaigns, newsletters, lifecycle sequences. Dedicated marketing skills own those lanes; this skill is for email from one human to another.
git clone https://github.com/Skillproofdev/email-discipline ~/.claude/skills/email-discipline
# restart Claude Code — triggers on "write an email", "reply to this", follow-ups, cold outreach draftsOne command — the repo IS the skill.
Two arms, same model (Claude Sonnet), identical prompts across 12 scenarios — the only
variable is this skill. Ground truth, word budgets, and the banned-fluff regex list were
frozen before either arm ran (bench/ground-truth/). Scored 2026-07-10.
| Metric | Baseline (no skill) | With email-discipline |
|---|---|---|
| Mechanical pass — budget + subject + fluff + format (scripted) | 2/12 | 12/12 |
| Within word budget | 2/12 (mean 163 words vs 130 budget) | 12/12 (mean 100) |
| Subject violations (>50 chars / generic) | 4 | 0 |
| Banned-fluff hits | 1 | 0 |
| Ground-truth coverage (questions answered, facts, asks) | 65/77 (84%) | 76/77 (99%) |
| Wrong premises corrected | 2/2 | 2/2 |
| Planted-trap violations (venting, credential dumps, buried asks…) | 9 | 4 |
| CTA with a real date where required | 11/12 | 12/12 |
| Blind paired preference (wins–ties–losses) | 1–1–10 | 10–1–1 |
| "Would reply now" reactions | 3/12 | 9/12 |
The headline: the baseline model writes polite, mostly fluff-free email that is too long
(over budget on 10 of 12 drafts) and buries or undates the ask — its follow-up said "whenever
you get a chance" and its escalation vented two grievances. The skill fixed exactly those:
every draft within budget, every required CTA dated, coverage up 84%→99%, and a 10–1–1 blind
preference. Negative results, published: the skill did NOT improve wrong-premise correction or
reply question-order (baseline already perfect at both), it lost one blind judgment (post-
incident update — the baseline's labeled sections scanned better), and both arms stacked two
consequences in the invoice chaser. Blind judging caveat: in-context judging by the scoring
agent with a pre-registered coin-flip blind (bench/runs/blind-key.json), not an independent
judge. Full protocol and per-scenario scores: bench/scoring.md, bench/results/verdict.md.
Twelve realistic scenarios, two per category. Each is a full self-contained prompt with
real-world texture — names, stakes, dates, and the complete inbound email for every reply
scenario, including planted traps (questions in odd order, a wrong premise to correct
politely, venting bait, credential-dump bait). All materials: bench/materials/, prompts:
bench/tasks.md.
| # | Category | Scenario |
|---|---|---|
| S01–S02 | Cold ask | (1) Ask a professor you met once for access to a non-public research dataset; (2) cold B2B email to a Head of Operations asking for a 25-min demo |
| S03–S04 | Tricky reply | (3) Client email with four questions scattered mid-paragraph, one premised on access you don't have; (4) VP status request with questions numbered 3-1-2, a buried unnumbered question, and a false "Security signed off" premise |
| S05–S06 | Escalation | (5) Vendor missed two deadlines — get an A/B decision from your director without venting; (6) recurring CI outages — get a P3 ticket raised to P1 before release freeze |
| S07–S08 | Follow-up after silence | (7) Second touch on a €24k proposal, 9 days of silence, a start slot expiring; (8) second nudge on a €3,800 invoice 15 days overdue |
| S09–S10 | Decline politely | (9) Turn down a keynote (and its panel) while keeping the door open; (10) decline a colleague's co-lead ask AND his fallback ask, offering one real alternative |
| S11–S12 | Internal update | (11) Weekly migration status, no decisions needed, a 12-row table tempting to paste; (12) post-incident summary with a blame trap |
Protocol. Two arms — base (Sonnet, no skill) vs skill (Sonnet + SKILL.md) —
identical prompts, 12 drafts each, generated before any scoring. Ground truth (per-scenario
question lists and order, wrong premises, required facts, CTA expectations, word budgets)
and the banned-fluff regex list were frozen in bench/ground-truth/ before either arm ran.
Arms are barred from reading ground truth or the scoring protocol.
Layer 1 — mechanical (scripted, zero judgment). bench/ground-truth/check_mechanical.py,
written and tested before the arms ran: body word count vs the Rule 3 budget (quoted thread
excluded), banned-fluff regex hits, subject ≤50 chars, generic-subject ban, thread-subject
keeping for replies, output-format compliance.
Layer 2 — content coverage vs pre-registered ground truth. Binary checklist items with quoted-fragment evidence: every inbound question answered in the asker's order, wrong premises corrected (corrected / silently accepted / dodged), required facts present, planted traps avoided, CTA-with-date where required.
Layer 3 — blind paired judging. Both drafts per scenario, A/B order set by recorded coin
flip before any draft was read (bench/runs/blind-key.json), judged on a five-second
"what does the sender want?" test, a reply-now/later/ignore call, and a forced pick;
unblinded only after all 12 judgments were recorded. As-run caveat: judged in-context by the
scoring agent, not by the pre-registered independent judge panel — labeled as such in
bench/results/verdict.md.
Negative results ship. If the skill loses a category, the table says so (it lost one blind judgment and tied another — see above).
"Write an email", "reply to this", "follow up with…", "draft a cold email to…", declines, escalations, internal updates — or pasting an email/thread and asking for a response. Explicitly excluded: marketing campaigns, newsletters, drip/lifecycle sequences, inbox triage, and email-sending integrations.
We test other people's skills for a living (public methodology). Our own skills get the same treatment — measured benchmarks, negative results included.
More skills from us: skillproof-skills index · Free tools for skill authors: SKILL.md validator, token calculator, Rules⇄SKILL.md converter.
MIT — use it, fork it, ship it.