Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

3 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

email-discipline

Emails people actually answer — 8 enforceable rules, benchmarked. Most published email skills are marketing machinery: cold sequences, drip campaigns, subject-line A/B tests. The email that actually eats your day — the ask to a colleague, the tricky multi-question reply, the follow-up after silence, the polite no — has almost no skill coverage, and what exists is either voice-mimicry that needs samples of your sent mail to work, or template structures that mandate the very warm-up paragraphs readers skip. This skill packages the discipline nobody else codifies: checkable rules a draft either passes or doesn't.

Built and bench-tested by SkillProof, the tested Claude skills directory.

What it does differently

We surveyed 87 email-writing skills (from a 16,682-skill dataset) plus the standalone field before writing this one. Four techniques showed up in none of them as enforceable rules — they are this skill's core:

  1. One goal per email, ask in the first two lines. Context comes after the ask, never before. Two unrelated asks → two emails.
  2. Word budgets by type, numeric. Cold ≤120 words, follow-up ≤60, reply ≤150, internal update ≤150. Over budget means cut, not apologize.
  3. Reply discipline. Read the whole thread, then answer their questions first, in their order, before adding your own ask. No skill we found codifies this at all.
  4. The recipient's-perspective pass. Reread as the recipient: "what do I do after reading this?" If the answer isn't one obvious low-effort action, rewrite.

Plus the rest: subject line = the ask (≤50 chars, deadline included), a mechanically scannable banned-fluff list ("I hope this email finds you well", "I wanted to reach out", "just checking in", the AI-tell vocabulary), exactly one CTA with a real date, and a strict output contract — subject + body, nothing else, variants only on request.

Explicitly out of scope: marketing campaigns, newsletters, lifecycle sequences. Dedicated marketing skills own those lanes; this skill is for email from one human to another.

Install

git clone https://github.com/Skillproofdev/email-discipline ~/.claude/skills/email-discipline
# restart Claude Code — triggers on "write an email", "reply to this", follow-ups, cold outreach drafts

One command — the repo IS the skill.

Benchmark results (measured, not estimated)

Two arms, same model (Claude Sonnet), identical prompts across 12 scenarios — the only variable is this skill. Ground truth, word budgets, and the banned-fluff regex list were frozen before either arm ran (bench/ground-truth/). Scored 2026-07-10.

Metric Baseline (no skill) With email-discipline
Mechanical pass — budget + subject + fluff + format (scripted) 2/12 12/12
Within word budget 2/12 (mean 163 words vs 130 budget) 12/12 (mean 100)
Subject violations (>50 chars / generic) 4 0
Banned-fluff hits 1 0
Ground-truth coverage (questions answered, facts, asks) 65/77 (84%) 76/77 (99%)
Wrong premises corrected 2/2 2/2
Planted-trap violations (venting, credential dumps, buried asks…) 9 4
CTA with a real date where required 11/12 12/12
Blind paired preference (wins–ties–losses) 1–1–10 10–1–1
"Would reply now" reactions 3/12 9/12

The headline: the baseline model writes polite, mostly fluff-free email that is too long (over budget on 10 of 12 drafts) and buries or undates the ask — its follow-up said "whenever you get a chance" and its escalation vented two grievances. The skill fixed exactly those: every draft within budget, every required CTA dated, coverage up 84%→99%, and a 10–1–1 blind preference. Negative results, published: the skill did NOT improve wrong-premise correction or reply question-order (baseline already perfect at both), it lost one blind judgment (post- incident update — the baseline's labeled sections scanned better), and both arms stacked two consequences in the invoice chaser. Blind judging caveat: in-context judging by the scoring agent with a pre-registered coin-flip blind (bench/runs/blind-key.json), not an independent judge. Full protocol and per-scenario scores: bench/scoring.md, bench/results/verdict.md.

Benchmark design

Twelve realistic scenarios, two per category. Each is a full self-contained prompt with real-world texture — names, stakes, dates, and the complete inbound email for every reply scenario, including planted traps (questions in odd order, a wrong premise to correct politely, venting bait, credential-dump bait). All materials: bench/materials/, prompts: bench/tasks.md.

# Category Scenario
S01–S02 Cold ask (1) Ask a professor you met once for access to a non-public research dataset; (2) cold B2B email to a Head of Operations asking for a 25-min demo
S03–S04 Tricky reply (3) Client email with four questions scattered mid-paragraph, one premised on access you don't have; (4) VP status request with questions numbered 3-1-2, a buried unnumbered question, and a false "Security signed off" premise
S05–S06 Escalation (5) Vendor missed two deadlines — get an A/B decision from your director without venting; (6) recurring CI outages — get a P3 ticket raised to P1 before release freeze
S07–S08 Follow-up after silence (7) Second touch on a €24k proposal, 9 days of silence, a start slot expiring; (8) second nudge on a €3,800 invoice 15 days overdue
S09–S10 Decline politely (9) Turn down a keynote (and its panel) while keeping the door open; (10) decline a colleague's co-lead ask AND his fallback ask, offering one real alternative
S11–S12 Internal update (11) Weekly migration status, no decisions needed, a 12-row table tempting to paste; (12) post-incident summary with a blame trap

Protocol. Two arms — base (Sonnet, no skill) vs skill (Sonnet + SKILL.md) — identical prompts, 12 drafts each, generated before any scoring. Ground truth (per-scenario question lists and order, wrong premises, required facts, CTA expectations, word budgets) and the banned-fluff regex list were frozen in bench/ground-truth/ before either arm ran. Arms are barred from reading ground truth or the scoring protocol.

Layer 1 — mechanical (scripted, zero judgment). bench/ground-truth/check_mechanical.py, written and tested before the arms ran: body word count vs the Rule 3 budget (quoted thread excluded), banned-fluff regex hits, subject ≤50 chars, generic-subject ban, thread-subject keeping for replies, output-format compliance.

Layer 2 — content coverage vs pre-registered ground truth. Binary checklist items with quoted-fragment evidence: every inbound question answered in the asker's order, wrong premises corrected (corrected / silently accepted / dodged), required facts present, planted traps avoided, CTA-with-date where required.

Layer 3 — blind paired judging. Both drafts per scenario, A/B order set by recorded coin flip before any draft was read (bench/runs/blind-key.json), judged on a five-second "what does the sender want?" test, a reply-now/later/ignore call, and a forced pick; unblinded only after all 12 judgments were recorded. As-run caveat: judged in-context by the scoring agent, not by the pre-registered independent judge panel — labeled as such in bench/results/verdict.md.

Negative results ship. If the skill loses a category, the table says so (it lost one blind judgment and tied another — see above).

When it triggers

"Write an email", "reply to this", "follow up with…", "draft a cold email to…", declines, escalations, internal updates — or pasting an email/thread and asking for a response. Explicitly excluded: marketing campaigns, newsletters, drip/lifecycle sequences, inbox triage, and email-sending integrations.

Why trust this

We test other people's skills for a living (public methodology). Our own skills get the same treatment — measured benchmarks, negative results included.

More skills from us: skillproof-skills index · Free tools for skill authors: SKILL.md validator, token calculator, Rules⇄SKILL.md converter.

License

MIT — use it, fork it, ship it.

About

Email skill for Claude: one goal per email, subject-as-ask, word budgets, no AI fluff. Benchmarked: 'would reply now' 9/12 vs 3/12 baseline.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages