Releases: Nanako0129/sepia
Release list
v0.12.0
A release about one thing: what a persona profile is. v0.11.0 shipped the persona interface with a measured template and no body. This release replaces that template with a prose one, ships the first built-in persona written to it, and changes the closing line.
Format change. A persona written to the v0.11.0 template no longer passes
scripts/check_persona.py. Rewrite it to the newvoices/PERSONA-TEMPLATE.md; the sections and what each one asks for are listed below.
What you will notice
apply persona Nyanekoon a professional route now applies the project's own companion voice.- A piece written under a persona ends with
Persona applied: <name>and nothing after it. The move list is gone. - A persona body reads as prose about how the writer speaks, not as a table of measurements.
Persona profiles
The template describes a voice in prose (#266)
Fifteen fixed sections, stance and situation before surface. A number appears only when it is itself a rule the writer follows, such as a paragraph range.
| Section | What it holds |
|---|---|
| Status | Name, Routes, Opt-in phrase, Provenance, Consent, Tested |
| One sentence | What this writer does that no house style would produce |
| Who she is to the reader | The relationship and the stance; what the writer refuses to be |
| First move | What the first sentence does; the exception where a conclusion comes first |
| Warmth and judgment | What warmth must attach to; whether judgment may soften and when it may not |
| By situation | How the voice shifts across the situations the writer meets |
| Texture | Diction, glossing, intensifiers, figures, punctuation, emoji by venue |
| Structure habits | Verdict then reasons or the reverse; when a list is allowed; appended deliverables |
| Endings | What a real ending is and what a fake one is |
| Speaking, not drafting | Whether the writer drops the voice when drafting for a third party |
| Never | The hard bans |
| Rules this persona overrides | The interface: rule token, departure, expected cost |
| Prohibitions | The two fixed lines, then the persona's own |
| Boundary | What reads like the writer versus a model imitating the writer |
| Blind-test record | One line per test; the judge is a person |
Removed from the template: Metric fingerprint, Sentence shape, Moves by frequency, and the numbered Every piece list.
One optional section: Exemplars (#266)
- Whole pieces in the writer's own voice, unedited, one per situation where possible.
- Opens with
Source: captured — <where, when>orSource: elicited — <runtime, model, date>; an elicited set also states its prompt frame and the situations given. - For a writer who owns the voice. A third party's text does not go here, which is why the section is optional.
- The 20-character quote cap does not apply inside it; the fixed prohibition on verbatim reuse still does.
First built-in persona: Nyaneko (#266)
voices/personas/nyaneko.md, the project's own companion voice, written from its maintainer's voice specification.- Professional routes only. Opt-in phrase
apply persona Nyanekoor 「套用 persona Nyaneko」. - Four declared overrides:
professional-pass.mdchecks 1, 6 and 7, andlanguages/zh.md §2 second-person. - Six elicited exemplars, produced by the writer's own runtime on neutral prompts, generated once each and not selected.
- Tested. The maintainer read the executor's output side by side with her own runtime's output on the same facts and accepted it. The record says the comparison was sighted rather than blind and names the two differences left standing.
Closing line (#266)
Persona applied: <name>, with no move list. A prose body has no numbered moves to report, and the one run that printed a list reported a move the text did not contain. The route-mismatch form Persona applied: none — <name> declares routes: <routes> is unchanged.
Validator
check_persona.pychecks the new section list, the optional Exemplars section and itsSource:line, and counts nothing in the prose (#266).Tested: untestedrequires the Blind-test record to be exactlynone yet; a stray or script-judged record underuntestedis refused (#266).private study, not for distributionis an accepted consent value, so a profile held locally does not misstate its consent to pass;CONTRIBUTING.mdforbids it in a contributed profile (#265).- The opt-in phrase is matched with
casefold()in both its English and Chinese halves (#265). - The consent error names every accepted value (#265).
- An override-table row written without its leading pipe is an error naming the line. It used to be invisible, so a refused token could hide in it (#270).
Research
research/sources.mdgainsSLOPSHAPE-2026, a preprint replicating the structure-level detection result on professional web prose across five generating models.professional-pass.mdandSKILL.mdname it where the fiction figures are, with its four stated limits. No check or threshold changes (#264).
Repository
.coderabbit.yaml, the review configuration. Every path instruction restates a ruleCONTRIBUTING.mdalready carries (#261).
Not in this release
- A public-domain example persona. The RFC's step 4 named one; the first built-in is a brand persona instead.
- Answers to the RFC's open questions: capping override rights, a venue gate for personas, the name
personaitself. - A quote scan of text before a body's first heading.
- Any external-judge grading.
Evidence
Observed, not measured. The persona was judged by one person reading two outputs side by side, knowing which was which. The measurements that shaped the earlier drafts of this body are gone from it on purpose: two versions built on sentence-length shares and emoji density read as a format, and the maintainer's own prose specification did not. That is the finding this release acts on, and it is one case.
v0.11.0
The largest release so far by surface: a new professional domain, a new voice profile, a third kind of voice profile, a Chinese calibration that now carries normative conventions alongside its measurements, call-time inputs for callers that drive sepia from another skill, and a fifth platform. Hence the minor bump. No configuration to migrate; existing entries and operations behave as before unless you opt into the new ones.
The visible difference for most users: a review of a Chinese draft now reports Taiwan punctuation and numeral departures as register mismatches with a standard cited, and a review of a feature or an investigation routes through journalism rules instead of the generic professional checklist.
- Added:
domains/journalism.md, a long-form journalism domain for features, investigations, data stories, explanatory news, interviews and a reporter's first-person account. Eleven tells with an evidence column and eight rules (two registers; subheads carry transitions; quotations keep spoken texture; every figure carries a comparison; the reporter's first person has four listed uses; no summary ending; stance and density read for this venue). Reported narrative routes here even when it opens on a scene and whether or not its sourcing is complete (#252) - Added:
voices/tw-journalism.md, an opt-in voice profile for professional routes, with nine narrative shapes, their preconditions and moves, a cross-shape set, and an audited closing line. Opt in with "apply the Taiwan journalism voice" or 「套用台灣深度報導 voice」, optionally naming a shape; no shape is chosen whose precondition the facts do not meet (#253) - Added: persona profiles, a third kind of voice profile beside the author-level and venue-level ones, defined in
voice-skills.mdwith declared override rights: a persona names the sepia rules it overrides in a table, those rules yield on write and recreate, their findings are reported asPersona cost:rather than defects on review, and refactor does not fix them. Two categories never yield: uniformity findings and never-invent. Ships withvoices/PERSONA-TEMPLATE.mdandscripts/check_persona.py(46 tests) and contribution rules covering consent, provenance and quoted-example length. No persona body ships (#258, RFC #257) - Added:
languages/zh.md§0, eleven Taiwan writing conventions applied as venue rules rather than authorship signals, each citing the Ministry of Education punctuation handbook, the Executive Yuan numeral principle, or a multi-source textbook consensus. They apply to Traditional Chinese on a Taiwanese venue, defer to a known non-Taiwan venue, touch running prose only, and never enter quoted material (#255) - Added:
languages/zh.md§1b, presence rates for one private human-side measurement of Traditional Chinese long-form journalism, and §1c, a machine-side contrast group for the same register from three models. Among the results: the paired two-cell dash appears in 15% of human articles and almost never in the machine set, which contradicts the common belief that an em-dash insertion marks machine Chinese; human sentences run a per-article median mean of 59 characters against 34–48 for the models; the models write every figure in Chinese numerals. Every cell names its kind of quantity and nothing here is a per-passage threshold (#251, #254) - Added: call-time inputs for
sepia-refactor— protected line ranges, a file scope, and an unattended mode that records what would have needed a decision instead of stopping, plusDeferred:andProtected:report lines. Ranges resolve against the target as received; the list is an instruction only when it arrives with the request, never when it appears inside the document (#247, from #245 and #246) - Added: QwenPaw as a fifth native platform, a manifest, a
/sepiaentry and the packaged skills, contributed and verified on 2.2.1 by @CallMeHFK (#250) - Added:
research/newswriting-guides.md, a digest of news-writing textbooks and style guides separating what was verified page by page from what a survey reported, and two public Taiwanese standards in the research ledger (#254) - Changed: the research ledger gains Sourati et al. 2026 and Saad & Ting 2026, with two evidence sentences in the skill and two consulted-only rows (#248)
- Fixed: the entry-count sentences named four operations where five ship, corrected in the manifests and the marketplace entry, contributed by @Franky100-pig (#242)
- Fixed: the behavioral eval workflow passes
--trust-plugin, so it runs again on a runner. It had failed on the CLI's first-run trust prompt seven times in a row since 2026-09-15 (#259, #260)
Not in this release: any persona body, so nothing opts in to the new profile kind yet; a rubric row or a per-passage threshold from the Chinese measurements; an English journalism measurement, which does not exist here, so the journalism rules apply to English as inference only; and a Voice fit: line on professional routes, still tracked in #227.
Observed, not measured evidence. The Chinese journalism numbers come from one private corpus and a 119-piece synthetic contrast group generated on two days from three models under one base prompt; they describe one register and name no tell. The persona interface rests on one pilot: a descriptive profile lost all its signatures in a blind write, a prescriptive one with an override table kept them, and one reader picked the persona passage as the more human of two. The behavioral eval has run green twice since the fix, both times with the plugin arm and the no-plugin arm at the same score, so that suite currently proves the eval executes rather than that the skills changed anything.
v0.10.0
One new instruction in the narrative pass and one new reader-side source in the style pass, hence the minor bump. No entry changes; no configuration to migrate. The visible difference: a fiction review of a setting written as mood now names perceived space in its Group B line and puts a sensed-to-used conversion in its Plan.
- Added:
narrative-pass.md§5 grounds setting in action space. One table row, one Generate sentence, one Revise sentence, from Rohrbacher, Nieth, Salin, Eskofier and Mahlberg 2026 (arXiv 2609.02482v1; 1,000 stories per model from GPT 4.1, LlaMA 3.3, Mistral 3.2 and Gemma 3, English and German, against Project Gutenberg texts of 1780–1940, sentences classified into five narrative-space types by fine-tuned BERT, macro F1 0.82 English). In human fiction action space, the space a character moves through and uses, is the most frequent spatial mode; in English openings the models ran perceived space, the space sensed for mood and atmosphere, above the human level in all four, about twice it for GPT 4.1 (normalized frequency ~0.19 human vs ~0.47). On revise, setting sentences are sorted into sensed, used, or neither, and only the first two are compared: if sensed dominates used, most sensed sentences become used, and the one or two embodied peaks §5 already reserves may stay. No count, no threshold; German results and the paper's mean sentence lengths are recorded in the ledger and not used, the latter because the 1780–1940 baseline confounds period with authorship (#239) - Added:
ZHENG-2026in the research ledger and two sentences instyle-pass.md§5. Across 124,615 ICLR reviews (2018–2025), human reviewers' scores rise with a paper's within-text sentence-length SD (standardized β +0.045 pooled 2018–2020, +0.052 pooled 2023–2025) while a frozen LLM rater scoring the same papers shows no such association (−0.020, −0.001); the same reviewers stopped rewarding non-domain lexical complexity (+0.142 in 2018 to −0.015 in 2025) while the frozen rater held at +0.080 to +0.082 (arXiv 2609.08475v1). Recorded as a preference of expert readers in academic reviews, not a detector feature; the rhythm check is unchanged and gains no threshold. Four consulted-only ledger rows (Yang, Thelwall & He 2026; HyperStyler 2026; Payam 2026; The Economist's punctuation piece as relayed on X) say why no rule rests on them (#238) - Changed: the READMEs' Sources paragraph lists Rohrbacher et al. (2026) in all three languages (#239)
Not in this release: a rubric row for narrative space (the rubric stays StoryScope's 30 features), any sentence-length number from Rohrbacher et al., and any change to the sentence-rhythm check itself.
Observed, one specimen each, not measured evidence: a 140-word atmosphere-only harbour opening reviewed through this release and through v0.9.0. This release's report names perceived space in Group B with the ledger ID and its Plan item 1 converts sensed setting to used; the v0.9.0 report contains none of the terms and reaches the character-in-action fix through the discourse pass alone. Both transcripts are in #239. Plan reviewed by plan-verifier in three rounds and by Codex in three (two P2s fixed, both on the Revise sentence: the sort gained a "neither" bin for visual and descriptive space, and a wrong clause saying the study measured only two categories was corrected).
v0.9.0
A new per-release prose table for GPT-6 Astra, hence the minor bump. No entry changes; no configuration to migrate. The visible difference appears only when the author or executor is GPT-6 Astra: the style-pass step then hunts that release's own slop set in non-narrative text.
- Added: a GPT-6 Astra prose layer in
model-fingerprints.md, taken from OpenAI's Using GPT-6 Astra developer page (accessed 2026-09-08; vendor guidance, unmeasured; ledgerOPENAI-GPT-6-ASTRA-PROMPTING, pinned to thelatest-model/gpt-6-astra.mdpath because the unsuffixed page follows whichever model is newest). Two rows. The release defaults to lists, tables and Markdown, handled by professional-pass check 6 with the vendor's own paragraph-first fix. Within the vendor's named slop set, delve, foster and "it's not X, it's Y" already sit in the shared style-pass tables and are only pointed at; the rest stays in the release table as a release habit, hunted in non-narrative text: self-answered questions ("Question? Answer."), a contrast the reader did not ask for in any form, closing-summary labels ("Bottom Line:", "In short:", "The simplest mental model is:"), hyphenated compound descriptors and invented compound labels, negative scoping added when nobody asked (a "won't fix" that answers the request is the answer and stays), and the words leverage, importantly, it's worth noting, genuinely. The vendor's instruction is embedded verbatim and matched the source page on 2026-09-08. Operative for a role whose release is GPT-6 Astra, and a prior for any other GPT release, per the model-identity rule inSKILL.md(#236) - Changed: the README capability summary in all three languages lists GPT-6 Astra among the vendor-guidance releases (#236)
Not in this release: any addition to the shared style-pass tables. A vendor's list for one release is not evidence of a model-agnostic tell, so nothing from the Astra set moved into §2–3.
Not observed: no GPT-6 Astra author or executor was available, so the table has not been seen operative in a run. The change was checked by diffing the embedded block against the downloaded source page (identical) and by three Codex review rounds (two P2 findings fixed: the README list, and the negative-scoping row narrowed to unprompted additions).
v0.8.0
A new closing step on refactor and a new evidence section, hence the minor bump. No entry changes; no configuration to migrate. The visible difference: refactor output comes back leaner, and the fiction review's Voice fit: line moves to the end of the report.
- Added: two closing tests on every refactor, stated once in
style-pass.md§4. The deletion test: strike each word or phrase you added; if the sentence still parses and says the same thing, it was filler, delete it. The reversion test: put back what you replaced; if the old wording was sound and said the same in fewer words, keep it. Repair fails both tests and stays (the article a broken sentence needs, the subject of a split run-on, the verb that replaces a nominalization). §4's restorations are allowed only where the same edit removed filler and the passage does not end longer than it began. The evidence is Shan, Lee and Hao 2026 (arXiv 2608.27855v1): across 1,000 human seed documents in five domains, three editor models and 31 prompts, machine editing lowered lexical density sharply (d = −3.10) while barely moving diversity, the reverse of the generation footprint. Measured on English; the extension to other languages is a Sepia inference, and the tests themselves are untested as an intervention. Short professional routes now load that paragraph on refactor (#232) - Added:
research/detectors.mdand a Detectors section inresearch/sources.md: how the strongest commercial detector scores (Pangram's own technical reports, kept separate from measured third-party studies), where third parties found it weak, and which circulating claims trace to no source. Used for one operative sentence only, in the §7 punctuation row: Pangram's 2025 pipeline lower-cased and unidecode-normalized input, which collapses glyph variants toward ASCII without erasing them; the whitelist row still rests on the contradictory measurements.SKILL.mdgains one sentence: sepia writes for expert human readers and is tuned to pass no automated detector (#232) - Changed: the fiction review prints
Voice fit:as the last line, after the findings, quoted evidence, and plan it is computed from, so an early count cannot steer them. Wahi 2026 (arXiv 2609.02246, an engineering report) measured the effect on a code-quality judge: rationale before score raised agreement with human experts from 42.6% to 51.9% with no rubric change (#231) - Changed: README prose in all three languages, each Chinese file written from the English and kept to its own locale's usage. The English text is the maintainer's, with the caveat about what was verified stated once instead of four times, the 95.5% → 93.9% figure scoped to StoryScope's LAMP-edited condition, and the voice-skills section stating external skills and the built-in profile separately; Shan et al. joins the Sources line. Sponsorship:
.github/FUNDING.ymland a Support section (#233, #234)
Not in this release: any content-word share, count, or threshold. The first draft of #232 had the executor compute and report one; eight Codex rounds and an architecture review showed a corpus-level effect size cannot be turned into a per-passage test, so the share, its definition and the Content words: line were removed and only the two tests shipped.
Observed, one specimen, not measured evidence: the same lightly AI-polished opinion paragraph refactored through v0.7.0 and through this release, then measured with a fixed English stoplist. Source paragraph content-word share 0.471; v0.7.0 output 0.505; this release 0.645. The new output is shorter with every point kept and applies the deletion test by name in its defect list. Transcripts and the stoplist are in #232.
v0.7.0
A new check in the style pass, a new load rule, and a new reference file, hence the minor bump. No entry changes; no configuration to migrate. The visible difference: the professional review's Style scan: line can now carry a rhythm finding, and Chinese text loads one more file.
- Added: a sentence-rhythm check in
style-pass.md§5. It looks for runs of three or more adjacent sentences of about the same length, in any language and unit of count, and reports them as a candidate signal that counts only alongside other hits. It sets no sentence-length cutoff: the studies behind it measure spread (human text varies more within a passage in every study that measured it, in English and Chinese), while mean length flips direction by model generation and by unit, so no length target exists to write down. Every non-fiction route now runs §5 with the §2–3 scan (#229) - Added:
references/languages/zh.md, loaded at the style-pass step when the target text is Chinese in any variant. One measured corpus (HC3-Chinese, 6,586 + 6,586 open-domain answers, GPT-3.5-era ChatGPT, 2023, 朱君輝 et al. CCL 2023): 語氣詞 density 0.003 vs 0.016 human, 連詞 0.036 vs 0.013, 「和」 11.76 vs 4.13 per answer, disyllabic share 0.532 vs 0.445, second person 0.021 vs 0.010. Hunt and restore lists are derived from that table and marked as inferences; the Taiwan editor observations (「不是…而是…」, 「○○性/感/化」, three-part parallelism, 「其實」 openers) are listed and marked unmeasured. No Taiwan academic study exists; the file says so (#229) - Added: three whitelist rows in
style-pass.md§7 for things readers keep asking about that the evidence does not support as signals. Punctuation density: two measures on the same Chinese corpus point opposite ways, and no per-type comma/period count exists for English or Chinese. Em dash: a per-release property (GPT-4.1 10.62 per 1,000 words, Claude Opus 4.6 9.09, GPT-5.4 1.43, Llama 0.00, human 3.23 from eight essays), so only clusters count and only the release's own prose layer applies. Paragraph count and length: direction contradicts across corpora (longer in how-to text, shorter in generated papers, more paragraphs in Chinese answers); only uniformity within one text remains a signal, asdiscourse-pass.md§3 already said (#229) - Changed: the Antigravity
/sepiaworkflow allowed only the reference files named in the routing table, which had excluded the model fingerprints and the voice registry as well as the new Chinese file; it now allows the conditional loads stated under the table (#229) - Research:
research/rhythm-syntax.md(zh-TW) and a new section ofresearch/sources.mdwith 16 ledger rows. Two findings worth knowing before you cite anything online: the "dispersion 16.4 vs 4.8" figures circulating for Shalevska 2025 are not in that paper, and GPTZero's "burstiness" has no published formula and was retired in 2023 (#229)
Not in this release: any per-passage rule based on short- or long-sentence counts. An earlier draft of §5 had one; six Codex review rounds showed that the corpus-level tail rates (measured in tokens on news leads, in words on science paragraphs) cannot yield a paragraph-level cutoff, so the detector was removed and only the run check shipped.
Observed, one run per specimen, not measured evidence: a Chinese and an English opinion paragraph reviewed through the working tree. The Chinese report loaded languages/zh.md and flagged four adjacent sentences of 35, 34, 36, and 32 characters plus seven connective openers; the English report flagged seven sentences of 18 to 21 words in a row. The same Chinese paragraph reviewed through v0.6.0 produced neither line. Transcripts in #229.
v0.6.0
A new experimental capability and a new entry, hence the minor bump. The three passes are unchanged.
- Added: a built-in voice profile, Hemingway (
references/voices/hemingway.md). Iceberg moves for fiction (leave the meaning out, emotion as action and speech, weather as weather, an occasional paratactic run against short declaratives) and the Kansas City Star rules for professional prose, each move traced to its source and to the StoryScope or LAMP axis it moves, with its known cost stated. Evidence inresearch/hemingway.md: the author's own statements, the 1915 style sheet, and criticism and corpus studies (Levin 1951, Smith 1983, Lamb 2010, Rice 2017, Ihrmark and Nilsson 2021, Lian et al. 2025), kept in separate tiers. Loads only when you opt in (#168) - Added:
/sepia-hemingway, a fifth entry on every platform: write or refactor fiction with the profile declared. On fiction, asking for strong de-AI also counts as opting in; sepia says which profile it is applying and that "no voice" runs plain sepia (#168) - Added: a
Voice fit:line in fiction review reports. It counts rubric rows already recorded against a profile's signature and, when three of five match with no over-correction advisory in the way, names the opt-in phrase. It never applies a voice. Professional routes do not get the line yet; #227 records what a closed report vocabulary needs first (#168) - Changed: the fiction report names every observed signal by its rubric row heading; the professional report gains a
Style scan:line for the §2–3 scan it already runs (#168) - Changed: refactor under a declared voice is defined: stage 1 unchanged, and in stage 2 a voice move may be applied only as the fix for an item on the defect list (#168)
- Docs: README badges for both workflows, latest release, and license (#84). Thanks @lunar-me for a grammar fix in discourse-pass.md (#147)
Observed, one run per language, not measured evidence: the same premise written twice by Gemini 3.8 Flash, once plain and once through /sepia-hemingway, then blind-reviewed. The profile version dropped the explained ending, the mirrored weather, and the unnamed places in favor of an open ending and real names and numbers, and paid the documented costs: interior access near the low pole and a uniform sentence recipe the uniformity finding still catches. Full pair and rubric tables: https://claude.ai/code/artifact/bb3a7bc0-df3d-4804-b58c-0fec6f094199
v0.5.0
Routing changes in this release, hence the minor bump.
- Added: a model-identity rule. sepia now resolves two identities before operating, the model that wrote the text and the model running the skill, each as family plus release or unknown, and loads the matching fingerprints. Identity comes from the user, from metadata, or from the executor's own system context, never from reading the prose: six-way attribution is a trained classifier at 68.4% macro-F1, and reading is not that classifier (#45)
- Added: a prose layer in
model-fingerprints.md, kept apart from the StoryScope-measured narrative layer and marked as vendor guidance, unmeasured. Each table is tagged with the exact release the vendor's prompting page names: Claude Fable 5.1, Fable 5, Opus 5, Opus 4.8; GPT-5.6; Gemini 3 series. A table is operative when the release matches, a prior otherwise. Anthropic's own mannered-prose instruction ("a dial worth turning" for "a parameter worth varying") is embedded verbatim and operative when Claude Fable 5.1 or Mythos 5.1 runs the skill (#45, #46) - Added: recorded absences. Sonnet 5, Opus 4.7, Opus 4.6, Sonnet 4.6, DeepSeek, and Kimi publish nothing about how their models write;
research/sources.mdrecords the pages consulted and the dates instead of leaving rows to be guessed at (#46) - Changed: review reports carry
Model:andProse layer:lines, one status per role; the fiction report also carriesNarrative layer:(#45) - Changed: the "Deletion beats addition" guardrail gains a register clause: a rewrite must not come out more promotional than its source (#45)
- Docs: Simplified Chinese README, thanks @javaht (#33). The English and Traditional Chinese READMEs describe the two-layer fingerprints
Observed on this tree, read-only: a Claude Fable 5.1 executor reported Prose layer: operative and replaced seeded metaphors with literal phrases; a Claude Opus 5 executor reported prior; Codex CLI reported executor=GPT version=unknown, prior, because its developer message names only the family. Antigravity and Grok identity exposure is untested and not claimed.
v0.4.1
The skill itself is unchanged: no routing, pass, or reference edits since v0.4.0. This release hardens the repository around it.
- Added: version-consistency check.
scripts/check_versions.pyverifies the three places declaring the skill version (SKILL.md frontmatter and both plugin manifests) agree, and runs in CI. Contributed by @AugustusW (#22) - Hardened: four rounds of review findings against that checker (duplicate keys, alternate YAML spellings, spellings one level up, escaped keys) ended with a structural fix: every scanned key must match one canonical grammar (
[A-Za-z0-9._-]+followed by:) instead of being tested against an enumerated blocklist, closing the class rather than the instance. 41 tests. Thanks @AugustusW for driving each round (#40, #41, #42, #43) - Added: behavioral eval in CI.
claude plugin evalruns a release-note de-AI-fication case with three graders (skill fired, no slop markers, reads human) on every push to main that touches the skill or the evals (#24, #27) - Docs: contribution requirements. Non-trivial changes need an approved issue first, with a private-reporting exception for security findings (#32, #38)
- Docs: install section claims any Agent Skills-compatible agent via
npx skills add Nanako0129/sepia; hands-on verification stays scoped to Claude Code, Codex, Grok, and Antigravity (#18)
v0.4.0
- Experimental: voice-skill composition. Opt-in interface for stacking sepia with a voice or style skill (minimalism methods, brand voices, persona guides). Declare the voice and sepia loads references/voice-skills.md on top of the normal route: architecture decisions first, 3–5 voice moves per piece with formula endings deliberately broken sometimes, review reporting the voice's known costs while uniformity findings keep full strength, and the venue keeping register precedence on professional routes (#15)
- Fixed: the voice-skills expectation table no longer hardens style-pass hits past style-pass's own clustering thresholds and whitelists (#16)
- Docs: both READMEs (English + zh-TW mirror) document the experimental feature
Grounding: one blind sepia review of a strict-minimalism specimen — a worked example, not measured evidence (see references/voice-skills.md).