Skip to content

Releases: HarperZ9/articulate

articulate-writing 0.5.0

Choose a tag to compare

@HarperZ9 HarperZ9 released this 26 Sep 16:39
d7d5244

A document folder can no longer run commands through judge, fix or
polish. Detector findings and the ruleset fingerprint do not change.

In 0.4.2 the claude -p child started in the caller's working directory. The
CLI reads .claude/settings.json from there, and -p skips the workspace trust
prompt. So a document folder that shipped a settings file ran its command hooks
on every editor call, on every platform, and its env block reached the
session. On Windows an npm claude.cmd shim added a second path: it runs
node by bare name, and cmd.exe looked for node in the working directory
before the PATH.

  • The CLI now starts in a new private folder that tempfile.mkdtemp creates.
    That folder holds the prompt file and an empty working folder for the child,
    and the editor removes it after the call.
  • Every call now passes --setting-sources user, so project and local
    settings never load. The full list is
    --setting-sources user --strict-mcp-config --tools "".
  • On Windows the child's environment sets NoDefaultCurrentDirectoryInExePath=1,
    so cmd.exe finds the shim's node on the PATH only.
  • For a batch shim, ) in an unquoted argument is now refused too. cmd.exe
    reads it as the end of a parenthesized block. A path with a space, such as
    one under Program Files (x86), is quoted and still runs. The check covers
    the prompt file's path as well as the CLI path, so a TEMP folder with &
    in its name is refused.
  • A temporary folder that cannot be written now raises ClaudeUnavailable,
    and the message leaves out the path. In 0.4.2 the OSError escaped as a
    traceback.
  • The editor searches the CLI's output for backend errors, such as a rate
    limit, only when the call fails. A fix of a document that mentions rate
    limits used to fail as ClaudeUnavailable. A call that exits nonzero now
    always fails, even when it printed output.
  • An older CLI that rejects one of the flags now raises ClaudeUnavailable
    with a message that says to upgrade. The flags are tested with CLI 2.1.251.
    The first version that accepts all of them is unknown.

Breaking for callers:

  • This release carries the minor version that 0.4.2 should have had. The
    0.4.2 changes for callers below were breaking, and a ~=0.4.1 pin picked
    them up as a patch.
  • The runner that claude_cli.run calls now receives cwd and env keyword
    arguments. A test double must accept them.

Tests: the new tests/test_claude_cli_batch.py, tests/test_claude_call.py
and tests/test_no_shell_calls.py and the updated CLI tests cover each change.
A unit test checks that the child's working folder is new, empty and apart
from the prompt file. A process test starts a stand-in from a folder that holds
.claude/settings.json and checks where it ran. A Windows test runs npm's own
claude.cmd template with a node.cmd planted in the caller's folder. A
Windows test runs every punctuation character through cmd.exe and checks that
the refused set matches what cmd.exe changes. The shell-call scan now also
catches ** keywords on process calls, a command whose first word names a
shell, COMSPEC and os.posix_spawn. Each of those has a sample it must catch.

Checked by hand, not in CI: the real CLI 2.1.251, started through
claude_cli.run from a folder with project SessionStart and
UserPromptSubmit hooks and with the API address on a closed local port, ran
both hooks under 0.4.2 and neither under this release. Not tested: whether a
folder's env block could have sent requests to another host, or whether a
CLAUDE.md there steered the judge. Both need files in the caller's folder,
which the CLI no longer reads.

articulate-writing 0.4.2

Choose a tag to compare

@HarperZ9 HarperZ9 released this 26 Sep 07:55
64cedf6

articulate-writing 0.4.2

The editor commands judge, fix and polish now find the claude CLI in
more setups, and the model call runs with no tools. Detector findings and the
ruleset fingerprint do not change.

The editor started the CLI by the bare name claude. That failed in two cases.
A process that an MCP host or a bundled app starts can inherit a PATH that
holds only System32. And on Windows, subprocess without a shell looks only for
claude.exe, so it never found the claude.cmd shim that an npm install puts
on the PATH. In both cases the start raised an uncaught FileNotFoundError.

  • New module articulate.claude_cli resolves the CLI. The environment variable
    ARTICULATE_CLAUDE_CLI names its path and wins when set. Its value must be an
    absolute path; a bare name or a relative path is refused.
  • Otherwise the resolver walks the absolute PATH entries itself. It takes
    claude.exe from any entry first, and a claude.cmd or claude.bat shim
    only when no entry holds claude.exe. So a setup that ran claude.exe
    before still runs it, whatever the PATH order. The current directory is never
    searched, and neither is a . or empty PATH entry. On Windows,
    shutil.which looks in the current directory first, so a document repo that
    shipped a file named claude.cmd would have run in place of the real CLI.
  • When nothing runnable is found, the editor raises ClaudeUnavailable. Its
    message names the variable and says the CLI must be installed and logged in.
    A start that fails with an OSError raises the same error. No message prints
    the value of the variable or the resolved path. A timeout message names only
    claude.
  • The prompt now travels in a temporary file passed with
    --append-system-prompt-file, on every path, and the file is removed after
    the call. The argument list holds a fixed line, that path and fixed flags. A
    prompt with many detector findings could pass the Windows command-line limit
    of about 32K characters, which failed with a misleading "could not be
    started" error. And a .cmd file runs under cmd.exe, which cuts an argument
    at the first newline and treats %, ^, &, |, <, > and ! as
    commands. For a batch file, a path with one of those characters is refused.
  • Every call passes --strict-mcp-config --tools "", so the model session has
    no built-in tool and loads no MCP server. The document is untrusted input,
    and before this change the claude -p child inherited the user's allow
    rules and MCP servers.
  • A timeout now stops the whole process tree: taskkill /F /T by its System32
    path on Windows, a process-group kill elsewhere. Before, a timeout on a batch
    shim killed only cmd.exe and then waited for the CLI to finish.
  • polish now catches a judge failure inside its rewrite loop, such as a rate
    limit halfway through, keeps the best version so far and exits cleanly.
    Before, that failure ended in a traceback.

Changed for callers:

  • The model's instructions now arrive as an appended system prompt, and the
    user turn holds a fixed line. The trust boundary text still ends the
    instructions, and the document stays on stdin.
  • editor.CLAUDE is removed, and setting it has no effect. The variable
    ARTICULATE_CLAUDE_CLI now names the CLI.
  • editor.run is removed. A test double that patched it would now reach the
    real CLI. A double now patches articulate.claude_cli.run.
  • A missing CLI raises ClaudeUnavailable, a RuntimeError, where it used to
    raise FileNotFoundError, an OSError. ClaudeUnavailable moved to
    articulate.claude_cli; articulate.editor still exports it.
  • claude_cli.resolve reads PATH from the environ mapping it is given.

Tests: tests/test_claude_cli.py and tests/test_claude_cli_process.py cover
the resolution order, a claude planted in the current directory, the closed
errors, the fixed argument list, every character cmd.exe reinterprets for both
.cmd and .bat, a 40,000-character prompt, and a timeout that must return
within 5 s while a grandchild sleeps. A parse of every source module checks
that no call asks for a shell. CI now also runs the suite on windows-latest,
where the stand-ins go through cmd.exe.

Not verified: a model reply through the batch path. A stand-in claude.cmd
around the real CLI 2.1.251 accepted the new arguments and reached the
backend, which answered with a credit error on the test machine.

articulate-writing 0.4.1

Choose a tag to compare

@HarperZ9 HarperZ9 released this 24 Sep 03:13
eefba5b

Fixed quadratic run time in the markup masks. Findings do not change, and the
ruleset fingerprint does not move.

Before the prose passes run, the detector blanks URLs, e-mail addresses, and HTML
tags on every line, plus quoted speech under a dialogue-exempt genre. It did this
with re.sub and patterns that backtrack quadratically on a long line where a
match never completes. strip_markup runs up to six times per line, so one such
line stalled check_text far past the 3 s budget the ReDoS test enforces.

Time for one check_text call, before and after:

  • "v1." * 10000 under flavored: 16.0 s before, 0.08 s after.
  • "1.1.1.1." * 5000 under flavored: 53.2 s before, 0.14 s after.
  • ("1.1.1.1." * 5000) + "@" under flavored: 47.8 s before, 0.14 s after.
  • "twenty-" * 8000 under flavored: 29.5 s before, 0.08 s after.
  • "<a " * 40000 under flavored: 23.5 s before, 0.33 s after.
  • "\u201ca " * 20000 under literary-fiction: 13.6 s before, 0.21 s after.

Before and after were measured on one machine, each pair in one session. That
machine was under load from other work, so the absolute numbers are noisy.

  • The e-mail part of the URL pattern retried every word boundary inside a long
    run of word characters, dots, or hyphens, and rescanned the rest of the run
    each time. The tag pattern rescanned to the end of the line from every
    unclosed <. The quote pattern did the same from every unclosed curly quote.
  • New module articulate.masking returns exactly what re.sub returns with
    each pattern, in linear time. The e-mail scan tests each run once, because
    every start inside one run succeeds or fails together. The tag mask stops at
    the last > in the line, since no tag can start after it. The quote mask
    skips an opening mark once one of its kind has failed to close before the
    next newline.
  • The patterns themselves are unchanged and remain the definition of record as
    detector.URL, detector.TAG, and detector.QUOTED. tests/test_masking.py
    compares each mask with re.sub on 4,000 generated lines and a set of hand
    cases. It also holds each mask, and the detector functions strip_markup and
    mask_quotes that apply them, to 0.25 s per call on long hostile lines, so a
    call site that goes back to re.sub fails as well.
  • tests/test_redos.py adds the inputs above, so the 3 s per-input budget now
    covers them. The tag and curly-quote inputs use 40000 and 20000 repeats. With
    10000 repeats the old code took 0.3 s to 0.6 s on the tag input and 3.3 s to
    5.0 s on the quote input, so a regression there could pass the budget.
  • Checked for identical output: full check_text results (rule ids, spans,
    labels, gate, texture, cadence) and per-line masks were compared before and
    after on every tracked file except the VS Code extension's lockfile, under no
    profile and under all 18 profiles, 6 genres, and 29 modes. The same check ran
    on 612 further documents (Markdown, HTML, SVG) under 9 configurations. No
    result differed. The benchmark output is byte-identical.
    The fingerprint hashes the tier patterns and leaves the mask patterns out. It
    reads sha256:9f78a7484bb20f84 before and after, so RULESET_SEMVER stays
    0.5.0.

Tests: 181 pass.

articulate-writing 0.4.0

Choose a tag to compare

@HarperZ9 HarperZ9 released this 23 Sep 04:08
7151097

The MCP surface now works from a bare install

articulate-writing 0.3.0 shipped an MCP entry point that could not start. It
imports fastmcp, which pyproject.toml declares under the mcp extra, so a
plain install produced a server that raised ModuleNotFoundError the moment a
host launched it. The install reported success and nothing failed until someone
tried to use it.

$ pip install articulate-writing        # 0.3.0
$ python -m articulate.mcp_server
ModuleNotFoundError: No module named 'fastmcp'

What 0.4.0 adds

articulate.local_mcp serves the same five tools over stdio JSON-RPC 2.0
(protocol 2025-06-18) using nothing but the standard library, plus
articulate.status and articulate.doctor. New console script:

pip install articulate-writing
articulate-mcp

doctor reports which tools run local (check, score) and which need an LLM
backend (judge, fix, polish), so a host with no backend knows what it still
gets rather than finding out per call.

The FastMCP server remains available under the extra:
pip install "articulate-writing[mcp]", then articulate.mcp_server.

One implementation, two transports

The tool bodies stay in mcp_server as do_check, do_score, do_judge,
do_fix and do_polish. Both transports call them, so behavior cannot diverge.
A test parses mcp_server.py for its @mcp.tool functions and asserts the stdio
server exposes every one, which runs in CI where the extra is absent.

Version alignment

The version lived in three places with nothing binding them, and the MCP server
reports the module value in serverInfo. A release could have shipped 0.4.0
while every client was told 0.3.0. tests/test_version_alignment.py now binds
pyproject.toml, articulate.__version__, the value initialize returns, and
the CHANGELOG heading.

Release gates

publish.yml pins every action by commit SHA, checks the release tag against the
declared version, records artifact digests, resolves every console script in a
clean venv, rebuilds a wheel from the sdist, and probes the stdio server from a
bare install before anything uploads. That last gate is the one that would have
caught this defect.

Verified on the runner before release: 3 entry points resolved,
stdio MCP ok: articulate 0.4.0, 7 tools.

articulate-writing 0.3.0

Choose a tag to compare

@HarperZ9 HarperZ9 released this 20 Sep 01:10

Articulate detects AI tells and prose devices, and rewrites AI-assisted writing to
a defensible, register-adaptive standard. The core detector runs
standard-library-only with no network call. A clean gate means the prose was
screened under a named ruleset. Articulate is not a detection-evader, and it makes
no claim that text is undetectable or human-authored. Disclosing AI assistance is
the user's responsibility.

What 0.3.0 adds

Three cadence detectors for prose that reads clean under the earlier ruleset yet
still carries a machine rhythm a reader hears.

  • Demonstrative summary-beat, "That is the ... half / side / piece / layer / angle
    / story / trick." Tier MEDIUM, so it gates under a strict profile. A compound noun
    like "That is the side effect" does not fire.
  • Two-imperative parallel slogan, "Verb the X, verb the Y." Tier LOW and advisory.
  • Evaluative fragment opener, "Strong foundation." Tier LOW and advisory,
    paragraph-initial only.

Detector ruleset semver moves 0.4.0 to 0.5.0, and FRAGMENT_OPENER folds into the
ruleset fingerprint, so a replay against an older receipt reports the change. A
verse genre suppresses the two new LOW categories. Every new pattern uses bounded
quantifiers and stays covered by the ReDoS safety test. The demonstrative extension
fires on none of a 26-snippet human control corpus.

Verification

151 tests pass. The build and the twine check pass. The shipped README and
CHANGELOG pass Articulate's own writing gate.

Install

pip install --upgrade articulate-writing

Source-available under FSL-1.1-MIT. Full notes are in CHANGELOG.md.

articulate 0.2.0

Choose a tag to compare

@HarperZ9 HarperZ9 released this 19 Sep 04:17

Comprehensive AI-tell detection.

This release widens the detector from banned constructions to the documented space of LLM writing tells: lexical, structural, cadence, delivery, and formatting. Tells are tiered HIGH, MEDIUM, and LOW, with stdlib statistical advisories for burstiness, opener diversity, n-gram repetition, and formatting density. It holds a strict false-positive line, verified against a human control corpus that stays clean.

  • 18 HIGH, 55 MEDIUM, 12 LOW tells, plus 7 statistical advisories.
  • 143 tests. Ruleset semver 0.4.0. Bounded-quantifier regexes, ReDoS-safe.

Install: pip install articulate-writing==0.2.0

articulate 0.1.0

Choose a tag to compare

@HarperZ9 HarperZ9 released this 18 Sep 13:06

Articulate is a local writing-quality and AI-tell detection and editing tool. It
flags the prose devices and machine-writing tells that make text read as
generated, scores how machine-textured a passage is, and rewrites prose to a
plain, skilled standard when an LLM backend is configured. The core runs on the
standard library with no network call. Detection quality and writing quality are
the goals. A detector score is a benchmark and a byproduct, never a target, and
this is not an evasion tool.

This is the first published release.

Install

pip install articulate-writing

The install adds two console scripts: articulate (the CLI) and articulate-lsp
(the language server). The detector, profiles, modes, genres, receipts, SARIF,
and LSP need no third-party package. The editor layer needs an LLM backend, and
the MCP server needs pip install "articulate-writing[mcp]".

What is in this release

Detect. Flags banned rhetorical devices (antithesis, including keyword-free
parallel-negation contrast pairs, corrective negation, rule-of-three, em-dashes,
filler intensifiers, corporate verbs), current frontier-model register, marketing
and email or blog tells, and Williams and Orwell signals (expletive openers,
nominalization density, passive voice, adverb density, cadence uniformity, opener
repetition). Emits span-level findings, a graded 0-100 texture score, and a
clean or flagged gate.

Adapt by register. A profile system (procedure, commit, research, readme,
essay, narrative, and more) sets which findings block. Fiction gates nothing.
Procedures and essays gate strictly. Profiles resolve from --profile, an
in-file writing-profile: tag, or the file path.

Modes and genres. A writing mode crosses a domain register with an
articulation need, such as memo/argue or technical-docs/explain. The genre
axis reads narrative and expressive prose by its own convention: literary-fiction,
genre-fiction, ya-fiction, memoir, screenplay, poetry. Under a fiction
genre, quoted speech is masked out of the device passes, the craft devices report
but never block, and a report-only lexicon flags generation artifacts such as the
somatic cliche. Screenplay classifies Fountain roles first, so only action lines
face the gate. Poetry reads by the line. Run articulate modes to list them.

Scientific and mathematical writing. academic/prove and
science-writing/explain target hard technical exposition: stating the idea
before the formalism, keeping a roadmap, defining each symbol once. The proof mode
does not rewrite by default, because a wrong change to a quantifier order or an
inequality direction changes a theorem. On a .tex file the editor masks every
math span before a rewrite and splices it back byte for byte.

Edit. judge reads the judgment-level failures a regex cannot see (confident
emptiness, vague abstraction, hedging with no position, weak verbs). fix
rewrites to the standard, self-checked against the detector. polish loops until
five qualities clear a bar, and a rewrite is kept only if it lowers no quality
score. The editor is gated on writing quality, never a detector score, and needs
an LLM backend.

Re-derivable receipts. receipt emits a verdict that replays to Match, Drift,
or Unverifiable against a pinned, versioned ruleset. A content-free audit receipt
keeps no verbatim text: it drops the matched substring and the exact offsets and
keeps only which rule fired, its tier and category, and the line, so a team can
retain and replay a record without storing the source. audit queries committed
receipts locally with no server and re-verifies that each source still holds.

CI and editor surfaces. SARIF 2.1.0 output for GitHub Code Scanning, Azure,
and reviewdog. A standard-library LSP server for VS Code, JetBrains via LSP4IJ,
and Neovim. A pre-commit hook, a GitHub Action, and an MCP server.

Boundaries

A clean gate, a low texture score, or a Match receipt means the prose was screened
under a named ruleset. It says nothing about whether a claim or a theorem is true.
A clearly written proof can still be false, and Articulate never checks the
mathematics. Correctness comes from referees and proof assistants, never from this
tool. A receipt attests that a ruleset ran and re-derives; it is not a compliance,
provenance, or verified-human claim. Content-free is not zero-leakage: which rules
fired and the line remain, which for a closed-vocabulary rule narrows the flagged
word to that rule's small public candidate set. Use drop when the flagged word
must stay secret.

Not in this release

A local-model editor backend and an --offline flag for air-gapped use are on the
roadmap. A labeled non-native-English corpus for a fairness check is not built, so
the tool makes no fairness claim yet. The self-hosted and hosted audit tiers, with
retention and access controls, come only after a paying signal.

Verified

121 tests pass on Python 3.9 through 3.12. The regression-gated benchmark holds at
100% recall, 100% specificity, and 0 regressions on its corpus.

Links