Releases: yottayoshida/modern-python-guidance
Release list
v1.3.0 — the hook says when it could not check a file
Summary: The PostToolUse hook was silent in exactly the same way for a clean file and for a file it could not check, so a broken hook read as a healthy one from outside (#209). It now says so — one line, shown to the user and not to Claude, the edit never blocked — and a clean file still writes nothing.
Added
- The PostToolUse hook now says when it could not check a file. A file over 2 MiB, one it could not read, a binary file, or a bundled catalog that loaded no guides used to produce exactly what a clean file produces — nothing on stdout, nothing on stderr, exit 0 — and a path it could not reach or a detector raising ended in a traceback of which Claude Code shows only the first line; either way a broken hook and a healthy one were indistinguishable from outside, and an installation could stay that way for months with nobody the wiser. The hook now writes a one-line
systemMessage, the field Claude Code shows to the user and not to Claude, naming the reason; the edit is never blocked, the exit status stays 0, and a clean file still writes nothing. Stderr was never an option: on exit 0 Claude Code sends it to the debug log only. Exiting non-zero was rejected because the transcript renders that ashook error … Failed with non-blocking status code:, which reads as a broken hook rather than an unchecked file, andadditionalContextwas rejected because the reader is the person, not Claude. Two things moved with it.Path.is_file(), which decided whether a path was worth checking, raised on a parent directory the hook cannot enter on 3.11-3.13 — a traceback, of which Claude Code shows only the first line — and answered False on 3.14, which is silence; the decision is now made withstat(), where only "nothing there" (ENOENT,ENOTDIR, and a name no filesystem can hold) stays silent, so a symlink loop now says so too. And the catalog load moved inside the same guard, which now also refuses an empty catalog —build_index()never raises; a missing or hollow guides directory is logged and comes back empty, and an empty catalog finds nothing in any file, which is the one shape of a broken install that reads exactly like a clean file — with the guard widened to any exception, since a hook must never break an edit and a detector's traceback told the user nothing; the message keeps a foreign exception's type. It passes through the same escaperdoctorwrites through since #245, and the JSON stays ASCII, so a path or an OS message carrying a newline or an escape sequence cannot add a line or redraw one.mpg check <file>is named in the message only for the failures it reproduces — too large, unreadable, binary — because for the otherscheckends in a traceback of its own. There is no per-file switch: the hook says so on every edit, andmpg setup --no-hookremoves it altogether. (closes #209)
v1.2.1 — doctor keeps its own report
Summary: mpg doctor is the command someone runs because they do not trust their tree, and until now the tree could take its report away: a matcher could make it stop answering (#242), a newline in a registered command could write a fabricated line into the report (#245), and an over-long command ended it in a traceback (#246). Those three, and twelve more of the same shape found while fixing them, now end in a report doctor wrote itself, with the exit status from the table.
Fixed
mpg doctorcould be stopped, crashed, or made to print lines it did not write by what a project holds — and a project can arrive bygit clone, since git carries symlinks andgit add -fputs even a git-ignoredsettings.local.jsonin a commit. The three reported cases, and nine more of the same shape found while fixing them, now all end in a report: one line per channel thatdoctorwrote, and the exit status from the table. Stopped: a matcher like".*" * 200 + "Z"backtracked inrefor more than twenty seconds, andre.compilealone is quadratic in branches sharing a prefix — thirty to sixty-eight seconds for a million characters. The portable subset has no groups, classes, or escapes, so its syntax check and its search are now written out directly and run in time proportional to the matcher; both were compared withreon every pattern of length five or less over an eleven-character alphabet — the subset's metacharacters, the newline, and letters from both tool names (177,156 patterns for syntax, 285,754 searches, no disagreement; a sixteen-character sweep during review, against seven subjects, found none either), and a test repeats that to length four. Asettings.local.jsonthat is a fifo made the read wait for a writer that never came; it is now opened non-blocking and refused as not a regular file (degraded). Crashed: acommandor a link destination the OS will not look up — 100,001 characters, a component past the name limit — raised out ofPath.exists()on 3.11-3.13; it now readsdegraded, since Claude Code meets the same wall, with the detail saying "does not exist or cannot be reached". That includes a link which compares equal to the bundled source and cannot be followed —<300 characters>/../<source>resolves to the source as a string while the OS refuses the path — which 3.11-3.13 answered with a traceback and 3.14 with "the link is right and there is nothing behind it" and a fix of reinstalling mpg. A--project-dirreached through a directorydoctorcannot enter now reads as nothing inspected (exit 2) instead of a traceback on 3.11-3.13. A settings file Python can open but not finish parsing — bytes that are not UTF-8, an integer past the digit limit, nesting deeper than the stack,NaNorInfinity, more than 1 MiB — now readsunknownrather than ending in a traceback, because Claude Code's parser may accept what Python's refuses; a JSON syntax error is stilldegraded. Read as healthy: a.claudethatdoctorcannot look into — an unreadable directory, a symlink loop — answeredabsenton 3.14, which exits 0, because 3.14'sPath.exists()folds every OSError into False; 3.11-3.13 raised instead. Both now readunknownon every interpreter, from a lookup that calls something absent only when the OS says nothing is there. The project root is found the same way, so an unreadable marker isunknownrather than skipped — skipping it haddoctordiagnose a directory further up,$HOMEwhen it was measured. Behind all of that, each channel is diagnosed inside its ownexcept Exception, so a failure nobody has found yet reports that channel asunknownwith the exception named and leaves the rest of the report intact. Lines it did not write (#245): a newline in a registeredcommandput a fabricatedmcp presentline on screen with no terminal involved, and an escape sequence could redraw one. Everythingdoctorprints now goes through one function that shows control, format, separator, and surrogate characters as escapes (\x0a,\u202e; an undecodable byte from a link as the byte,\xff), and output to a stream that is not UTF-8 is written as ASCII — under cp1251 an ellipsis is byte 0x85, which is NEL. A consequence: a multi-line error fromclaude mcp getnow reads as one line with\x0ain it. What--run-interpreterruns, and whateverclaude mcp getstarts to answer for the MCP channel, are outside this; whichclauderuns, and what it runs with, is not —claudeis now looked up and run with the relativePATHentries dropped, because those resolve against the current directory and so may be the project's: aclaudefound only through one is not run (the MCP channel readsunknown), and an npmclaude, a#!/usr/bin/env nodescript, no longer picks up a project'snode_modules/.bin/node. (closes #242, #245, #246)- The README now says that the MCP server
mpg doctorstarts, throughclaude mcp get, may be the project's own: in a folder trusted in Claude Code, a project whose.claude/settings.jsonorsettings.local.jsonapproves its.mcp.jsonservers gets its server started — and its command run, as the user — whendoctorasks. Measured with Claude Code 2.1.268, in a throwaway home directory: without trust or without that approval the server stays pending and is not started. It is Claude Code's behaviour, whichdoctorinherits; the MCP paragraph used to say only thatclaude mcp get"starts the server". mpg setup,mpg setup --no-hook, andmpg uninstall— every command that readssettings.local.jsonto edit it — waited forever on one that is a fifo, and ended in a traceback on one they could open but not parse; they now refuse both, like any other settings file they cannot safely edit, and refuse one larger than 1 MiB for the same reason. A file holdingNaNorInfinity, which they used to read and write back, is refused as well: neither is JSON. This covers whatjsonitself rejects: nesting shallow enough forjsonand deep enough to exhaust the stack when the settings are copied for merging is not covered.
v1.2.0 — Checks that exercise what they report
Summary: doctor reported on things it had not looked at, and three of those gaps close here. The bundled assets are opened before setup links them, a hook matcher is only answered from a subset both engines were measured to read alike, and — behind a flag, because it means running a string out of a settings file — the registered interpreter can be asked whether mpg actually loads in it.
Added
mpg doctor --run-interpreterruns the hook's registered interpreter and reports whether mpg actually loads in it. Without the flag nothing is executed, which is the same behaviour as before and the reason the flag exists: closing this gap means running a string out of a settings file, and a settings file can arrive with a cloned repository. What is checked is the output rather than the exit status — a script that ignores its arguments and exits 0 passes every shape check and every exit-status probe while loading nothing, measured before the check was designed around it — sopresentrequires a line naming mpg and a version. Any version counts: a hook wired to an interpreter holding an older mpg is working, and comparing against the running installation's own version would repeat the mistake the symlink and MCP channels each avoid deliberately. A non-zero exit, output that is not a version line, or an interpreter that cannot be started at all readsdegraded— the last one because Claude Code spawning the same hook meets the same wall, so it is a measured failure rather than an unmeasured one; a timeout, a failure to start for reasons of the running process, or more than four distinct commands to try, readunknown. The execution is bounded: no stdin, four kilobytes of output read, five seconds per interpreter, its own process group killed on timeout, an environment holding onlyPATHandHOME(enough for a pyenv or conda shim to start, and nothing that carries a secret), a scratch working directory, and four interpreters at most. Three of those bounds exist because review broke the first version without them: draining the pipe after a timeout waits on an escaped grandchild that still holds it, and the command never returned; draining it at all reached 2.3 GB resident againstcat /dev/zero; and accepting any space-free token as a version let a child print an erase-line escape and rewrite doctor's own verdict on the terminal. What the bounds do not cover is written beside them rather than left to be discovered — a child that callssetsid()outlives the group kill, the network is open to it, and it can write anything the invoking user can write. (closes #236)
Fixed
mpg doctorreportedpresentfrom evidence that did not establish it, and now measures what it claims. The two symlink channels were judged by the name of the target: a directory calledmodern-python-guidancewith nothing in it, or an emptymodern-python.md, read as healthy — the decay the module's own docstring opens by naming, a link whose target moved still resolving as a name. Both paths topresentnow open the target and read a byte through a non-blocking descriptor confirmed to be a regular file, so aSKILL.mdsymlinked to a fifo cannot hang the command instead of answering it. The hook channel looked at whether the interpreter path existed and at nothing else: it never readargs, never readtype, and never read thematcherthat decides whether the hook fires at all, so a registration pointing atBash— one that cannot run on an edit — reported healthy, and so did/, which exists and is not a file. It also read only the first mpg entry in the file, whilemerge_hookhas always promised to converge "from ANY starting state", meaning a second, broken registration sat behind a healthy verdict. Every mpg entry is now examined, and the count that matters is per hooked tool rather than per group:matcher: "Edit"besidematcher: "Write"covers what mpg's own matcher covers and is not a duplicate, while two entries inside one group are. Matchers follow the documented rule, where simple characters mean an exact name or a|/,-separated list of exact names and anything else means an unanchored regular expression; what this process cannot evaluate is reportedunknownand neverdegraded, because Claude Code evaluates the regular-expression form in JavaScript and calling a matcher broken on the strength of Python's disagreement would report a working registration as broken. Failing to compile turned out to be only half of that problem, and the narrowing that closes the other half is the entry below.doctorexecuted nothing at all when this shipped — the interpreter path in a settings file is an arbitrary string, and running it to find out whether the hook works would make a read-only diagnostic a way to run whatever a settings file names, including one that arrived with a cloned repository. The gap that left was stated in the README rather than hidden, and the entry below closes it behind a flag rather than by default. (closes #231, #232, #233, #234)link_statepromised in its docstring that a link it cannot walk is answered as stale, andos.readlinksat outside thetrythat made that true — a second syscall against a pathis_symlink()no longer owns, which took the caller down instead of classifying. Moving it inside widened whatstalecovers, andsetupdeletes on that verdict, so the fix had to travel: a path that stopped being a symlink is now refused the way the flattened branch already refuses it, rather than unlinked. Measured against the previous code, which printedAgent Skills linkedand returned success after deleting the real file it had been pointed at.mpg setuplinked a project at a bundled source it had never opened._find_skills_diraccepted the directory onis_dir()and_find_rule_sourceaccepted the rule onis_file(), so a packaging accident shipping an emptyskills/modern-python-guidance/satisfied both: setup printedAgent Skills linked …and exited 0, and onlydoctor— run later, if at all — reported that the installation delivers nothing. Both linking call sites now open the source and read a byte before creating the link, using the same content predicatedoctormeasures links with, and refuse with the hollow path named;--dry-runrefuses too, rather than promising a link the real run would decline. The locators themselves stay shape-only, which is the opposite of what the issue proposed: a content predicate inside them stopsdoctorlocating a hollow source at all, so the channel answersunknown("cannot locate") before reaching the branch that says the link is right and there is nothing behind it — turning a measured breakage into an unmeasured one, anddoctor's exit from 1 into 2, against what this file's own[Unreleased]entry above and the README both promise. That direction is now held by a test walking the real locator rather than a patched one; every existing test monkeypatches the locators and none would have caught the regression. The predicate moved tosetup_cmd, besidelink_state, the other classification setup and doctor share, and opens withgetattr(os, "O_NONBLOCK", 0): the bare attribute raisesAttributeErroron Windows, past everyexcept OSErrorin the callers, which would have takenmpg setupdown before it reached MCP or hook registration — a platform the README already documents as losing the two symlink steps and keeping the rest.scripts/verify_wheel_assets.pygained the same check on both assets; the rule file had none at all, so an emptymodern-python.mdshipped undetected. Three test fixtures createdSKILL.mdwithtouch(), pinning "an installation that delivers nothing sets up fine" — the bug itself — and now write content. (closes #238)mpg doctoranswered for an engine it does not run. The regular-expression form of a hook matcher was evaluated with Python'sreand the result reported as what Claude Code would do, but Claude Code evaluates it with JavaScript'sRegExp, and the two languages differ at the edges:(?P<x>Edit)|Writefires in Python and is a syntax error in JavaScript, so the hook never runs while the channel readpresent— a silently wrong answer, which is the one failure mode a diagnostic cannot afford. That form is now evaluated only when the pattern is written in a subset measured to mean the same thing in both engines, and only against a plain tool name; anything else readsunknown. The subset was drawn by exhaustive sweep against node and CPython rather than by reasoning about which constructs differ, and the difference mattered: a first subset chosen by hand admitted( ) + ?and let through every possessive quantifier —Edit*+fires in Python, which added them in 3.11, while JavaScript has no such syntax — disagreeing on 555 of 69,104 admitted patterns, every disagreement a possessive, on every interpreterrequires-pythonallows. Dropping+and?makes those unspellable and dropping(alongside keeps the test single rather than a set plus a(?exception: 271,452 patterns of length five or less, no disagreement. The cost runs one way and is deliberate — matchers that do work, like(Edit|Write)or\d, now readunknown(exit 2, visible) rather thanpresent(silent when wrong), and the channel's detail says whatmpg setup --with-hookwould do, since a matcher whose evaluation is what failed carries nofix. Two limits are stated rather than implied: the simple form (Edit|Write,Edit, Write) reaches no engine on either side and is still answered from the documented rule, and the subset does not make evaluation terminate —".*" * 200 + "Z"is admitted and does not answer within 20 seconds, which predates this change and is tracked in #242. (closes #237)
v1.1.0 — Doctor for the channels that fail silently
Summary: mpg setup writes four delivery channels and every way they decay is silent, so the only symptom is that guidance stops arriving. mpg doctor reads each one back and says which are working, which are broken, and which could not be determined — reporting only, never repairing.
Added
mpg doctorreports the state of each delivery channelsetupwrites — the MCP registration, the Agent Skills symlink, the rule symlink, and the PostToolUse hook — and repairs none of them. Every way those decay is silent: a flattened symlink still has content, a link whose target moved still reads as a name, an unregistered hook simply never fires. Nothing raises, so the only visible symptom is that guidance stops arriving, with no signal pointing at the cause — the workspace this was measured against had all four broken for two and a half months while looking, from the outside, like a tool that did not work. Channels read aspresent,degraded,absent, orunknown.absentis deliberately not a failure:--mcp-only,--skills-only, and--no-hookeach make an unwritten channel a configuration someone chose, and counting it as breakage would make a red result mean nothing.unknownis deliberately not health: an empty report set exits 2 rather than 0, because a run that evaluated nothing has not established that anything works. The MCP verdict is read fromclaude mcp get'sStatus:line rather than by comparing the registered command against this installation's paths. That comparison was written first and measured wrong — it describes whichever interpreter invokeddoctor, so a user who installs with uv and runsdoctorout of a source checkout sees healthy links reported as degraded, which is what the first run against an already-repaired workspace did. Whether the registration connects is a property of the registration, so the answer does not depend on who asks. The cost is thatclaude mcp getstarts the server to answer (1.5s, against 0.05s forclaude --version);claude mcp listwould connect to every configured server and took 30s, so it is not used. (closes #211)
Changed
- The predicate
setupuses to decide whether a delivery symlink is already the one it would write is now a named function shared withdoctor, rather than an inline test duplicated acrosssetup_skillsandsetup_rules. The two readers ask different questions of the same classification —setupreplaces a link pointing anywhere else,doctoraccepts one pointing into another working installation — and having that difference is only safe because the classification itself lives in one place. One behaviour changed in the extraction: the inline form calledPath.resolve()without catching anything, so a symlink loop raised out ofmpg setuprather than replacing the link. How that surfaces depends on the interpreter — measured, 3.12.12 raises RuntimeError while 3.13.14 and 3.14.6 return the input unchanged and reach the same verdict without an exception — so the regression test asserts the outcome rather than the mechanism, and only demonstrates the fix on the versions that raise. That case, and the OSError a hostile or racing tree can raise, now classify as stale and are replaced. No existing test covered the path.
v1.0.1 — Holders for the last three surfaces
Added
-
The three surfaces 1.0 left out of the freeze now have something holding them, and are frozen:
check --format json,detect-version --format json(with the MCP tooldetect_python_version), and exit codes. They were excluded for one reason — nothing compared them against a documented shape — and the fix had to start with the comparison rather than the declaration.checkanddetect-versionare held by a recursive field-path comparison against the examples in design.md, using==rather than the<=the three existing JSON surfaces use. That direction matters: deleting a field from design.md shrinks the documented set, and a smaller subset still fits, so the older comparisons hold the serializer to the document but not the document to the serializer — measured, not assumed. Walking the whole value rather than the top-level names is what putstarget_python.sourceandmatches[].dependency_compatibility.statusunder the freeze; a comparison of outermost keys would let either vanish. Thesourcelabels are frozen as a vocabulary too, compared againstPythonVersionSourcefrom the schema section alone — the precedence section already lists all five, so a file-wide search would have reported agreement no matter what the schema section said — and separately checked by running one input per label, because aLiteralenforces nothing at runtime and a documented label no input produces would otherwise pass. Exit codes are frozen as the rows of a table, not as semantics in general. Comparing a table against a set of scenarios shows the document and the tests agree; it cannot show no other exit exists, and reading the exits out of the source would not help either, sincesetupanduninstallexit with a variable, argparse produces its own, and a signal never reaches the interpreter. What falls outside is written beside the table: argparse's exits, termination by signal, uncaught exceptions, andhook, whose status the PostToolUse contract already holds.BrokenPipeError → 0is not among the guarantees —main()restores the defaultSIGPIPEdisposition, so a closed pipe terminates the process by signal (141 as a shell reports it,-SIGPIPEin a parent reading the raw status) and theexceptclause is not what a caller observes. Adding an exit condition is not treated as additive the way a new JSON field is: an old client ignores a field, but a new condition can change what an existing input returns. (closes #224) -
The two places the version is written —
pyproject.tomlandsrc/modern_python_guidance/__init__.py— are compared. They are edited by hand and nothing read them back against each other; a bump that updated only the first has shipped before, putting a wheel on PyPI whose--versionreported the previous release and spending a version number to correct it. Every other packaging check derives frompyproject.tomlalone, so all of them stayed green while the two disagreed. The comparison uses the imported value rather than the file's text, since what a caller sees is whatimportproduced.
Changed
-
design.md's rule that the CLI "defaults to JSON when piped" now names its exception.
detect-versiontakes--format json|plainrather thanjson|humanand defaults toplainwhether or not a pipe is attached, because the plain version string is what scripts read. The behaviour is unchanged; what changes is that the general rule no longer contradicts it — freezing the surface while the document described it wrongly would have frozen the contradiction. -
The README listed five frozen surfaces; it lists six, and points at the two distinctions VERSIONING now draws — which side of each JSON surface is actually compared, and why exit codes are frozen row by row rather than as semantics. A list kept in two places drifts in one of them, which is what it did between the change that added the sixth surface and this release.
No behaviour changed. The only difference the wheel carries over 1.0.0 is the version string itself — skills/ and rules/ are untouched, and the sole edit under src/ is __version__. What actually moved is the documentation of what the version number promises, which the source distribution carries and the wheel does not.
v1.0.0 — Frozen surfaces, each with a holder
Summary: 1.0 states what the version number promises. docs/VERSIONING.md names five frozen surfaces — the CLI, the MCP tool schemas, the JSON output field sets, the guide frontmatter schema, and the PostToolUse hook stdout contract — and, for each, the test that holds it. Reaching that point meant repairing what the packaging had been claiming falsely: the Typing :: Typed classifier had shipped for 35 releases without its PEP 561 marker, no platform was declared anywhere, and four frequency: high guides reached no session at all because their only route was a catalog that measurement found unused. Nothing here changes the CLI or the MCP surface for existing callers; what changes is that those surfaces are now promises with something behind them.
Added
docs/VERSIONING.mdstates what a version number promises. Five surfaces are frozen from 1.0 onward — the CLI surface, the MCP tool schemas, the JSON output field sets (additive only), the guide frontmatter schema, and the PostToolUse hook stdout contract — and each names what holds it in place, because a frozen surface nothing reads back is the failure this project already made withTyping :: Typed. Three were held already: by the design.md schema comparison, by the hook's stdout tests, and by the frontmatter parser refusing violations. The CLI surface and the MCP schemas were not, and each gained a test — command names, positionals and option names read back frombuild_parser(), and every tool's wholeinputSchemacompared against a snapshot, so alayerchanging from integer to string or alimit.maximumcut from 50 to 5 fails rather than slipping past a comparison of parameter names. Three things are deliberately outside the freeze because nothing compares them today:checkanddetect-versionJSON output, and exit-code semantics; naming them would have made the document's own claim false in its opening paragraph, and widening a freeze later is not a breaking change.maxItemsis excluded for a different reason — derived from the catalog size, so freezing it would freeze the catalog — with a control test pinning the premise of that exclusion. The Python import API and the guide catalog are named as not frozen, ids included, sinceretrieveselects by them and a rename is a breaking change recorded here rather than made quietly. (closes #203)
Changed
- The
Development Statusclassifier moved from3 - Alphato4 - Beta, which 35 tagged releases, a four-version CI matrix and a 92% coverage gate already described; a release that self-describes as Alpha contradicts itself on its own PyPI page. A test now derives the required value from the version inpyproject.toml, so the eventual bump to 1.0 cannot land while the classifier still says Beta. Prereleases of 1.0 are exempt, since1.0.0rc1is where a project finds out whether it is stable and requiring the stable classifier there would force the claim ahead of the evidence. The issue asked for a line item in a v1.0 release checklist; no such checklist exists in this repository, and a test fires in the same commit as the version bump, which a checklist line can be skipped in. (closes #207)
Removed
docs/superpowers/and the two planning documents inside it. They were the plan and design used to file five issues on 2026-08-06; those issues were filed, implemented in #184–#189, and closed, and the spec's text is the issue bodies verbatim — #179 and its siblings carry it, so removing the drafts loses nothing the project relies on. No document in the repository referenced the directory, it never reached the wheel, and the plan opened with an instruction addressed to whatever agent read it next, which a public repository has no reason to keep offering. Recover withgit log --all --full-history -- docs/superpowers/. (closes #217)
Fixed
-
Four
frequency: highguides now reach a session.dataclass-modern,pytest-parametrize,ruff-over-flake8, anduv-over-piphad no detector — so the hook could never surface them — and no entry in the embedded-patterns section, which left the MCP catalog as their only route, and #152 measured that route as unused. All four are now carried by the Rules file, which loads on Python files and on project config (pyproject.toml,requirements*.txt,setup.cfg,.python-version,Pipfile), so the two toolchain guides arrive exactly when their own trigger files are being edited. The always-loaded body grew from 674 to 771 tokens; that is paid on every matching edit, which is why the routes were recorded on the issue before implementation. Adding detectors was considered and rejected: the guides presentfrozen/slots/kw_onlyas decisions rather than corrections, so an automatic finding would fire on correct mutable dataclasses. Each embedded line is now checked against the wording of its own guide, after a draft recommendeduv sync— a commanduv-over-pipnever mentions, and whose GOOD section keepsuv pip install -r requirements.txtrather than asking anyone to abandonrequirements.txt. (closes #208) -
The README no longer leaves out the delivery path that actually works. Its first line offered "MCP, CLI, or Agent Skills" without naming the Rules file, and the Highlights entry said Rules auto-inject on
.pyfile touch — both inaccurate, since Rules also load on project config and a scan of session logs found the MCP catalog essentially unused. The ordering now follows what reaches a session first, and a test ties the claim toRULE_FRONTMATTERso the prose cannot drift from the paths again.search_guideshad its own version of the problem: its description told agents to search for anything outside "~5" embedded patterns and named pytest as an example, which this change made false in the same commit that moved pytest into the rules body. (closes #152) -
The Python versions the classifiers advertise are now checked, and against two different facts because they answer to two. The lowest must agree with
requires-pythonin both directions — those two tell installers the same thing or one of them is wrong. The set must be a subset of what CI tests, in one direction only:Programming Language :: Python :: 3.14has been inpyproject.tomlsince the initial scaffolding while CI gained 3.14 in #105, a PR that changedci.ymlalone, so equality would have made every commit before that one a violation and would forbid trying a Python in CI before committing to support it.Environment :: Consoleis checked in the same one direction: the claim requires a console script, a script does not require the claim. Reading the CI matrix needs no new dependency — the pattern is scoped tomatrix:and to thetestjob, and it raises rather than comparing against an empty set when it cannot find the list, since that is the failure that reports agreement exactly when the check has lost its footing. The floor is derived by asking the specifier which minors it admits rather than by reading its written bounds, because>=3.11,>=3.12has two and only the higher one is real. License classifiers are deliberately left alone: #213 proposes dropping them as deprecated under PEP 639, and a test defending something slated for removal points the wrong way. (closes #220) -
CONTRIBUTING no longer states a test count or a guide count. The suite count was stale by a factor of three, so a contributor checking their environment against it would read a third of the suite failing to collect as a healthy setup. Pinning the numbers with tests was the alternative and was rejected: the guide count already has three places checking it against the catalog, and a fourth would be the duplication #219 removed. The one number kept is the dependency audit's, which says "as of" and means it. (closes #216)
-
The package ships the PEP 561 marker that its
Typing :: Typedclassifier has been promising since the first release. Withoutpy.typed, mypy and pyright treat an installed package as untyped no matter how annotated its source is, so every consumer silently lost the type information the classifier advertised — and the published wheel, downloaded and inspected, contained no marker at any path. The wheel verification now asserts the marker against the installed package rather than against the checkout, since the source tree holding the file proves nothing about what was packaged; the check was falsified by hand, because it runs only in the build job where a green result would otherwise be its own first evidence. (closes #204) -
The supported platforms are stated, and
mpg setupexplains itself when Windows refuses to create a symlink. No platform was declared anywhere before: noOperating Systemclassifier, no mention in README, and CI running Linux alone. Both classifiers and a README section now name the two platforms with evidence behind them — Linux from CI, macOS from development — and Windows is deliberately absent from the metadata while README states what fails there and why.os.symlinkraisesWinError 1314on Windows without Developer Mode or elevation, and the bare error text read like a defect in mpg rather than a privilege the OS withholds, leaving two of the advertised delivery methods missing with no indication of what to change. The new hint is scoped to that error rather than to the platform, because Windows also raisesOSErrorhere for path lengths and read-only volumes, and answering those with a privileges setting would send the reader after a fix that does not apply. Whether to fall back to copying on Windows is left open. (closes #206)
v0.6.0 — Catalog selection filters + grouped --help
Summary: The catalog can now be sliced the way its metadata always allowed: --layer and --frequency filter search and list on both the CLI and MCP surfaces, and mpg list --with-content emits the selected guides in full — the shape to pipe into a generated system prompt. mpg --help groups the nine commands by purpose and shows runnable examples. A contributor's first commands now work as documented: ruff check . passes from the repository root, and an ordinary session no longer leaves the working tree dirty. One behavioural cost is recorded under Changed: the new option names make the --f and --l abbreviations ambiguous.
Added
searchandlistaccept--layerand--frequency, and thesearch_guides/list_guidesMCP tools accept the same two arguments. Filters are conjunctive, so--layer 1 --frequency highreturns the intersection.retrieveis unchanged on both surfaces: it selects by explicit ID, anddocs/design.mdsays so in two places. (closes #30, #31)mpg list --with-contentemits each selected guide's full body alongside its metadata, which is the shape to pipe into a generated system prompt or rules file. This is where bulk retrieval belongs — the requestedretrieve --allwould have maderetrievea selection command and contradicted the documented guarantee that it preserves explicitly requested IDs. The body is the same oneretrieveserves, and the field is absent unless the flag is passed, so the defaultlistschema is unchanged. (closes #29)mpg --helpgroups the commands by what they are for and shows four runnable examples, instead of listing nine flat with no indication of how to invoke them. The grouping is composed by hand because argparse has no notion of it; the same table registers the subparsers, so a command cannot appear in the listing without existing, and adding one to the listing without wiring it up is now rejected rather than silently exiting 0. Tests compare the rendered help against the parser's own commands and parse every example shown, so a listing that drifts from the CLI fails rather than misleading. (closes #36)
Changed
- Adding
--layerand--frequencymade two option abbreviations ambiguous that previously resolved:--fno longer selects--format, and--lno longer selects--limit, on bothsearchandlist.argparseaccepts unambiguous prefixes by default, and the new options collide with the old ones; both now exit 2 withambiguous option. Nothing inside this repository used the short forms, but a script that did will need them spelled out. Turning prefix matching off entirely would break every other abbreviation, so the behaviour stands and is documented instead.
Fixed
-
ruff check .from the repository root passes. It reported 34 violations, all of them inbench/fixtures, where outdated patterns are the point — so a contributor following the obvious command met a wall of failures that were never theirs to fix. Those fixtures are now excluded, and only those: the runner and scorers elsewhere underbench/are still linted, which a test verifies by reading back the file list ruff says it will check. Excludingbench/wholesale would have been shorter and would have quietly dropped the tooling too. CONTRIBUTING states the scope CI uses rather than leaving the two commands to disagree. (closes #123) -
.coverage,uv.lock, and the pathsmpg setupwrites under.claude/are now ignored, so an ordinary development session no longer leaves a dirty working tree.uv.lockis ignored rather than tracked because CI installs withuv pip install --system -eand never reads a lock file, so committing one would pin nothing; the choice is recorded next to the rule rather than left implicit. The.claude/entries name individual paths instead of the directory: a directory-wide rule would silently swallow anything the project later decides to track there, whereas an unlisted path shows up ingit statusand gets a line added deliberately. A test asksgit check-ignorefor each path rather than grepping.gitignore, since a pattern being present and a path being ignored are different claims, and it pins in both directions — the generated paths stay hidden and project content stays visible. (closes #125) -
Catalog selection by category now runs through a single predicate rather than four independent implementations of the same comparison — the scored search, the fuzzy fallback,
mpg list, and thelist_guidesMCP tool each had their own. Nothing was broken beforehand; the point is that adding two filters to four separate sites is how a filter comes to be applied on three paths and quietly ignored on the fourth, and the MCP tool was the copy easiest to overlook because its siblingsearch_guidesalready delegates to the shared search. The accepted values for both new filters come from the frontmatter vocabulary the parser validates against, so the CLI choices, the MCP schema, and the parser cannot drift apart. The MCP server checks them in code as well: anenumininputSchemais advertised to the client, not enforced by the server, and without the checklayer: 99would return an empty array — which reads as "no such guides" when it means "no such layer".
v0.5.12 — Coverage-blind check fixes + symlink write disclosure
Summary: Several checks reported healthy without covering what they claimed — an identifier collision made the wheel-asset gate blind rather than noisy, the MCP stdout-purity test only stayed green because it exercised one request shape, and the release checker's permissions block silently dropped the scope its own checkout needs. Each now establishes its coverage or fails, and a weekly dependency audit is held to the same standard. Separately, mpg setup and mpg uninstall now disclose every symlinked directory a run writes through; the traversal is unchanged and is still not confinement.
Added
- A weekly dependency audit (
.github/workflows/audit-dependencies.yml) runsuv auditover the dependencies this project actually resolves and opens an issue keyed on the advisory set — so a closed issue about old advisories cannot suppress a new one, and a repeat of the same set cannot open a duplicate. It does not run on pull requests: an advisory published against a transitive dependency has nothing to do with the change under review. The verdict is decided byscripts/check_dependency_audit.py, which fails rather than reporting "nothing found" when it cannot establish what the audit covered — a missing subcommand, an unrecognized JSON shape, or an implausibly small audited set. That guard exists because a clean verdict over the wrong corpus is exactly how this check would fail silently: while it was being designed, a barepip-auditreported no vulnerabilities against its own dependencies rather than this project's. CONTRIBUTING documents how to run and triage it locally. (closes #165)
Fixed
-
The wheel-asset gate compared guide identifiers as bare filename stems, which is only correct while no two categories share one. Measured, the consequence runs the opposite way from what the issue anticipated: a collision does not make the gate noisy, it makes it blind. The expected set shrinks along with whatever the index dropped, so the two agree and the check passes — a guide missing from the shipped wheel would be reported as healthy. A guide moved between categories was invisible to it for the same reason. It now compares
(category, id)pairs derived from the wheel's own paths, which removes the dependency on globally unique stems rather than documenting it. (closes #141) -
The scheduled Python release checker files its follow-up issues with
tier:4-extendandarea:contentalongsideenhancement, so automatically created work lands in the same triage as everything else instead of outside the classification scheme. A test reads the--labelarguments specifically; matching anywhere in the workflow text would have been satisfied by the issue body, which names the same words. Duplicate detection is unchanged. (closes #160) -
The scheduled Python release checker declared only
issues: write, and declaringpermissionsat all drops every scope not listed — including thecontents: readitsactions/checkoutstep needs. Public repositories allow the checkout regardless, which is why ten consecutive weekly runs succeeded and the omission never surfaced; it would have failed the moment this repository went private. Both scopes are now declared, and a test asserts the pair without admitting any other write scope. (closes #163) -
The MCP stdout-purity test judged pollution by recomputing the expected byte count with
json.dumps's defaultensure_ascii=True, while the server serializes withensure_ascii=False. Any non-ASCII character reaching stdout made the two counts diverge and failed the test with no stray output present at all — not a hypothetical, sinceguide_index.pycomposes BAD/GOOD summaries with a→and a realsearch_guidesresponse measures 12 bytes "short" under the old comparison. The test only stayed green because it exercisedtools/listalone. The check now judges the stream's structure directly — every line non-empty, free of surrounding whitespace, a standalone JSON-RPC object carrying aresult/error/methodbody, and the stream terminated by a newline — so it no longer depends on the server's serialization settings, and the test additionally pins which response ids belong on stdout so that a well-formed but extra message is still caught. Falsification tests pin that every pollution shape the byte comparison used to catch still fails, and that a genuinely non-ASCII payload passes. (closes #173) -
mpg setupandmpg uninstallnow announce, once per run before writing anything, when.claudeis a symlink and which directory the writes actually land in. Previously the per-file guards refused a symlinkedsettings.local.jsonbut said nothing about a symlinked.claudedirectory, so the hook settings and the Skills/Rules symlinks all silently went to the link target. The symlink is still followed — refusing would break deliberate "config lives elsewhere" layouts — so this closes the silence, not the traversal; SECURITY.md states plainly that neither the per-file checks nor this disclosure confine the.claudetree.--mcp-only, which writes nothing under.claude, prints nothing. (closes #170) -
The same disclosure now covers every directory a run writes through, not just
.claude. A symlinked.claude/skillsor.claude/ruleswas followed with no note at all — the case that motivated the boundary #170 had to document.mpg setupandmpg uninstallnow walk each write target below the project root, report the outermost symlinked directory on the way, and de-duplicate: a symlinked.claudestill yields one note covering all three targets, whileskillsandrulespointing at different trees yield two, because there genuinely are two destinations. The final path component is skipped — those are mpg's own symlinks, so including them would make a secondmpg setupannounce mpg's own links back. The write targets are supplied by the callers rather than duplicated in the settings module, so there is no second copy of the layout to drift, and a source scan pins that no fourth target has appeared unannounced. Traversal is still not refused, and SECURITY.md still says the disclosure is not confinement. (closes #192)
v0.5.11
Highlights
- Resolve target Python and dependency applicability consistently across CLI, MCP, checks, and hooks.
- Make detection coverage and benchmark claims explicit, traceable, and auditable.
- Make project-scoped MCP setup and uninstall reliably target the requested project directory.
Fixed
- Local MCP setup now creates an explicit project directory before registration and runs every
claude mcpmutation from that directory. - Local MCP uninstall honors an explicit project directory and fails closed when the target does not exist.
- Corrected the documented PostToolUse matcher and clarified symlink-flattening diagnostics.
Documentation
- Expanded security, trust-boundary, dependency-applicability, and hook-behavior documentation.
See CHANGELOG.md for the complete release notes.
v0.5.10 — Benchmark credit guard + CI invariant tests
[0.5.10] — 2026-06-27
Summary: V5 benchmark runner now requires explicit --allow-credit-use opt-in for non-dry-run execution, and CI release-artifact invariants (build-verify-upload order, OIDC scope, no-rebuild-in-publish) are pinned by regression tests.
Added
- Regression tests pin the verified-artifact release flow in
ci.yml: build job verifies wheel assets before uploading, publish job reuses that verified artifact instead of rebuilding,id-token: writestays scoped to the publish job, and artifact upload is limited to release/workflow_dispatch events. Includes a negative-test fixture proving the pre-hardening workflow shape would fail the invariant. (closes #140) bench/run-v5.shnow requires--allow-credit-usefor non-dry-run execution;claude -pbenchmark sessions may consume credits depending on the account type.--dry-runoutput shows the session count and directs users to opt in explicitly. (closes #153)- Tests for the credit-use guard: non-dry-run without
--allow-credit-useexits 2, dry-run warns without requiring opt-in.
Docs
- V5 benchmark reproduction docs restructured into cost/credit safety guidance, a low-cost manual path, and an automated path with dated budget wording.
- V1/V2 benchmark procedure (
docs/benchmark-procedure.md) marked as historical with a cross-link to V5 and issue #124.