Repository navigation
Authoring a prompt-scrub rule pack from scratch #100
will-lamerton
announced in
Articles
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Built by the Nano Collective, a community collective building AI tooling not for profit, but for the community.
The release announcement walked through the headline shape of
prompt-scrub: whatscrub()andrehydrate()do, why determinism matters for prompt cache prefixes, and where the threat model draws the line. A sibling post went deep on the secret detector and how its three layers stack. This post is for the other side of the same coin: how to add a detector that does not live inside the main package. Rule packs are the extension point, and the contract is small enough that you can write, test, and publish one in a sitting.What follows is a walkthrough aimed at a developer publishing their first custom detector pack: the
Detectorinterface as it actually appears insrc/types/index.ts, the three export shapes the loader accepts, why detectors deliberately do not return a replacement string, and the four integration points you have to wire up (the package, the config, therules listoutput, and the collision resolver that runs underneath).The contract in one screen
The whole rule-pack contract fits on one screen. Two interfaces, both in
src/types/index.ts:That is it. A rule pack is an npm package that exports one or more objects implementing
Detector. Eachdetect()call walks the text, finds zero or more matches, and returns them as an array ofFinding. There is no constructor signature the loader requires, no base class you have to extend, no registry to register with. The interface is structural: if your object has anameand adetect, it slots in.A worked example, trimmed from the existing
EmailDetectorso the shape is concrete:Note what the detector returns and what it does not. It returns
value(the matched substring) andplaceholderPrefix(the category namespace, here'Email'). It does not return a placeholder likeEmail_2, and the reason is structural, not stylistic.Why detectors never return a
replacementThis is the part of the contract that surprises first-time authors, because every tutorial they have ever read on "how to write a redaction detector" hands them a
Findingshape with areplacementfield. The whitepaper forprompt-scrubdid the same thing. The runtime contract deliberately omits it.The reason is that the final placeholder has to be deterministic across the whole text, and the number suffix (the
2inEmail_2) depends on the session, not on the detector. If a detector returnedreplacement: 'Email_1'for its first match, and then a second detector also returned a match whose placeholder should beEmail_2, the engine would have a choice between re-numbering detector A's output (which breaks the deterministic-prefix guarantee for prompt caches) or letting collisions sit across detectors. Both options are bad.The runtime shape splits the responsibility. Detectors say two things: this is the match (
valueplusspan), and this is what category namespace it belongs to (placeholderPrefix, e.g.'Email'). The core engine takes that, walks the resolved findings, assigns the numeric suffix per category in span order, and writesEmail_1,Email_2, etc. into the session map. Rehydration gets the mapping back and swaps the placeholders for the originals.For a rule-pack author this means two practical things:
detect(). The engine will overwrite them or, worse, leave them in place and confuse rehydration.placeholderPrefixthat uniquely identifies your category in the user-visible output.Codenameworks for a project codename detector;'ApiKey'works for a third-party API detector;'Ssn'works for a US Social Security number detector. Two detectors that pick the same prefix will produce interleaved numbering, which is what you want.The three export shapes
src/core/rule-packs.tsaccepts a rule pack if its top-level module looks like any of these:The loader walks them in order: it prefers the named
detectorsexport, falls back to adefaultarray, falls back to adefault.detectorsarray. Anything else is treated as "this package did not export any detectors" and a warning is printed to stderr:The same warning fires if your object is missing
nameordetect. Validation per-item is the checktypeof detector.detect === 'function' && typeof detector.name === 'string'. There is no required inheritance, no decorator, no factory function. A detector is anything with the right shape.Practical implications:
defaultarray keeps the import line tidy:import myPack from 'my-pack'.defaultto be the package object itself, the third shape lets you add more fields later without breaking importers.All three shapes are stable. We have not seen a strong reason to deprecate any of them, and the loader explicitly tries them all before giving up.
What the loader actually does
loadConfiguredRulePacks()insrc/core/rule-packs.tsis short. Walking it line by line is a fast way to see what guarantees you get as a rule-pack author:loadConfig()first. Ifconfig.rulePacksis empty or undefined, it returns{ detectors: [], metadata: [] }with no side effects.await import(packName). This is bare Node dynamic import, so the package name has to resolve via the normal Node resolution rules: either an installed npm dependency, a path-style fallback used in tests, or a workspace package.nameis paired with a metadata record{ name, source: 'rule-pack: <packName>', defaultState: 'on' }.Two consequences worth knowing:
scrub()orrules listpath. If your module has top-level side effects, those side effects will run on first use, not on import ofprompt-scrubitself.'on'. There is no built-in switch to mark one off by default. If you want to ship a noisy detector behind opt-in, you have to either name it in a way that flags it as opt-in and let consumers turn it off viaScrubOptions.disabledDetectors, or document it loudly in your README.Wiring it up: the four integration points
A rule pack is nothing without the four integration points that actually make it run. Each is small; together they form the boundary between your package and the rest of the world.
1. The npm package
Your detector has to be reachable by Node's resolver. Publish to the registry, install in your project (
npm install my-pack), or for local development usenpm install /path/to/my-pack. The test suite uses the same path. Whatever you do, make sure a plainrequire('my-pack')from insidenode_modules/@nanocollective/prompt-scrubwould resolve.2. Global config
In your project's config directory,
~/.config/prompt-scrub/config.jsonon Linux (~/Library/Application Support/prompt-scrub/config.jsonon macOS,%APPDATA%\prompt-scrub\config.jsonon Windows), therulePacksarray is where you list the packages you want loaded:{ "rulePacks": ["my-pack", "another-pack"] }The loader reads this file via
loadConfig()and unions itsrulePackswith theprompt-scrub.rulePacksfield from your projectpackage.json.3. Project package.json
For per-project configuration, the same key lives under a
prompt-scrubnamespace in yourpackage.json:{ "prompt-scrub": { "rulePacks": ["my-pack"] } }src/core/config.tsreadspackage.jsonfromprocess.cwd(), so it picks up whatever project you happen to be running from. The loader de-duplicates between the two sources via aSet, so a package listed in both ends up loaded exactly once.4. The
rules listCLI commandOnce configured, your detectors show up alongside the built-ins when a user runs
prompt-scrub rules list:This command is what people use to check that a pack is actually loaded. If your detector does not show up in this output, the loader silently skipped it and the most likely cause is a bad export shape. Check for the warning printed at startup first.
Collision behaviour: where your detector slots in
The reason the contract is small is that the heavy lifting lives in
resolveCollisions()insrc/core/collision-resolver.ts. Your detector's findings are pooled with the built-ins, then a single pass resolves overlaps. The rules:span[0]ascending and walked left to right.DETECTOR_PRIORITY:SecretDetectoris 1,EmailDetectoris 2,UrlDetectoris 3, and so on throughCodeTellDetectorat 8.99, including rule-pack detectors. If your finding overlaps withSecretDetector, the secret wins, every time. This is the "missing a credential is worse than missing a name" policy, applied mechanically.The implication for rule-pack authors is that your detector's findings will survive only when they do not conflict with a higher-priority built-in. Two patterns show up in practice:
docs/features/authoring-rule-packs.mdcalls out the priority table verbatim, including the note that there is no public mechanism for a custom detector to overrideSecretDetectorand there will not be one. That policy is what keeps the secret detector useful for the cases it is tuned for, and it is not negotiable through the rule-pack API.A minimal end-to-end pack
To make the moving parts concrete, this is a pack that detects internal project codenames for two strings,
ApolloandZeus. It uses the named-export shape and TypeScript:Build and publish the package (
prompt-scrub-projectxhere), then list it in the user'spackage.json:{ "prompt-scrub": { "rulePacks": ["prompt-scrub-projectx"] } }Run
prompt-scrub rules list. The new detector shows up withSource: rule-pack: prompt-scrub-projectxandDefault State: on. From that point on, any text scrubbed via the library or CLI gets every occurrence ofApolloandZeusreplaced withCodename_1,Codename_2, etc., in source order, in the same pass that catches emails and paths.The packing itself is unremarkable. What is worth taking away is how few places in the system know about your detector: the loader, the metadata entry for
rules list, and the collision resolver. Everything else (the placeholder numbering, the session map, the rehydration pass) does not need to know that your detector exists. That is the shape of a contract designed to be small on purpose.What is stable and what is not
prompt-scrubis at 1.0.0 and the rule-pack contract shipped in this release. The two interfaces reproduced above (DetectorandFinding) are documented as the runtime contract for everything outside the package. The general approach is conservative: any change to a public interface is a breaking change for rule packs and would land in a major version bump.What is reasonable to expect to grow (additive, non-breaking):
defaultStateis the only knob exposed today. If a rule-pack author needs to mark a detector as off by default, that would land as a new optional field rather than as a replacement.What we do not plan to change:
docs/features/detectors.md. There are no plans to expose priority to rule-pack authors, because the secret-priority behaviour is a deliberate trade the whole product depends on.Until something in
src/types/index.tsmoves, the contract you see there is the contract your pack will run against.Where to look in the repo
src/types/index.tsfor the canonicalDetectorandFindingshapes.src/core/rule-packs.tsfor the loader (about 70 lines including comments).src/core/config.tsfor how global config andpackage.jsonare read and merged.src/core/collision-resolver.tsfor the priority table and the overlap pass.docs/features/authoring-rule-packs.mdfor the user-facing version of the same contract.Issues, PRs, and rule-pack proposals are welcome at the repo. If you write one, please include a minimal test that exercises each of the three export shapes, plus a
rules listline so others can see what they get when they install it.Source, tests, and the full rule-pack contract: https://github.com/Nano-Collective/prompt-scrubber
All reactions