Skip to content

Documentation Process

ProxyPrints Docs Bot edited this page Jul 29, 2026 · 11 revisions

Documentation process

The standing system this repo's documentation runs on, in one page. If you're wondering "where does X go" or "why does the wiki look like that," this is the answer.

The core rule: docs/ is source of truth, the wiki is generated

Everything durable about this fork lives in docs/ and is committed alongside the code it describes. The GitHub wiki (ProxyPrints.github.io.wiki) is a generated view of a curated subset of docs/ — never a second place to write documentation.

Never hand-edit a wiki page this system manages. Every generated page carries a <!-- GENERATED PAGE --> marker at the top saying so, and naming its docs/ source. Edit the source instead; the page regenerates on the next push to master that touches docs/** (or on-demand via workflow_dispatch). A hand-edit to a generated page survives only until the next regeneration, then silently vanishes — not a data-loss risk (the source is what's real) but a wasted edit.

  • What's published: exactly the files listed in .github/wiki-publish-map.json, a hand-curated mapping mirroring docs/README.md's own audience groups ("Understanding the system" / "Using it" / "Operating it"). Update the mapping file whenever docs/README.md's index changes — the publish workflow does not parse docs/README.md's prose, by design (parsing markdown structure at runtime is fragile; a hand-maintained mapping is not).
  • Renames update the mapping in the same PR. If a published doc's source path or its wiki page name changes, update wiki-publish-map.json in that same PR — a wiki page's URL is its identity, and GitHub wikis have no redirects: renaming a page in the mapping doesn't move the old URL, it abandons it. Prefer a stable wiki name once picked, even across a source-file rename or migration — e.g. docs/self-hosting.md still publishes as the wiki's existing Instance-Admin-Guide page (not a new Self-Hosting page) specifically so nothing that already links to it breaks.
  • What's excluded, always: docs/proposals/ (drafts/HOLD specs — not yet real, shouldn't read as if they are), docs/reports/ (relayed session artifacts, not reference material), docs/audits/ (point-in-time findings, same reasoning).
  • pointer_pages: a page whose content has fully moved elsewhere (e.g. Research-and-Proofs, once its own placeholder, now superseded by docs/theory.md) becomes a small generated stub linking to the page that replaced it, rather than being deleted — same reasoning as the rename rule above: the URL stays alive even though the content doesn't live there anymore.
  • legacy_pages: any wiki page that predates this system and has no docs/ source yet goes here so it stays linked from the generated Home/Sidebar instead of going dark. Empty as of this writing — the 3 pages that used to sit here (User-Guide, Instance-Admin-Guide, Research-and-Proofs) are now either migrated into docs/ (the first two) or a pointer_pages entry (the third). The publish workflow never touches a page outside all three lists — it only ever deletes a page it can prove it generated, by checking for that page's own marker.
  • Mechanism: .github/workflows/docs-wiki-publish.yml together with .github/scripts/publish_wiki.py. Requires a WIKI_PUSH_TOKEN repo secret (a classic PAT — GitHub wikis are a separate git repo the default GITHUB_TOKEN can't push to); see that workflow's own header for exact setup steps.

The site is a second generated target, same source, ONE transform

Per proposals/proposal-i-docs-as-site-source.md (APPROVED, staged build — PR-I-1 shipped, PR-I-2+ not yet started), the same docs/ files can also publish as real pages on the site itself, at /guide, via a second per-page target ("site" in wiki-publish-map.json's targets array) alongside — not instead of — the existing wiki target.

Single-transform architecture: .github/scripts/publish_wiki.py is the ONLY place link-rewrite logic exists, for both outputs. Its transform_links()/rewrite_link() take an optional repo_to_site map — absent (the default), it's exactly the original wiki-only 2-way resolution (same-wiki link or GitHub blob URL); given a real map, it adds a 3rd resolution branch (a link to a "site"-targeted page becomes that page's own route) and changes how a wiki-only target resolves (an ABSOLUTE wiki URL rather than a same-repo link, since the site itself doesn't host that page). .github/scripts/publish_site.py, a thin sibling script, imports this shared logic (no reimplementation) and writes pre-transformed markdown for every "site"-targeted page into frontend/generated-docs/ (gitignored). frontend/src/pages/guide/[[...slug]].tsx has no transform logic of its own — it only reads that markdown and renders it to HTML via a JS markdown library at Next.js build time (a rendering concern, kept separate from the transform concern above). Run locally via npm run docs:generate (from frontend/); next build/next dev degrade gracefully with a console warning, not a crash, if that hasn't been run yet — /guide simply has no pages until it has.

Live for two docs so far — docs/overview.md (/guide) and docs/user-guide.md (/guide/using-it, added as a widening pass after PR-I-1); docs/self-hosting.md and docs/theory.md remain per this proposal's proposed initial mapping. Widening the list is normal, low-risk work (add "site" + sitePath to a wiki-publish-map.json entry). A Nav.Link to /guide in frontend/src/features/ui/Navbar.tsx ships alongside this widening (ungated, same as Download/guide has no backend dependency), closing proposal-i's "not yet built" item 3.

Known gap, deliberately deferred (2026-07-25): mermaid diagrams are wiki-only for now. docs/identification-pipeline.md (FIG-1) and docs/pipeline-fidelity-gate.md (FIG-2) each carry a committed ```mermaid fence. GitHub and the wiki both render it natively; the site's guide page does not — frontend/src/pages/guide/[[...slug]].tsx renders markdown via marked v18, which has no mermaid support, so a fence would show as raw fenced source text on proxyprints.ca. Neither page is in this file's "site" target list for exactly that reason. Three ways to close the gap, owner sign-off needed before any of them: (a) add a mermaid renderer to the guide page (new frontend dependency); (b) commit an SVG beside each fence and have the site prefer it (no new dependency, but a second artifact that can drift from the mermaid source); (c) leave it wiki-only indefinitely. Nothing about this gap blocks the docs//wiki commit itself.

Link-rewrite parity fixtures: since one function now serves two output shapes (wiki mode vs. site mode) rather than two separate implementations, .github/scripts/tests/test_publish_wiki_link_rewrite.py runs a shared fixture set (.github/scripts/testdata/link_rewrite/cases.json) through both modes and pins which cases are expected to produce identical output versus legitimately diverge by mode — a future edge-case fix that doesn't also update the fixture fails this test, rather than silently diverging. Same rationale as the federation hash tool's permanent parity test.

readme.md is a third generated target — assembled, not transformed

Per proposals/proposal-i-readme-pipeline.md (SHIPPED), the repo-root readme.md is generated by a third script, .github/scripts/publish_readme.py — but it's a different kind of generation from the wiki/site targets above. Those two TRANSFORM whole docs/ pages (link rewriting, one page in → one page out). readme.md ASSEMBLES one output file from several small, hand-authored prose regions scattered across docs/, each marked:

<!-- README-REGION: license -->

...

<!-- END README-REGION -->

Deliberately a different marker name from DATA-EXTRACT (the site pipeline's own extraction contract, table-only by design) — these regions are prose, and reusing the table-only contract's name would misdescribe what gets parsed. A marked region is copied verbatim, never link-rewritten: readme.md lives at the repo root, not in docs/, so a relative link correct for a region's actual location in docs/ would resolve to something else entirely once copied into the root file. Every region therefore uses an absolute GitHub URL for any file reference — guarded by a dedicated test (.github/scripts/tests/test_publish_readme.py) — rather than teaching the script a second, output-relative link-resolution mode.

Generated AND COMMITTED, not gitignored — the one real structural difference from the wiki/site outputs. GitHub renders readme.md directly from the default branch with no build step in between, so the committed file has to actually be current at all times. The correctness gate is therefore a CI parity check, not a build-time regenerate: docs-lint.yml's readme-parity job reruns publish_readme.py into a scratch copy and fails the PR if it differs from the committed file — same generate-and-diff shape as the link-rewrite parity job above, applied to a whole generated file instead of a fixture set. Regenerating, mechanically, means running the script and committing the result — no different from hand-editing any other tracked file.

Lint catches mechanical rot

docs-lint.yml runs docs_lint.py on every PR touching docs/** and weekly regardless. Its original, always-hard checks are two, mechanical (the interconnection rules below are newer, and hard-fail as of 2026-07-23 — see that section below):

  1. Every [[wiki-link]] and markdown [text](path) link resolves to a real file.
  2. Every backtick-quoted, path-shaped reference in prose (e.g. `frontend/src/features/card/Card.tsx`) exists in the repo (or is a known-gitignored file, or an explicit allowlist entry with a stated reason).

Failures annotate the PR diff. It never auto-fixes anything — a broken link/path is a fact worth a human's one-line correction, not a silent rewrite.

Known limitation, stated plainly: this can only catch broken links/paths. It cannot tell a stale status claim ("not yet built" for something now shipped, a "current stage" claim a later section already contradicts, a date stamp that's drifted) from a true one — resolving that requires reading the doc's actual content and cross-checking it against reality, which is exactly what the next section is for.

Roster tethers: an enumerated set is tethered to its definition in code

The rule. Any document that ENUMERATES a set which is actually defined in code must be tethered to that code by a lint rule, with code as the source of truth. A hand-maintained list of things-that-exist will drift, and the drift is silent precisely where it matters most — the item nobody remembered to add is the item nobody is checking.

Where this was earned. pipeline-fidelity-gate.md enumerated the calculators whose output the gate audits. It listed eleven identities; code declared fourteen. The gate performed perfectly on all eleven it listed — and two of the three it omitted were entirely dead in production (stage-d-illustration-v1 had cast 3 votes in its whole existence; local-name-frequency-v1 had produced no output at all, not even a skip log). The audit was not wrong about anything it looked at. It simply never looked, because nothing bound its list to reality. Note the shape of this: dormancy is exactly what a coverage-based audit misses, because a calculator producing nothing generates no divergence to explain and so reads as clean by being invisible. A roster tether catches it, because it checks against what is DECLARED, not against what showed up in the output.

How to write one — the two shipped examples, check_extractable_primitives_tether() and check_calculator_roster_tether() in docs_lint.py, are the pattern to copy:

  1. Derive the roster from code, never restate it in the linter. A second hand-maintained list inside the check reproduces the exact drift the check exists to prevent; it only moves the stale list from the doc into the lint.
  2. Key on the declared symbol, not on the shape of its value. The calculator tether matches module-level *_ANONYMOUS_ID = "..." bindings rather than regexing for <name>-vN strings, because test fixtures and extractor version keys share that shape and are not roster members.
  3. Match the full identity, not a normalised family. calculator_family strips the -vN suffix; matching on it would let a version bump ride on a stale entry, leaving a doc paragraph describing an engine that no longer runs while CI stayed green. Full-identity matching makes the bump fail loudly. A version bump SHOULD be a two-file change.
  4. Exclusions are an explicit, per-entry-justified allowlist. Same discipline as ALLOWLIST above: an exclusion has to be a visible decision with a stated reason (evidence-transfer-v1 casts no votes; question-feed-hypothetical-vote is a UI construct), never "it was failing so I dropped it."
  5. Annotate the doc, not the code. The finding is "the doc is behind code"; the fix goes in the doc, so that is where the ::error:: points.

What an entry must say. Being on the list is not enough — an entry that describes a calculator's intent while it is silently dead is the same blind spot in prose form. State what the thing does and its current status, dormancy and no-op-by-design included, plainly and without papering over: "casts no votes by construction, routes only" and "dormant, 3 votes ever, root cause diagnosed, fix in flight" are both entries that do their job.

Other hand-maintained rosters across docs/ are candidates for the same treatment; tether them as their code-side definition becomes unambiguous, and prefer no tether over a tether whose mapping needs judgment to evaluate — a rule that fires on honest content teaches people to ignore the linter.

Interconnection lint

The 2026-07-23 owner ruling abolished the letter/number decision-labeling convention outright ("kill the lettering convention all together ... each subject should have one document or they should at least reference each other"). There is no central decisions register and no label grammar: a decision lives written out in prose in its own subject doc, and the subject doc is the source of truth. docs_lint.py enforces that model with four rules (hard-fail as of 2026-07-23) on top of the mechanical link/path checks:

  1. No new D-number decision labels. The abolished convention wrote a decision as a bold D-number marker or a "decision D-number" phrase (the sibling vote-weight decision stream used the same D-ledger plus VW-style refs). The lint flags any such definition-style label reintroduced in a living doc — write the decision out in prose instead. Structural enumerations the sweep deliberately kept are NOT decision labels and are not flagged: funnel steps (F), requirements (R), test scenarios (T), mappings (M/C/S/I/L), editor-spec items (E), file-change rows (XF), and license/PR tokens (GPL-3.0, PR-5). A historical aside in the (formerly '...') form is allowed; the archive buckets below are exempt; and a verbatim, self-declared decision-record doc that keeps its labels intentionally is allowlisted (today: reference/funnel-spec.md — the living, de-lettered authority for that surface is features/grid-selector.md).
  2. Orphan check. Every docs/*.md must be reachable from README.md or MANIFEST.md, transitively, via markdown/wiki links or backtick-path references. The record/archive buckets index themselves by their own dated-listing convention and are excluded: reports/, data/, audits/, proposals/mockups/.
  3. Supersession pointer. A line carrying a SUPERSEDED/SUPERSEDES marker must name what supersedes it (a link, a backtick path, a §section, a #ref, or a self-describing SUPERSEDED-BY-X status) somewhere in its own paragraph or the next one — so a multi-line "HISTORICAL — SUPERSEDED" banner that names its superseder a few lines away still passes.
  4. Same-subject cross-reference. Two proposals/proposal-<letter>-*.md docs covering the same subject must reference each other — at least one direction, so a reader can navigate between them. This is the anti-fragmentation guarantee the "one doc per subject, or they reference each other" ruling asks for, replacing the old shared-label heuristic.

Hard-fail (flipped 2026-07-23). These four rules now print as ::error:: and count toward the exit code, same as the original link/path/tether checks — docs-lint.yml's Run docs lint step runs python3 .github/scripts/docs_lint.py --strict. The de-lettering sweep (PR #357) had already left the whole corpus clean under --strict (exit 0) before this flip, so promoting the four rules from warn-only to blocking changed no doc content — only what CI enforces going forward. DOCS_LINT_STRICT=1 in the job's env is the equivalent alternate trigger, documented here in case a future workflow edit prefers the env-var form over the CLI flag.

The judgment coherence pass (quarterly)

Lint catches what's broken. It cannot catch what's wrong but still resolves — a Part header claiming "in progress" for something the same file's own status section says merged, a cross-reference pointing at a real file that no longer contains the claimed content, a "known gap" that's actually long since fixed. That needs a human (or an agent session) actually reading every doc and checking its claims, on a cadence loose enough that rot has time to accumulate into something worth a dedicated pass, but tight enough that it doesn't compound for years. Quarterly is that cadence.

This checklist is derived directly from the first such pass (2026-07-18):

  1. Inventory every file in docs/. For each: is its status language (dates, "planned"/"in progress"/"not yet built", "HOLD"/"DRAFT" markers) still accurate? Does it have cross-references, and do they still point somewhere real? Is it reachable from CLAUDE.md's index or docs/README.md — or orphaned?
  2. Cross-check every status claim against something real, not against another equally-stale doc: the same file's other sections, a sibling doc's own status section, the actual current codebase (grep/ find for a referenced path or symbol), or git log for a cited PR number or date. A claim that can't be cheaply verified this way (an external PR's live open/closed state, a third party's stated intentions) gets flagged for a session with the access to check it — never guessed at.
  3. Watch for duplicate or colliding headings within a long file (an anchor collision is a real navigation hazard, easy to miss by reading linearly).
  4. Watch for miscategorized content — an item sitting under "Known gaps" that's actually resolved, a bullet under the wrong header entirely.
  5. Check docs/README.md's own index for completeness (every doc in a bucket it should be in) and .github/wiki-publish-map.json for drift from it (the mapping is hand-maintained specifically so this check has to be a deliberate step, not something that happens for free).
  6. Check CLAUDE.md's flat docs index for parity with docs/README.md — both should list the same "current" reference docs; a doc missing from one but not the other is exactly the kind of gap this pass exists to catch.
  7. Compile a migration/decision list before touching anything non-mechanical. If a wiki-only page has no docs/ source, or a fix would mean rewriting a doc's actual technical content (not just its status line or a cross-reference), that's a decision for the owner, not something to resolve unilaterally mid-pass.
  8. Fix only what's mechanical: status lines, cross-references, stale paths, miscategorized bullets, duplicate headings. Leave technical content — and always leave docs/theory.md's substance — untouched unless the fix is explicitly approved as its own, separate piece of work.

Upstream wiki: linked and attributed, never mirrored

docs-upstream-wiki-drift.yml together with upstream_wiki_drift.py check weekly whether chilli-axe/mpc-autofill's wiki has changed since the last check, and update docs/upstreaming/upstream-wiki-drift.md's table in place — which page changed, when, at which upstream commit.

This is detection only, deliberately: upstream's wiki text has no clear license, so nothing from it is ever copied into this repo, in this workflow or anywhere else. The correct response to a drift-log entry is always a human decision — read the upstream page, and if it's worth reflecting here, write our own words, linking to and crediting the upstream page as the source, on review. A drift-log row is a prompt to look, never content to paste.

Clone this wiki locally