Skip to content

How this was written

github-actions[bot] edited this page Oct 9, 2026 · 109 revisions

AI disclosure

Tippani was written with AI assistance, and not marginally — essentially all of it. This file says exactly what that means, because "AI-assisted" is used to describe everything from a tab-completed variable name to a wholly generated codebase, and those are not the same claim.

Two separate questions get conflated whenever AI comes up in a repository. They are answered separately below:

  1. Was the code written with AI? Yes, throughout. See How this repo was written.
  2. Does the app use AI — does it send your library anywhere, call a model, run inference? No. Nothing. See The app itself contains no AI.

The app itself contains no AI

This matters more than the first question for anyone deciding whether to run Tippani, so it comes first.

  • No model calls at runtime. There is no OpenAI, Anthropic, or local-inference code path in internal/ or the frontend. Not disabled by default — not present.
  • No model ships with the binary. Tippani is one static Go binary with an embedded SPA and a SQLite file.
  • Your highlights are never sent to a model. They are never sent anywhere.
  • The only outbound calls are metadata lookups, to the sources named in the README (Google Books, Open Library, TMDB, TheTVDB, Wikidata), that somebody asked for or that a screen somebody opened makes for the pictures it is about to draw; the Pushover messages a reader set up; and a GitHub release check that runs only when an admin presses the button. Cover and portrait fetches go through a host allowlist with an SSRF guard. Single sign-on, where it is configured, talks to the operator's own provider.
  • Every one of them is written down. Since 3.1.0 each outward call is a line in the log of the job that made it — method, host, path, what came back — kept for 30 days where an admin can read and export it, with every key and credential in the address replaced by … before it is kept.
  • And you can switch all of it off. TIPPANI_OFFLINE=1 refuses every one of those before it is dialled, sign-in to your own provider excepted — internal/outbound is a transport wrapped around each of the four HTTP clients that reach the internet, and a test names every &http.Client{} in the tree so a fifth cannot appear ungated. The library itself needs no network to read, so nothing you already have stops working.
  • Nothing is sent anywhere to be encrypted either. Backup archives are sealed (AES-256-GCM, Argon2id, both from Go's standard library and golang.org/x/crypto) entirely in-process. No key service, no escrow, no network call. Since 1.4.2 there is a recovery key, and it is 32 bytes in your own data directory — nothing holds a copy but you.

There is one AI feature under consideration, and it is in the roadmap under Later / maybe — not built, not started: opt-in digest summaries against an OpenAI-compatible endpoint you configure with your own key, off unless you turn it on. If it is ever built it will be off by default and will say plainly what leaves the machine. Until then, the honest summary is: an AI wrote this app; the app does not use AI.


How this repo was written

Practically all of it was written in Claude Code — design discussion, schema, Go, React, CSS, SQL migrations, tests, and the documentation including this file — with me directing the work, deciding what gets built, and reviewing what comes back.

The audit trail is the git history itself: nearly every commit carries a Co-Authored-By: trailer naming the model that worked on it. The block below is the measurement, stamped with the commit it was counted at and re-counted by the kit's ai_census.py --check. The breakdown by model, and the commits that carry no trailer, follow it under By model.

Measured as of 996cae99dc04b0069536890ff23256cd967b4ca3 (2026-10-09)

Figure Value
Commits, no merges, exclusions applied 2,107
AI-assisted commits 2,094 — 99.4%
Lines added / removed, AI-assisted commits +913,425 / -385,181
Lines added / removed, all commits +923,022 / -393,664
Surviving lines from AI-assisted commits 525,091 of 529,243 — 99.2%

These figures describe the tree at that sha and change with every commit, which is why they are stamped: an unstamped percentage cannot be checked against anything, and so is not a disclosure.

Verify these numbers

# The marker, matched case-insensitively against the whole commit message. The census
# applies it as a Python regex; -E below is the closest git has.
AI='^[ \t]*co-authored-by:.*(claude|anthropic|copilot|chatgpt|openai|gemini|cursor|codex|aider|devin|\[bot\])'

# The exclusions, in the order the census applied them.
EX=(':(exclude)*.lock' ':(exclude)**/*.lock' ':(exclude)package-lock.json'
    ':(exclude)**/package-lock.json' ':(exclude)yarn.lock' ':(exclude)**/yarn.lock'
    ':(exclude)pnpm-lock.yaml' ':(exclude)**/pnpm-lock.yaml' ':(exclude)go.sum'
    ':(exclude)**/go.sum' ':(exclude)vendor/**' ':(exclude)third_party/**'
    ':(exclude)node_modules/**' ':(exclude)dist/**' ':(exclude)build/**'
    ':(exclude)target/**' ':(exclude)*.min.js' ':(exclude)**/*.min.js'
    ':(exclude)*.min.css' ':(exclude)**/*.min.css' ':(exclude)*.map'
    ':(exclude)**/*.map' ':(exclude)**/*.generated.*' ':(exclude)**/generated/**'
    ':(exclude)**/__snapshots__/**')

# 1. Commits, then AI-assisted commits.
git rev-list --count --no-merges HEAD -- . "${EX[@]}"
git rev-list --count --no-merges -E -i --grep="$AI" HEAD -- . "${EX[@]}"

# 2. Lines added and removed -- AI-assisted commits, then all commits.
git log --no-merges -E -i --grep="$AI" --numstat --format='' -- . "${EX[@]}" \
  | awk '$1 ~ /^[0-9]+$/ { a += $1; d += $2 } END { print "+" a, "-" d }'
git log --no-merges --numstat --format='' -- . "${EX[@]}" \
  | awk '$1 ~ /^[0-9]+$/ { a += $1; d += $2 } END { print "+" a, "-" d }'

# 3. Surviving-blame share, at the coverage the census used: text files only, under the
#    blame cap, blamed at HEAD. Then look each line's commit up in the AI-assisted set.
#    Drop the two skip tests and this over-counts by every binary and generated blob.
shas=$(mktemp); lines=$(mktemp)
git log --no-merges -E -i --grep="$AI" --format=%H | sort -u > "$shas"
git ls-files -- . "${EX[@]}" | while IFS= read -r f; do
  [ "$(wc -c < "$f" 2>/dev/null || echo 0)" -le 2000000 ] || continue
  [ -s "$f" ] && ! grep -Iq '' "$f" 2>/dev/null && continue
  git blame --line-porcelain --no-progress HEAD -- "$f" 2>/dev/null
done | awk '(length($1) == 40 || length($1) == 64) && $2 ~ /^[0-9]+$/ { print $1 }' > "$lines"
awk 'NR == FNR { ai[$1]; next }
     { total++ } $1 in ai { hit++ }
     END { printf "%d of %d lines (%.1f%%)\n", hit, total, 100 * hit / total }' \
  "$shas" "$lines"

What the count covers

  • Marker. A commit counts as AI-assisted when its message matches ^[ \t]*co-authored- by:.*(claude|anthropic|copilot|chatgpt|openai|gemini|cursor|codex|aider|devin|\[bot\]), case-insensitively. Merge commits are excluded throughout, because a merge authors no content.
  • Unmeasured history. None. The first commit, 6f8f7033 (2026-07-02), already carries the trailer, so no history predates the convention. The thirteen commits with no trailer are named under By model.
  • Exclusions. *.lock, package-lock.json, yarn.lock, pnpm-lock.yaml, go.sum, vendor/**, third_party/**, node_modules/**, dist/**, build/**, target/**, *.min.js, *.min.css, *.map, **/*.generated.*, **/generated/**, **/__snapshots__/**.
  • Blame flags. git blame with no -M and no -C. Move and copy detection reattributes moved lines to their original commit and can shift the surviving share by tens of points, so the flags used are declared rather than left for the reader to guess.
  • Coverage. Tracked files: 1,584 blamed, 4 removed by the exclusion list, 460 skipped as binary, 1 skipped over the 2,000,000-byte blame cap, 0 staged but absent at the stamped sha, 0 whose blame failed and were counted as unattributed. Blame is taken at that sha, so uncommitted working-tree edits sit outside every figure. Nothing was sampled: every other line was counted.
  • Floor, not measurement. Untagged AI-assisted commits, if any exist, make every figure above a lower bound.
Exclusion Files removed
package-lock.json 3
go.sum 1

By model

Models used, by commit count, at the same commit. Merges are excluded, and the block's exclusion list is not applied, so 2,104 commits carry a trailer against the block's 2,094. One of them, f5000428 (3.1.5's release commit), carries two trailers and is counted once under each of two models, so the rows sum to 2,105. One commit, b77a22e5, carries its line outside git's trailer block, so the count and the breakdown below grep the message. The per-commit listing asks git for trailers, and shows b77a22e5 with none:

Model Commits
Claude Opus 5 1,169
Claude Opus 5.5 696
Claude Opus 4.8 150
Claude Fable 5 54
Claude Sonnet 5 19
Claude Sonnet 5.5 7
Claude Haiku 4.5 5
Claude Fable 5.1 4
Claude Sonnet 4.6 1

Some of those trailers carry a (1M context) suffix naming the long-context variant: 450 of the Opus 5 commits, 466 of the Opus 5.5 ones and 146 of the Opus 4.8 ones. It is the same model with a larger window, so the table folds them; the last command below prints them unfolded if you would rather see it raw.

Thirteen commits carry no trailer, and they divide four ways. Five are github-actions[bot] regenerating the roadmap's known-bugs block from the issue tracker — machine-written but not AI-written, and that distinction is the point of this file: a script rendering a JSON file into HTML is not a model making choices. One (16879616) adds attribution URLs for Bookcision, Readest and pretext to the README, typed by hand. Six are oversights rather than categories, each AI-written and each of which should have said so: the 2.0.0 and 2.1.0 release stamps, one Go test refactor (087b2dc0), and three commits whose author is recorded as Claude with no trailer beside it (faa0ec36, 4da5b829, 3121bf7e). One (ae85064b, a CI fix) is recorded under the owner's name with no trailer, and the history does not say which it was. They are named here instead of being quietly fixed, because a disclosure that rounds its own gaps away is not one.

To see it yourself:

SHA=996cae99   # the commit the block above is stamped with

# every commit, with the model that co-authored it
git log --no-merges --date=short "$SHA" \
  --format='%h %ad %s — %(trailers:key=Co-Authored-By,valueonly,separator=%x2C)'

# the count, and the breakdown above (1M-context variants unfolded)
git log --no-merges --format='%b' "$SHA" | grep -c '^Co-Authored-By: Claude'
git log --no-merges --format='%b' "$SHA" | grep '^Co-Authored-By' | sort | uniq -c | sort -rn

The agent configuration used to do the work — session skills and subagent definitions under .claude/ — is gitignored and not part of this repository, so what you see here is the output, not the toolchain.

Increasingly this is several agents at once, not one conversation in sequence

Later releases were built by fanning a task out across parallel subagents under my direction and then reconciling what came back, rather than by one session doing everything in order. Three examples from 2.1.x, all of them in the history:

  • The Bengali interface (2.1.1, 2,446 strings) was written by six agents in two passes, each working from a committed style sheet — docs/wiki/Bengali-style.md — rather than from the English alone, because six writers with no shared register produce six registers in one interface. The merge was then checked mechanically: key set, placeholder parity, nothing lost, and on the 442 keys where writers disagreed, that the file holds one of their readings rather than an invented third. The register checks are in docs/wiki/Bengali-style.md. In 2.2.x one agent rewrote every string again from a per-key dossier of the code that renders it, and a second, independent agent rated the result against the brief before it was committed.
  • The roadmap audit ran one reader per section and then four sceptic agents whose instruction was to refute the first pass, not to agree with it. They overturned or amended ten of thirty-three findings. Acting on the first pass alone would have deleted real backlog items and kept stale ones, which is the whole argument for the second pass.
  • Two decision entries in docs/wiki/Design-decisions.md were drafted by agents from the plan documents and then verified line by line against the code before being inserted. One of them corrected a figure I had written and repeated: "1,299 of 2,446 keys carry a comment" had counted comment lines.

The rule that makes this safe is that an agent's finding is worth nothing until something executes or cites. More agents means more plausible-looking output per unit of my attention, which is a hazard and not a benefit unless the verification scales with it. So a finding carries a file:line or it is not a finding; a test is broken on purpose to watch it go red before it is trusted; and where a claim matters, a second agent is pointed at it with instructions to break it. This session alone, that pass caught two of my own tests asserting the wrong thing — one that a comma stays inside a tag name when the vocabulary deliberately splits on it, and one that a Bengali abbreviation may not end in a vowel sign when that is exactly how Bengali abbreviates.


What is actually checked, and what that does not cover

AI-written code fails differently from hand-written code. It compiles, it reads well, it is plausibly commented, and it can still be wrong — so plausibility is worth nothing here and only execution counts. What the repo actually runs:

  • 2,107 Go test functions and 5,067 frontend tests, across 911 test files — the Go half over real HTTP handlers against a real SQLite database, not mocks. Counted, not estimated, and every number here has a command that reproduces it:

    grep -rhoE '^func Test[A-Za-z0-9_]+' --include='*_test.go' . | wc -l   # Go functions
    cd web/frontend && npx vitest run                                      # 5,067 of them
    cd web/frontend && npm run journeys                                    # + 154 in the browser
    find . -name '*_test.go' -not -path './node_modules/*' | wc -l         # 349 Go files
    find ./web/frontend -path '*/node_modules' -prune -o -type f \
         \( -name '*.test.*' -o -name '*.spec.*' -o -name '*.journey.*' \) \
         -print | wc -l                                                    # 562 frontend

    npm test NO LONGER RUNS ALL OF THEM, AND THAT IS THE POINT. 5,067 is what npx vitest run reports across the three vitest projects, and the browser tier is not among them — it has its own config, because it needs a globalSetup that builds the binary and seeds a library. npm test runs two projects — 4,131 tests over 356 files; npm run lint:rules runs the third, 936 assertions over 100 files; and npm run journeys runs 154 tests over 106 files against a real server in a real browser, which is the tier that would have caught the bug all this is named after. Those 100 READ THE SOURCE TEXT and assert how it is spelled: never truncate a name, spacing is a constant, no emoji glyphs, the typescale. They are worth keeping and they were never tests, because the app can be entirely broken and all 100 of them still pass — none of them runs it. A suite let a feature ship 100% dead that way. CI runs lint:rules as its own step, so a broken design rule still fails the build; it just stops being counted as evidence that anything works.

    THREE OF THE FOUR ARE NOW CHECKED RATHER THAN TRUSTED. This paragraph has said "recount rather than trust it" for six recounts and gone stale after five of them, twice inside the session that recounted it — most recently between the commit that fixed the numbers and the commit two after it, which added cases. A habit that fails that reliably is a guard's job: test/rules/ai-counts.test.js counts the Go test functions, the Go test files and the frontend test files the way the commands below do, and fails when this line disagrees. The frontend TEST total is the one it cannot check — it.each expands at run time, so the only honest count is the runner's — and it is the one to distrust here. A number in a file like this one is stale the moment it is written, so recount rather than trust it — these three had drifted from 645 / 1,293 / 180 before anyone checked them at 1.12.0, from 725 / 1,581 / 203 before they were recounted at 1.14.2, from 765 / 1,672 / 217 by 1.15.0, from 807 / 1,759 / 233 by 2.1.1, from 924 / 1,771 / 284 by 2.2.0, from 1,085 / 1,844 / 320 by 2.3.0, from 1,100 / 1,853 / 323 when they were recounted for 2.2.3, and most recently from 1,153 / 1,977 / 338, from 1,336 / 2,218 / 394, from 1,357 / 2,223 / 398, from 1,360 / 2,245 / 401, from 1,380 / 2,358 / 418, from 1,391 / 2,366 / 419, from 1,466 / 2,772 / 471, from 1,493 / 3,041 / 520, from 1,493 / 3,071 / 521, from 1,493 / 3,083 / 522, from 1,493 / 3,111 / 523, from 1,493 / 3,324 / 533, from 1,494 / 3,350 / 535, from 1,494 / 3,416 / 540, from 1,494 / 3,426 / 541, from 1,494 / 3,431 / 542, from 1,494 / 3,434 / 543, from 1,494 / 3,435 / 543, from 1,494 / 3,436 / 543, from 1,494 / 3,439 / 543, from 1,494 / 3,448 / 543, from 1,494 / 3,449 / 543, from 1,494 / 3,450 / 543, from 1,494 / 3,562 / 551, from 1,508 / 3,590 / 555, from 1,508 / 3,595 / 556, from 1,508 / 3,610 / 558, from 1,508 / 3,616 / 559, from 1,508 / 3,624 / 561, from 1,509 / 3,626 / 562, from 1,509 / 3,631 / 563, from 1,512 / 3,633 / 563, from 1,513 / 3,642 / 564, from 1,514 / 3,645 / 565, from 1,522 / 3,645 / 565, from 1,524 / 3,646 / 565, from 1,533 / 3,648 / 566, from 1,538 / 3,652 / 567, from 1,542 / 3,652 / 567, from 1,685 / 4,102 / 633, from 1,685 / 4,135 / 636, from 1,687 / 4,137 / 636, from 1,690 / 4,142 / 636, from 1,690 / 4,151 / 636, from 1,692 / 4,154 / 636, from 1,692 / 4,156 / 636, from 1,692 / 4,157 / 636, from 1,692 / 4,158 / 636, from 1,703 / 4,188 / 638, from 1,703 / 4,191 / 638, from 1,705 / 4,198 / 638, from 1,707 / 4,201 / 638, from 1,708 / 4,202 / 638, from 1,708 / 4,209 / 639, from 1,708 / 4,219 / 640, from 1,708 / 4,234 / 641, from 1,708 / 4,245 / 642, from 1,708 / 4,248 / 643, from 1,708 / 4,257 / 644, from 1,708 / 4,265 / 644, from 1,708 / 4,272 / 645, from 1,708 / 4,281 / 646, from 1,710 / 4,292 / 646, from 1,713 / 4,296 / 647, from 1,713 / 4,299 / 647, from 1,713 / 4,312 / 648, from 1,731 / 4,430 / 660, from 1,738 / 4,430 / 664, from 1,738 / 4,435 / 665, from 1,738 / 4,455 / 681, from 1,738 / 4,470 / 695, from 1,804 / 4,784 / 792, from 1,804 / 4,787 / 794, from 1,805 / 4,787 / 796, from 1,806 / 4,789 / 796, from 1,806 / 4,793 / 797, from 1,806 / 4,794 / 797, from 1,815 / 4,794 / 799, from 1,816 / 4,794 / 800, from 1,818 / 4,794 / 800, from 1,818 / 4,797 / 801, from 1,819 / 4,797 / 801, from 1,819 / 4,801 / 802, from 1,819 / 4,802 / 802, from 1,822 / 4,827 / 810, from 1,828 / 4,828 / 811, from 1,828 / 4,834 / 812, from 1,828 / 4,898 / 817, from 1,958 / 4,834 / 839, from 1,893 / 4,954 / 835, from 1,998 / 4,954 / 856, from 2,000 / 4,957 / 863, from 2,010 / 4,968 / 866, from 2,009 / 4,961 / 864, from 2,013 / 4,972 / 868, from 2,006 / 4,974 / 866, from 2,023 / 4,976 / 869, from 2,029 / 4,998 / 872, from 2,031 / 4,998 / 872, from 2,036 / 5,004 / 875, from 2,037 / 5,006 / 875, from 2,039 / 5,006 / 875, from 2,040 / 5,006 / 875, and from 2,040 / 5,006 / 877, from 2,051 / 5,016 / 887, from 2,051 / 5,016 / 888, from 2,051 / 5,016 / 889, from 2,052 / 5,015 / 889, from 2,055 / 5,015 / 890, from 2,055 / 5,015 / 891, from 2,055 / 5,018 / 892 before the latest recount — which is why each one now sits beside the command that produces it. The last of those drifts is worth naming because it was one work session: a number recounted honestly at the start of a stretch is stale by the end of it. The 2.2.4, 2.2.5 and 2.2.6 passes added forty-four cases between them, every one for a defect a release review found rather than for a feature — and six of those defects were introduced by the pass before. That is the number worth reading here rather than the total: the two habits named below are what stops a fix pass being a source of work, and they did not stop it three times running.

    THE ONE THAT GOT PAST THREE REVIEWS is worth naming, because it is a gap in the habits rather than a lapse. A panel called a callback it did not define with no argument; the prop it had been handed was the host's record setter, so the host was set to undefined and the page unmounted. Twenty tests around it asserted requests, payloads and absences — and every one of them stubbed that callback as () => {}, which cannot see what it was given. The repair for THAT then passed the host its own record back — never undefined and never anything either, since setting React state to the same reference is a bail-out — and the new test could not see it, because it asserted what the panel passes and stopped there.

    A callback crossing a component boundary has three things worth asserting: that it is called, what it is called WITH, and what the other side does with it. A stub answers only the first, and the second missed a release.

    The frontend count went DOWN while its file count went up, which is the sort of number that ought to be explained rather than reported. 5751757 collapsed per-datum tests into tables and dropped eleven that asserted nothing the neighbouring case did not already assert. A suite is not better for containing eleven tests that cannot fail alone, and a total that only ever climbs is a total nobody is reading.

    Two habits behind those numbers are worth naming, because they are what stops a plausible test from being a useless one. Assert on values, never on counts: "got 3, wanted 3" passes happily while the three are the wrong three, which is the entire failure mode of a filter or a facet. And check the test by breaking the code: the search facets landed green on the first run, which is not reassuring on a change that touches fifteen queries — so the predicate was neutered to confirm seventeen tests noticed, and again to confirm two more noticed a subtler reversal. A test written after the code, by the thing that wrote the code, is worth exactly what its failure proves.

    A FOURTH HABIT, and the newest: open the thing in a browser before saying it works. The "add a link from an id" popup shipped its logic with thirteen unit tests over the URL it builds, every provider pattern pinned against the server's own — and the popup, rendered, had three defects at once. It drew no ✓ to press, because the form registered with the surface OUTSIDE the dialog rather than the dialog itself, and a modal with nothing registered draws no tick at all. Its four provider pills all looked identical, the chosen one included, because the class the click set (is-on) is not the class the stylesheet styles (active). And the link it saved drew a pill reading vocab.source.tmdb.label, because the provider table's middle column is a locale key and the pill builder used it as a name. All three passed 2,745 tests. None of them is subtle in a screenshot, and none of them is visible to a test that asks a function what it returns: a helper can be right in every case while the component that calls it renders a control that does nothing. The three DOM tests that now cover them were each watched to fail against the broken code before being kept, per the habit above.

    A SIXTH, and it is the fourth habit sharpened: the stylesheet a guard reads is not the stylesheet a reader gets. Every CSS guard in this repo reads web/frontend/src/index.css, which is the author's intent; web/dist/assets/*.css is what a browser is served. .tp-scrim declared backdrop-filter and its -webkit- twin, the minifier's collapse kept only the twin, and the focus blur the owner had asked for was in the source and had never once reached a screen — with every guard green, because every guard was reading the wrong file. It was found by getComputedStyle in a real browser on a real library and is now held by prefixed-pairs-survive.test.js, which asks the BUILT file whether each twin's standard property survived, and by panel-depth.mjs, which asks the browser whether the scrim actually blurs. The class is silent by construction: nothing errors, nothing warns, and the rule is simply absent.

    A SEVENTH, from the same measurement: "the hook is called" is not "the thing happens". PanelHost had called useBodyScrollLock since it was written, and the page behind a panel went on scrolling, because the lock hid the overflow of <body> and the element that scrolls a standards-mode document is <html>. A grep for the call site said the rule was obeyed. Reading getComputedStyle( document.scrollingElement).overflow said it was not.

    A FIFTH, from the same feature: a fixture that invents its input invents the answer. The publisher's record page came with a DOM test asserting a studio gets a company's page and a company's id spaces — passing on a record shaped role: 'studio', which the server cannot send. work_person.role for a studio is director: movies.director holds a film's director AND a game's studio, and media_type is the only thing separating them, a split six comments in this repo are written about. So the screen said "The person" over Electronic Arts and offered it an IMDb /name/ page while its own test went green. The fixtures carry the wire's shape now, and seven cases fail if the media-type arm is removed. Where a test hand-writes a payload, check the payload against the handler that builds it — an assertion is only as true as the shape it asserts on.

    1.15.0 added a third habit, for a case the other two cannot reach. Seven features in that release all rewrote the same two functions, so instead of building them in sequence I had seven implementation specifications written against the tree, each blind to the others, and then reconciled. It found three defects live in the shipped app before a line of the feature was written — a render branch that would print the answer above the options, a switch whose default returned the correct quote among the distractors, and a badge that counted cards the deck then refused to serve, whose test asserted the empty deck as correct. It also found two bugs that lived between two features and belonged to neither, which is the class no single spec and no single test was ever going to reach. The reconciliation is folded into docs/wiki/Design-decisions.md §8.

    The same release is also the clearest case for writing the plan first: the retired plan specified cloze grading word by word and said why in as many words — "a whole-string budget earned by long neighbours will hide a wholly missing short word" — and I built the whole-string version anyway. Three commits later a documentation pass compared the two and found that "want of a wife" was accepting "want of a life". The plan was right and the code was not, and only reading them side by side said so.

  • CI on every push: go vet ./..., gofmt -l . (red when it lists a file), go test ./..., a smoke test that boots the server and health-checks it, a frontend build, a check that the roadmap's generated regions still match the data files they come from, a check that the UI glossary's inlined stylesheet matches the one the app actually ships, and a check over the Home greeting's fixed-date tables — 129,210 greetings across 59 regions, every day of a year and every hour bucket. That last one exists because every way it can break is silent: a greeting rendering {name} literally, a commemoration wishing you a happy one, or a country resolving to its neighbour's time zone. None of those throw, and none of them fail a build.

  • Ten harnesses run a real browser rather than a DOM emulator, because what they measure does not exist in jsdom: make perf, make typescale, make frame-scroll, make panel-depth, make hero-control, make glyph-align, make controls, make sheet-drag, make metadata-layout and make overlay-scroll. (This line said eight while it listed nine.) make metadata-layout measures the Metadata and Settings desk layout (tab height, underline overhang, portrait size, two-line records, masonry holes), and each of its five checks was mutation-verified. The last is the newest and the clearest case for the whole category: it asks whether dismissing an overlay leaves the page where the reader left it, and NEITHER of the two ways that broke is observable in jsdom — it has no layout, and focus() there scrolls nothing whatever you pass it, so a restore that moves the page and one that does not are the same call. It is also the harness that found the second half of its own bug: the panel fix made those numbers right and the popover's stayed at 0. make perf measures how long the app holds the main thread per action; make typescale turns every type dial to 200%, sets the root font size to 24px, and fails when a screen clips something it did not clip before. The second is a DIFFERENCE and not a threshold on purpose: plenty of this app clips deliberately — a line-clamped intro, an ellipsised path — so a check that failed on all clipping would fail on the design and would be switched off within a week. Its first run found 59 elements, twelve of them the user avatar's initial cut off on every screen in the app at once, because a 38px badge was inheriting the page's 1.55 leading around a letter that grew with the type dial. The first repair froze the letter's size instead, which fixed the clipping by taking it off the dial — and typescale.test.js, a guard that predates all this, failed it in the same run. That is the check working: two rules met, and the older one was right. The remaining 47 sit in scripts/screenshots/typescale-baseline.json as a ratchet that may fall and never rise. Note what the LITERAL version of the pack's test would have reported here: nothing. Tippani never sets a root font size — applyTypeScale writes finished pixels into --type-* — so setting the root to 24px alone leaves the app untouched and would have returned a clean bill of health for a stylesheet full of px boxes.

  • A BRANCH THAT ONLY RUNS WHEN A PICTURE EXISTS IS THE BLIND SPOT BOTH LAYERS SHARE. screens-mount.test.jsx mounts every screen with every request REFUSED — deliberately, and its own note argues the case well: a refusal needs no invented payload shape, so the file cannot rot into eleven guessed formats. make controls runs against seed.mjs, which has no artwork, because this container cannot fetch any. Between them, no layer ever executes the arm of a conditional that draws a picture.

    So a free identifier inside one of those arms compiles, bundles, passes 3,400 tests and throws the first time a reader with a real library opens the screen. Three shipped in three consecutive commits, all in one span of StatsPage.jsx, each found by a rater reading the diff rather than by anything that runs.

    This note used to end by saying the class could not be mechanised, and that was wrong — wrong in the one direction the owner's standing instruction forbids ("create mechanical controls … don't depend on your prompting"). It weighed two ways of RUNNING the code, a smoke pass against per-branch rendering, and never considered READING it. The next rater pass found two more of the same defect live, in code nobody had touched: the Delete key on a row of the annotations table and of the dialogue table each called a setter bound in their parent. Both were reachable buttons; both threw when pressed.

    test/rules/no-free-names.test.js is the mechanical control. Babel is already in this project's node_modules — Vite's React plugin brings it — so the scope analysis costs no download and about a second, and it names the file and the line. It is declared in devDependencies anyway, because a transitive dependency is a fact about somebody else's package.json rather than a promise. It asks one question and is not a type checker: whether every identifier a module reads is bound somewhere it can see. It catches both the shipped StatsPage crash and the two the rater found, each by name.

    Its allow-list is the part that can quietly stop working. A browser's globals include plain English words this app also uses for its own things — open is its commonest prop name — and excusing one blinds the check to exactly the parent/child shape it was written for, worse than a crash, because window.open is truthy and the screen renders wrong instead of throwing. Eight such names were on the list and none of them was carrying any code; they are off it, and a synthetic case now shows the check a child reading its parent's open. A window API is written window.open.

    And that pruning is no longer a judgement, because the judgement was wrong. Removing those eight by hand left location behind — the app binds it in two files — and a rater found it the same afternoon. So a case asserts the list is EXACTLY the globals the tree reads: 57 of its 106 names were excusing nothing at all. The rule costs one line of upkeep the first time this app uses crypto, and that is what it buys — the name gets on the list on the day there is code to check it against, rather than years before.

  • A GUARD THAT WALKS THE SOURCE TREE PASSES WHEN THE WALK FINDS NOTHING, and two dozen of them do. Most are a list compared against [], so an extension narrowed by one character, a moved tree, or a wrong TIPPANI_SRC turns the guard green while it checks nothing. Not all of them, and an earlier version of this paragraph said "every one", which is not true — most carried a floor of their own, which is what made the ones that did not so easy to miss: the habit existed, it was just not enforced. Measured by disabling the shared floor and running the suite, seven guards stay green over an empty tree: confirm-finality, nested-dismiss, person-router, infodot-copy, one-stand-in, pack-citations and typescale. (An earlier count of four here was taken before the conversions and never re-measured; this one is from the run.) That is what the floor in src-files.js is for — none of the seven can be reached now without it throwing first. Two were caught this way — by a rater, not by a run — and the second was the sibling of the first, in the commit that fixed the first. A per-file assertion is the wrong repair for exactly the reason this repo already states about screens: it is a line each, and that is how one of them goes on being right while the other quietly stops. test/src-files.js is the shared walk, and it THROWS below its floor rather than returning a short list — a tree floor for a wrong root, a per-guard floor for a predicate that stopped matching.

    Twenty-eight guards read a directory themselves; twenty-two have converted. Most were a plain readdirSync(SRC).filter(...) — non-recursive — so each had ALSO been silently skipping src/demo/install.js for as long as that directory has existed, and every one still passes with it in scope. The six that remain, and one added at 3.0.2, read a directory that is not the source tree at all — web/dist, src/textures, docs/plans, the repo, .github/workflows — and test/rules/one-walk.test.js names each with its reason. So the count is a floor being held rather than a debt being paid: what it stops now is a NEW guard walking the source tree by hand, which is the case that matters, because a new guard is written by whoever has just been bitten by the thing it checks and is not thinking about whether its own walk can come back empty.

    The per-branch rendering stayed as well, because the two answer different questions: the scope check knows a name is missing, and only a render knows the arm draws the right thing. Neither presses anything, and the owner's standard for a test is "is the button clickable (for all buttons)?" — so test/dom/table-row-delete.test.jsx renders both tables and presses the Delete key on a row: two cases per screen, and reverting one screen's row fails that screen's two. It also asserts the half that would be easy to break while fixing the first half: the press ASKS rather than deletes, the same question the card view puts. Note that a throw inside a React handler does not come back out of fireEvent — the synthetic event system reports it to the window — so a case that only presses asserts nothing, and this one listens for the error event.

    AND A PROPERTY CAN BE HELD REDUNDANTLY, WHICH BREAKS THE MUTATION TEST WITHOUT BREAKING THE CODE. Mutation is how every case in this file was checked — delete the fix, watch the case fail, put it back — and it assumes one line holds the property. A phone sheet interrupted mid-landing is guarded three ways on purpose, because a gesture arriving during the settle animation can reach that code by three paths and the cheapest correct answer is for each path to check: so test/dom/sheet-from-the-bottom.test.jsx's double-drag case stays green when any ONE of the three is deleted, and fails only on all three. That is not a sleeping case, and the case says so in its own comment so the next reader spends no time on it. The alternative — funnelling the three paths through one checked function so a single mutation would fail — trades a real property held three cheap ways for a testability property held one way, and the gesture is no more correct for it. Where mutation cannot settle a case, the comment has to.

    AND A PROPERTY CAN BE HELD IN TWO TIERS AT ONCE, WHICH TAKES TWO MUTATIONS. Practice stopped being a "seeing" event in two places, because it was one in two: the server bumped the half-life after a Practice answer, and the client reported a Practice card's three other options to the same route. One mutation cannot reach both — deleting the server's leg leaves the client's intact and the library still drifts — so each has its own, and each was run: re-adding the server bump fails TestReviewSeen at the practice leg, and removing the client's mode === 'practice' fails reports nothing as seen in Practice. A third covers the shape a plausible fix would have taken: gating the bump on srPracticeCounts rather than removing it, which compounds 2.5 × 1.05 and fails the same Go case on the NUMBER rather than on which line ran.

    AND WHAT IS NOT COVERED, SAID RATHER THAN IMPLIED. There is no browser case for it. The effect is invisible at the seeing multiplier's 1.0 default, so a journey would have to move a RANGE first — and this harness has no verb for a range, by design: type clicks and types, choose drives a <select>, and neither operates a slider. Adding one is a vocabulary decision rather than a test, so the schedule arithmetic stays where it is observable as a number, in the Go tier over real HTTP against a real SQLite file; this paragraph is the record that the gap was chosen rather than missed.

  • make controls asks a question of every control instead of asserting a fix. It presses everything a reader can press on fifteen surfaces — twelve screens, both work details, and the character panel, which is reached through a DOOR the run opens first and refuses to continue past unless a panel actually opened — and asks two things of each: did anything at all change — a dialog, a panel, the route, focus, the scroll position, the surface's own text — and if not, did the control SAY it was disabled. A control that answers no to both is a lie to the reader whatever the reason, and the reason is never visible from the outside. It also counts a menu's rows, and flags one that opens empty.

    It is written that way because the alternative does not work. A test that knows the fix can be written to fail without it and still guard nothing — assert the exact handler that was repaired and the next dead control beside it passes. This asks the property, so a control nobody has looked at has to answer too. Its first run found the ⋯ menu opening an EMPTY floating card on six of twelve screens: the desktop bar drops its own Help row because a ? stands two controls away, and Home, Quotes, Search, People, Stats and Settings publish no actions on a desktop at all — their useScreenBar rows are mobile && gated. Every screen rendered, the button opened, and the defect was the ABSENCE of rows in a box one line tall. No screenshot review and no unit test in this repo was ever going to see that.

    Three of its own failure modes are failures rather than notes, which took three separate lessons in one session to get right: a surface that draws almost nothing did not render and is reported instead of scoring "0 controls, 0 dead"; a control that moved between enumeration and the press was not tested and is reported instead of skipped; and the run exits non-zero on a finding rather than on a warning.

    And one surface may not cost the other fourteen. The run stopped on Settings with a detached frame and reported 412 presses: every surface after it, and the whole 390 pass, went untested while the exit code said only "stopped early". A probe one screen can silence does not guard the rest. Each surface now runs behind its own fence, and one that throws is filed as a surface that did not render — a FAIL, because "nothing on it was tested" is the same fact whether the screen came up blank or the run fell over on it. What threw was a class the probe could not see: a button whose handler assigns location.href (Settings.jsx:2407, Download, which streams the archive when there is one and navigates to the server's answer when there is not), where the check for a control that leaves the app reads an anchor's href and this is a button. That press is caught where it happens, recorded report-only because the control DID something, and the surface re-opened so the controls after it are still pressed. Not a skip list keyed to the control's name: this file has twice learned that a probe keyed to a spelling stops guarding the class the moment the spelling changes.

    AND THE TOUCH-FLOOR CEILING IS A FACT ABOUT A LIBRARY, NOT ONLY ABOUT A WIDTH. controls-baseline.json was keyed by width alone and its 390 number — 326 controls under 44px — was measured against the owner's restored backup, while make controls seeds public-domain titles and a cast of three. A bigger library draws more controls, so the seeded run measured 187 against 326 and printed ok with 139 of slack: a hundred new sub-44px controls could have landed under it. The key is [fixture][width] now, the harness names its shelf, and --update-baseline refuses to write a ceiling under no name at all.

    The part worth keeping is WHY it was invisible. The probe deliberately does not fail on a missing ceiling — failing there is how a ratchet gets deleted rather than filled in — so a fixture name that misses the baseline turns the ratchet off and the run still exits 0. test/rules/controls-ratchet.test.js is the second control: it asserts, in 200ms, that the shelf the harness names has a ceiling at every width the harness runs. Two fifty-minute runs cannot notice what a file-shape check finds instantly.

    AND HALF A SHELF PASSES THAT CHECK, which is the version of it that shipped. The seeded shelf was required at every width and the backup shelf was exempted, on sound reasoning — a CI machine has no archive, so demanding that shelf would fail a check nobody can satisfy. The unstated consequence: the backup shelf had a 390 ceiling and no 1280 one, so make controls on the machine that HAS the archive could not exit 0 at all, because an unrecorded width exits 3 — and the case that would have said so was the one that had been narrowed away. The rule is about COMPLETENESS now: nobody has to record a shelf, and anybody who records one records every width, because a shelf with one width is a gate with a hole in it that never stops exiting 3 to say so.

    THE HARNESSES TAKE THE OWNER'S ARCHIVE BY DEFAULT, AND FOR A DAY ONLY ONE OF THEM DID. The owner's ruling was "use it for all tests"; the branch was written into run-controls.sh and CLAUDE.md was written as though all seven had it, while five went on calling seed.mjs unconditionally. Wiring them up then failed on the layer underneath: HARNESS_ACCOUNT is a hard-coded screenshot-bot and the TIPPANI_USER override was a line in controls.mjs, so that probe reached a restored library and the other six got a 401 and thirty seconds of waitForFunction before dying with a timeout that named no account. And under THAT, a hard-coded --movie-id 2 — a fact about the seeded fixture — asked the archive for a page that is not a film. Three layers, one shape: a rule written once, in one of the places that needs it.

    test/rules/harness-archive.test.js is the guard, and it checks the shape rather than the run: every harness that seeds calls the one shared decision, calls it before anything is built or booted, and keeps no copy of the branch; the account is decided in exactly one file and that file lets the environment win; and the backup.env parser is run in a real shell against a file with no trailing newline, an export prefix, an indent, quotes and CRLF — every one of which used to drop a variable silently and fall back to seeding, which prints the same first line as a machine with no archive at all.

    The proof that the wiring works is runs rather than a claim, and the count is exactly six: make sheet-drag (ten ok), make panel-depth (seven, having opened Geralt of Rivia in The Witcher 3), make typescale (no new clips), make frame-scroll, make hero-control, and run-with-server.sh --seed --screens home (one capture, of the owner's own Home). Each against the owner's restored archive, each exiting 0.

    AND THREE OF THEM WERE RE-RUN ON eb37313 rather than left quoted from an earlier build, because that commit changed code they exercise: typescale because the clamp exemption narrowed from skipping the element to skipping only its vertical check (which can newly count a clamped box clipping SIDEWAYS), and panel-depth and sheet-drag because both now choose their subject through pickFilm. typescale 0 with no new clips, panel-depth 0 with seven, sheet-drag 0 with ten. Quoting the earlier runs would have been the mistake this file has now recorded three times.

    make controls IS THE SEVENTH AND HAS NOT DONE SO YET, and this paragraph said "seven runs… each exiting 0" for half an hour while it hadn't. It reaches the archive and walks every surface clean, and it exits 3 — "the app came back clean AND the touch floor was measured against nothing" — because the backup shelf's 1280 ceiling has never been recorded. A --update-baseline run is what records it, and that run takes seventy minutes against a real library. Until it finishes and a plain run confirms the numbers, the honest sentence is this one. A rater caught the earlier version by reading the commit message next to it, which said in as many words that make controls could not exit 0.

    AND A PROBE RUN IS A CLAIM ABOUT A BUILD. The first two sheet-drag runs quoted in the register predated part of the change they were quoted for — offsetNow's matrix3d arm landed after them — so it was re-run on the commit itself. A build that has moved since the run is a different claim, and the gap is invisible in the log.

    The seeded ceilings are 0 / 0 at 1280 and 187 / 9 at 390, and they were measured twice: a second full run returned the same two numbers across all thirty surfaces. An exact ceiling is only worth having if the count is stable — a second run of 186 or 188 would have meant recording a range and saying so.

    And the ratchet judges both directions now. A count that ROSE is the easy half; a ceiling the app has LEFT BEHIND is the other, and it is the half that let 139 controls of room sit unnoticed — the thing gets better, the number stays, and the gate has space in it until the next regression spends it. Both fail. That arithmetic moved out of controls.mjs into scripts/screenshots/ratchet.mjs for a reason worth stating: while it lived after a browser walk of thirty surfaces, the only way to ask whether the rule was right was to spend an hour producing an input for it, and it went wrong twice without anyone noticing. controls-ratchet.test.js imports the same module and asks it both ways in a millisecond.

    And the fingerprint had no geometry, so one control's whole effect was invisible. A phone sheet's grab bar moves the sheet between its anchors and does nothing else — same panel, same rows, same text, same scroll — so the probe pressed it, saw every field it reads unchanged, and reported it dead. That press was ALSO genuinely dead for a reader, for an unrelated reason, and the two hid each other: fixing the bar would have left the probe still reporting it, and trusting the probe would have hidden that there was anything to fix. The fingerprint now carries the sheet's height — the sheet's and not the body's, because a body grows when a lazy cover arrives and a dead control must not be able to borrow a change it had nothing to do with.

    Its buckets are split three ways, and that is what makes the gate reachable. Six FAIL, because each is the app lying to a reader and each has one right answer: a control that does nothing and does not say so, a menu that opens empty, a header that is not the pack's, a screen that scrolls sideways, a route that is not a screen, a surface that did not render. Two RATCHET against controls-baseline.json PER WIDTH — controls under the 44px touch floor, and controls drawing a glyph beside their words on a phone — because both are real debt that no one commit can zero, and a count allowed to grow is a count nobody reads. One is a report only: a control the harness could not reach is as often a fact about the harness as about the app, and failing on it would make the gate a measure of the run's luck. It used to fail on ANY non-empty bucket, which made a green run impossible by construction — every run read FAIL, so the FAILs that mattered stopped being read. panel-depth.mjs and hero-control.mjs each shipped with a silent SKIP that exited 0, and each was cited in prose as a guard while a run of it could prove nothing at all.

  • Two guards added in 2.2.0 were each watched to fail before being kept, which is the same standard the rest of this list is held to and is worth naming because both protect something invisible. CASE WHEN ? <> '' in the cast merge is what stops a TMDB refetch blanking the character art a TheTVDB fetch found — replaced with plain assignment, the test reports both an untouched and a corrected row losing it. And OneTimeEnv.FreshInstall, which stops a new install being told about a change it never lived through, needed a SECOND test: on a genuinely fresh database the pass writes nothing whether the guard is there or not, so inverting the guard left the obvious test green and only a direct call with a populated database caught it.

  • A fix can be present, commented, and inert. Three handlers cleared the server's 60s write deadline for a long job with http.NewResponseController(w).SetWriteDeadline, each with a comment explaining why it was needed. All three did nothing: the controller walks the response-writer chain by calling Unwrap(), this package wraps every response twice, neither wrapper implemented it, and the http.ErrNotSupported that came back was discarded with _ = at every call site — because there is nothing useful a handler can do with it. Nothing logged, nothing failed, and the source read as though the problem had been handled. It was found only by writing a throwaway probe that built the real middleware chain over a real connection and printed what the call returned, which is the general lesson: when a fix cannot fail loudly, the only way to know it works is to ask it. The permanent version of that probe is now a test, and it asks both with and without gzip, because a fix to one wrapper and not the other passes for browsers and fails for everything else.

  • A green suite is evidence about whatever the suite can see, and no more. A screen was built to lock the document and scroll two columns inside it, marked done on 2,238 passing frontend tests, and did not scroll at all: one link in its chain of heights was missing, so both columns grew to their content and the locked body cut off everything past the first screen. jsdom has no layout — scrollHeight is a constant 0 there — so not one of those 2,238 tests could have failed on it, and none of them was about it. The sentence "done, suite green" was true and irrelevant in the same breath. The repair is two guards rather than one, because a single one cannot cover this: screen-scroll-chain.test.js reads the stylesheet and fails when a link is deleted, and make frame-scroll opens the page in Firefox and fails on a clipped page or a column that cannot scroll. Before claiming a layout works, measure the layout — the browser harness exists for exactly the class of failure the unit suite is blind to.

  • A test that catches its bug in six runs of eight has found the bug, and does not guard against it. 3.1.0's queue and log writer are concurrent by design, and two of their tests first caught their own named defect only by chance: the wait for a job's last lines before its finishing write (one failing run in three of -count=5 with the wait removed), and the order of shutdown's steps (six or seven runs in eight with the log pool closed first). Each was rewritten to hold the race open instead of hoping to land in it — a hook that parks the log writer, a switch that holds every line until shutdown's last flush — and was then red on every run with the defect put back. Two such switches are in the shipped binary, TIPPANI_JOBS_HOLD (the queue claims nothing, so a journey can see a job wait, or, set to running, holds the job it claims at its start, so a journey can see one run) and TIPPANI_LOG_HOLD. Both are honoured only while TIPPANI_OFFLINE is on, which a test proves, and every file that uses one names it in its header. internal/jobs runs raced every night with the other packages; TestEveryTestedPackageIsInTheNightlySweep is what made it join.

  • Two numbers that must agree read each other rather than a copy. The threshold that decides whether a cover is worth replacing is the server's (lowResCoverWidth); the client draws the same fact in red. Written as a comment saying "mirrors the server", it drifted for as long as nobody looked, and a design handoff then proposed a third value for the same question — at which point the interface could call a cover unusable while the one button offered to repair it declined. cover-floor.test.js parses the Go constant out of its own source file and asserts the JS export equals it, which is bin-kinds.test.js's shape applied to a scalar: read the authority, never a transcription of it. The same test also asserts the client keeps no second copy of the number, because the drift surface a check like this closes is the copy, not the disagreement.

  • The committed SPA is checked against the sources it was built from, by web/dist_inputs_test.go against the hashes in web/dist-inputs.json that npm run build writes. This one is here because the guard it backs up was correct and still let a stale bundle onto main: CI rebuilt web/dist and diffed it, which works, but only speaks after the push — and the commit it caught had not touched web/frontend at all. It edited internal/i18n/en.txt and bn.txt, which src/i18n.js imports with Vite's ?raw, so every string in the interface comes from them and editing one changes the bundle. Nothing about a .txt file inside a Go package says "you have just changed the frontend". Moving the check into go test is the whole point: it needs no Node, no six-second build and nothing installed in the clone, so it fails in the working tree rather than in a workflow. The input set outside web/frontend/ is derived from the imports that escape it, because a hand-kept list is the same blind spot one level up.

  • docs/wiki/Design-decisions.md is a decision log: every design decision, the reasoning that produced it, the alternative turned down, and — where it applies — the part I got wrong and what changed my mind. It used to be the design document I wrote before building, and it had drifted into describing a system that was partly never built and partly rebuilt underneath it. A log fixes that structurally: an entry that records a decision and records that the decision was wrong stays true forever, where a design document that is 80% accurate gives you no way to tell which 80%. Comments in the code explain why rather than what, for the same reason.

  • Migrations are numbered, embedded and append-only, each in its own transaction — and the runner refuses to open a database newer than itself. Forward-only means the downgrade failure mode is a success: an old binary finds all its own migrations applied, skips them, returns nil, and serves an app missing every table added since, with nothing in the log. It happened, from a stray tag winning :latest in CI. Stopping with both version numbers turns a four-migration audit into one line.

What that honestly does not cover:

  • Passing tests are not proof. A real example from this repo: the parity test whose entire job is to fail when a field is added to one quote kind and not the other skipped embedded structs, so two new fields rode past it while the suite stayed green. It was found by reading the test, not by running it. Tests can pass for the wrong reason, and AI-written tests can be confidently wrong about their own coverage.

  • Confident documentation is not verified documentation. Where this file makes a claim, it was checked against the tree; treat prose elsewhere in the repo as a strong hint and the code as the truth.

  • The demo was the one thing here with no test at all, and now has one. web/frontend/src/demo/install.js is a fetch shim that answers the API with dummy data so the Pages demo can run with no server, and nothing checked it against the handlers it is imitating. In 1.4.1 its backup response was found returning created_at where the server returns created — so the demo's Settings screen had been rendering "Invalid Date" for as long as that card existed. Nobody's data is at risk from a shim, which is exactly why it drifts: a fake that is close but not identical fails in the one place no test looks. 1.5.0 exported its route function and started asserting shapes, which caught two of the same class before they shipped. The cover is partial — it asserts the answers the newest screen reads, not every route — so the drift risk is reduced rather than closed.

  • A bug report is a report of a symptom, and the symptom is often not the bug. 1.7.0 opened with "quotes should be included in the daily quiz". They already were — the deck has drawn them since the medium existed, and two tests written to check that passed against the code as it stood. What was broken was the Settings control, which offered Books, Films & shows and a third option labelled Both that silently meant all three: the word undercounted what it did, and there was no way to ask for quotes at all. Building what the report literally asked for would have meant changing a deck that was correct. The first move on any report like this is a test that tries to reproduce it, and the useful outcome is as often "this passes, so look upstream" as a red bar. Neither of those two tests existed before, either: every deck test seeded a single medium, so "the deck serves standalone quotes" had only ever been asserted for a library containing nothing else, which is not a library anyone has.

  • Consistency is not something you can review for. 1.6.0's whole subject was the app agreeing with itself, and almost nothing in it was found by looking at a screen. Four icons were another icon — Share and Upload were both a tray with an arrow in it a pixel and a half apart, Export and Metadata were the same three strokes at coordinates half a unit apart — and the nav kept its own copy of the set at a different stroke weight, so the Library tab was the identical open book the "currently reading" badge wears. The same action was delete on a card and del in a table. A window could be dismissed five different ways. Every one of those is invisible in a screenshot of any one screen, because each screen is internally fine. The tests that now hold the line compare every exported glyph with every other one — both exactly and with all coordinates stripped, so a near-miss cannot hide behind a rounding nudge — and read the source for control labels rather than the help file, because a doc test that only reads the docs agrees with itself forever. Three of them earned it immediately by failing on the tree as found.

  • The frontend had no test runner, so anything it parsed was parsed on trust. 1.5.0 added one (Vitest, dev-only — the three runtime packages below are unchanged) and moved the two bespoke check scripts into it. What it covers is the pure logic: routing, credit splitting, recall status, grouping, the share formats, the quote-card wrap engine, the partial-date rules, the demo shapes. Rendering is still largely unchecked, so this list has shrunk rather than emptied. What prompted it was this, from 1.4.2: web/frontend/src/secret.js reads the backup archive's binary header in the browser, by fixed byte offsets into a format defined in Go. Nothing checked the two agreed. A check script did, in CI — it is a Vitest case since 1.5.0 — and it earned itself immediately by failing on the first run, for a bug I had written into the parser minutes earlier: the read window covered a maximal account name but stopped a few bytes short of the field after it, so an archive's recoverability read as absent for exactly the accounts with long names. That is the shape of every bug in this class. It does not throw, it does not look wrong, and it is only ever wrong for inputs nobody happened to try.

  • docs/ui-glossary.html is generated now, and the hand-written half is what had gone wrong. Its oldest failure mode was mechanical: the page inlines the built stylesheet so its samples are styled by real app rules, and every frontend build renames index-<hash>.css, so the snapshot rotted silently. A generator regenerates it and CI fails when it is stale, which ends that half — though only after 1.5.2 found that it had been regenerating 140KB of stylesheet inside an HTML comment for two releases, because the page's opening comment named the <style> tag in its own prose and never closed, and the generator finds its block by searching for that tag. Every sample rendered unstyled, which is the one thing the page exists not to do, and --check passed throughout: the bytes it compared were exactly the bytes it had written. A gate that only reads its own output cannot fail. It refuses now rather than writing into a comment.

    The entries used to be hand-written too, and they had lagged a full release. The page went on offering a Paper / Film aesthetic toggle after v3 replaced aesthetics with material sets — a control driving an attribute that appears zero times in the stylesheet — and drew four CSS classes the app had deleted. They are now built by web/frontend/scripts/glossary-build.mjs from a catalogue module, from glossary declarations beside the components (which render the real component, so a sample cannot quietly lose a class the way the buttons had lost tactile), and from web/frontend/src/tokens.js for the constants. glossary-registry.test.js fails when an entry names a class or an identifier the source no longer has, and it carries two ratchets — 57 components with no entry (37 of them icon glyphs), and 141 of the 171 entries still on carried markup — that may fall and may not rise. Documenting the filled icon set took the glyph figure from 49 to 37 in one pass, which is the ratchet working as intended rather than a number being relaxed. Those two numbers are the honest measure of what is still undone here, and they are asserted rather than described. Feeding the prose from web/frontend/src/help.jsx remains the roadmap's help & density section. What 1.6.0 added is the cheaper half of that: web/frontend/test/rules/help.test.jsx asserts that every screen a nav list can reach has an entry, and that a control the app labels is a control the help names. It found three gaps on the first run — the whole Quotes filter row, the Catalogue's group-by, and an "Export all" button that had stopped meaning "all" when the list screens started exporting the filtered view. All three had been read past repeatedly. The test reads the source for those labels, not the help file, for the reason above: a doc test that only reads the docs agrees with itself forever.

  • Known bugs are recorded, not hidden — see Known bugs, not yet fixed on the roadmap. That list is generated from the issues I have accepted, so it cannot quietly go stale in either direction: an accepted report lists itself without me writing it up, and a fix removes it only by the issue actually being closed. Accepting is a human step on purpose — publishing a stranger's text to a public page automatically is a different risk from the ones on this page, and no amount of escaping makes a wrong report right.

  • AI review is worth more than AI code, and it is the same model. 1.4.2's design went to three adversarial reviewers before a line was written — one asked to attack the cryptography, one disaster recovery, one the Go implementation — and between them they killed the design. The recovery key was to live in a column of the users table; a restore replaces that table wholesale, so restoring any archive, resetting the instance, or deleting the account would have destroyed the key silently, and the only surviving copy sat in a directory the next restore deletes. Two of the most ordinary operations there are, in order, no error at any point. The same reviews disproved a claim I had already put in the 1.4.1 release notes — that renaming an account orphaned its archives — by pointing at the two lines that make it false. The fix for both was to make the design smaller. Worth being precise about what happened, though: the reviewers are the same model that wrote the design, given a different instruction. What changed was not intelligence but stance — "find what is wrong with this" is a different question from "build this", and it is the one that was not being asked.

  • A confident diagnosis is worth no more than confident code. The concurrency defect that sat in that list for two releases came with a written-up cause — the connection pool allows four writers where the plan specified one — and a written-up fix, a mutex. Both were wrong. The real fault was the lock order: a DEFERRED transaction that reads before it writes must upgrade its read lock, and SQLite fails that upgrade instantly rather than waiting, so the 5000ms busy_timeout was never consulted. It was caught by making the test fail on purpose and reading the error code — 517, SQLITE_BUSY_SNAPSHOT, which names the upgrade — not by re-reading the explanation, which was fluent and had been sitting there being fluent for months.

  • Changing a default changes nothing, if the default was ever written down. 1.7.2 turned the share image's colour switch off by default. The one-character version of that change — flip true to false — is correct, reviews clean, and alters the behaviour of no device that has ever opened the panel: the hook behind it persists on mount, so the old default had already been stamped into local storage by the first render, and a stored value beats a default every time. What makes this worth writing down is that the obvious test passes. Clear the storage, render, assert the switch reads Off — green, and green for a change nobody would experience. The test that means anything is the one that seeds the retired key with the old value and asserts the switch still reads Off, which is the only version that can tell a default from a decision somebody made.

  • The safest-looking way to walk a schema was the one that lost data. The bin snapshots a deleted thing and its whole subtree, and "follow the foreign keys" is obviously the robust way to find that subtree — it cannot go stale, it needs no list to maintain. Except that in this schema two of the tables that travel with a quote have no foreign key at all: item_reviews is polymorphic (kind, item_id) and work_reads is (kind, work_id), and both are cleared by AFTER DELETE triggers instead. An FK walk finds neither. The restore then works perfectly: the book comes back, the quotes come back, and a year of spaced-repetition history is silently a new card. Nothing throws, nothing logs, and the person who notices is the person who had that history. The fix was to declare the table list and write down, in the migration, why the obvious approach is wrong — because the next person to touch it will reach for the FK walk for exactly the same good reasons.

  • A feature request names a mechanism; the thing worth building is the need underneath it. "Show the changelog by fetching it from git" cannot be done by the shipped artifact at all — the image is distroless, there is no git and no shell in it, the docs are outside the build context, and the CSP stops the browser calling GitHub. Taking the mechanism literally leads to a dead end; taking the NEED — "stop sending me to a website to find out what changed" — leads to embedding the file, which is better on the hardware this runs on because it also works with the network off. Worth saying out loud to whoever asked, though, rather than quietly building something else.

  • "No importer can fill this column" is not a reason to leave it out of the import path. Adding a translator to books made the staging queue look irrelevant: no third-party format carries one, so a column on the approval queue could only ever move an empty string. The step that argument skips is that the app's OWN Markdown export is an importer's source, and every import is staged — so the field survived the export, survived the parse, and was dropped on the way into the queue. Exporting a library and importing it back would have lost every translator in it, with a successful import and matching counts saying nothing had happened. The question is never "can a source fill this", it is "does anything on the round trip have to carry it".

  • A switch written as a default plus overrides is a bomb with the fuse in the next feature. handlePeopleNames read q := <the books.author query> and then overrode q for actor, director and speaker. Correct for exactly as long as author was the only book-side role — and the moment translator became valid, asking for translators answered with every AUTHOR in the library, tallied, named as translators and offered for renaming. The same file already carried a twenty-line comment explaining this hazard about two OTHER functions, which had been fixed; this third one had been missed. Knowing the shape is not the same as having swept for it.

  • A control strip that grows one button per release has no moment where it breaks. The selection bar shipped with four word-buttons and left 1.11.1 with eleven controls, every one of them added for a good reason and none of them the one that broke it — because nothing broke. It simply became, on a phone, a strip wider than the phone, one release at a time. The fix was not smaller buttons but a decision the bar had never been asked to make: WHICH THREE matter, with the rest behind a ⋯. Worth noticing that the failure had no error, no test and no screenshot — the only signal was a reader saying it looked crowded.

  • Turning words into glyphs takes the state off the screen with them. The quiz toggle reads “Skip in quiz” or “Add to quiz”, and that label was doing two jobs: naming the action and reporting which way round the selection currently is. A single glyph would have kept the first and silently dropped the second, and the loss is invisible in review because the button still works. So the picture flips with the label — two glyphs drawn as a pair, which is the only reason the set has two drawings that are nearly the same picture on purpose.

  • A rule that five queries splice belongs in the one string they share. Keeping a quote out of the Daily Quiz is a single condition, and it has to reach the three candidate fetches, the count behind the cards-left badge, and the breakdown behind "where you stand". Written into four of the five, the failure is not an error: it is a badge counting a card the deck will never serve, which reads as the quiz being broken rather than as a filter being inconsistent. reviewSource.where() already existed as that shared string, and the whole feature is one clause added to it — which is only obvious once you have gone looking for every place the condition would otherwise have to be repeated.

  • The obvious home for a flag can invent history. "Not in the quiz" is a scheduling fact, so item_reviews looks like where it belongs. But that table has no row at all for a quote that has never been reviewed, and four separate queries read "a row exists" as "this card has been seen" — so excluding a quote and putting it back would quietly promote it from never-seen to seen-and-overdue. A lie about the reader's own history, told by a preference they set for something else entirely. The column went on the quote instead, where it also travels for free through the bin, the account backup and the export, none of which had to learn it exists.

  • A presence flag is a function waiting to be called. The selection bar draws some of its own controls (the tag field, the sticker dialog, the delete confirm), so for those it passed a bare true into the action registry purely to mean "available". That read fine for two releases and broke the moment a new action — a shelf dropdown — was one whose run the bar actually invoked: the action appeared, the control rendered, and choosing a shelf threw. Nothing about the flag was wrong; it was that a value in a slot typed as a callback will eventually be called.

  • A gesture bound to closest('button') breaks on the day the card IS a button. Long-press-to-select ignores presses that land on one of the card's own controls, which is right, and it did it by asking whether the target had a button ancestor. Quote cards are divs, so it worked. A cover tile is a <button> — the whole cover is the thing you click to open the book — so every press on it matched, and the gesture did nothing at all on two entire screens. Invisible to inspection, invisible on a desktop, and caught only because the new board got its own test.

  • A bug of omission needs a sweep, not a case. Scrolling inside a popup also scrolled the page behind it. The fix is one CSS property, and the temptation is to add it to the popup that was reported — but the defect is not a wrong value anywhere, it is that nobody thought about scroll chaining at all, so the test is an invariant over the stylesheet: every scroll container declares overscroll-behavior, with a named exemption list and a guard that the sweep still matches something. Widening it from vertical to sideways immediately found one nobody had reported — the top navigation, where running off the end is the browser's back gesture, so a nav that navigates away. The same sweep showed the other half was worse than the report: eleven full-viewport overlays, and seven of them never froze the page behind them at all.

  • scripts/claude-kit-setup.sh is checked by scripts/claude-kit-setup-check.sh, and the one thing it is for has never been run. The check is a sandbox, not a stub suite: its own HOME holding a copy of the installed kit's real guard, a real git init, and stubs only for claude and npm, the two commands that would reach the network, both logging their arguments so the check sees the right plugin installed and npm ci run in both packages. It runs the setup script through the cases its history broke on — a hook that is not the kit's, the kit's hook naming a version that is gone, one whose text has drifted, the kit's text under someone's own commands, a plugin record naming an older version than the newest cached or a directory that is not there, a settings.json that is not JSON, an exclude file with no final newline or one that only mentions visual-verify in a comment, a linked worktree, core.hooksPath, a HOME with a space, no clone — and asserts the last line of each, so no run over a hook it did not write ends in "set up". The hook is exercised through real git commits, which run it only if it is executable: one with a copied kit file staged must be refused, one with an ordinary file must go through, for the written hook and for the line it offers (with a command after it, and after the kit moves to a new version). It is run by hand; CI cannot fetch the private kit the guard comes from. Mutation-checked: thirteen edits to the setup script each turned it red — overwriting any hook, dropping the audit, exiting with the failure count, dropping the newline repair, preferring the newest cache to the record, not checking the record's directory, dropping the chmod, not counting a foreign hook or a missing clone, dropping one npm ci, misspelling the plugin, dropping || exit 1 from the offered line, loosening the exclude match.

    The script's first version was also run for real in the cloud container on 25 September, against a fresh state made by moving the plugin records, the plugin cache, the marketplace clone and the user settings aside: it installed the kit, wrote the thresholds, the hook, the exclude line and both npm cis, a nested claude -p started after it fired the kit's hooks, and a second run changed nothing. What else was measured there, since CLAUDE.md now carries only the instructions: claude plugin install before a marketplace add answers "not found in marketplace", and the add fails on HTTPS authentication until the private kit is attached to the session; the session that installed the kit then ran Bash for an hour without the kit's activity log ever being written and with none of its skills in its list, while claude plugin list showed it enabled; with the install record emptied and the cache moved aside, a nested start installed nothing; a nested claude -p started with every CLAUDE_KIT_* stripped from its environment read the project settings' env block and the user settings' alike; and a staged copy of the kit's work-rating/SKILL.md was only flagged "review only" by the guard, whose built-in names predate it. What has not been run is the one thing the script is for: executing AS a cloud environment's setup script, before Claude Code launches, in a session started with the kit attached. Whether its per-clone half survives into later sessions, which each start from a fresh clone, is unknown with it.


Third-party code

A fair question about AI-written code is whether any of it is somebody else's. What can be stated from the tree:

  • The dependency surface is deliberately tiny and fully declared — three direct Go modules (modernc.org/sqlite, golang.org/x/crypto, golang.org/x/time) in go.mod, and three runtime npm packages (react, react-dom, @chenglou/pretext) in web/frontend/package.json. Everything else in the binary is standard library.
  • What was borrowed is credited by name in the README's Attribution section — the pretext text-reflow library, the CC0 texture packs behind the paper/film skins, the metadata sources, and the apps whose export formats are read as import sources.
  • Eighteen type families ship in the build, all @fontsource packages and all OFL-1.1 — free to use, embed, modify and redistribute. Twelve of them arrived in 1.15.0 as the alternates behind Settings → Type. They are bundled rather than fetched, which is the same rule as everything else here: this app makes no network request the reader did not ask for, and a type picker that phoned a font CDN would have been the first exception, on a screen about how your own words look.
  • Design influences are named where they apply, in the code and in docs/wiki/Design-decisions.md — the Radarr-style status bar, the *arr-style cover folder — so an idea taken from elsewhere is attributed rather than passed off.

If you spot something in here that belongs to someone else and is not credited, that is a bug worth an issue.


Responsibility, and license

I am responsible for this code — for what it does, for its bugs, and for the decision to ship it. "An AI wrote it" is an explanation of method, never an excuse, and it does not transfer to whoever runs the software.

Tippani is MIT licensed (see LICENSE) and I hold the copyright, on the same terms as any other MIT project. If you find something wrong, open an issue — a bug report is as useful here as anywhere, and arguably more so.

Last verified against the tree at v1.4.2.