Skip to content

Releases: ifrit98/Explainer

v0.7.0 — Reviews that run themselves

Choose a tag to compare

@ifrit98 ifrit98 released this 04 Oct 10:25

Reviews that run themselves, and a record of which findings matter.

  • --run for the probe, the cold read, and the blind test. One fresh claude -p call reads only the explainer's folder (no settings, no MCP servers, no CLAUDE.md, no stdin) and the reply is saved as review/runs/<time>-<tool>[-<rendering>].json with numbered findings. A blind test is scored by a second fresh call against the rubric; --regrade scores saved answers again.
  • Adoption log. explainer decide records what the author did with each finding (review/decisions.yaml); explainer findings --stats reports adoption per tool and severity. First measurement: cold-read blocking findings 100% adopted, the probe's main gaps 83%, cold-read edge findings 3%, blind-test audit gaps 9%. The prompts changed in response (edge findings only where the reader would stop; at most five audit items per list), and so did the grader's rubric, which had failed correct answers for not repeating the rendering's own example.
  • Diagram budget: explainer check warns when the first diagram has more than 9 nodes (budget.diagram declares more, with a reason).
  • Review sheets take each bookmark frame after its animation has finished, so readers no longer report mid-animation overlaps.
  • New example: rates-inflation. How a higher policy rate lowers inflation, on a Stage 3 page with a toy model and a toggle for each assumption: anchored expectations, a supply shock, the disputed cost channel (the price puzzle), and the disputed neo-Fisherian view. Every claim is tagged established, estimate, disputed, or constructed. Its probe found a claim the toy contradicted ("anchoring makes disinflation cheaper") and its cold read found the toy's timing outside the real-world estimate; both fixed.
  • Verification gaps closed: claims for git-objects, cdn-request, and ste-80; second-round cold reads of the six renderings fixed in the previous round (git-bisect, git-objects, cdn-request, ste-80, attention, the sky-blue diagram); blind tests of the four backfilled ones, all passing after fixes. ste-80 went through five rounds: rule three is now "use concrete words" (it changed a noun too), the word count is honest, and each swap lands as it is spoken.
  • Since 0.6.0, every example has a narrative.md and a v0.6.0 cold read. The backfill for git-bisect, git-objects, cdn-request, and ste-80 found blocking problems in all four, two of them errors: the git-bisect remedy for a bug that comes and goes could not find the bug in its own example, and the CDN page said "faster" over bars that showed the CDN slower. Fixed, with records in each review/understanding.md. ste-80 re-rendered (16 → 15 → 13 → 9 words, with the count now honest).
  • GitHub page (since 0.6.0): the README and the landing page lead with first-principles answers and the chat eval, with an interactive page next to a video at the top; repository description and topics updated.

Upgrade: plugin users get v0.7.0 on the next /plugin update; the explainer wrapper now pins @v0.7.0. --run needs the claude CLI on PATH.

v0.6.0 — Rebalance: omission and excess both cost

Choose a tag to compare

@ifrit98 ifrit98 released this 04 Oct 07:23

Rebalance: the reader's effort has two sources, omission and excess, and the checks now push against both. From v0.3 to v0.5 every check caught omission and none caught excess, so every review round added words (principles 947 → 1,802 words, softmax prose 412 → 1,234, odd-squares video 59 s → 4 min 21 s).

  • Principles cut to about 1,340 words; the detail of §10 and §11 moved to references/completeness.md and references/narrative.md. New: effort has two sources; explain from first principles; calibrate to the reader (default: a technical reader new to the subject); "could anything be cut?" in the understanding test; tag only claims whose status a reader could mistake; §12, three tiers (answer, quick artifact, published), and "a chat answer is not a small artifact".
  • Where STE-80 gives way (writing.md): a term of art, a cause and its effect in one sentence, a labeled analogy, a formula, no definitions for what the reader knows. Plus "calibrate to the reader".
  • Cold read reads as the audience in model.md, reports excess (already known, repeated, a detour) and marks blocking findings. Pass rule: nothing blocks; fix edge findings only when short; cut excess; prefer fixes that replace words; stop when nothing blocks.
  • Probe and blind test read as the audience: the probe marks gaps main or edge and lists model entries the audience does not need; the blind-test audit gains excess.
  • Length budget in explainer check: a warning (!) above 600 words of prose or 150 s of video; budget in model.yaml declares another length with a reason and fails the check when exceeded.
  • Tiers in the explain skill (Step 0) and verify; explainer new --quick skips narrative.md. Templates drop "Stage 1: controlled prose" from visible titles and default the audience.
  • explainer eval chat: six questions about phenomena answered by fresh claude -p calls under several system prompts and graded blind. It found that the principles made chat answers 50% longer, and that v0.6's first changes did not help chat. A chat rule fixed it: the current principles now give the best-scored and shortest answers (7.8/10, 292 words, against 6.3 and 399 with no system prompt). docs/evals.md.
  • New examples: sky-blue (a physical phenomenon from first principles, prose and diagram, 650-word budget) and attention (a Stage 3 page with L1–L4 and a toggle that removes each part of attention).
  • Fixes found by the new cold read: the dijkstra finality argument named D and E as the only exits and said "reaching D costs at least ten" before D became eight through B (a flaw the v0.4.0 fix introduced); the argument now follows a path to its first node outside the settled set, in the model and the re-rendered video (203 s). The softmax prose leaned on a table that came after it and used undefined symbols; reordered and cut (924 → 830 words).

Upgrade: plugin users get v0.6.0 on the next /plugin update; the explainer wrapper now pins @v0.6.0.

v0.5.0 — the narrative pass

Choose a tag to compare

@ifrit98 ifrit98 released this 04 Oct 04:02

The narrative pass: an explanation is a model plus the path a first-time reader takes through it.

  • narrative.md (Pass 3, principles §11), scaffolded by explainer new: the reader before and after; the question and the result in words; a motive and a reason for the approach; an introduction ledger of every term, symbol, name, and visual convention; the beats in order, each with its bridge, what is shown, and what is said; concrete to symbol; links between representations; the close.
  • explainer coldread <slug> --rendering R: a fresh agent meets the narrative or one rendering for the first time and reports, in order, every reference it was not given, everything shown but unsaid, every leap, and whether the question, motive, and close are there. --rubric prints the pass rule. The verify skill runs it before the blind test.
  • Ledger check: with a narrative.md, explainer check fails a scene symbol (on-screen math, a math label, "the n-th" in speech) that the ledger does not introduce. check -v notes explainers with no claims or no narrative.
  • Layout: a crowded issue for text closer than 0.1 units to other text. stagger_labels spaces its rows by label height (it fixes two touching labels in the softmax video).
  • odd-squares, rebuilt from a narrative: the old video passed its blind test and failed its cold read (n never introduced, the result never said aloud, two formulas never read). The narrative's own cold read, before rendering, changed the argument: each L is two bigger than the one before, and the first L is one tile, so the Ls are the odd numbers in order. The Ls are now real L shapes, the next L is drawn, the full sum to 19 is shown in the L colors, and the video ends on the answer. Claims and terms added.
  • Docs: principles §11, a narrative step in the explain and video skills, a cold-read step in verify, the authoring guide's narrative section with the odd-squares case, concepts, CLI, and model reference.

Full history: CHANGELOG.md

v0.4.0 — review findings become rules

Choose a tag to compare

@ifrit98 ifrit98 released this 04 Oct 03:54

Review findings become rules. Each finding from the v0.3.0 blind tests is fixed in its example and turned into a check, a probe rule, or a template line that applies to every explanation.

  • Terms: terms in model.yaml lists the phrases a rendering must not use for a concept. explainer check fails on one. Probe rule 3 now asks for one name per concept.
  • Meaning of quantities: probe rule 6 asks what each number shown means for the reader. check -v lists each accept_unjustified function as a debt.
  • Pace: render and review report a claim with less than 1 s of pause after its line, or more than 10 s of narration over one picture after its mark.
  • Blind-test audit: the reviewer also lists concepts with two names, numbers without a meaning, steps without a why, and (video) rushed points. The rubric treats each as a gap.
  • Pages keep predict pauses: explainer check fails a page that plays a video with predict pauses through a plain <video>. The landing page now uses Explainer.video.
  • Mermaid: explainer check --diagrams renders every Mermaid block with the pinned Mermaid CLI (mmdc or npx). CI runs it.
  • Examples: dijkstra uses "estimate" for the tentative value and "distance" only for the final one; the finality argument is three lines, one step per picture, with pauses; relax is shown as a test. softmax-temperature replaces the entropy definition with a why-entropy claim (surprise in bits, why a log and not a count, live per-token surprise on the page), and pauses after why exp.
  • Second-round fixes from the audit: softmax "scaled gap" (a terms entry, which also caught the page); Dijkstra's finality argument now draws the settled region as A and C, exits through D and E, and says why reaching D costs at least its estimate; the narration says each relaxation's sum. Probe rule 2 checks each step of an argument as written. Review-sheet headings show event times.
  • Principles: §10 adds one name per concept, a meaning for each quantity, time for each claim, and "turn each review finding into a rule".

Full history: CHANGELOG.md

v0.3.0 — complete the chain: claims and the probe

Choose a tag to compare

@ifrit98 ifrit98 released this 04 Oct 03:54

Completeness: explanations must not omit the connections a learner needs.

  • Claims in model.yaml: why, guarantee, mechanism, definition, limit, misconception, each with required fields (a why needs the simplest alternative and its failure; a guarantee needs an example and a counterexample). Renderings mark where they cover each claim (<!-- claim: id -->, data-claim, self.claim()), and explainer check fails on incomplete or uncovered claims.
  • Function rule: a function named in model.md (exp, log, sqrt, sigmoid, …) with no why claim fails the check, unless accept_unjustified gives a reason.
  • explainer probe: a prompt for a fresh agent that reads only model.md and lists omissions before rendering. On the existing examples it found "why exp" and the Dijkstra negative-edge counterexample.
  • Model template: new sections Why this form, Concrete cases, Terms, Scope. Principle 10, "Complete the chain".
  • Blind test: claims with ask become quiz questions; the rubric expects their cases.
  • Examples: softmax-temperature now explains why exp (with the divide-by-sum and squares counterexamples, live on the page and in the video), why divide by T, shift versus scale, and a worked T = 2 step. dijkstra adds why the smallest estimate, finality with this run's numbers, path reconstruction, and a negative-edge chapter. git-bisect adds the halving trace and a case where bisect names the wrong commit.
  • Checks: scientific notation is one number (6.0 × 10⁻⁶); hash fragments like 4e10 stay identifiers; render reports repeated narration lines.
  • Docs: authoring guide, model reference, web toolkit reference, docs index, contributing guide.

Full history: CHANGELOG.md

v0.2.0 — plugin, model checks, review tooling

Choose a tag to compare

@ifrit98 ifrit98 released this 04 Oct 03:54
  • Claude Code plugin (explain, video, verify skills; bin/explainer) and marketplace.
  • model.yaml and explainer check; load_model and explainer sync.
  • Review tooling: timelines, layout check, explainer review, explainer quiz.
  • Stage 3 web toolkit and page template; predict-first gates, video predict pauses, MP4 chapters.
  • Video components, LaTeX via TinyTeX, sentence-timed captions, fail-fast scenes.
  • Examples: odd-squares, dijkstra, cdn-request. CI.

Full history: CHANGELOG.md