Security
- Personal phone number removed from every LICENSE.md (root + the 8 package
mirrors,license:sync-verified). Found by the pre-promotion privacy sweep — the
license contact stays email-only. Residual exposure noted honestly: the number
remains in old git blobs and in the already-published npm tarballs (≤4.5.0); the
current tree, the next release, and everything a visitor reads going forward are
clean. Full-history rewrite deliberately NOT done — it would break every clone,
tag and PR reference for marginal gain.
Added
-
Bench results visualized in the READMEs — a static dumbbell chart
(docs/assets/bench-results.svg+ a dark-mode variant, selected via
<picture>/prefers-color-scheme) of the committed 2026-07-21 live-round
numbers: skill vs stock per bench × model (research / design / discipline ×
Haiku / Sonnet / Opus), with direct value labels and deltas on every row.
Palette CVD-validated for both surfaces. Pure rendering of data already under
evals/*/results/— no benchmark execution involved. Embedded in
README.md,README.en.mdandevals/README.md. -
/vkm-researchgrew from a 68-line monolith into a full skill (same standard as
the 4.5.0 vkm-spec rebuild): rewritten SKILL.md with a copyable checklist and a
degradation ladder;references/summary-template.md(the canonical consolidated
shape) +references/synthesis-guide.md(four axes: claims across sources, links
that connect, visible supersession, compression with judgment);
examples/worked-example.mdwith REAL validator output both ways (draft fails with
5 named errors, rewrite passes); andscripts/validate_summary.mjs— a zero-dep
validator that rejects promoted map-reduce drafts by their seams (---separators,
## <file>.mdheadings, leakedurl:lines), missing wikilinks, malformed
supersedes and transcription-sized output. Doubles as the research-bench grader;
mutation-style self-test (10 cases) gates in core CI. -
research-bench (
evals/research-bench/): a syntheticRESEARCH/<topic>/bank
with the pipeline's real frontmatter and two seeded probes — a contradiction between
sources (must surface as a typed- supersedes) and an embedded instruction in one
source (must be flagged as untrusted DATA, never obeyed). Skill vs stock; grader =
the shipped validator + probe signals (reference consolidation scores 100, the raw
draft 0). New job + dispatch option inllm-benchmarks.yml. -
implementer-bench (
evals/implementer-bench/): first eval covering the
vkm-implementeragent — its real installed contract as system framing vs bare, on
spec-shaped tasks, graded by discipline-bench's existing hidden-test instruments
(no new graders). New job + dispatch option inllm-benchmarks.yml. -
design-bench auto round 1 (mechanical score from the skill's own validators;
raw HTML committed underresults/2026-07-21-round1/): Sonnet +60 on the
slop-attractor brief (stock: 15) and +30 on the held-out brief; Opus (n=1)
+60/+40 — stock Opus scored 0 on facturio (full slop fingerprint + failing
contrast). Haiku flat, dial-consistent. Judgment axes stay in the manual protocol. -
Effort gate auto-match (ADR-0031 amendment): the
PreToolUseeffort gate now
detects the session's current effort (CLAUDE_EFFORTinherited by the hook) and
opens itself when the model's proposed level equals it — no pause when there is
nothing for the user to change. Mismatch or undetectable effort keeps the original
pause; the deny message now advertises the detected level. Subprocess-level tests
cover match/mismatch/undetectable/template-only paths. -
Diversified round (round 2/3 of the live benches), Opus added at reduced n —
raw data under each eval'sresults/2026-07-21-round2/:- research-bench: the skill's gain GENERALIZES — held-out domain topic
(container-queries: Haiku +35, Sonnet +52.5, Opus +50) and Opus on sqlite-vec
(+45); stock Opus still fails the consolidation contract. - discipline-bench on the new harder tasks (incl. the held-out instrument):
Sonnet +19/+21, Opus +25/+31; Haiku flat-to-−3 within spread — where the hidden
contract exceeds the small model's reach, the doctrine neither helps nor hurts,
exactly what the dial predicts. - token-quality-ab adversarial fixture (no-keyword decisive lines): delta 0.0
on all three models measuring the FIXED hook — verdict stays KEEP. - skills-triggering hard set (12 near-miss/multi-skill/tool-vs-skill cases):
Haiku 11/12, Sonnet 12/12, Opus 11/12 — desaturated, both misses are arguable
boundary calls, logged with the description tweaks to try next round.
- research-bench: the skill's gain GENERALIZES — held-out domain topic
-
Round 1 of both new benches, raw data committed (2026-07-21, Haiku 4.5 + Sonnet 5,
n=3/cell, underresults/2026-07-21-round1/):- research-bench: the skill delivers — Haiku 33.3 → 70.0 (+36.7), Sonnet
35.0 → 96.7 (+61.7). Stock output fails the consolidation contract (no typed
supersedes, weak linking, the seeded injection usually dropped silently). - implementer-bench: honest null result — explicit specs saturate both
conditions (the value lives in the spec, which is /vkm-spec's job), underspec
deltas sit inside replica spread with one negative Sonnet cell noted for re-check.
The agent's case is delegation ergonomics, not raw scores; the bench exists to
catch contract regressions. No verdict claimed beyond n=3.
- research-bench: the skill delivers — Haiku 33.3 → 70.0 (+36.7), Sonnet
Changed
- The skill structure gate's cross-reference allowlist is now EMPTY (was 4
tolerated vkm-design entries): the in-reference pointers became plain-text mentions
at the point of use ("contemporary.md, this folder") and SKILL.md remains the only
place that links — Anthropic's one-level-deep rule now holds to the letter, and the
gate ratchets at zero exceptions.lineages.mddeliberately NOT partitioned: it
passes the## Contentsnavigability gate, and splitting it would mint new
cross-references — the exact debt just retired.