Skip to content

Releases: gitayg/productizer

v4.60.0 — CI drafts the release a person publishes, and the dashboard stops promising a page a tag never made

Choose a tag to compare

@github-actions github-actions released this 15 Sep 07:53

This page covers v4.58.0, v4.59.0 and v4.60.0. None of them had a release
page: v4.58.0 was never tagged because its CI run failed, and v4.59.0 was
tagged before anything here could create a page.

Security checks that could not fail now can

  • The shipped sast checks could never fail. The checks.yaml template
    every scaffolded repository receives ran semgrep without --error, and
    semgrep exits 0 even when it finds something — measured: 4 findings, exit 0.
    So sast, sast-auth and secure-coding-controls passed in every repo that
    used them. Their rule count also called a semgrep option that does not exist,
    so it read 0 on every run. Both are fixed in the template. If you scaffolded
    from an earlier version, run /productizer:upgrade to see the difference.
  • secret-scan runs gitleaks for real. gitleaks 8.30.1, pinned, with its
    download checked against a SHA-256 before extraction. On six planted
    credentials in real issuer formats it caught 6. Running the actual binary
    found four ways a scan can pass without scanning, each now guarded: an
    unreadable file is skipped with exit 0, so coverage is read from gitleaks' own
    log rather than its exit code; and the repository being scanned can switch
    the scan off with .gitleaksignore, a root .gitleaks.toml, or an empty
    config, each of which now fails the check instead of passing it.

/productizer:upgrade can only report

It promised to change nothing, but it granted unrestricted shell access, so
that promise rested on the model behaving. The model can no longer invoke it on
its own, file writes and edits are denied, and shell access is limited to
re-running its own report. Measured: five calls outside that scope were denied,
and the same five succeeded with the scope widened back. It also passes the
report's exit code to the model, which it silently dropped before.

A tag push drafts the release, and a person publishes it

Before v4.60.0, pushing a version tag published nothing: 22 of 81 version tags
had a release page, and no workflow had ever created one. Now a v* tag push
runs the checks on that tag first, and only then creates the GitHub release as
a draft. The job reads the release back and fails if it is anything but a
draft, never edits an existing release, and is the only job allowed to write.
This page is the first draft it produced.

The architecture view shows what a change did to the spec

The Visualizer compares the spec against the commit a change was measured
from, and marks each requirement NEW, CHANGED, SUPERSEDED, WITHDRAWN or
REMOVED. It does not claim a requirement's coverage changed, because that needs
a check run at the earlier commit too. Checks that take no file list now read
n/a instead of ?, which wrongly implied their count could not be read.

Also

  • contradiction-check.py reports "not measured" and exits 4 when a file holds
    no requirement it can read. It previously reported "0 requirements, 0 pairs"
    and exited 0 on this repository's own spec.
  • /productizer:help lists commands as well as skills.
  • build-release-notes.sh has a self-test covering its exit codes. Writing it
    fixed paths that printed "no requirement changed" or "none found" before
    failing.
  • CI actions moved to Node 24 releases, each pinned by a verified commit.

What this release does not do

  • It does not backfill the 59 older tags that have no release page.
  • Only sast and secret-scan run here. This repository declares the
    other shipped security checks but leaves them disabled or untriggered, each
    with a written reason: dependency-audit has no package manifest to scan, two
    checks need tags nothing passes yet, and one depends on a script that does not
    exist.
  • secret-scan does not scan git history, only the working tree.
  • The upgrade command's permission limits were measured in headless runs.
    In an interactive session an out-of-scope call should ask you rather than be
    denied; that was not observed.
  • The generated notes' pull-request section lists the last 50 merged pull
    requests rather than those in the release's range.

Evidence: commits 23f9402 (v4.58.0), fa8dace (v4.59.0), 5117f48 (v4.60.0).

v4.7.3 — R1's acceptance row says nothing asserts it

Choose a tag to compare

@gitayg gitayg released this 29 Aug 22:13

The first answer from walking the eleven requirements that have no acceptance
row. R1 — one living spec per product — is asserted by nothing: no declared
check, no script that refuses on an ambiguous or unreachable spec home, and
none of validate-spec.py's 53 diagnostic codes. product.spec_home is declared
in config.json and never read back, which makes R1 the requirement most exposed
to the thing it exists to prevent — two allocators both handing out R42.

The row records that, rather than leaving the requirement absent from the table
and indistinguishable from one nobody has looked at yet. The work to fix it is
B14; the row points there.

Recording it exposed a defect in the view, not fixed here: "Requirements with
no test" counts requirements with a ROW, not requirements with a VERIFIER, so
writing an honest "nothing asserts this" moved the count 11 to 10 while nothing
became verified. Telling the truth made the dashboard look better, which is
backwards, and the tile should count what is asserted rather than what is
written down.

v4.7.2 — the hygiene check was case-sensitive

Choose a tag to compare

@gitayg gitayg released this 29 Aug 19:22

The pattern list is lower case. The check was not case-folding. So a
capitalised spelling of a forbidden name walked straight past a rule that
names it — and did, into the previous commit and onto the public remote.

Verified before fixing: the shipped pattern matched 1 of 3 test lines
case-sensitively and 2 of 3 with -i. The names in question are written
CamelCase everywhere they actually appear, so the rule was close to inert for
its main purpose.

Fixed with -i, three names added, and the offending text removed from the
backlog. The first version of this fix wrote the two spellings into a comment
as an illustration, and the file promptly flagged itself — this file is checked
like every other, which is the property that caught it. The comment now says so
instead of demonstrating it.

Not fixed here: the previous commit's message and its release note still carry
the names. A commit message cannot be scrubbed without rewriting published
history, which is not something to do unasked.

v4.7.1 — B12 was the wrong shape

Choose a tag to compare

@gitayg gitayg released this 29 Aug 19:18

B12 said "add a SessionStart hook". Two things were wrong with that.

The model it pointed at was the wrong one. That store serves its memory alphabetically because it keeps no timestamps — no ranking, no recency, no selection at all.

And the gap was misdiagnosed as retrieval. Not everything learned is a requirement: "the build breaks unless you run X first" is nobody's obligation, and forcing it into spec.md corrupts the one thing that file is for. There is nowhere for it to live. That is a storage gap.

So: a committed store in the repo, not a per-user memory — every objection to memory (per-user, uncommitted, dies on clone, disagrees silently) is an objection to it being outside the repo, and putting it inside removes all four.

What that then requires is the part worth writing down. A store that can contradict the spec without going through intake is a side channel into the spec. So it sits below the spec and cannot outrank it: a learning informs, never obligates; one that contradicts an active requirement is a finding routed through intake, the same road Stage 9 uses for code-versus-spec drift; and one that turns out to be a real obligation graduates to an R id and stops being a learning.

Two rules, both about not trusting yourself: unverified until independently corroborated — the same dispatch can never confirm its own learning — and dormancy that stops serving without deleting, with a missing timestamp never read as infinitely old.

One thing Productizer can do that a free-text store cannot: a learning may cite R14, and because ids are permanent, "this learning is about a requirement that has since been superseded" is mechanically detectable by the check drift-reverse.sh already runs over code.

v4.7.0 — thirteen pending items opened

Choose a tag to compare

@gitayg gitayg released this 29 Aug 17:51

B11-B23 record what this session found and did not fix: the agent halts on a
contradiction without asking for the ruling; there is no cross-session memory
at all; a check has no rendering after a human overrides it; nothing asserts
one living spec per product; eleven requirements have no acceptance row; the
runner format ships with no executor; the measurement suite has never met a
real model. B9 is closed - --strict is clean - and B10 carries the graded
result.

Making a never-triggered check amber, which shipped in 4.5.0, was most of a
good idea and one bad one. A check scoped always that did not run is broken.
A check scoped by paths that did not match is not applicable to this change,
and counting it turned the tile amber on every commit that touched no shell
script - which is how an amber signal stops meaning anything, three commits
after being built to mean something.

run-checks.sh now records trigger_scope, and the view reads it. Found by
falsifying rather than by looking: the first attempt failed because the view's
check-row normaliser built a fixed dict and dropped the new key, so every check
read as scoped and the real repo happened to render correctly for the wrong
reason. Two arithmetic errors in the same neighbourhood were caught the same
way - "3 check(s), all passing · 1 not applicable" of three checks is two
claims that cannot both be true, and the partial denominator counted a check
that never ran among those that had.

Four states, falsified: scoped miss alone is green and says what did not apply;
an always check that did not run is amber; both together are amber and count
each separately; a real failure is red.

v4.6.0 — if it needs a person, it is a link

Choose a tag to compare

@gitayg gitayg released this 29 Aug 13:15

Anything drawn at att or warn is now an anchor to where it can be acted on:
tiles to their banner or their tab, board cards to the row or check behind
them, a failed Stage 5 row to the checks banner.

Calm things stay plain, deliberately. Making everything clickable destroys the
signal the loud things carry, so a tile at level '' renders as a div and a
linked one renders as an — the level decides, not the author.

Cross-panel links open the target's panel before jumping. A fragment that
changes the hash without switching the tab lands the reader on a hidden
section, which is a worse dead end than no link: proven by removing the handler
and watching the hash move to #stage5 while the panel stayed on p-dash with the
target at display:none, height 0.

Two things deliberately left unlinked, both recorded in views.md. A blocked
backlog row is already where it is acted on, so a link there would be a 2-cycle
rather than a route; the link runs board card to row instead. And a check that
passed while covering less than it declared has no banner and nothing on the
page that says more than its own row does — a link landing somewhere silent
about what was clicked is the dead end this change removes.

The banners block now builds before the stat tiles, so a tile can read the id
of the banner that was actually emitted rather than guessing it. Proven a pure
move first: byte-identical output on four page shapes before any substantive
edit.

v4.5.0 — a red tile that said PASS, and a check nobody asked

Choose a tag to compare

@gitayg gitayg released this 29 Aug 12:46

The Checks tile rendered class att — red — with the headline word PASS. Both
halves were separately true: PASS was the run verdict, and the red came from
bad_checks being non-empty. Together they were unreadable, and the reason was
that bad_checks counted not_triggered alongside fail and hollow.

A check that was never triggered and a check that failed are different facts.
Now: green PASS when everything ran and passed, amber PARTIAL when something
never ran ("2 of 3 ran and passed · 1 never triggered, so it was not checked"),
red FAIL for a real failure. hollow stays red — a check that exits clean
having examined less than it declared is a failure, not an absence.

The check in question was solver-corpus, triggered on
paths: ["**/contradiction-check.py"], so it only ran when the solver's own
source changed and reported not_triggered on every other commit. Its recorded
state was exit_code: None, tool: None, files_in_scope: 0 — it had no
exit code key at all, because it was never started.

That is the check whose why says precision must stay at 1.00. The property
most worth knowing on every run was the one measured least often. It is a 0.3s
stdlib self-test with no external dependency, so it now runs always. The gate
went from "3 declared, 2 triggered" to "3 declared, 3 triggered", and
solver-corpus reports PASS with 7 true positives covered.

Also adds a refresh button to the header. It copies a prompt, not a reload: the
page is a snapshot, and re-serving identical bytes teaches the reader that
nothing changed when nothing was re-measured. The prompt carries the exact
command rebuilt from the argv this run was given — so a non-default --out
survives — and the one sentence needed to not misread the result.

v4.4.0 — finding a leak was creating one

Choose a tag to compare

@gitayg gitayg released this 29 Aug 12:08

check-hygiene.sh printed the offending line. run-checks.sh stores each
check's output in checks-result.json, which is committed. So a detected leak
was written verbatim into the repo, and the next run found its own report and
recorded that too — a loop that fed itself.

It now reports by location and pattern class and never quotes the match: the
same rule the delegated review agents already follow for injected text, which
had not been applied to the check that needed it most. It still detects; it
just names the file and line and lets the reader open it.

evals/results/ is gitignored. Run artifacts are regenerated per run and carry
the absolute paths of whoever ran them — the leak that shipped in v4.2.0, in a
second costume. The corpus stays tracked; the runs do not.

The measured result they produced is recorded instead, on B10: end to end the
classifier catches 16 of 16 on the 26-case corpus at 0.94 precision, against 13
of 16 for the same model with no plugin. The three it uniquely catches are
domain entailment, an unquantified adjective against a numeric bound, and
vocabulary drift — the semantic tail the deterministic checker cannot reach.
n is small, the corpus is ours, and the harness is early-access.

The view:

  • "N things need you" links to the banners it is counting.
  • The acceptance-row banner's copy-prompt used to tell an agent to add the
    missing rows, which it would have done by inventing test names — turning
    "unknown" into a confident "yes" in the table whose whole job is recording
    whether a test really asserts a requirement. It now interrogates the
    maintainer instead, writes a row only from an answer, records "nothing
    asserts this yet" as a real answer, and names the ids left unanswered.
  • Opt-in --stale-after: the page notices it is old and prints the command to
    regenerate itself, rebuilt from real argv. It does not reload — re-serving
    identical bytes teaches the reader that nothing changed when nothing was
    re-measured. It says the page is old, never that it is wrong: an old page
    over an untouched repo is exactly right. Off by default, because embedding a
    generation time trades away byte-identical output.

v4.3.0 — the hygiene check was excluded from its own gate

Choose a tag to compare

@gitayg gitayg released this 29 Aug 09:32

A home path shipped to a public repo, and the check that exists to stop
exactly that was passed a file list with the offending file removed from it.

.claude/productizer/checks-result.json is committed, and the runner wrote
its own absolute config and root into it. Two occurrences of
/Users/<name> went out in v4.2.0. check-hygiene.sh detects them
correctly; the pre-release gate excluded that path from the manifest to stop
the run churning its own output, and so excluded the only file with a leak.
"PASS hygiene 245/245" was true and worthless.

The runner now emits config repo-relative and root as .. The absolute
values stay in memory for path resolution and never reach disk. Nothing
outside the runner read either field.

Also in this release:

  • Stage 5's EARS rules no longer judge superseded requirements. Two of this
    skill's own rules were in direct collision: a superseded requirement keeps
    its original text verbatim (an ERROR to edit), while the EARS check ran on
    every requirement regardless of status. The warning could not be cleared
    without breaking the retention rule. Found by splitting R14/R16/R21, which
    could never have silenced it. --strict is now clean on this repo's spec.
  • R14, R16 and R21 each stated two obligations under one id. Superseded, text
    retained, replaced by R23-R28, one obligation each. No id reused, none
    renumbered; verified against a baseline.
  • CI, which this repo has never had: claude plugin validate, the Stage 5
    gate, and both self-tests, on every PR and every push to main. Fails closed,
    least privilege, and no github.event.* interpolated into a run: block.
  • The view gains a Board tab; the kanban moves off the Dashboard. Setup
    becomes a checklist that appears only while something needs a person, and
    says plainly which items will never tick and why.
  • The untagged-versions banner said "so CI builds the releases" in a repo with
    no CI. It now reads the workflow triggers and says what is actually true.
    n/a and "not run" had been rendering identically; they no longer do.

v4.2.0 — the spec diff reaches Build

Choose a tag to compare

@gitayg gitayg released this 29 Aug 08:47

Eleven new scripts, seven new references, and four path bugs that all
produced the same failure: a confident false negative.

Stage 3 now receives the spec diff, not just the spec. A superseded
requirement keeps its original text, so a removal was invisible to a run
reading only the current file — the dropped behaviour survived in the code.

Stage 5 derives its coverage denominator from the spec rather than from the
check's own declaration, so a check can no longer shrink its own scope and
pass. Team-level settings are honoured only from the committed config; a
local override of one is ignored with a warning, and disabling every check
is a load error rather than a clean run over nothing.

New:

  • spec-diff.sh the delta handed to Build, with the no-baseline cases kept apart
  • drift-reverse.sh code citing a requirement that no longer stands
  • validate-spec.py EARS and id permanence, executable: 61 codes, ERROR/WARN
  • signals.sh/score.sh evidence and judgment split, judgment keyed to an evidence hash
  • req-trailer.sh Productizer-Req: provenance, orphans, COV_ coverage
  • init.sh all of Stage 0 in one command, with an interview fallback
  • graduate.sh repeated corrections into durable guidance
  • retrieval-budget.sh / ab-harness.sh / stage-snapshot.sh the measurement suite
  • usage-audit.sh our own scripts' error rate and option contract

Onboarding no longer dead-ends. The survey forks to a weaker evidence tier
instead of refusing, and gained probes for CLI surface, public API, config
keys, CI jobs and skill inventories. Two of four real repositories tested
could not be imported before this, including this one.

Fixed, all four the same class — a relative path anchored to the wrong
directory, reported as an absence:

  • spec-diff.sh reported a tracked spec "absent" from any subdirectory
  • run-checks.sh resolved ROOT, the default config and --changed against
    three different wrong places; it also wrote a nested shadow .claude tree
  • stage-status.sh counted with grep -c || echo 0, which emits two zeros
  • import-survey.sh --help printed bash's cd builtin help

Docs corrected against the code: GUIDE.md and README.md still named
.claude/sdlc/ and the lettered stages; SKILL.md claimed six stages while
defining nine, and stage 5 had no heading at all. The gitignore rule in
SKILL.md took its verdict from git check-ignore -v, which exits 0 on a
negation and so refused exactly the repos that had already fixed theirs.

Measured, and it is worse than we said: the symbolic cross-check scores
1.00/0.70 on the 21 cases we wrote first and 0.67/0.12 on a harder 26-case
corpus. The figure tracked the corpus, not the checker. B3 records both.