Skip to content

Releases: MichaelYcJo/SpecSeal

0.9.0 — a record is written by a machine and trusted like one, and now something reads it

Choose a tag to compare

@MichaelYcJo MichaelYcJo released this 07 Sep 13:33
22662f1

Four work items, six issues. One sentence in four sets of clothes: a record is written by a machine and then trusted like one, while nothing checks what it says.

What ships

A loaded file may no longer name a version at or above the running one (#179, #98). The check read one number — the version in plugin.json — so a version written ahead of its release was green every day until the day it shipped, and red on that release's own preparation commit, hours in, after the broad gate had already run. Alongside it, three places credited -z alone with turning git's escaping of non-ASCII paths off; re-measured on git 2.50.1 over four quoting variants plus a control, either argument does it by itself.

The round record carries the reviewer's paste-ready fix, and a | inside a cell no longer truncates the row (#187, #189). The review skill requires a paste-ready fix for every 🔴 and every 🟡 and spends four paragraphs on what makes one paste-ready. round_record.py new copied the report's four tables and dropped everything else — so not one of those blocks reached the file a fix pass is told to open instead of the report.

evidence-check reads what a record says about the tree (#190). A ledger row is a claim about the tree that something reads; a spec.md, a plan.md, a round record or a phase record states the same kind of thing and nothing read it. A record of a work item that has not shipped is now refused when it names a backticked compound identifier nothing outside the records carries, and a path#unit@hash stamp in a record is resolved by the ledger's own reader. The boundary is a file the release already removes — the work item's ledger fragment — so there is no new state to maintain and no list to keep.

round_record.py new says what bound the next round is under, as it writes the record (#207). The chain bounds a run one step earlier than the cap: after a record whose floor row reads no, at most one later record may close on a fix, and the record that reads its fixes ends the run whatever it finds. That was enforced only at the broad gate — after every round had already been spawned. #179's own chain overran it by three rounds, and the gate refused 37.9 minutes and 180 calls of agent work that were then reverted.

What the release measured about itself

The #190 · #207 run was capped, and the cap did its job. Three review rounds and two fix passes. Round 2 reopened; round 3 read the fixes and ended the run whatever it found. It found seven more things and not one of them blocks — they were filed rather than fixed, which is exactly what docs/review-chain-spec.md §The reopening exists to produce: #217 through #222, beside #215 and #216 for the two answered questions.

Twice, a fix closed the coordinate and left the class one step over. Round 1's 🔴 2 became round 2's 🟡 1 became round 3's 🟡 2 — the printed bound disagreeing with the gate it exists to predict, in three different places, each time inside the fix for the one before. Round 3 settled it by construction rather than by reading: a differential of bound_line against chain_check.stopping_floor over all 584 record sequences of length ≤ 3, which found exactly one disagreeing class, 16 sequences, all permissive.

The Windows leg had been red since the commit that added the records arm — through three review rounds and two broad gates. A coordinate the arm built printed with the platform separator, so the same file read seal\specs\… from a record and seal/specs/… from a ledger row naming it. Every round and every gate ran on macOS, where the fix is a no-op. That is agent-contract §13 in its plainest form — a defence resting on a platform guarantee nobody removed. The case passes ntpath rather than skipping off Windows.

One rider had to be re-stamped at the release, and that is the squash's own cost. The stamp named a feature-branch commit the squash into the release branch discarded, so the rider-reachability check went red at the release rather than on the branch that wrote it — the class docs/branch-and-release.md records for release → main, arriving one direction earlier.

Every segment was measured at its boundary and posted to the flow log: 413 tool calls over roughly 100 minutes across four measured segments. Two implementer segments at 1.00–1.01 tools per turn against two reviewer segments never flagged for batching; the reviewers cost less than half of either fix pass and found eleven things the fix passes had not. The log also carries an instrument correction: the output-token undercount is not a scale factor — 3.2× on one segment and 334× on another, the same day — so no published reading involving a reviewer's output tokens is usable until that ships.

The gate

./bin/test 2541 passed, 2 skipped · uvx ruff check . and uvx ruff format --check . clean over 110 files · ./bin/evidence-check . exit 0 with 764 ok · 0 drifted · 0 broken · 0 external · 0 old-format · gather_changelog.py --check and fold_ledger.py --check both exit 0, no ledger fragment left and no open evidence-todo row. Every feature pull request was green on all six legs at the commit that merged it, Windows included.

Full changelog: CHANGELOG.md · Compare: v0.8.3...v0.9.0

0.8.3 — three items, and the rounds that measured their own chains

Choose a tag to compare

@MichaelYcJo MichaelYcJo released this 06 Sep 21:36
3dccf04

Three items, out of eight the release was cut for. Stopping there was the
call rather than the failure: #175 alone ran four rounds and three of its
findings were blocking. The five that did not ship are carried into 0.9.0 with
what each of them had already learned.

All three are the same shape from different directions — a program that had
already decided the right answer and then said something else.

The evidence check printed the name of a file it had not read (#163)

Where a --ledger pattern crosses a symlink before a ..--ledger 'x/lnk/../ledger.md', with x/lnk pointing at y — the checker opened the
file the pattern actually names and printed a header naming a different one.
os.path.relpath folds .. the way normpath does, by rewriting the string
rather than by asking the filesystem, so it answered x/ledger.md: a real
file, usually holding different rows.

The exit code was right and the rows were right. What was wrong was the name
a person reads and then goes and opens.

A ledger path rendered for a person now goes through one helper that drops the
root's own leading segments and touches nothing else about the spelling. A
path that is not under the root prints in full — longer than before, and
naming the file that was read. That covers a local-mode root seen from a
linked worktree, where the ledger genuinely does sit outside the tree.

An unwritable virtualenv turned a refusal into a traceback (#177)

The ignore that keeps .venv out of git status is written from a finally
in ensure, which is what makes it an exit-level guarantee rather than a list
of remembered paths. It is also what puts that write on the two exits whose
entire product is a sentence.

write_text was unguarded. So on a .venv the process cannot write to,
bin/test's refusal reached stderr and a PermissionError followed it —
the answer, and then a traceback printed underneath the answer.

The write is guarded now and says what it could not do: it names the ignore,
the reason, and that the virtualenv stays visible to git status until either
that write succeeds or the reader removes the directory. The remedy names no
cause on purpose
— the guard fires on more than one, and a remedy that
guesses which is a remedy that is sometimes wrong.

Two odd rows still ended the cost report, and both were values (#175)

parse_time states this file's rule — one odd row must not end the report —
and count had applied it to whether a value is a number at all. Neither
reached a value of the type a field already carries:

  • a transcript mixing a zone-aware stamp with a naive one exited 1 with stdout
    empty, on the report and on --json alike;
  • a transcript whose only paired call begins and ends on one timestamp printed
    the span line and then lost the token block and the family table behind a
    ZeroDivisionError.

Both are closed at a funnel rather than at the sites that consumed them. A
stamp carrying no zone is read as UTC at parse_time — the assumption that
same line already made when it rewrote a trailing Z — and a share of a span
is taken through a new share, which prints a dash and one line saying why
when the span is not positive.

Two smaller things, both about a record that had stopped being true

A ledger row whose guarantee a change makes conditional gains the condition;
it is not removed.
Two rows were in that position and both are repaired in
place, because a row is removed when a change takes away the code it cites and
this one took nothing away. What was wrong was an unstated precondition rather
than the mechanism — and the mechanism is exactly what the guard was written to
keep.

A case that reads three named constants catches a rename and cannot see a
fourth constant added beside them.
That is three ways of catching an edit to
a value that exists and no way at all of catching one being added — the drift
where the document goes stale and the suite stays green. The set is now
derived from the function's own return statements, so a value it gains
has to reach the document before the suite is green again. Measured rather
than argued: with a sixth value added, the derived case exits 1 naming it and
the named case passes. Both are kept, because they fail on different edits.


Installed sessions are told about this release by hooks/version-check.py,
which reads git tags and nothing else. Run /specseal:update to take it, then
reload — a change to skills/, agents/ or templates/ binds no session
until the plugin cache refreshes.

0.8.2 — the ruler, and the three readings taken under it

Choose a tag to compare

@MichaelYcJo MichaelYcJo released this 06 Sep 09:13
39f5cad

The ruler first, then the three readings taken under it. docs/flow.md set
that order and said why: the first item is what measures the other three, so
anything built beside it is built without it.

A run's report carries one comparison table (#170)

session-cost prints a token line — turns, output, cache write, cache read
— summed over a run's main transcript and every segment beside it, in one
command:

tokens        18 transcripts, 1,617 turns
  output              792,908
  cache write       7,688,414
  cache read      480,245,207

Every way it degrades prints a smaller number rather than raising, so the
line also says how many transcripts it covered — that count is what a reader
holds it against.

skills/verify/SKILL.md now states the nine-row table a run's report carries
and where each row comes from. The two documents that carry it point at the
owner instead of copying the rows, because two copies are two things to keep in
step.

The repository ships a command for its own suite (#156)

bin/test builds a virtualenv on first use and reuses it: 5.24 s cold,
0.60 s warm.
The claim is not that the first call is cheap — it is that the
expensive call is only the first.

CONTRIBUTING.md names it first. The old uvx --with pytest form stays as a
labelled no-write fallback, carrying the 55–58 s per call that demoted it,
because a fallback named without its cost gets promoted back. agents/smith.md
and agents/warden.md now tell a spawned segment to look for a runner the
repository ships instead of assembling one — in language that names no
repository's own command, since those files ship to every install.

The flow log's roll fires when a version has shipped (#155)

It guessed the next version by bumping a minor. At 0.8.1 that closed a log
opened for 0.9.0 that nothing had used, and opened a second issue with the same
title.

Now a log is titled by the version it rolls fromchore: flow measurement — after 0.8.2 — which is a fact at the moment it is written rather than a
prediction, and a roll happens only when a new version has shipped. This
release is the migration
: the old-convention log read as naming no version,
rolled, and the first log under the new convention is open.

A fence may not close past a section the record needs (#169)

round_record.py new accepted a fenced block that closed after a later
heading, and wrote a record with that section silently gone. The rule is one
sentence — a fence must close, and its span must not cross a line the generator
reads out of the report — and the guard derives that line list from two
constants the module already used to declare what it looks up.

Upgrading

Nothing to do. bin/test is new and optional; every other change is in the
skills, agents and documents the plugin ships, and takes effect at your next
plugin update.

One thing to know if you write review records by hand: round_record.py
now refuses a report whose fenced block, or whose HTML comment, would take a
section out of the record. The refusal names the heading it would have lost and
what to do about it.

Known, and scheduled

Four issues came out of this release's own runs and are on 0.8.3 and 0.9.0:

  • #177 — an unwritable .venv turns bin/test's refusal sentence into a
    traceback.
  • #179 — a loaded file naming a real version is a timer: green today, red
    at the release that ships it.
  • #182 — the fence guard's enumeration names three copies where the
    property is every copy taken out of raw, and there are four.
  • #180 — three rules that were already written down were each re-broken
    inside one run. Written down and arriving at the moment of the act are
    different states, and a fourth copy of the rule is the failure mode rather
    than the fix.

Full changelog: https://github.com/MichaelYcJo/SpecSeal/blob/v0.8.2/CHANGELOG.md

v0.8.0 — the chain's own machinery

Choose a tag to compare

@MichaelYcJo MichaelYcJo released this 05 Sep 08:12

A review run now knows when to stop, and a check can no longer report clean while something is missing

Two failures shaped this release, and both were silent. A review run had a ceiling — three rounds, five while something blocking was open — and nothing underneath it, so the ceiling got spent like a budget: one run went seven rounds and the last three found nothing that mattered. And a ledger check narrowed to one work item's own file reported everything fine while the shared ledger rotted beneath it — three review rounds and two fix passes all said ok, and the unscoped read at the pull request found fifteen stale rows and one false claim, every one in a file that branch had touched.

A run now has a floor. Stop when a round finds nothing that leaves the root and nothing that crashes. Whatever else that round found is deferred to a named answerer or becomes an issue. The reviewer answers it in a line of its own, the round record carries a row for it, and chain_check.py reads that row at the pull request rather than leaving it to whoever is awake.

A fix pass may add a unit; that unit's fix may not. The rounds a floor removes are exactly the rounds that were reading what the previous fix pass had just created — measured across four rounds of one work item, three consecutive rounds found their finding inside the unit the previous fix added rather than in the fix itself. By construction the fix ships reviewed and the unit it added ships unreviewed, in one commit. The record now declares the depth of each new unit, and a unit added to answer a finding inside another new unit is refused, with the two places it goes instead named in the refusal.

A narrowed ledger check announces what it did not read. Guidance binds only a session that reads it, and the session this trap was sprung on had narrowed the command on its own initiative. So the tool says it now: a scoped run opens by naming every ledger it skipped, one per line, and how to read them. A run that narrowed to exactly what the defaults would have opened says nothing.

Two names for one file are one ledger, matched by inode rather than by a spelling of the path — so a case variant on a case-insensitive filesystem, a hard link and a symlink all count as read. Comparing paths had put a platform inside the answer.

A round record written after the fixes it commissioned no longer looks like one written before them. By the time a late record is committed its verdict cells read fixed at <sha>, which is exactly what a correct record looks like after its own update pass — so lateness left no trace, and the reviewer's drafted text died in a report while the next segment rebuilt it from scratch. The checker now refuses a record whose adding commit descends from a commit its own verdicts name as the fix.

What else moved

  • Every segment's record now says what ran it — the agent and the model, joined by a word rather than a punctuation mark. Two work items were metered segment by segment before this, and not one of those readings can be attributed afterwards, because the model lived only in a session transcript.
  • A measurement meant to span versions was being written to the issue the next release deletes. The rolling log and the durable ledger are now told apart, the durable one found by a label rather than hardcoded. And a rolling log used to be born empty; it now opens carrying what it rolls from, the version it closes on, and where the cross-version readings go.
  • Older work items are not made red. Every new rule is keyed to the id of the work item that wrote it: a record from before the cutoff prints instead of failing, because a merged record has no honest repair. A row that is present and malformed is refused at any age, since formatting is always the author's.

Where the cost went, honestly

The last work item of this release took fifteen review rounds where three is the rule. Two of them reviewed the feature. The other thirteen were the tool reviewing its own fixes and its own paperwork, and half of every finding was located in a record rather than in code. That is measured in #161, which is the whole of 0.8.1 — alone in its release, so that everything after it is the first work run under a chain that stops.

Full changelog: v0.6.0...v0.8.0

v0.7.0 — a phase hands the next one a record

Choose a tag to compare

@MichaelYcJo MichaelYcJo released this 05 Sep 08:11

What a build phase learned now survives the phase, and every segment records what it was asked to do

Half of the work in a run is what a segment discovered while doing it. Until this release that half reached the next segment only if the orchestrating session retyped it into the next prompt, and when it didn't, nothing recorded that anything had gone missing. It happened: one phase moved a rule out of an agent's definition into a temporary home, and the next phase deleted that temporary home before the rule had reached anywhere permanent. The rule left the repository and no record said so.

A build phase now leaves a committed record, the same way a review round already does. A new phase template mirrors the round's shape for the build side — what the phase was asked, what it found, and a table of what left the tree and where it must land. It is wired into the plan template, the implementation skill and the implementing agent, so a spawned session writes the record at each phase's close without being told twice.

And both records now say what the segment was asked to do, not only what it found. The cheapest review round ever measured here — under eight minutes, one blocking finding and four smaller ones — was cheap because its prompt named eight specific things to try to break, in order. That fact was recoverable only from a session transcript. Both templates now carry a section for it, filled in right after the segment posts.

Enforcement here is a template blank and a skill instruction, not a checker. A refusal for a missing section was considered and set aside: it would need a red test, a stated failure direction and a cutoff to measure against, for a mechanism that has shipped zero records so far. It is revisitable once real phase records exist to learn from.

What else moved

  • Measuring a segment and logging what it found is now automatic in every repository this plugin installs into. The instruction used to live in one operator's own memory file, naming a fixed issue number that went stale twice. It is now part of the verification skill: after every segment, find this repository's open measurement log and post the numbers. Where no such log is open — nearly every installed repository — the step does nothing, asks nothing, and measures nothing.
  • The log rolls itself forward. A new script, wired into the release workflow, closes the current log and opens the next one every time a release reaches the default branch, so it keeps growing without anyone opening an issue by hand.
  • The rollover retries once on "no log open", and never on "more than one". A search-index lag right after a label write can only undercount what is actually open, never overcount, so only a zero reading gets a second look. Found while building this same change, when a listing came back empty immediately after a label was applied and a direct read in the same breath showed it already there.

Full changelog: v0.6.0...v0.7.0

v0.6.0 — the agent contract

Choose a tag to compare

@MichaelYcJo MichaelYcJo released this 03 Sep 14:21

Every agent now reads the same rules from one file, instead of an orchestrator retyping them into every prompt

Half of every spawn prompt sent to smith, warden and scribe used to be identical, hand-typed from memory each time — how to read an exit code, what an agent must not run, how a probe is written. A rule that lived only in a human's memory of what to type was a rule that could go missing with nothing recording that it had, and it happened: a review round ran without a rule that had already been established two rounds earlier, and a reviewer's own forbidden full-suite run once produced a review's best finding, because nobody was assigned to run it.

agent-contract is a new skill carrying sixteen numbered rules that apply to every agent this plugin spawns. It is injected automatically — listed in each agent's definition, it arrives before the agent's first action, with nothing typed and no file to open. It appears in the skill listing because that is the mechanism the harness uses to deliver it, but it is not something to run yourself.

Each agent keeps only what is genuinely its own. warden keeps how it works in isolation and its report format; smith keeps its build procedure and mutation-testing discipline; scribe keeps its fact-finding rules. The session driving the work is bound by the same contract too — the rule that went missing before was broken by the orchestrator itself, not by an agent, so binding only the agents would have missed the party that needed it most.

Where this feature lives was decided by testing it, not by guessing. An earlier plan put it in a plain documentation file; six small experiments (in docs/experiments/) showed that location silently fails, and found the one that actually works.

What else moved

  • Three GitHub milestones were reshaped so each carries one theme: 0.7.0 is now about giving agents a way to hand findings to the next step of their own work (a record of what a segment learned, and the measurement that watches it); 0.8.0 is the review-process bounds those measurements make possible; 0.9.0 adds two new agents.
  • Two smaller issues were opened from what this work surfaced: whether the shared rule-file should split further once more agents exist, and whether a repository can tell "never set up" apart from "was set up and lost its records."

Full changelog: v0.5.0...v0.6.0