Skip to content

Releases: aarontzeng/dev-lead

dev-lead v0.6.14

Choose a tag to compare

@github-actions github-actions released this 16 Sep 23:10

0.6.14 — a flaky link turns a leg into pure burn, and that is not a calibration row

An attempt to calibrate glm-5.3 spent 35 minutes reaching step 8 with zero
completed tool calls and 159 bytes of output, against three connect
failures on a host whose links to opencode.ai and cursor were both
unstable that day. Each retry re-sent the same ~130 K input, so a quarter
of the model's five-hour bucket went on nothing before the operator
killed it.

Recorded as NOT MEASURED rather than as a slow or expensive result. A
number produced by a degraded network says nothing about the model and
would poison the table it sat in -- and the $0.25 is a network artefact,
not a per-leg price. The operational half is the transferable part: a
leg's cost is not bounded by one pass, so watch the step counter rather
than the clock and kill a leg that is not advancing. A run that has read
nothing by step 8 is not slow, it is looping.

dev-lead v0.6.13

Choose a tag to compare

@github-actions github-actions released this 16 Sep 23:10

0.6.13 — the Go bucket split is per model, not per family

0.6.6 of the runtime file put the whole glm-5.x family on the $60 monthly
bucket. It is per model: glm-5.3 sits on the $15 tier with kimi-k3 and
qwen3.8-max, while glm-5.1, glm-5.2 and glm-5.3-flash have the larger
one. The consequence is not cosmetic -- it turns glm-5.3 from roughly
sixty legs a window into roughly ten, which is the difference between a
model you can lean on and one you ration.

A peer caught it from the dashboard and the reading checks out
arithmetically: $0.2454 spent showing as 8.2 % is only consistent with a
$3.00 five-hour cap, since against $12 the same spend reads 2.0 %. Noted
in the file as the general trap -- sibling version numbers are not a
bucket, so read the model's own row rather than inferring from the family
name, which is how this got shipped wrong in the first place.

dev-lead v0.6.12

Choose a tag to compare

@github-actions github-actions released this 16 Sep 23:10

0.6.12 — criteria for what may go to a contributor tier, checkable rather than felt

The owner delegated the disclosure judgement to this layer, so it needs
rules that can be checked. Three: decide per repository once and record
it, because a repo's class is stable while a change's is not knowable at
a glance and a judgement re-made every round decays; a stop list naming
credential material, third-party confidential material, personal data
beyond work identifiers, and unannounced-product specifications; and a
screen run against the frozen target before the first dispatch rather
than from memory, because the tree is what the leg actually receives.

Run on quanta-mcp-gateway the screen found no credential material, 153
files and 6.7 MB handed to each leg, and colleague work addresses and
internal project codenames present. That residual is named here rather
than assumed absent -- it is what the owner accepted when he accepted the
contributor tiers.

The half a repository check cannot see is recorded with it: the claims
file and the lens text are written by the lead and travel with the tree,
so they can disclose material the repo does not contain. Screen what you
write by the same rules. When anything trips, the fourth leg moves to
glm-5.3, and three legs with a stated reason beat four with a disclosure
nobody decided.

dev-lead v0.6.11

Choose a tag to compare

@github-actions github-actions released this 16 Sep 23:10

0.6.11 — the fourth-leg rule is the owner's now, and the terms question is answered

0.6.10 carried the lens-based leg selection as guidance with the terms
question left open, because a plugin should not quietly decide what may be
disclosed. Both were settled the same day: the owner adopted the rule into
their CLAUDE.md, which is the roster, and ruled on the terms.

Contributor tiers put the prompt into a training set and a review leg
receives the whole frozen tree. Accepted for review legs, with the
exception carried per change: material that must not reach a training set
means the fourth leg runs on glm-5.3 instead. That keeps the judgement
with whoever knows what is in the diff, rather than in a default nobody
revisits.

The division of labour is now explicit in both directions. The roster
names which leg; this repo holds the prices, the buckets and the
calibration rows, and the owner's file points here for them.

dev-lead v0.6.10

Choose a tag to compare

@github-actions github-actions released this 16 Sep 23:10

0.6.10 — releases do not reach readers until the cache moves; pick the leg by lens

The day this repo shipped four releases, the maintainer's own session was
reading cache 0.6.5 and a peer was on 0.6.4 while master stood at 0.6.9.
The opencode-go guidance written that morning had eight occurrences in the
checkout and four in the cache both sessions actually read, so half a
day's work was invisible to everyone including its author. The scripts
were byte-identical across the gap, so nothing ran stale -- luck, not
design. Tooling should resolve from the git checkout, which moves with
the release, and a release note is a claim about what readers will see
that is wrong by default.

Second, leg selection guidance from a peer, recorded as guidance and not
as a default: mechanical lenses to muse-spark, free pool first; judgement
lenses to glm-5.3, explicitly unproven until it has a calibration row;
kimi-k3 and qwen3.8-max only for top-tier judgement at one or two legs a
window, probed immediately before dispatch because the buckets are
shared. Effort follows the lens, since reasoning bills as output and
xhigh costs nothing on a cheap model and doubles a dear one.

The decision that actually matters is left to the owner and marked as
open, because it is about terms rather than cost: both muse-spark
contributor variants put the prompt into a training set, and a review leg
receives the whole frozen tree -- source, design documents, sometimes a
spec. If that is unacceptable for this material then muse-spark leaves
the roster and glm-5.3 is the cheapest independent family. A plugin must
not settle that quietly.

dev-lead v0.6.9

Choose a tag to compare

@github-actions github-actions released this 16 Sep 23:10

0.6.9 — the journal says where its rows are, because their absence misled twice

The calibration journal holds no rows by design -- it is the method, and
the tables live in each family's runtime reference file, which one
sentence in the format section already said. Nothing pointed at the
actual files, so grepping this document for a model and finding nothing
looked like evidence that the model was unmeasured.

That misread happened twice on 2026-09-16. A session searched here for
opencode-go and reported the paid pool uncalibrated, when the kimi-k3 row
was sitting in the opencode runtime file. And the maintainer made the
same assumption in the other direction, writing a release edit that
targeted this file's non-existent table -- the assert fired, the edit did
nothing, and 0.6.6 shipped claiming a row it had not added.

So the format section now carries an index of the five tables and says
plainly that an absence here is not a measurement. Same class as the
gateway's own lesson about signals that cannot distinguish two cases:
this document could not distinguish "not measured" from "measured
elsewhere", and both readers took the first.

dev-lead v0.6.8

Choose a tag to compare

@github-actions github-actions released this 16 Sep 23:10

0.6.8 — the go buckets are shared across sessions, and a roster default is not mine to set

Two sessions on two machines measured the same thing hours apart: this
host spent kimi-k3's 5-hour window on one review leg, and a session
elsewhere then got Insufficient balance from glm-5.3 and minimax-m3 while
the free pool still answered. Same key, one pool. Neither observation
proves it alone -- "my machine ran out" reads the same either way -- and
only putting them side by side shows the buckets are shared. So a go leg
must be probed immediately before dispatch, because another session may
have just emptied the window; a multi-leg round can starve its own later
legs; and Insufficient balance decodes as this model's 5-hour window
being full, possibly by somebody else, rather than the model being down
or the month being spent.

0.6.6 of the runtime file wrote "default to opencode-go/glm-5.3". That
took a peer's recommendation and shipped it as policy. The roster is the
owner's and lives in their CLAUDE.md, changing only on their word for a
given round; the peer who proposed it has since said the same. The
passage is now cost information, with the reliability argument recorded
beside it: a paid leg that another session can drain is a worse default
than a free one that stayed answerable throughout, while the capability
half has no calibration data yet.

Methodology gains the shape behind 0.6.6's false commit message: a guard
fired, the edit did not apply, and git add -A staged whatever was there
-- severing "what I changed" from "what I verified" exactly when an edit
silently did nothing. A commit message that asserts an outcome is a claim
and belongs to the verify-before-you-say rule.

dev-lead v0.6.7

Choose a tag to compare

@github-actions github-actions released this 16 Sep 23:10

0.6.7 — correct 0.6.6's false claim, and family is the model you dispatched

0.6.6's commit message ended by saying kimi-k3 "earns its first journal
row". It did not. The edit that would have added it asserted on a table
shape that does not exist -- rows live in each family's runtime file, not
in calibration-journal.md, which is the methodology -- the assert fired,
and the release was committed with git add -A anyway, so a true message
shipped with a false paragraph on the end. Recorded here rather than
fixed quietly: a commit message is a claim, and this one was wrong. The
row is added now, in the family table where it belongs.

The substance: family is the model you actually dispatched this round --
not the provider path (0.6.5) and not the adapter's catalogue either. The
same fact about multi-family adapters supports two inferences and only
one is sound. cursor alone can serve xAI, OpenAI, Anthropic, Google,
Moonshot and Composer, so scoring by what an adapter COULD serve marks
nearly every candidate a clash and collapses the rule into refusing
everything. A peer read a catalogue overlap that way and would have ruled
out a leg that was in fact cross-family under the running roster. The
wrong reading is the tempting one, because a rule that refuses more feels
safer -- which is why both directions are now written down together.

Corollary: "the cursor leg" and "the grok leg" are shorthand that stops
being true the moment the model changes. Account the round's model id.

dev-lead v0.6.6

Choose a tag to compare

@github-actions github-actions released this 16 Sep 23:10

0.6.6 — the go pool is metered per model per 5 hours, and one leg can drain it

Measured the hard way. A single opencode-go/kimi-k3 review leg on a plan
document took 62 % of that model's 5-hour bucket, and four hours later
every go model answered "Insufficient balance". That error is the WINDOW,
not the month: each model has its own dollar bucket, 5 h is 20 % of the
monthly cap and a week is 50 %, so it clears on its own and topping up is
the wrong reaction. Two sessions read the dashboard independently and
agree. Buckets differ fourfold -- kimi-k3, qwen3.8-max and gpt-5.6-luna
sit at $3 per 5 h while the glm-5.x family has $12 -- so the paid go leg
now defaults to opencode-go/glm-5.3, a genuinely new family on the large
bucket, with kimi-k3 and qwen3.8-max reserved for judgement lenses at one
or two legs per window. Effort is billed as output tokens, so raising it
costs most on the dearest models: pick it from the lens, not the model.
On a balance error, fall back to the free muse-spark leg and do not retry
the same model.

Also a one-character money trap: free and paid twins differ only by a
-free or :free suffix, three copies of muse-spark exist across providers,
and choosing wrong raises no error. Full ids with provider prefixes are
now required in prose as well as tables, and the two places this repo
itself used a bare name are fixed -- that is the exact shape where the
table is right and the sentence people remember is wrong.

Corrected: DeepSeek is no longer a free-pool family on this account
(absent from the listing, consent-gated on the paid pool), so the
reviewer table stops offering it as one, and the historical
deepseek-v4-flash-free incident is annotated as an id that no longer
resolves -- the failure shape is what that story is for. The cause of its
disappearance is deliberately not guessed at.

kimi-k3 earns its first journal row from the round that cost the window:
four sole findings, all lead-verified, with honest NOT REACHED gating.

dev-lead v0.6.5

Choose a tag to compare

@github-actions github-actions released this 16 Sep 23:10

0.6.5 — family is the model, not the path; a probe is a call, not a listing

A peer session measured the newly enabled opencode go plan on 2026-09-15 by
sending every one of its 27 models a throwaway prompt, and four things from
that round belong in the doctrine rather than in a chat log.

  • methodology §1: the cross-family corollary now cuts both ways. The
    existing bullet said one CLI serves many families; the inverse is the
    trap that actually bites at accounting time — the same family reached by
    a different provider path (opencode-go/grok-4.6 beside cursor's grok,
    opencode-go/gpt-5.6-luna beside codex's terra, paid muse-spark beside the
    free one) looks like a third family and is not. Family is the model
    lineage; the path is irrelevant. cursor-agent already being a
    multi-family adapter is why this is a general rule, not a go-plan note.
  • dev-lead SKILL Phase 1: "probe availability cheaply" now says cheap means
    one real call, never a listing. opencode models kept listing models the
    workspace's permission switches had made uncallable.
  • methodology §4: the positive-side twin of "a refusal is not evidence
    until you know why it refused" — a success is not evidence until you know
    which layer produced it. A catalogue says exists; only a completed call
    says runs.
  • opencode-runtime: a first-probe section for the go pool — 22 callable
    across seven new families, 5 blocked by the owner's deliberate workspace
    switches (a decision, not a defect), and a new silent-death shape:
    glm-5.3-flash returns an empty string with no error and no counter, which
    is not the broken tokens.output=0 case. Only a one-shot probe catches it.

Named as not done: no full review leg has been driven end-to-end through
opencode-go/ on this account (the emitter is verified two ways — run, and
read), and none of the 22 has a calibration row; each returned one READY.
Capability is the next measurement, against a review problem with a known
answer, and is not claimed here.