Releases: qualitymd/quality.md
Releases · qualitymd/quality.md
Release list
v0.35.2
v0.35.2 - 2026-07-15
/quality skill
- Automatic evaluation selection now discovers ready SDK evaluators before
considering current-session harness judgment. A usable SDK matching the
invoking Codex or Claude agent is preferred; when no provider match is usable,
the CLI's deterministic automatic winner is used, and harness judgment is the
no-SDK fallback. - Provider-named requests now map directly to that provider's independent SDK
evaluator without a harness-versus-SDK question. Before writing evaluation
artifacts, the workflow reports the selected method and whether explicit
intent, configuration, provider affinity, CLI discovery, or harness fallback
determined it. - Skill metadata advances to
0.35.2and retains
qualitymd >=0.35.0 <0.36.0compatibility because this patch changes the
agent-mediated selection policy without changing the skill-facing CLI
surface.
Compatibility / migration
- No CLI command, automatic-discovery ordering, JSON receipt, persisted
evaluation artifact, or QUALITY.md format behavior changes. Existing
0.35.xCLI installations remain compatible with the updated skill.
Compatibility:
- CLI:
v0.35.2 - QUALITY.md specification:
0.12 (Draft) - /quality skill:
0.35.2, requiresqualitymd >=0.35.0 <0.36.0
v0.35.1
v0.35.1 - 2026-07-15
/quality skill
- Evaluation progress now reports preflight, evidence review, report
generation, meaningful model coverage, and whether attention is needed.
Request windows, work units, payload schemas, concurrency caps, worker or
subagent orchestration, and resume loops stay behind the agent interface
unless a decision or recovery step requires them. - Default in-session evaluator selection is now concise information. An
independent evaluator remains discoverable for future invocations; any
offered current-run change is a real wait-for-answer choice. Evaluator
selection now precedes the first feedback-log write, so the pre-mutation
progress boundary matches the actual workflow. - Skill metadata advances to
0.35.1and retains
qualitymd >=0.35.0 <0.36.0compatibility because this patch changes the
agent-mediated interface without changing the skill-facing CLI surface.
Documentation
- The agent-mediated UX guide now defines the implementation boundary and shows
how to translate protocol state into task phases and meaningful coverage.
Compatibility / migration
- No CLI command, receipt, persisted evaluation artifact, or QUALITY.md format
behavior changes. Existing0.35.xCLI installations remain compatible with
the updated skill.
Compatibility:
- CLI:
v0.35.1 - QUALITY.md specification:
0.12 (Draft) - /quality skill:
0.35.1, requiresqualitymd >=0.35.0 <0.36.0
v0.35.0
v0.35.0 - 2026-07-15
CLI
- Automatic evaluator discovery now probes and reports both Codex and Claude
candidates even when the first is ready. Dry-run and run receipts carry each
candidate's structuredverified,assumed, oruncheckedauthentication
basis; when multiple runtimes are usable, the selection reason names the
deterministic ordering decision and every usable runtime not selected. - Claude readiness now uses the documented non-interactive authentication
status probe. A verified logged-out runtime is skipped, while an unavailable
probe remains an explicit authentication assumption; credential values are
never reported.
/quality skill
- Provider-named evaluator requests that could mean either the current harness
or a same-provider SDK subprocess now prompt for that transport choice. When
default precedence selects in-sessionharnessjudgment, the explanation
names the fresh independent SDK alternative and its one-run and durable
configuration paths. - Skill metadata advances to
0.35.0and requires the matching
qualitymd >=0.35.0 <0.36.0CLI line for intent-faithful evaluator selection
and receipt interpretation.
Compatibility / migration
- Automatic evaluator receipts add
authenticationBasisto every candidate
and now report every built-in candidate, including usable runtimes not
selected by deterministic ordering. Consumers of JSON receipts should accept
the expanded candidate array and additive field; no persisted evaluation
artifact migration is required.
Compatibility:
- CLI:
v0.35.0 - QUALITY.md specification:
0.12 (Draft) - /quality skill:
0.35.0, requiresqualitymd >=0.35.0 <0.36.0
v0.34.0
v0.34.0 - 2026-07-15
CLI
- Atomic evaluation, report, lint-repair, and cancellation writes no longer
leave orphaned temporary directories behind. Provider result files are now
scoped to each completed evaluator call and cleaned up immediately after the
checkpoint consumes them. - Evaluation advice now synthesizes one stakeholder-facing executive summary
after finding and recommendation ranking. The run report## Summaryrenders
its bottom-line headline, concise narrative, and concrete key points instead
of exposing the scoped root's aggregation trace. evaluation.jsonadvances to schema version 10 and
EvaluationOutputResultreferences the persistedEvaluationSummaryResult.
Report build and status require that advice payload before treating a run as
reportable.
/quality skill
- Skill metadata advances to
0.34.0and requires the matching
qualitymd >=0.34.0 <0.35.0CLI line for the schema-10 evaluation artifact
and executive-summary work unit.
Compatibility / migration
- In-flight schema-version-9 evaluation runs cannot resume under this release
and must be started again. Completed version-9 artifacts remain historical
data; there is no migration or dual reader.
Compatibility:
- CLI:
v0.34.0 - QUALITY.md specification:
0.12 (Draft) - /quality skill:
0.34.0, requiresqualitymd >=0.34.0 <0.35.0
v0.33.0
v0.33.0 - 2026-07-15
CLI
- Evaluation concurrency now comes from the selected evaluator transport rather
than host CPU count: harness and Codex default to four, while Claude remains
sequential. Workspaceevaluation.concurrencystays the only configuration
override and is clamped by any evaluator maximum; no CLI flag is added. - Direct evaluator runs now create their immutable manifest with the selected
evaluator before issuing work and use a completion-driven bounded pool. Each
accepted result is persisted before its slot is refilled, so fast calls no
longer wait behind the slowest sibling and accepted work survives
interruption. Harness runs retain their bounded rolling request window. evaluation.jsonadvances to schema version 9 with structured dispatch
capabilities and concurrency provenance. In-flight version-8 runs must be
started again; completed historical artifacts remain historical records.
/quality skill
- Harness evaluation may assign one self-contained outstanding request to each
native worker while the parent keeps graph, artifact, quality-control, and
retry authority. Progress distinguishes the runner's outstanding cap from
requests actually dispatched.
Compatibility / migration
- In-flight schema-version-8 evaluation runs cannot resume under v0.33.0 and
must be started again. Completed version-8 artifacts remain historical data;
there is no migration or dual reader. - Automatic evaluation concurrency no longer follows host CPU count. Harness
and Codex default to four; Claude is sequential even when a higher workspace
cap is configured. Set positiveevaluation.concurrencyin workspace
configuration when a lower explicit cap is required.
Compatibility:
- CLI:
v0.33.0 - QUALITY.md specification:
0.12 (Draft) - /quality skill:
0.33.0, requiresqualitymd >=0.33.0 <0.34.0
v0.32.2
v0.32.2 - 2026-07-15
Internal
- Evaluator prompts now place canonical shared policy, task, model, and
inspection context before each work-unit delta so provider prompt caching can
reuse the longest safe prefix across independent sessions. Claude uses its
supported cache-stable system-prompt preset, and run-local evaluator-call
logs preserve separately reported cache-read and cache-creation input tokens.
Evaluator sessions remain fresh and are never resumed or forked for reuse.
/quality skill
- Skill metadata advances to
0.32.2; its compatible CLI line remains
qualitymd >=0.32.0 <0.33.0because this patch preserves the skill-facing
command and evaluation artifact contract.
Compatibility:
- CLI:
v0.32.2 - QUALITY.md specification:
0.12 (Draft) - /quality skill:
0.32.2, requiresqualitymd >=0.32.0 <0.33.0
v0.32.1
v0.32.1 - 2026-07-14
CLI
- Evaluation history and next-run numbering now share one recognition rule:
currentevaluation.jsonmetadata takes precedence over the historical data
manifest and folder prefix, unreadable or invalid metadata falls through, and
manifest-less run folders whose slug containsqualityare no longer
silently excluded.
Internal
- Evaluation execution now derives frames, work scheduling, protocol requests,
artifacts, and summaries through pure domain functions and Effect-native
ordered traversal. Source collections across the runtime are expression-built
and exported collection types are readonly.
/quality skill
- Skill metadata advances to
0.32.1; its compatible CLI line remains
qualitymd >=0.32.0 <0.33.0because this patch preserves the skill-facing
command and artifact contract.
Compatibility:
- CLI:
v0.32.1 - QUALITY.md specification:
0.12 (Draft) - /quality skill:
0.32.1, requiresqualitymd >=0.32.0 <0.33.0
v0.32.0
v0.32.0 - 2026-07-14
Specification
- QUALITY.md specification
0.12 (Draft)clarifies that an area'ssource
identifies the subject or starting boundary being judged, not a precomputed
prompt payload or permission boundary for every supporting file. Evaluation
methods may inspect separately classified supporting context without silently
widening the area or requirement.
CLI
- Requirement judgment now runs in a fresh coding-agent inspection session.
Codex, Claude, or the invoking harness searches and reads requirement-specific
context inside a neutral, read-only workspace policy with network, approval
escalation, writes, repository-instruction authority, and executable
verification disabled. - The runner no longer walks, truncates, or packages one area-wide source bundle
and no longer schedulesresolveSourcework. Each requirement returns its
assessment, rating, and evidence proposal together; downstream factor/area
synthesis, advice, ranking, and reports remain tools-off over accepted data. evaluation.jsonadvances to schema version 8. Per-areasourcesbundles are
replaced by sealed per-requirementevidencemanifests containing evaluated
or supporting roles, validated workspace-relative file locators,
runner-computed bytes and SHA-256 digests, capture times, limits, and a
canonical manifest hash—never file bodies or tool transcripts.- Evidence acceptance now rejects lexical and symlink escapes, missing or
non-text files, evaluated files outside concrete source selectors, invalid
line/heading locators, and unbound finding references. Invalid proposals retry
asevidence_invalid; honestly insufficient evidence remains a partial,
blocked, or non-rated judgment outcome. - Runnable evaluator methods are now exactly
harness,codex, andclaude.
autotries a ready Codex runtime and then Claude. Directopenaiand
anthropicHTTP evaluators, inactiveshellandmanualnames, API-profile
fallback,apiKeyEnv,baseUrl, and API-specific failures are removed.
Authentication belongs to the selected coding-agent runtime and may use its
documented login, subscription, or API-key mechanism. - Dry-run output now reports evaluator isolation capabilities, the requirement
inspection policy, work counts, concurrency, and effective source selectors
without claiming resolver choice, bundle size, or a static evidence package.
Determinism now explicitly covers runner mechanics and artifact projection,
not identical evidence or ratings across agent runtimes or repeated runs.
/quality skill
- Skill metadata is
0.32.0and requires
qualitymd >=0.32.0 <0.33.0. Harness checkpoints now inspect each
requirement's authorized workspace directly, keepsourceas the judged
subject, classify other context as supporting, treat repository instructions
as untrusted data, and submit judgment plus evidence for runner validation. - Codex and Claude are the only standalone evaluator fallbacks. The workflow no
longer presents direct API evaluators or CLI-managed credential fields; login,
subscription, and API-key authentication are all runtime concerns.
Documentation
- README and Codex/Claude automation guidance now distinguish deterministic
runner orchestration and provenance from agentic evidence selection and
judgment, and document agent-runtime authentication without API evaluator
methods.
Packaging
- The Windows
install.cmdshim prefers PowerShell 7 when available, avoiding
legacy module-discovery failures inherited from a PowerShell 7 parent while
retaining Windows PowerShell as its fallback.
Compatibility / migration
- In-flight schema-version-7 evaluation runs cannot resume under v0.32.0 and
must be started again. There is no migration, dual reader, or source-bundle
compatibility shim; historical completed artifacts remain historical data. - Replace configured
openaioranthropicevaluator profiles with acodex
orclaudeagent-runtime profile (kind, optionalmodel, optional
command), or select explicitharness. RemoveapiKeyEnvandbaseUrl;
configure authentication in the chosen runtime. - Harness clients must service requirement inspection requests directly and
return the combinedassessment,rating, andevidencepayload. They no
longer receive source bundles orresolveSourcerequests.
Compatibility:
- CLI:
v0.32.0 - QUALITY.md specification:
0.12 (Draft) - /quality skill:
0.32.0, requiresqualitymd >=0.32.0 <0.33.0
v0.31.0
v0.31.0 - 2026-07-14
Specification
- Source selectors containing glob metacharacters are treated as globs only
when they parse as supported glob syntax. Malformed globs fall through to an
existing relative path and then prose inference; absolute or escaping paths,
empty glob results, unreadable paths, and collection failures remain explicit
filesystem-source failures.
CLI
- Replaced the Go implementation and GoReleaser pipeline with one Effect v4
TypeScript application compiled by pinned Bun into standalone executables.
The command tree, JSON and exit contracts, evaluation artifacts, and generated
reports remain compatible; the repository no longer ships Go source,
go install, a Go fallback, or a Node sidecar. - Codex and Claude evaluations now use their supported TypeScript agent SDKs;
direct OpenAI and Anthropic API profiles and harness checkpoints remain
available. Evaluator capabilities, readiness evidence, bounded source access,
cancellation, retry, and call metadata are explicit rather than normalized
across providers. - Each area gets one immutable, hashed context assembled from its source bundle,
frame, rating criteria, and body guidance. Every requirement judgment runs in
a fresh provider session or thread so sibling transcripts cannot leak into one
another, and resume reconstructs context without provider-retained sessions. - Release archives and npm platform packages now cover Darwin, Windows, Linux
glibc, and Linux musl on arm64 and x64. Checksums are mandatory, installers
detect libc, and release repair preserves already-published Bun artifacts
because compiled output is not byte-reproducible. - Harness-backed evaluation runs are no longer pinned to concurrency 1: the
runner now keeps a rolling window of dependency-ready work requests
outstanding — up to the run's resolvedevaluation.concurrency— topping
the window up as results are accepted, so the invoking harness can judge
independent requests in parallel or fan them out to subagents.
Omitted-concurrency harness runs now use the shared automatic default
(max(2, NumCPU*2)) instead of forced 1; receipts, dry-run previews, and
progress output state the resolved concurrency and window width. awaiting_evaluatorreceipts carryevaluatorRequests— the outstanding
bounded work-request set — replacing the singularevaluatorRequest, and
--evaluator-resultaccepts one result envelope or a JSON array covering
any subset of the outstanding requests. Not-yet-submitted requests stay
outstanding at no retry cost; a schema-invalid or failed member re-emits
for its retry attempt (with itslastFailurenamed) without touching other
members' accepted results.evaluation.jsonbumps to schema version 7: the harness checkpoint state
is now the pluralpendingEvaluatorCalls. Clean break: an in-flight
awaiting run created before the upgrade cannot resume across it — re-run
it; completed runs are unaffected.qualitymd evaluation statuslists
every outstanding request for an awaiting run.
/quality skill
- Skill metadata is now
0.31.0and requiresqualitymd >=0.31.0 <0.32.0.
Evaluation guidance distinguishes SDK agent runtimes from direct API
evaluators, reports capability remedies, and treats area context and
requirement session isolation as runner-owned boundaries. - The evaluate workflow's checkpoint loop services the outstanding request
set: it may delegate independent requests to subagents, submits results as
they become ready (one envelope or several per call), and names the window
width on the first windowed receipt.
Compatibility / migration
- The Go implementation,
go installpath, and GoReleaser assets are removed.
Install the standalone archive, npm launcher, or Homebrew cask instead. - In-flight runner artifacts awaiting evaluator input use a clean schema break
from version 6 to version 7 and should be started fresh; completed evaluation
artifacts remain readable. - Harness clients should read
evaluatorRequestsand
pendingEvaluatorCalls, and may submit one result envelope or an array. The
former singular fields are no longer emitted. - The QUALITY.md specification stays at
0.11 (Draft); malformed-glob handling
clarifies source-selection behavior without changing the document schema.
Compatibility:
- CLI:
v0.31.0 - QUALITY.md specification:
0.11 (Draft) - /quality skill:
0.31.0, requiresqualitymd >=0.31.0 <0.32.0
v0.30.0
v0.30.0 - 2026-07-11
Specification
- The format now commits to non-filesystem source selectors:
sourcestays a
single string, and its kind is detected from the selector itself — glob
metacharacters make a glob, an existing filesystem entry is a path, and any
other selector describes the evaluated material in prose (for example
"open tickets in the support queue"), resolved by evaluating tools. The
frontmatter shape andquality.schema.jsonare unchanged. The specification
version moves to0.11 (Draft): what asourceselector means in a
conforming document has changed.
CLI
- Evaluation runs resolve every source selector through a per-kind resolver.
Path and glob selectors keep the deterministic workspace walk; a prose
selector is dispatched to the invoking harness as aresolveSourcework
request on the existing checkpoint transport, and the returned material is
validated, capped, hashed, and captured into the run artifact before any
dependent judgment. Selector kinds are detected and pinned at run creation
and honored on resume. - A selector kind the selected evaluator cannot resolve now fails the run at
plan time with the newselector_unsupportedcategory — naming the
selector, its detected kind, and the remedy — distinct from
source_unavailable, which keeps meaning the named material is missing. evaluation.jsonbumps to schema version 6 and gains a per-areasources
provenance record: each area's selector, detected kind, serving resolver
(and harness runtime when harness-resolved), bundle hash, capture time, and
per-file hashes — with captured content kept for harness-gathered evidence
so resume re-judges against the exact evidence of record.qualitymd evaluation run --dry-run --jsonand run receipts now include
the per-area source dispatch plan (selector, detected kind, resolver).- Evaluation runs now resolve area sources as the format specification
defines: a source-less root area evaluates the QUALITY.md file's directory,
a source-less child area inherits the nearest ancestor's source, and glob
selectors (for exampledocs/**/*.md) are expanded. Previously these were
silently packaged as empty evidence. - An area source that resolves to no readable files now fails the run loudly
withsource_unavailable, naming the unresolved selector, instead of
judging the area against an empty source bundle. - Source packaging skips symlinked directories and files, sockets, and
devices instead of crashing on them (is a directory), so repositories with
committed symlinks evaluate and resume cleanly. qualitymd lintno longer reports spec-permitted extension frontmatter as
invalid-frontmatter: a key that names no model property is now the
warning-severityunknown-keyadvisory, so conforming documents carrying
extension properties lint valid and load. Misplaced root keys on nested
areas remain errors.quality.schema.jsonnow accepts any non-empty scalar (string, number, or
boolean) for content scalars such asassessment, rating-level
criterion, andratingsoverride values, matchinglint; its$comment
no longer claims a rating-level ordering check that no tool performs.
/quality skill
- The harness checkpoint loop now serves
resolveSourceresolution requests
alongside judgment requests: the skill gathers exactly the material the
request's selector describes and returns it verbatim, or reports
source_unavailablewhen it does not exist. /qualityskill metadata now declares version0.30.0and requires the
qualitymdCLI0.30.xline.
Compatibility / migration
- The QUALITY.md specification moves to
0.11 (Draft). Asourceselector is
now defined for non-filesystem material: a selector that is neither a glob
nor an existing filesystem path is prose describing the evaluated material,
where before only filesystem selectors were defined. Existing path and glob
selectors are unchanged. - Runner
evaluation.jsonartifacts now use schema version6and add the
per-areasourcesprovenance record; in-flight version 5 runner runs should
be started fresh. - A selector kind the selected evaluator cannot resolve now fails the run at
plan time with the newselector_unsupportedcategory.source_unavailable
keeps its narrower meaning: the named material is missing.
Compatibility:
- CLI:
v0.30.0 - QUALITY.md specification:
0.11 (Draft) - /quality skill:
0.30.0, requiresqualitymd >=0.30.0 <0.31.0