Skip to content

Releases: qualitymd/quality.md

v0.35.2

Choose a tag to compare

@github-actions github-actions released this 15 Jul 23:18

v0.35.2 - 2026-07-15

/quality skill

  • Automatic evaluation selection now discovers ready SDK evaluators before
    considering current-session harness judgment. A usable SDK matching the
    invoking Codex or Claude agent is preferred; when no provider match is usable,
    the CLI's deterministic automatic winner is used, and harness judgment is the
    no-SDK fallback.
  • Provider-named requests now map directly to that provider's independent SDK
    evaluator without a harness-versus-SDK question. Before writing evaluation
    artifacts, the workflow reports the selected method and whether explicit
    intent, configuration, provider affinity, CLI discovery, or harness fallback
    determined it.
  • Skill metadata advances to 0.35.2 and retains
    qualitymd >=0.35.0 <0.36.0 compatibility because this patch changes the
    agent-mediated selection policy without changing the skill-facing CLI
    surface.

Compatibility / migration

  • No CLI command, automatic-discovery ordering, JSON receipt, persisted
    evaluation artifact, or QUALITY.md format behavior changes. Existing
    0.35.x CLI installations remain compatible with the updated skill.

Compatibility:

  • CLI: v0.35.2
  • QUALITY.md specification: 0.12 (Draft)
  • /quality skill: 0.35.2, requires qualitymd >=0.35.0 <0.36.0

v0.35.1

Choose a tag to compare

@github-actions github-actions released this 15 Jul 23:01

v0.35.1 - 2026-07-15

/quality skill

  • Evaluation progress now reports preflight, evidence review, report
    generation, meaningful model coverage, and whether attention is needed.
    Request windows, work units, payload schemas, concurrency caps, worker or
    subagent orchestration, and resume loops stay behind the agent interface
    unless a decision or recovery step requires them.
  • Default in-session evaluator selection is now concise information. An
    independent evaluator remains discoverable for future invocations; any
    offered current-run change is a real wait-for-answer choice. Evaluator
    selection now precedes the first feedback-log write, so the pre-mutation
    progress boundary matches the actual workflow.
  • Skill metadata advances to 0.35.1 and retains
    qualitymd >=0.35.0 <0.36.0 compatibility because this patch changes the
    agent-mediated interface without changing the skill-facing CLI surface.

Documentation

  • The agent-mediated UX guide now defines the implementation boundary and shows
    how to translate protocol state into task phases and meaningful coverage.

Compatibility / migration

  • No CLI command, receipt, persisted evaluation artifact, or QUALITY.md format
    behavior changes. Existing 0.35.x CLI installations remain compatible with
    the updated skill.

Compatibility:

  • CLI: v0.35.1
  • QUALITY.md specification: 0.12 (Draft)
  • /quality skill: 0.35.1, requires qualitymd >=0.35.0 <0.36.0

v0.35.0

Choose a tag to compare

@github-actions github-actions released this 15 Jul 21:43

v0.35.0 - 2026-07-15

CLI

  • Automatic evaluator discovery now probes and reports both Codex and Claude
    candidates even when the first is ready. Dry-run and run receipts carry each
    candidate's structured verified, assumed, or unchecked authentication
    basis; when multiple runtimes are usable, the selection reason names the
    deterministic ordering decision and every usable runtime not selected.
  • Claude readiness now uses the documented non-interactive authentication
    status probe. A verified logged-out runtime is skipped, while an unavailable
    probe remains an explicit authentication assumption; credential values are
    never reported.

/quality skill

  • Provider-named evaluator requests that could mean either the current harness
    or a same-provider SDK subprocess now prompt for that transport choice. When
    default precedence selects in-session harness judgment, the explanation
    names the fresh independent SDK alternative and its one-run and durable
    configuration paths.
  • Skill metadata advances to 0.35.0 and requires the matching
    qualitymd >=0.35.0 <0.36.0 CLI line for intent-faithful evaluator selection
    and receipt interpretation.

Compatibility / migration

  • Automatic evaluator receipts add authenticationBasis to every candidate
    and now report every built-in candidate, including usable runtimes not
    selected by deterministic ordering. Consumers of JSON receipts should accept
    the expanded candidate array and additive field; no persisted evaluation
    artifact migration is required.

Compatibility:

  • CLI: v0.35.0
  • QUALITY.md specification: 0.12 (Draft)
  • /quality skill: 0.35.0, requires qualitymd >=0.35.0 <0.36.0

v0.34.0

Choose a tag to compare

@github-actions github-actions released this 15 Jul 19:11

v0.34.0 - 2026-07-15

CLI

  • Atomic evaluation, report, lint-repair, and cancellation writes no longer
    leave orphaned temporary directories behind. Provider result files are now
    scoped to each completed evaluator call and cleaned up immediately after the
    checkpoint consumes them.
  • Evaluation advice now synthesizes one stakeholder-facing executive summary
    after finding and recommendation ranking. The run report ## Summary renders
    its bottom-line headline, concise narrative, and concrete key points instead
    of exposing the scoped root's aggregation trace.
  • evaluation.json advances to schema version 10 and
    EvaluationOutputResult references the persisted EvaluationSummaryResult.
    Report build and status require that advice payload before treating a run as
    reportable.

/quality skill

  • Skill metadata advances to 0.34.0 and requires the matching
    qualitymd >=0.34.0 <0.35.0 CLI line for the schema-10 evaluation artifact
    and executive-summary work unit.

Compatibility / migration

  • In-flight schema-version-9 evaluation runs cannot resume under this release
    and must be started again. Completed version-9 artifacts remain historical
    data; there is no migration or dual reader.

Compatibility:

  • CLI: v0.34.0
  • QUALITY.md specification: 0.12 (Draft)
  • /quality skill: 0.34.0, requires qualitymd >=0.34.0 <0.35.0

v0.33.0

Choose a tag to compare

@github-actions github-actions released this 15 Jul 12:36

v0.33.0 - 2026-07-15

CLI

  • Evaluation concurrency now comes from the selected evaluator transport rather
    than host CPU count: harness and Codex default to four, while Claude remains
    sequential. Workspace evaluation.concurrency stays the only configuration
    override and is clamped by any evaluator maximum; no CLI flag is added.
  • Direct evaluator runs now create their immutable manifest with the selected
    evaluator before issuing work and use a completion-driven bounded pool. Each
    accepted result is persisted before its slot is refilled, so fast calls no
    longer wait behind the slowest sibling and accepted work survives
    interruption. Harness runs retain their bounded rolling request window.
  • evaluation.json advances to schema version 9 with structured dispatch
    capabilities and concurrency provenance. In-flight version-8 runs must be
    started again; completed historical artifacts remain historical records.

/quality skill

  • Harness evaluation may assign one self-contained outstanding request to each
    native worker while the parent keeps graph, artifact, quality-control, and
    retry authority. Progress distinguishes the runner's outstanding cap from
    requests actually dispatched.

Compatibility / migration

  • In-flight schema-version-8 evaluation runs cannot resume under v0.33.0 and
    must be started again. Completed version-8 artifacts remain historical data;
    there is no migration or dual reader.
  • Automatic evaluation concurrency no longer follows host CPU count. Harness
    and Codex default to four; Claude is sequential even when a higher workspace
    cap is configured. Set positive evaluation.concurrency in workspace
    configuration when a lower explicit cap is required.

Compatibility:

  • CLI: v0.33.0
  • QUALITY.md specification: 0.12 (Draft)
  • /quality skill: 0.33.0, requires qualitymd >=0.33.0 <0.34.0

v0.32.2

Choose a tag to compare

@github-actions github-actions released this 15 Jul 12:02

v0.32.2 - 2026-07-15

Internal

  • Evaluator prompts now place canonical shared policy, task, model, and
    inspection context before each work-unit delta so provider prompt caching can
    reuse the longest safe prefix across independent sessions. Claude uses its
    supported cache-stable system-prompt preset, and run-local evaluator-call
    logs preserve separately reported cache-read and cache-creation input tokens.
    Evaluator sessions remain fresh and are never resumed or forked for reuse.

/quality skill

  • Skill metadata advances to 0.32.2; its compatible CLI line remains
    qualitymd >=0.32.0 <0.33.0 because this patch preserves the skill-facing
    command and evaluation artifact contract.

Compatibility:

  • CLI: v0.32.2
  • QUALITY.md specification: 0.12 (Draft)
  • /quality skill: 0.32.2, requires qualitymd >=0.32.0 <0.33.0

v0.32.1

Choose a tag to compare

@github-actions github-actions released this 15 Jul 03:28

v0.32.1 - 2026-07-14

CLI

  • Evaluation history and next-run numbering now share one recognition rule:
    current evaluation.json metadata takes precedence over the historical data
    manifest and folder prefix, unreadable or invalid metadata falls through, and
    manifest-less run folders whose slug contains quality are no longer
    silently excluded.

Internal

  • Evaluation execution now derives frames, work scheduling, protocol requests,
    artifacts, and summaries through pure domain functions and Effect-native
    ordered traversal. Source collections across the runtime are expression-built
    and exported collection types are readonly.

/quality skill

  • Skill metadata advances to 0.32.1; its compatible CLI line remains
    qualitymd >=0.32.0 <0.33.0 because this patch preserves the skill-facing
    command and artifact contract.

Compatibility:

  • CLI: v0.32.1
  • QUALITY.md specification: 0.12 (Draft)
  • /quality skill: 0.32.1, requires qualitymd >=0.32.0 <0.33.0

v0.32.0

Choose a tag to compare

@github-actions github-actions released this 15 Jul 02:21

v0.32.0 - 2026-07-14

Specification

  • QUALITY.md specification 0.12 (Draft) clarifies that an area's source
    identifies the subject or starting boundary being judged, not a precomputed
    prompt payload or permission boundary for every supporting file. Evaluation
    methods may inspect separately classified supporting context without silently
    widening the area or requirement.

CLI

  • Requirement judgment now runs in a fresh coding-agent inspection session.
    Codex, Claude, or the invoking harness searches and reads requirement-specific
    context inside a neutral, read-only workspace policy with network, approval
    escalation, writes, repository-instruction authority, and executable
    verification disabled.
  • The runner no longer walks, truncates, or packages one area-wide source bundle
    and no longer schedules resolveSource work. Each requirement returns its
    assessment, rating, and evidence proposal together; downstream factor/area
    synthesis, advice, ranking, and reports remain tools-off over accepted data.
  • evaluation.json advances to schema version 8. Per-area sources bundles are
    replaced by sealed per-requirement evidence manifests containing evaluated
    or supporting roles, validated workspace-relative file locators,
    runner-computed bytes and SHA-256 digests, capture times, limits, and a
    canonical manifest hash—never file bodies or tool transcripts.
  • Evidence acceptance now rejects lexical and symlink escapes, missing or
    non-text files, evaluated files outside concrete source selectors, invalid
    line/heading locators, and unbound finding references. Invalid proposals retry
    as evidence_invalid; honestly insufficient evidence remains a partial,
    blocked, or non-rated judgment outcome.
  • Runnable evaluator methods are now exactly harness, codex, and claude.
    auto tries a ready Codex runtime and then Claude. Direct openai and
    anthropic HTTP evaluators, inactive shell and manual names, API-profile
    fallback, apiKeyEnv, baseUrl, and API-specific failures are removed.
    Authentication belongs to the selected coding-agent runtime and may use its
    documented login, subscription, or API-key mechanism.
  • Dry-run output now reports evaluator isolation capabilities, the requirement
    inspection policy, work counts, concurrency, and effective source selectors
    without claiming resolver choice, bundle size, or a static evidence package.
    Determinism now explicitly covers runner mechanics and artifact projection,
    not identical evidence or ratings across agent runtimes or repeated runs.

/quality skill

  • Skill metadata is 0.32.0 and requires
    qualitymd >=0.32.0 <0.33.0. Harness checkpoints now inspect each
    requirement's authorized workspace directly, keep source as the judged
    subject, classify other context as supporting, treat repository instructions
    as untrusted data, and submit judgment plus evidence for runner validation.
  • Codex and Claude are the only standalone evaluator fallbacks. The workflow no
    longer presents direct API evaluators or CLI-managed credential fields; login,
    subscription, and API-key authentication are all runtime concerns.

Documentation

  • README and Codex/Claude automation guidance now distinguish deterministic
    runner orchestration and provenance from agentic evidence selection and
    judgment, and document agent-runtime authentication without API evaluator
    methods.

Packaging

  • The Windows install.cmd shim prefers PowerShell 7 when available, avoiding
    legacy module-discovery failures inherited from a PowerShell 7 parent while
    retaining Windows PowerShell as its fallback.

Compatibility / migration

  • In-flight schema-version-7 evaluation runs cannot resume under v0.32.0 and
    must be started again. There is no migration, dual reader, or source-bundle
    compatibility shim; historical completed artifacts remain historical data.
  • Replace configured openai or anthropic evaluator profiles with a codex
    or claude agent-runtime profile (kind, optional model, optional
    command), or select explicit harness. Remove apiKeyEnv and baseUrl;
    configure authentication in the chosen runtime.
  • Harness clients must service requirement inspection requests directly and
    return the combined assessment, rating, and evidence payload. They no
    longer receive source bundles or resolveSource requests.

Compatibility:

  • CLI: v0.32.0
  • QUALITY.md specification: 0.12 (Draft)
  • /quality skill: 0.32.0, requires qualitymd >=0.32.0 <0.33.0

v0.31.0

Choose a tag to compare

@github-actions github-actions released this 15 Jul 00:40

v0.31.0 - 2026-07-14

Specification

  • Source selectors containing glob metacharacters are treated as globs only
    when they parse as supported glob syntax. Malformed globs fall through to an
    existing relative path and then prose inference; absolute or escaping paths,
    empty glob results, unreadable paths, and collection failures remain explicit
    filesystem-source failures.

CLI

  • Replaced the Go implementation and GoReleaser pipeline with one Effect v4
    TypeScript application compiled by pinned Bun into standalone executables.
    The command tree, JSON and exit contracts, evaluation artifacts, and generated
    reports remain compatible; the repository no longer ships Go source,
    go install, a Go fallback, or a Node sidecar.
  • Codex and Claude evaluations now use their supported TypeScript agent SDKs;
    direct OpenAI and Anthropic API profiles and harness checkpoints remain
    available. Evaluator capabilities, readiness evidence, bounded source access,
    cancellation, retry, and call metadata are explicit rather than normalized
    across providers.
  • Each area gets one immutable, hashed context assembled from its source bundle,
    frame, rating criteria, and body guidance. Every requirement judgment runs in
    a fresh provider session or thread so sibling transcripts cannot leak into one
    another, and resume reconstructs context without provider-retained sessions.
  • Release archives and npm platform packages now cover Darwin, Windows, Linux
    glibc, and Linux musl on arm64 and x64. Checksums are mandatory, installers
    detect libc, and release repair preserves already-published Bun artifacts
    because compiled output is not byte-reproducible.
  • Harness-backed evaluation runs are no longer pinned to concurrency 1: the
    runner now keeps a rolling window of dependency-ready work requests
    outstanding — up to the run's resolved evaluation.concurrency — topping
    the window up as results are accepted, so the invoking harness can judge
    independent requests in parallel or fan them out to subagents.
    Omitted-concurrency harness runs now use the shared automatic default
    (max(2, NumCPU*2)) instead of forced 1; receipts, dry-run previews, and
    progress output state the resolved concurrency and window width.
  • awaiting_evaluator receipts carry evaluatorRequests — the outstanding
    bounded work-request set — replacing the singular evaluatorRequest, and
    --evaluator-result accepts one result envelope or a JSON array covering
    any subset of the outstanding requests. Not-yet-submitted requests stay
    outstanding at no retry cost; a schema-invalid or failed member re-emits
    for its retry attempt (with its lastFailure named) without touching other
    members' accepted results.
  • evaluation.json bumps to schema version 7: the harness checkpoint state
    is now the plural pendingEvaluatorCalls. Clean break: an in-flight
    awaiting run created before the upgrade cannot resume across it — re-run
    it; completed runs are unaffected. qualitymd evaluation status lists
    every outstanding request for an awaiting run.

/quality skill

  • Skill metadata is now 0.31.0 and requires qualitymd >=0.31.0 <0.32.0.
    Evaluation guidance distinguishes SDK agent runtimes from direct API
    evaluators, reports capability remedies, and treats area context and
    requirement session isolation as runner-owned boundaries.
  • The evaluate workflow's checkpoint loop services the outstanding request
    set: it may delegate independent requests to subagents, submits results as
    they become ready (one envelope or several per call), and names the window
    width on the first windowed receipt.

Compatibility / migration

  • The Go implementation, go install path, and GoReleaser assets are removed.
    Install the standalone archive, npm launcher, or Homebrew cask instead.
  • In-flight runner artifacts awaiting evaluator input use a clean schema break
    from version 6 to version 7 and should be started fresh; completed evaluation
    artifacts remain readable.
  • Harness clients should read evaluatorRequests and
    pendingEvaluatorCalls, and may submit one result envelope or an array. The
    former singular fields are no longer emitted.
  • The QUALITY.md specification stays at 0.11 (Draft); malformed-glob handling
    clarifies source-selection behavior without changing the document schema.

Compatibility:

  • CLI: v0.31.0
  • QUALITY.md specification: 0.11 (Draft)
  • /quality skill: 0.31.0, requires qualitymd >=0.31.0 <0.32.0

v0.30.0

Choose a tag to compare

@github-actions github-actions released this 11 Jul 16:00

v0.30.0 - 2026-07-11

Specification

  • The format now commits to non-filesystem source selectors: source stays a
    single string, and its kind is detected from the selector itself — glob
    metacharacters make a glob, an existing filesystem entry is a path, and any
    other selector describes the evaluated material in prose (for example
    "open tickets in the support queue"), resolved by evaluating tools. The
    frontmatter shape and quality.schema.json are unchanged. The specification
    version moves to 0.11 (Draft): what a source selector means in a
    conforming document has changed.

CLI

  • Evaluation runs resolve every source selector through a per-kind resolver.
    Path and glob selectors keep the deterministic workspace walk; a prose
    selector is dispatched to the invoking harness as a resolveSource work
    request on the existing checkpoint transport, and the returned material is
    validated, capped, hashed, and captured into the run artifact before any
    dependent judgment. Selector kinds are detected and pinned at run creation
    and honored on resume.
  • A selector kind the selected evaluator cannot resolve now fails the run at
    plan time with the new selector_unsupported category — naming the
    selector, its detected kind, and the remedy — distinct from
    source_unavailable, which keeps meaning the named material is missing.
  • evaluation.json bumps to schema version 6 and gains a per-area sources
    provenance record: each area's selector, detected kind, serving resolver
    (and harness runtime when harness-resolved), bundle hash, capture time, and
    per-file hashes — with captured content kept for harness-gathered evidence
    so resume re-judges against the exact evidence of record.
  • qualitymd evaluation run --dry-run --json and run receipts now include
    the per-area source dispatch plan (selector, detected kind, resolver).
  • Evaluation runs now resolve area sources as the format specification
    defines: a source-less root area evaluates the QUALITY.md file's directory,
    a source-less child area inherits the nearest ancestor's source, and glob
    selectors (for example docs/**/*.md) are expanded. Previously these were
    silently packaged as empty evidence.
  • An area source that resolves to no readable files now fails the run loudly
    with source_unavailable, naming the unresolved selector, instead of
    judging the area against an empty source bundle.
  • Source packaging skips symlinked directories and files, sockets, and
    devices instead of crashing on them (is a directory), so repositories with
    committed symlinks evaluate and resume cleanly.
  • qualitymd lint no longer reports spec-permitted extension frontmatter as
    invalid-frontmatter: a key that names no model property is now the
    warning-severity unknown-key advisory, so conforming documents carrying
    extension properties lint valid and load. Misplaced root keys on nested
    areas remain errors.
  • quality.schema.json now accepts any non-empty scalar (string, number, or
    boolean) for content scalars such as assessment, rating-level
    criterion, and ratings override values, matching lint; its $comment
    no longer claims a rating-level ordering check that no tool performs.

/quality skill

  • The harness checkpoint loop now serves resolveSource resolution requests
    alongside judgment requests: the skill gathers exactly the material the
    request's selector describes and returns it verbatim, or reports
    source_unavailable when it does not exist.
  • /quality skill metadata now declares version 0.30.0 and requires the
    qualitymd CLI 0.30.x line.

Compatibility / migration

  • The QUALITY.md specification moves to 0.11 (Draft). A source selector is
    now defined for non-filesystem material: a selector that is neither a glob
    nor an existing filesystem path is prose describing the evaluated material,
    where before only filesystem selectors were defined. Existing path and glob
    selectors are unchanged.
  • Runner evaluation.json artifacts now use schema version 6 and add the
    per-area sources provenance record; in-flight version 5 runner runs should
    be started fresh.
  • A selector kind the selected evaluator cannot resolve now fails the run at
    plan time with the new selector_unsupported category. source_unavailable
    keeps its narrower meaning: the named material is missing.

Compatibility:

  • CLI: v0.30.0
  • QUALITY.md specification: 0.11 (Draft)
  • /quality skill: 0.30.0, requires qualitymd >=0.30.0 <0.31.0