-
Notifications
You must be signed in to change notification settings - Fork 0
Skill Quality
skills quality record explicitly records one quality receipt for one selected
skill revision. skills quality inspect reads the recorded receipts from the
existing evidence database. The equivalent MCP tools are skill_quality_record
and skill_quality_inspect. They share the same validation and storage behavior.
Use Node.js 24+ and a reviewed checkout or prepared installed CLI containing these
commands. Replace i9-skills with node bin/index.mjs in a prepared source checkout.
Published npm version 0.3.4 does not include these commands or the observation
flags below; use the reviewed candidate until a separately authorized release.
Neither quality command executes a skill, benchmark case, evaluator or official validator,
installs an integration, changes a package or publishes a result.
A successful record means valid metadata was stored or an identical event was
retried. It can contain fail, blocked or not-run; exit status zero does not
mean that a skill passed. Reading instructions, reporting activation or completion,
and observing a catalog change never manufacture quality receipts.
| Recording input | Stored assurance | What remains unverified |
|---|---|---|
Closed metadata event, with optional --package-root byte checking |
caller_assertion |
Whether the declared validation or evaluation happened, the artifact claims, grading, metrics, executor and source Git identity. |
Closed behavioral event plus an explicit retained --benchmark directory |
verified_retained_benchmark |
Executor/model authenticity, grading, metrics, declared fresh contexts, source Git identity, causal improvement and general effectiveness. |
A supplied assurance field is rejected. Hashing a package or importing JSON
containing an observed-process label cannot select another assurance. The
official_validation kind is available for caller-reported metadata, but this
recording interface does not run or observe the official process. An official
receipt imported from elsewhere remains a caller assertion.
verified_retained_benchmark means the server verified the complete frozen
suite/package inventories, run JSON, run receipts and retained artifact bytes.
It does not authenticate a coherently rewritten collection of files or the
assertions inside those files. No assurance is an official conformance,
publication-readiness or statistical-effectiveness guarantee.
Save the following as receipt.json. It describes a synthetic evaluation that
was not run, with both expected run arms missing. The repeated b and c digests
are synthetic caller assertions, not measured package or artifact hashes. The
artifact locator is inert; metadata-only recording does not open or verify it.
{
"schema_version": 2,
"event_type": "skill.quality.recorded",
"event_id": "10000000-0000-4000-8000-000000000001",
"correlation_id": "10000000-0000-4000-8000-000000000002",
"occurred_at": "2026-09-19T12:00:00.000Z",
"source_host": "manual",
"source_adapter": "example",
"session": null,
"payload": {
"collection": "demo",
"skill": "example-skill",
"source": {
"repository": null,
"source_ref": null,
"resolved_git_sha": null,
"package_path": "packages/example-skill",
"package_sha256": "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb"
},
"kind": "behavioral_evaluation",
"method": {
"name": "external-evaluation",
"version": null,
"revision": null,
"source_sha256": null
},
"result": "not-run",
"coverage": {
"expected_pairs": 1,
"imported_runs": 0,
"missing_runs": 2,
"paired_cases": 0,
"declared_agent_pairs": 0,
"declared_controlled_agent_pairs": 0,
"fixture_runs": 0,
"manual_runs": 0,
"unexecuted_runs": 0,
"critical_failures": 0,
"treatment_critical_failures": 0,
"comparison_status": "incomplete",
"selected_cases": [
{
"case_id": "synthetic-case",
"phase": "exploratory",
"attempts": 1,
"joint_workflow": false
}
],
"treatment": {
"expected": 1,
"executed": 0,
"blocked": 0,
"not_run": 0,
"missing": 1,
"passed": 0,
"failed": 0,
"unresolved": 0
}
},
"artifacts": [
{
"locator": "evaluation/result.json",
"sha256": "cccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccccc",
"scope": "evaluation_artifact"
}
],
"limitations": []
}
}Choose an absolute dedicated database outside the installed collection, the command's caller workspace and any selected package/benchmark tree:
i9-skills skills quality record \
--db /data/evidence.db --file receipt.json
# The same event can be supplied on stdin:
i9-skills skills quality record \
--db /data/evidence.db --file - < receipt.jsonThe response contains recorded, event_id, event_type, assurance and
identity_key. Repeating this exact canonical event returns recorded: false.
Changed content under the same event ID, or an ID already used for a read,
lifecycle or catalog occurrence, returns evidence_conflict. A new observation
needs a new UUID, even when it concerns the same revision. Earlier failed,
blocked and missing evidence stays recorded.
File-backed input must be a stable canonical regular file with one link. Its
identity is rechecked before storage preparation, and it cannot occupy or alias
the database or any SQLite -journal, -wal or -shm path. Safe siblings and
relative input names remain valid; stdin has no selected file. Valid recording
can create safe missing database parents without changing the input file.
The event and payload are closed: unknown fields fail. Collection, skill,
source host/adapter and method names are bounded lowercase slugs. Use canonical
UTC timestamps with milliseconds and opaque UUIDs, with session: null.
Unknown source/method information is null, not an invented revision.
A non-null source repository must be a public HTTPS URL without credentials,
query or fragment. A source ref requires a declared resolved Git SHA and
repository. The package path and artifact locators are portable relative paths.
Method fields are name, version, revision and source_sha256. Artifact
fields are locator, sha256 and scope; supported scopes are official_result,
benchmark_manifest, benchmark_run, evaluation_artifact and
comparison_summary. Metadata-only scope labels do not verify those artifacts.
limitations uses the documented codes returned by the interface; arbitrary
prose is rejected. The server adds mandatory assertion and inference limits.
Supported codes are caller_assertion, source_identity_asserted,
artifact_locator_inert, validator_authenticity_unverified,
behavior_not_assessed, grading_asserted, executor_asserted,
metrics_asserted, joint_workflow, comparison_incomplete, fixture_evidence,
manual_evidence, acceptance_is_phase_label, historical_revision,
reported_budget_overrun and retained_run_limitations. These codes describe
limits; supplying one does not establish an observation or quality tier.
For official_validation metadata, replace behavioral coverage with the closed
counts selected_packages: 1, executed, blocked, not_run, passed and
failed. All other counts are 0 or 1; executed = passed + failed and
executed + blocked + not_run = 1. The result must agree with those counts.
This consistency check is not observation of an official execution.
source.package_sha256 identifies the full selected package tree, including
references, assets, scripts and other regular resources. The shared benchmark
representation sorts portable relative paths by UTF-8 bytes and hashes each
path NUL file-sha256 LF entry. It does not execute scripts or interpret the
resources as instructions.
When recording an actual metadata observation, --package-root optionally
checks that this full inert inventory matches the declared package digest:
i9-skills skills quality record \
--db /data/evidence.db --file actual-receipt.json \
--package-root /inputs/example-skillThe synthetic b digest above will not pass this check for an unrelated real
package. Unsafe links, hardlinks, special entries, oversized trees or detected
changes fail before database selection. Keep selected inputs stable; the
inspection does not promise confinement against hostile concurrent mutation.
Successful byte checking still records caller_assertion.
The normalized identity formulas match existing lifecycle evidence:
source_key = SHA256(JSON([repository, package_path]))
identity_key = SHA256(JSON([source_key, resolved_git_sha, package_sha256]))
Use normalized source fields returned by inspection when computing a source key.
source_ref remains stored metadata; a mutable ref is not a resolved revision.
Repository, ref and Git SHA are caller assertions even when package bytes are
verified. Two skills with the same name can have different sources or revisions;
name alone does not combine their evidence.
A changed resource yields a different package digest and identity. A historical
receipt continues to describe its recorded revision; it is not silently moved
to the current package. No latest-revision alias, supersession or automatic
invalidation erases earlier receipts. Query the exact identity_key to inspect
one revision and retain negative history when evaluating a replacement.
Prepare/import evidence using the
behavioral benchmark guide.
Select the actual retained freeze, not a detached comparison JSON or public
summary that cannot replay its original retained artifacts. The selected package
name and tree_sha256 must match the request's skill and package digest.
The following helper builds a complete request from an explicitly selected
freeze. Save it as make-quality-request.mjs. It only reads metadata and writes
one new request file; the quality recorder subsequently performs the full
verification. Its neutral not-run coverage is request scaffolding, replaced by
the server's retained coverage. Always submit this request with --benchmark;
do not record the scaffolding alone as an execution report.
import fs from 'node:fs';
import path from 'node:path';
import { createHash, randomUUID } from 'node:crypto';
const [directory, output, skill] = process.argv.slice(2);
const root = path.resolve(directory);
const manifestBytes = fs.readFileSync(path.join(root, 'manifest.json'));
const manifest = JSON.parse(manifestBytes);
const suite = JSON.parse(fs.readFileSync(path.join(root, 'suite.json')));
const subject = manifest.packages.find((item) => item.name === skill);
const selected = suite.cases.filter((item) => item.skills.includes(skill));
if (!subject || selected.length === 0) throw new Error('Select a frozen package.');
const expected = selected.reduce((sum, item) => sum + item.limits.attempts, 0);
const receipt = {
schema_version: 2,
event_type: 'skill.quality.recorded',
event_id: randomUUID(),
correlation_id: randomUUID(),
occurred_at: '2026-09-19T12:00:00.000Z',
source_host: 'manual',
source_adapter: 'example',
session: null,
payload: {
collection: 'demo',
skill,
source: {
repository: null,
source_ref: null,
resolved_git_sha: null,
package_path: `packages/${skill}`,
package_sha256: subject.tree_sha256,
},
kind: 'behavioral_evaluation',
method: {
name: 'retained-benchmark-comparison',
version: '1',
revision: null,
source_sha256: null,
},
result: 'not-run',
coverage: {
expected_pairs: expected,
imported_runs: 0,
missing_runs: expected * 2,
paired_cases: 0,
declared_agent_pairs: 0,
declared_controlled_agent_pairs: 0,
fixture_runs: 0,
manual_runs: 0,
unexecuted_runs: 0,
critical_failures: 0,
treatment_critical_failures: 0,
comparison_status: 'incomplete',
selected_cases: selected.map((item) => ({
case_id: item.id,
phase: item.phase,
attempts: item.limits.attempts,
joint_workflow: item.skills.length > 1,
})),
treatment: {
expected,
executed: 0,
blocked: 0,
not_run: 0,
missing: expected,
passed: 0,
failed: 0,
unresolved: 0,
},
},
artifacts: [{
locator: 'manifest.json',
sha256: createHash('sha256').update(manifestBytes).digest('hex'),
scope: 'benchmark_manifest',
}],
limitations: [],
},
};
fs.writeFileSync(output, `${JSON.stringify(receipt, null, 4)}\n`, { flag: 'wx' });Run it with caller-owned paths. The helper refuses an existing output file:
node make-quality-request.mjs \
/inputs/frozen-benchmark ./benchmark-request.json example-skill
i9-skills skills quality record \
--db /data/evidence.db --file ./benchmark-request.json \
--benchmark /inputs/frozen-benchmarkThe recorder verifies the entire freeze and every imported run/artifact, including cases unrelated to the selected package. Missing/changed artifacts, changed package or freeze identity, invalid run receipts and a stale requested digest block recording before a database is resolved or created. It reads retained evidence without executing cases, installing tools or copying private outputs into SQLite.
The server derives method, result, coverage and artifact references from those
verified bytes. A stored retained receipt references the actual manifest and
selected run JSON hashes, which bind the verified artifact inventory and full
asserted run limitations. It retains the suite digest and separately scoped
suite comparison status/result. The selected receipt's result applies only to
cases whose frozen skills list includes that package.
Read these fields together:
-
coverageretains expected/imported/missing arms, treatment execution and results, critical failures in both arms, fixture/manual/declared-agent counts, comparison completeness, and selected case phases/attempts. -
benchmark.result_basisiscritical_treatment_criteria. A result ofpasscan coexist with noncritical failures;baseline_failed_criteriaandtreatment_failed_criteriaretain those counts too. -
benchmarkretains reported baseline/treatment overruns, comparability reason codes and the number of runs declaring limitations. Overruns and unknown or differing contexts prevent a controlled comparison; they do not erase runs. - Mandatory
limitationsretainfixture_evidence,manual_evidence,comparison_incomplete,reported_budget_overrun,retained_run_limitationsandjoint_workflowwhen applicable. Multiple selected skills remain a joint workflow, not causal credit for an individual package.final-acceptanceremains a frozen phase label; untouched execution is an operator assertion.
Missing or unexecuted treatment yields honest not-run or blocked coverage;
a failed critical treatment remains fail. A passing critical treatment can
still have an incomplete comparison. Reading a receipt does not rerun the
benchmark or authenticate the executor, grader, metrics or source Git identity.
Choose an existing database and explicit collection/period:
i9-skills skills quality inspect \
--db /data/evidence.db --collection demo --skill example-skill \
--from 2026-09-01T00:00:00.000Z --until 2026-10-01T00:00:00.000Z \
--kind behavioral_evaluation --limit 20The matching MCP query is a complete inert input:
{
"collection": "demo",
"skill": "example-skill",
"from": "2026-09-01T00:00:00.000Z",
"until": "2026-10-01T00:00:00.000Z",
"kind": "behavioral_evaluation",
"limit": 20
}CLI filters --source-key and --identity-key map to MCP source_key and
identity_key. --identity-key selects one exact revision; --source-key selects
its normalized source family across revisions. Both take lowercase SHA-256 values.
Skill, kind and identity filters are optional; collection, from and until are
required. The window includes from, excludes until and uses the receipt's
asserted occurred_at, not recorded_at. No caller-supplied timestamp is
independently authenticated.
Inspection returns scope, complete bounded matching_receipts and kind counts,
rows containing normalized receipt plus recorded_at, truncated,
next_cursor and inference limits. Rows sort by (occurred_at, event_id).
When truncated is true, pass its exact next_cursor as CLI --after or MCP
after, keeping all period/collection/skill/kind/source/identity filters unchanged.
The display limit may change. A cursor cannot be reused for another scope and
is not a mutation token. Counts describe the full matching window, not just the
remaining page. Inspection never dereferences artifact locators.
For MCP recording, wrap the complete event from this page in
{"receipt": <event>}. Optional server-local selections are package_root and
benchmark; there is no database or assurance field in tool arguments. Select
the database when starting the server:
i9-skills mcp serve --db /data/evidence.dbUse the same normalized event UUID for a retry. The two tools expose closed schemas through discovery. All selected filesystem paths refer to the server's machine and require a stable owned input tree. The server rejects database paths inside the installed collection, its known caller root or selected package and benchmark roots. A plugin-root process cannot infer an otherwise unknown consumer workspace; select external state deliberately. See Skill MCP for transport and recovery.
| Boundary | Limit and behavior |
|---|---|
| Caller receipt | 16 KiB; up to 16 inert artifact references; arbitrary fields and free-form limitations fail. |
| Verified retained receipt | Manifest plus up to 160 selected run references, still within the same 16 KiB total receipt cap. A selection too large for this cap fails before storage; references are never silently dropped. |
| Query | 4 KiB; canonical positive UTC period of at most 366 days; opaque cursor at most 256 characters. |
| Scan work | At most 5,000 matching receipts before cursor or display filtering. |
| Display/output | Default 20, maximum 100 entries; serialized JSON at most 64 KiB. |
| Package inspection | At most 2,048 total filesystem entries, including directories, 4 MiB per file and 32 MiB total; every regular resource participates. |
| Benchmark verification | Existing freeze/run/receipt/artifact limits in the behavioral guide; byte integrity and execution assertions remain separate. |
Only an explicit valid writer creates/upgrades the dedicated evidence database.
Migration 4 adds the quality projection while preserving issued migration 1–3
checksums, old reads, lifecycle/catalog rows and negative history. An existing
schema 1–3 reader returns schema_upgrade_required without modifying storage.
An older consumer rejects schema 4; future/altered/untracked schemas and unrelated
SQLite databases fail safely. No command downgrades, resets, copies, prunes or
deletes evidence to recover.
Before an intentional upgrade of valuable history, retain an owner-selected
verified backup. A code rollback does not downgrade the database; use a compatible
consumer or separately authorized backup restoration. invalid_input requires
correcting the closed metadata or selected evidence; diagnostics omit raw input
and private paths. storage_unavailable requires checking the selected external
file/parent/schema. evidence_conflict requires preserving the earlier event and
using a new ID for a genuinely new observation. query_limit_exceeded requires
a narrower matching period/scope; lowering the page limit does not lower scan
work. response_too_large requires a smaller display or selection. Inspection
never creates a missing directory, database or upgrade.
Store portable metadata only: logical names, source assertions, method identity, counts, known limitation codes and relative artifact locators/hashes. Keep raw traces, prompts, private package copies, credentials, participant information and unsanitized diagnostics in separately controlled evidence, outside SQLite and public docs. The selected absolute package/benchmark paths are runtime inputs and are not stored in the receipt. A syntactically accepted public source URL is not a privacy review; the caller owns what they disclose.
Skill Memory presents
quality_receipts separately from lifecycle, catalog and weaker read evidence.
It retains kind/result/assurance and exact revision metadata with explicit
truncation. Its retention inspection counts one skill.quality.recorded
occurrence, not an additional occurrence for the mirrored envelope. Neither
receipt inspection nor memory decides promotion, retention/deletion policy,
activation or publication.
The source contracts are SkillQualityValidator, SkillQualityService and SkillQualityRepository. See Validation for separately executed structural conformance and Lifecycle Evidence for existing event meanings and denominators.
Ordinary repo validate-official continues checking every canonical package and
the temporary scaffold without recording a quality receipt. In a prepared CI
environment, the operator can select one canonical skill and explicitly request
an observation using all three flags together:
node bin/index.mjs repo validate-official \
--project ./collection \
--quality-request /tmp/owned-quality/request.json \
--quality-db /tmp/owned-quality/evidence.db \
--quality-output /tmp/owned-quality/new-observationThe parent directories must already exist at canonical absolute paths. The output directory must be new. Keep the selected request, database and output outside the installed package, caller repository and selected skill. Request, DB and sidecar collisions are rejected before setup; valid siblings are allowed. Those reserved destination names also reject normalized case-only variants, conservatively including on a case-sensitive filesystem. Do not select a consumer database for the disposable CI demonstration.
The closed request contains only collection, skill, and source. Its
source.package_path is .agents/skills/<skill>, and its package_sha256 must
match every regular file in that package. Source repository/ref/Git revision
remain assertions even when the CI configuration supplies the checked-out head.
The observer derives tool version, immutable source revision and archive digest
from the repository's reviewed configuration, checks the actual version response,
and fingerprints the selected package immediately before and after its real
validation call. It retains official-quality.json before opening the database.
Only unchanged bytes after completed setup and a matched version can receive
locally_observed_official_process. Conformance failure remains a failed receipt;
timeout, interruption, unavailable process or execution error remains blocked.
Setup/version failure or changed/unavailable post-validation bytes retain a
blocked artifact without an observed receipt. A late database failure preserves
the retained artifact and returns failure; no automatic reset or evidence deletion
occurs. CLI failure still reflects other canonical/scaffold failures independently
of the selected receipt.
The artifact and observation JSON written to the CI log contain compact method, fingerprint, status and result metadata, not stdout/stderr, source text or raw error messages. Ordinary validator diagnostics remain separate. Inspection treats relative artifact locators as inert. Copying that JSON into public record ingress cannot recreate an observed assurance tier. Correct-version official conformance does not prove behavior, rights, safety, native installation or publication readiness.
The repository validation workflow demonstrates this path for skill-authoring
using runner-owned disposable request/DB/output, then performs the same read-only
quality query exposed to CLI/MCP consumers when a receipt database exists. A
successful observation without its database fails the workflow. Setup/version
failure does not create a database merely to query it. A passing workflow is evidence only
for its exact head. Local tests use fake process seams and synthetic workspaces;
they neither install Python nor provide actual official-process evidence.