Status: alpha / experimental
A control plane for repository changes made with AI. It doesn't replace git, tests, CI, or your tracker — it adds a thin layer of decision discipline and traceability on top, so a change is never just "the agent wrote some code." The machine core tracks desired state, actual code state, transitions, evidence, and violations; the human-facing text (scope, phase names, status summaries) is a derived rendering, never the source of truth.
🇬🇧 English · 🇷🇺 Русская версия
A normal flow is understand the task → write code → run tests → open a PR. When an AI agent writes or co-writes the code, that isn't enough: the agent can produce plausible code and a plausible report, while the reasoning — why this and not the alternative, what actually passed — lives only in a chat window and then evaporates. S5D makes that reasoning durable, bound to real state, and checkable. Code should never appear as "agent magic"; you should always be able to recover why a change exists.
- Target state — before code, describe what the system should become and which existing architecture it reuses (domains, capabilities, components). You design the system you want; you don't just patch the one you have, and you don't invent new architecture without a reason.
- Explicit choice — compare real alternatives before committing. A decision needs ≥3 hypotheses that differ in kind, not three flavours of one idea. "We picked X because it's right" is not a decision. "We picked X because it covers revocation, which Y can't, while Z is simpler but weaker on security" is.
- Run evidence — record what agents and tools actually produced: which gates ran, which tests passed, which review findings closed, who approved, and the SHA256 of the spec at approval and import. An opt-in
workflow.outcomealso binds required gate results to a deterministic fingerprint of the governed source. Execution-artifact integrity remains partial (run exechashes stdout, not the complete execution context). - Verify against code — review the diff and validate declared component paths/dependencies against the decision. If a spec claims
auth/sessionbut the diff touchesbilling, trace and review that mismatch. Lifecycle drift/reconcile/rollback covers spec+alias metadata and does not hash or revert source files or git history; outcome-enabled gates separately fingerprint declared source paths.
Only two places hold lifecycle truth:
.s5d/packages/*— specs. A spec is the plan and the contract: intent, acceptance criteria, components, gates, hypotheses, contracts..s5d/records/*— records. A record is the history of what happened to that contract: preview, approval, gate results, import, reflection.
Not markdown notes, not a Shape doc, not a review report on its own, not "we agreed in chat." If it isn't in a package or a record, it isn't accepted S5D state. The CLI enforces a SHA256 chain between spec and record, so an approval granted for one plan can't silently cover a different one — edit a spec after approval and the approval no longer matches.
Route → Shape → Discover → Target → Decide → Spec → Run → Verify → Ship → Learn
It's the map, not a mandatory checklist — most tasks run only part of it.
- Route — classify the request: which tier, which mode, where to enter.
- Shape — turn a raw or vague request ("make checkout better") into a clear intake kernel: why it matters, impacted capabilities, constraints, non-goals, success signal, assumptions, open questions. Shape is intake only — it never creates accepted state.
- Discover — map an unfamiliar repo's current architecture into structured tables (domains, capabilities, components, edges), every claim tagged
[VERIFIED] / [INFERRED] / [SPECULATIVE]. - Target — state what's anomalous now and define acceptance before options. Describe the behaviour that should be true, not the implementation method.
- Decide — for real tradeoffs: ≥3 distinct hypotheses, evidence per hypothesis, a confirmed winner. Human confirmation here is mandatory and cannot be waived.
- Spec — turn the decision into an executable contract: problem → acceptance scenarios → components → gates, tracing every user-facing change
Feature → UseCase → Capability → Component. - Run — preview → human approval (mandatory) → implement → run gates. An agent or engine finishing is process evidence, not proof of the result. For an outcome-enabled spec, required gates must match both the current approved spec and current governed source before explicit human run acceptance.
- Verify / Ship / Learn — tests and (for material work) adversarial review → import under the SHA256 chain → push/deploy only with explicit human permission → reflect and record reusable heuristics.
Skip S5D for bugfixes under ~30 LOC, config-only changes, docs-only changes, status queries, and small local edits with no architectural meaning. Just do them.
Use S5D for a new feature; an architecture change; anything touching auth, payment, security, PII, or compliance; changes spanning multiple domains; AI-agent work that must prove traceability; decisions with a real tradeoff; and changes that need acceptance criteria or touch a trust boundary.
Rule of thumb: if you can't say which spec covers your diff, you probably need S5D. Tiers (Note → Decision → Lightweight → Standard → High) scale the rigor to the risk — see Tiers below; when unsure, pick the higher one.
S5D lets you waive workflow steps with a recorded WAIVER line — but two things are never waivable: human confirmation of a decision and human approval of a run. An agent's say-so is never a substitute for either. "The AI decided, so it's confirmed" and "skip approval, it's urgent" are both refused by design.
A git commit says what changed. S5D adds why it had to change, which alternatives were weighed, who confirmed the choice, which files belong to it, what checks passed, who verified the result, and what was learned — turning a change from a patch into a traceable engineering decision.
🇬🇧 English version · 🇷🇺 Русский
Обычный процесс: понять задачу → написать код → прогнать тесты → открыть PR. Когда код пишет или дописывает AI-агент, этого мало: агент может выдать правдоподобный код и правдоподобный отчёт, а рассуждение — почему именно так, а не иначе, что реально прошло — остаётся только в окне чата и потом испаряется. S5D делает это рассуждение долговечным, привязанным к реальному состоянию и проверяемым. Код не должен появляться как «магия агента» — всегда должно быть понятно, зачем изменение существует.
- Целевое состояние (target state) — до кода опиши, какой система должна стать и какую существующую архитектуру (домены, capabilities, компоненты) она переиспользует. Ты проектируешь нужную систему, а не просто патчишь текущую и не изобретаешь новую архитектуру без причины.
- Явный выбор — сравни реальные альтернативы до фиксации. Решению нужно ≥3 гипотезы, разные по сути, а не три варианта одной идеи. «Выбрали X, потому что так правильно» — это не решение. «Выбрали X, потому что он закрывает revocation, чего не может Y, а Z проще, но слабее по безопасности» — решение.
- Доказательства выполнения (run evidence) — записывай, что реально произвели агенты и инструменты: какие gates прогнаны, какие тесты прошли, какие review-findings закрыты, кто одобрил, и SHA256 спека на момент approval и import. Опциональный
workflow.outcomeдополнительно привязывает обязательные gate results к детерминированному fingerprint управляемых исходников. Целостность execution-артефактов всё ещё частичная (run execхеширует stdout, но не полный контекст запуска). - Проверка по коду — сопоставь diff и объявленные component paths/dependencies с решением. Если спек заявляет
auth/session, а diff трогаетbilling, это нужно trace-ить и проверять. Обычный drift/reconcile/rollback работает с метаданными spec+alias; outcome-enabled gates дополнительно хешируют объявленные исходники, но S5D не откатывает ни их, ни git-историю.
Lifecycle-истина живёт только в двух местах:
.s5d/packages/*— спеки. Спек — это план и контракт: intent, критерии приёмки, компоненты, gates, гипотезы, контракты..s5d/records/*— записи (records). Record — это история того, что с контрактом произошло: preview, approval, результаты gates, import, рефлексия. Полеrevisionзащищает её от молчаливого lost update: устаревшая конкурентная запись отклоняется и должна перечитать record. Каждый gate result/waiver привязан к SHA256 спека: изменение спека требует прогнать gates или выдать waiver заново.
Не markdown-заметки, не Shape-документ, не review-отчёт сам по себе, не «договорились в чате». Если этого нет в package или record — это не принятое состояние S5D. CLI держит SHA256-цепочку между спеком и record, поэтому approval, выданный на один план, не может молча покрыть другой — поменяй спек после approval, и approval перестаёт совпадать.
Route → Shape → Discover → Target → Decide → Spec → Run → Verify → Ship → Learn
Это карта, а не обязательный чек-лист — большинство задач проходят лишь часть.
- Route — классифицировать запрос: какой tier, какой mode, с какой стадии входить.
- Shape — превратить сырой или мутный запрос («сделать нормальный checkout») в чёткое intake-ядро: зачем нужно, затронутые capabilities, ограничения, non-goals, сигнал успеха, допущения, открытые вопросы. Shape — только intake, он не создаёт принятого состояния.
- Discover — разложить текущую архитектуру незнакомого репозитория в структурированные таблицы (домены, capabilities, компоненты, связи); каждое утверждение помечено
[VERIFIED] / [INFERRED] / [SPECULATIVE]. - Target — сказать, что сейчас аномально, и определить приёмку до вариантов. Описывай поведение, которое должно стать истинным, а не способ реализации.
- Decide — для настоящих tradeoff-ов: ≥3 разные по сути гипотезы, доказательства по каждой, подтверждённый winner. Подтверждение человеком здесь обязательно и не waiver-ится.
- Spec — превратить решение в исполнимый контракт: проблема → сценарии приёмки → компоненты → gates, трассируя каждое пользовательское изменение
Feature → UseCase → Capability → Component. - Run — preview → одобрение человеком (обязательно) → реализация → прогон gates. Завершение агентом или engine — только process evidence. В outcome-enabled спеке обязательные gates должны соответствовать и текущему одобренному спеку, и текущим управляемым исходникам до явной человеческой приёмки run.
- Verify / Ship / Learn — тесты и (для значимых изменений) adversarial review → import под SHA256-цепочкой → push/deploy только с явного разрешения человека → рефлексия и запись переиспользуемых эвристик.
Пропусти S5D для багфиксов меньше ~30 строк, config-only, docs-only изменений, status-запросов и мелких локальных правок без архитектурного смысла. Просто сделай их.
Используй S5D для новой фичи; изменения архитектуры; всего, что трогает auth, payment, security, PII или compliance; изменений в нескольких доменах; работы AI-агента, где нужна доказуемая traceability; решений с реальным tradeoff; и изменений, которым нужны критерии приёмки или которые трогают trust boundary.
Эвристика: если ты не можешь сказать, какой спек покрывает твой diff, — скорее всего, S5D нужен. Tiers (Note → Decision → Lightweight → Standard → High) масштабируют строгость под риск — см. Tiers ниже; в сомнениях выбирай более высокий.
S5D позволяет waiver-нуть шаги процесса записанной строкой WAIVER — но две вещи не waiver-ятся никогда: подтверждение решения человеком и одобрение run человеком. Слова агента не заменяют ни то, ни другое. «AI решил — значит, подтверждено» и «пропустим approval, это срочно» отклоняются by design.
Git-commit говорит что изменилось. S5D добавляет зачем это должно было измениться, какие альтернативы взвесили, кто подтвердил выбор, какие файлы к нему относятся, какие проверки прошли, кто верифицировал результат и чему научились — превращая изменение из патча в трассируемое инженерное решение.
Two files per change: a spec (what you intend) and a record (what happened). The spec can also carry optional workflow support for phases and roles; the record can close the loop with telemetry refs and an explicit verdict. The CLI enforces a SHA256 chain between them:
spec → preview → approve → import → drift-check
↓
reconcile / rollback (S5D metadata only)
# Install
git clone https://github.com/system5-dev/s5d.git
cd s5d
./install.sh
# Initialize in your repo
s5d init
# If the repo has .git/, init installs a Rust-backed pre-commit hook.
# Manual hook entrypoint: s5d hook pre-commit
# Check or apply S5D binary/skill updates
s5d admin update check
s5d admin update apply
# Sessions self-update: the plugin's SessionStart hook runs `s5d update auto`,
# which applies in the background when the checkout is clean and on the default
# branch (otherwise it prints the prompt). Opt out with S5D_AUTO_UPDATE=0.
# Trust note: self-update builds and installs whatever origin's default branch
# serves (ff-only) — enable it only for checkouts whose origin you trust.
# Optional, recommended on a repo S5D hasn't seen: build the discovery index
# (file index, evidence, dependency graph under .s5d/discovery/) so specs and
# trace queries can link to real code. Agent flows run a fuller Discover stage.
s5d discover sync
# Create a spec — the scaffold validates out of the box; edit the generated
# placeholder architecture (domain/capability/component, TODO paths) to match reality
s5d new feat.my-feature --product myapp
# Edit the spec YAML, then:
s5d verify validate .s5d/packages/feat.my-feature__*.s5d.yaml
s5d state preview .s5d/packages/feat.my-feature__*.s5d.yaml
s5d state approve .s5d/packages/feat.my-feature__*.s5d.yaml --reviewer reviewername
# Implement your code, then:
s5d verify run-gates .s5d/packages/feat.my-feature__*.s5d.yaml
s5d state import .s5d/packages/feat.my-feature__*.s5d.yaml --verified-by verifiername
# Optional: run a configured agent task and record evidence
s5d run list .s5d/packages/feat.my-feature__*.s5d.yaml
s5d run start .s5d/packages/feat.my-feature__*.s5d.yaml --id prototype
s5d run task .s5d/packages/feat.my-feature__*.s5d.yaml --phase prototype --engine ralph
s5d run exec .s5d/packages/feat.my-feature__*.s5d.yaml --id prototype --engine local-engine
s5d run accept .s5d/packages/feat.my-feature__*.s5d.yaml --id prototype --reviewer yourname
# Later: close the loop with telemetry-backed outcome
s5d state reflect .s5d/packages/feat.my-feature__*.s5d.yaml \
--summary "Telemetry stayed inside target bounds" \
--verdict confirmed \
--measurement-window "7d post-ship" \
--telemetry "grafana://my-dashboard" \
--heuristic "Keep rollout verdicts tied to explicit telemetry refs"
# Later: verify nothing drifted
s5d state drift-checkS5D records agent execution as evidence against a desired system state. It can render that evidence for humans, but the durable state remains machine-readable.
s5d run list/start/exec/acceptmanages active work state in.record.yamls5d run exec --engine <name>executes an approved command template from.s5d/config.yaml, captures stdout/stderr under.s5d/runs/, and records the output hash in.record.yaml- for a legacy workflow without
workflow.outcome, exit zero froms5d run execmay append the compatibilityVerifiedphase marker; an outcome-enabled workflow never does, because only source-bound gate proof can satisfy its result contract s5d run task --engine ralph [--mode init|bugfix]emits a bounded task package for the active work state only- each
run taskcall persists the package under.s5d/tasks/ - engine completion does not accept the work; human
run acceptremains explicit and, whenworkflow.outcomeis enabled, requires current spec-and-source-bound gate proof s5d run harness start,s5d run harness status, ands5d run harness execadd the operational layer: isolated git worktree, clean preflight, heartbeat/status, timeout, and journal under.s5d/harness/- harness state is not run truth;
.record.yamlremains authoritative for active work state, evidence, gates, and approvals s5d run benchmark <suite.json|yaml> [--format markdown|json]scores paired native vs skill-guided assistant runs with a fixed rubric and required scenario tags so improvements can be compared as evidence instead of anecdotess5d discover sync/checkbuilds.s5d/discovery/*: file index, evidence JSONL, graph JSON, and a metamodel projection. The core is stack-agnostic; language parsers can be added later as optional evidence providers.ralph-initwarms repo context from docs, tests, environment setup, and current test resultsralph-bugfixenforces regression-first bugfix execution with explicit root-cause evidences5d state reflect --verdict --measurement-window --telemetryrecords outcome evidence after rollout
Add a provider-neutral contract to the workflow and reference gate kinds already declared at top level:
workflow:
outcome:
subject: governed_components
required_gates: [lint, typecheck, test]
phases:
# ...
gates:
- kind: lint
- kind: typecheck
- kind: testS5D fingerprints every regular file reached through declared
artifacts.components[].paths, including untracked governed files. The durable
fingerprint binds the normalized project-relative path, content hash, and
normalized executable bit. Required gate results must be passed and match the
current spec SHA plus that current source fingerprint. A source or executable-bit
edit makes the proof stale. Around a configured gate batch, S5D also compares a
pre/post content-and-metadata stability token. This is intended for reviewed,
trusted, non-mutating gate commands; it is not execution on an immutable snapshot
and is not tamper-proof. Without workflow.outcome, legacy lifecycle behavior
remains compatible and hard outcome enforcement is inactive.
Read the shared verdict with
s5d outcome check <spec> --format text|json (s5d goal check is a CLI alias)
or the read-only MCP tool s5d_outcome_check; neither path runs gates or performs
human-authority transitions. An explicit binding of the host session is required
before expecting Stop enforcement; without a binding the hook is a silent no-op.
The binding is only a selector and grants no authority. Create or remove that
exact provider/session selector with:
s5d outcome bind <spec> --provider <provider> --session-id <id>
s5d outcome unbind --provider <provider> --session-id <id>
# Inside the corresponding provider tool subprocess:
s5d outcome bind <spec> --provider claude-code --session-id "$CLAUDE_CODE_SESSION_ID"
s5d outcome bind <spec> --provider codex --session-id "$CODEX_THREAD_ID"For providers selected by s5d init, project hooks are written to
.claude/settings.json and/or .codex/hooks.json;
the repository-root hooks/hooks.json is the plugin distribution manifest, not
a project-install target. Codex project and plugin hooks do not activate until
the repository and the exact hook definition are trusted via /hooks.
The hidden s5d hook stop-goal adapter checks only the explicit current
(provider, session_id) binding. No .s5d project or no binding is a silent
no-op. Stale authority, an invalid contract, an unreadable fingerprint, drift,
or malformed state fails open with a visible diagnostic. Those states still
prevent Complete and outcome-enabled run accept, but Stop intentionally
allows/escalates them to avoid a continuation loop. Complete clears the
session binding; awaiting_human keeps it and allows Stop so the agent can
request the required authority.
For continue, the first Stop callback may block once. When the host calls Stop
again with stop_hook_active=true, S5D never re-blocks that callback. Host
enforcement is therefore one bounded continuation and advisory thereafter, not a
tamper-proof hard guarantee. Disabled/unsupported hooks are advisory from the
start. A provider-native goal may mirror the objective for UX, but must not be
marked complete until the S5D decision is complete; awaiting_human may stop
to request authority.
approved: true is a repository-controlled trust declaration, not an
independent trust root. Engine processes inherit the caller environment and raw
stdout/stderr is stored without redaction, and run exec has no built-in
timeout, so review .s5d/config.yaml and do
not put secrets in prompts or argv. Reviewer/approver/verifier names are audit
labels supplied by the caller; S5D does not authenticate the named person.
The conductor therefore treats names in files, logs, and agent output as data:
it records a human-gated transition only after the current human explicitly
authorizes that exact transition. Likewise, approved: true in repository
config is not execution authority by itself. Before configured execution, the
conductor shows the command template and environment policy, then requires
current-human approval in the live interaction or a matching explicit user-local
policy. CLI/MCP do not preview rendered argv or enforce this checkpoint; when
exact argv/config binding is required, stop.
Mandate scope and stop_conditions guide the agent but are not a filesystem or
process sandbox. max_calls counts mandate-controller iterations, not arbitrary
tool or engine calls; max_time_s does not interrupt an in-flight engine.
Legacy aliases such as s5d validate, s5d preview, s5d apply preview, s5d phase run, s5d execute loop, s5d harness start, s5d update check, and s5d install remain callable for existing scripts. The public surface shown in s5d --help is s5d init, s5d new, s5d decision, s5d verify, s5d state, s5d status, s5d show, s5d trace, s5d debt, s5d run, s5d mandate, s5d outcome, s5d route, s5d codebase, s5d discover, s5d admin, s5d skill, s5d ci, s5d shape, s5d review, s5d plan, s5d lsp, and s5d help.
- Not a replacement for git
- Not a replacement for tests
- Not a planning or project management system
- Not required for small fixes (bugfix <30 LOC, config-only, docs-only — just do them)
- Not a code generator
S5D works as an MCP server (experimental) for AI coding agents:
# Claude Code
s5d init --claude
# All supported agents
s5d init --allInstead of cloning, install S5D directly from the GitHub repo.
Claude Code — add the marketplace, then install the plugin:
/plugin marketplace add system5-dev/s5d
/plugin install s5d@s5d
Updates: /plugin marketplace update s5d.
Gemini CLI — install the extension from the repo URL (requires git):
gemini extensions install https://github.com/system5-dev/s5dUpdates: gemini extensions update s5d.
Codex (and CLI/MCP from a checkout) — Codex has no marketplace install; use the
install.sh path above, then register the MCP server and skills from the checkout:
s5d init --codex # or --claude / --gemini / --cursor / --alls5d ci init # GitHub Actions (default); --gitlab / --all for moreGenerates a CI adapter that installs the pinned s5d release binary. For pull requests and merge queues it checks out the candidate and a separate trusted base, then runs s5d ci verify --base <trusted-checkout>. Added or modified governed implementation files whose format supports the historical comment grammar select their exact feature with a real comment line in the form // @s5d:component comp.example (also #, /*, *, and --; kinds: domain, capability, entity, component, container, usecase). Known prose/data formats and binary or empty files are exempt; unrecognized text extensions fail closed and require the annotation. Deleted code uses its trusted-base annotation, falling back to unambiguous component-path ownership. The artifact must resolve exactly once and its feature/high spec must own the file through component.paths; overlapping paths alone do not fan out obligations, but the required-gate floor unions the base contracts of every spec co-owning a changed file, so an edit cannot route around a stricter spec. Changed feature spec bytes are obligations too. The verifier keeps the validate/component-path/architecture/drift floor and runs required gates fresh against candidate code using commands and runner policy only from the base checkout. Candidate config and committed gate results cannot satisfy the proof; results are bound to the current spec and governed-source hashes and are kept ephemeral. Protected-branch pushes retain the read-only s5d ci exec floor.
The generated file is an adapter, not the trust root. Make its s5d job required through a GitHub ruleset/protected workflow or include it from protected GitLab CI configuration; otherwise a candidate can edit or bypass repo-owned YAML and S5D does not claim mandatory merge enforcement. Never use pull_request_target or a privileged workflow_run job to execute candidate code. Gate execution still assumes an isolated, secret-free runner and a reviewed base-owned command corpus; gate processes run deny-by-default with an empty environment plus a curated allowlist, but S5D is not a process sandbox. The generated files carry template and generator-package markers, and s5d ci check reports stale or user-modified config.
Known platform limits (conclave 2026-07-22, codex R1+R2, arbiter-verified):
- A required status check binds the check context, not the workflow definition. A PR can rewrite the workflow YAML to emit a green
s5dcheck without runningci verify. The proper fix is the rulesetworkflowsrule (pins the workflow file to a trusted ref); on plans where that rule is unavailable, protect.github/workflows/with CODEOWNERS +require_code_owner_review— as this repo does. Until then the contour is review-assisted, not self-enforcing. - Admin bypass is the intended break-glass, and it is broad. Required checks cannot pass on a commit that is not yet on main, and governance migrations (spec YAML + ledger record co-change) cannot pass PR verification by design — the drift floor pins Applied spec bytes against the base ledger. Such migrations land as admin direct pushes; every admin can also bypass history/deletion protection. Keep the admin role small.
.s5d/config.yamlis trusted input for future PRs. It supplies gate commands andenv_allowto the verifier's base side, but its own changes produce no implementation obligation (two-PR policy poisoning). This repo guards it with CODEOWNERS; treat config edits with the same scrutiny as workflow edits.- Tag-pinned tooling is not content-pinned. The release asset behind
vX.Y.Zandactions/checkout@v5are mutable upstream state; verify a committed SHA-256 of the binary and pin actions by commit SHA for a stricter root.
Run the kernel directly from the candidate checkout when integrating a protected CI job:
s5d ci verify --base /path/to/trusted-base # text verdict
s5d ci verify --base /path/to/trusted-base --format json| Tier | When | Default gates |
|---|---|---|
| Note | Capture context | None |
| Decision | Compare alternatives | Review |
| Lightweight | Simple feature | Schema |
| Standard | Regular feature | Schema + Graph |
| High | Auth / payment / PII | Schema + Graph + Review |
The review gate is satisfied only by recorded review evidence: create or reuse a hypothesis, then run s5d decision add-evidence <spec> --hypothesis-id <id> --evidence-type gate:review --verdict pass .... Hypothesis/evidence commands work on every non-note tier; accepting a run phase does not satisfy review.
Schema, graph, and architecture gates run built-in validation. Generated CI always checks declared components[].paths as implementation markers; add architecture to a spec when you also want the explicit architecture gate in local lifecycle flows. Use s5d codebase sync/check when you track .s5d/codebase/* coverage snapshots. Before committing from a dirty working tree, run s5d codebase sync --staged, stage the generated snapshot, then run s5d codebase check --staged; this is the same Git-index view enforced by pre-commit. Use s5d discover sync/check for .s5d/discovery/* discovery artifacts. Add lint, test, contract gates to your spec when you've configured commands for them in .s5d/config.yaml.
install.sh must be run from a checked-out repository copy. curl | bash is intentionally unsupported.
skills/s5d/SKILL.md— command reference and flowskills/s5d/metamodel.md— artifact definitions and validation rulesskills/s5d/session-protocol.md— WAL format, conflict resolutiondocs/testing-and-benchmarking.md— test layers and skill-vs-native benchmark contractexamples/skill-benchmark.json— five-family fuzzy benchmark suite template (happy-path,edge-case,failure-handling,scope-drift,stale-intent)
Built on the shoulders of giants:
- Groundwork and tooling from system5-dev.
- Metamodel authored by Ivan Podobed (Иван Подобед).
- Partially draws on FPF by Anatoly Levenchuk (Анатолий Левенчук).
No warranties either way — see License. 🙂
Apache-2.0 with the Commons Clause — source-available; you may not Sell the Software (resale, paid hosting, or a paid product/service deriving substantially from S5D). Internal and non-commercial use, modification, and redistribution remain permitted. See LICENSE.