Agent Skill Pattern is repository-local terminology for composing established agent mechanisms:
- a small, progressively disclosed Agent Skill recognizes a bounded subtask;
- the parent delegates that subtask to a separate, usually cheaper custom agent;
- the worker receives a narrow tool surface and isolated context;
- the worker writes the artifact directly and returns compact status.
The name describes this wiring, not a novel primitive. See
docs/agent-skill-pattern.md for the detailed
architecture and docs/research/prior-art.md for the
evidence and terminology review.
Use the pattern only for a repeatable, well-bounded task that does not need the parent's full history and can be evaluated independently. Delegation can isolate working context and move work to a cheaper model, but it adds routing, tool, latency, and reliability costs. Model price alone is not evidence that the complete workflow is cheaper or good enough.
| Role | Responsibility |
|---|---|
| Harness | Exposes Skills and agents and performs delegation. |
| Parent | Owns the main task and decides whether the Skill applies. |
| Skill | Provides progressive discovery and routing instructions only. |
| Worker | Performs the bounded subtask with minimal tools, writes the artifact, and returns compact status. |
Editable architecture diagrams and rendered PNGs are in
docs/diagrams/.
| Use case | Components | Status |
|---|---|---|
| ASCII art | ascii-art Skill, agent, notes |
Implemented; controlled benchmark was faster but costlier and lower quality. |
| Semantic test corpus | semantic-test-corpus Skill, agents, confined MCP, notes |
Implemented; controlled benchmark saved AI credits and parent context but lost held-out quality and reliability. |
| Release-note synthesis | release-note-synthesis Skill, agent, two-tool MCP |
Feasibility only; runtime wiring failed before semantic quality could be tested. |
| Action-item extraction | retained v3 Skill and agent | Feasibility only; context isolation worked, but held-out tuple quality and grounding failed. |
| Feature documentation | feature-documentation Skill, fixed-Haiku agent, excluded-pilot report |
Excluded pilot NO-GO; delegated docs and treatment adherence failed, so main was not authorized. |
| Feature documentation v2 | feature-documentation-sonnet-v2 Skill, fixed-Sonnet agent, preregistration |
Preregistered design only; mandatory Skill routing, six-pair excluded pilot, 24-pair held-out main, zero observations. |
| Unit-test authoring | v1 result, Sonnet 4.6 v2 result | Both excluded pilots stopped on frozen integrity boundaries and returned NO-GO; held-out main remains forbidden. |
The compact experiment index is the canonical cross-study summary. Each study retains one concise protocol and one canonical report.
- Semantic corpus: the target GPT-parent-to-Haiku-worker arm used 38.8% fewer combined AI credits and 58.0% less parent cumulative input than GPT inline, but used 88.0% more total model tokens, took 72.0% longer, missed held-out path and mutant-quality floors, and was treatment-adherent in only 1/12 units.
- ASCII: treatment finished 54.9% faster among 20 complete pairs, but used 69.7% more combined AI credits and 53.9% more parent cumulative input, with deterministic pass 25 percentage points lower and blinded overall quality 0.883 points lower.
- Release notes: the Skill and worker were reached, but canonical MCP tools never executed and no draft artifact was produced. Semantic quality was never tested.
- Action items: the v3 mechanics isolated transcript/file work from the parent, but the three excluded pilots averaged 0.462 tuple F1 and failed the 100% grounding requirement.
- Feature documentation: the two-block excluded pilot used 17.4% fewer combined credits but 52.0% more parent cumulative input and 57.3% more total model tokens. Both delegated guides scored zero and both treatment runs failed routing/adherence, producing a frozen NO-GO; the main study was not run.
Quality was assessed by repository-owned deterministic external evaluators and, for ASCII, blinded artifact judging. The parent model did not grade its own output. Release-note and action-item runs were excluded feasibility probes without valid control arms; they are not controlled evidence. No main study followed any probe.
AI credits are the runtime's available usage measure, not a portable dollar cost. Combined values include parent and worker usage and exclude evaluator/judge usage unless explicitly stated.
| Path | Purpose |
|---|---|
.github/skills/ and .github/agents/ |
Live routing Skills and custom agents |
tools/ and tests/ |
Confined MCP implementations and tests |
docs/ |
Pattern, diagrams, prior art, and implementation notes |
experiments/ |
Protocols, canonical results, and retained reproducibility source |
The repository intentionally excludes raw event streams, copied runtime payloads, per-run evidence bundles, generated dashboards, and obsolete one-shot harnesses. Negative findings, limitations, preregistered dispositions, and the Git history that contains the original evidence packages remain explicit.
The pattern remains a research pattern, not a validated optimization. The controlled studies did not establish a positive quality-adjusted efficiency result, the feasibility probes did not authorize confirmation, and no main study was run.
Released under the MIT License.