Skip to content

Upstream Research

github-actions[bot] edited this page Sep 30, 2026 · 1 revision

Upstream benchmark

Observed on 2026-09-12. This is a static comparison of public primary sources, not a production certification. The source snapshots were inspected without executing their scripts or installing their skills into a consumer.

Discovery

The cited skills.sh listings were checked directly. Additional indexed searches covered skills-merger, skill-merger, merge-skills, merge skills, skill-creator eval, skill evaluation, skill-evaluator, brainstorming, and writing-skills.

The published Skills CLI version 1.5.26 was also used for two queries, with npm lifecycle scripts disabled: skills find 'skill creator' and skills find 'skill merger'. The creator query returned Anthropic and other authoring candidates. The merger query returned unrelated memory, code, spreadsheet, specification, and model merging packages; those tasks do not match skill synthesis. No suitable dedicated skill merger was found within this search budget. This is limited search evidence, not proof that none exists.

Install counts, badges, and search order were discovery signals only. No package was labeled production approved based on those signals.

Pinned sources and decisions

Source ID Listing / primary package Commit Scoped license Decision
vercel-find-skills skills.sh, package d667282815248da03a08a18272b5d2eef9caf77c MIT, root Reuse discovery ideas; author a bounded qualification procedure
anthropic-skill-creator skills.sh, package 34040c9c568585f6929bedeaad110ad08f079624 Apache-2.0, package override Reuse authoring and evaluation principles; implement portable helpers independently
superpowers-brainstorming skills.sh, package b36e0829c6d0140e93cfef2ca599b1b07d4a7797 MIT, root Reuse design and decomposition ideas, scoped to skill design
superpowers-writing-skills skills.sh, package b36e0829c6d0140e93cfef2ca599b1b07d4a7797 MIT, root Reuse baseline and pressure-case ideas
skillport-skill-evaluator skills.sh, experimental package 51334ae94919fc1c261673ffefdce9176c9094ef MIT, root Secondary rubric reference; reject as a production-approved base because it declares WIP

The source lock records all 41 files across these five packages, full file inventories and hashes, applicable license hashes, and local consumers. Downloaded bytes were checked against the pinned Git blob identities before SHA-256 capture. Root licenses outside a package are separate evidence; the package digest includes only its files. A hash records byte identity and does not imply that tests ran.

The machine algorithm is sha256. Hash each file's exact raw bytes without newline or encoding normalization. Sort package-relative POSIX paths by UTF-8 byte order; concatenate each path, NUL (0x00), its lowercase hexadecimal digest, and newline (0x0A), then hash those combined bytes for package_sha256. Directory entries and file modes are excluded. Licenses outside the package have their own repository-relative path and digest. The offline validator recomputes the aggregate from recorded digests; source-byte verification occurred during capture and requires a fresh fetch to repeat.

Contribution decisions

Local consumer Useful idea Adaptation or rejection Local evidence/consumer
Skills discovery Task/synonym discovery and primary source links Replace installation-oriented output with a qualification report and explicit rights/security review Source review
Creator Intent, progressive disclosure, realistic outputs Keep orchestration separate from specialist procedures and require a self-contained license Authoring guide
Evaluator Baselines, observable grading, comparison, iteration Use runtime-neutral execution and separate untouched acceptance cases Evaluation rubric
Skill design Clarify outcomes, compare approaches, reduce scope Preserve prior authorization; remove mandatory provider dispatch and unrelated global activation Decision guide
Synthesis Combine traceable source contributions Original synthesis procedure; reject concatenation, duplicate scripts, unrelated responsibilities, and automatic adoption Synthesis guide

The package instructions and helpers are independently authored implementations of selected methodological ideas. No upstream code, UI server, complete prompt, or asset is vendored. Apache-2.0 licenses the original work; a future actual copy/adaptation must preserve its source's required license and notices rather than inheriting this statement automatically.

Naming basis

Searches on 2026-09-12 covered skill naming, skill-naming, and naming-conventions on skills.sh. The general naming listing led to a pinned entrypoint oriented toward brands, products, characters, places, and titles. Its entrypoint scope was inspected for discovery; its full resource tree was not qualified and no material was adopted. Other surfaced platform-specific naming results did not establish a reusable skill-identifier specialist.

skill-naming is an original, narrower procedure based on the Agent Skills name constraints and the requested domain-affinity/cardinality convention. It returns one naming decision and scoped collision evidence. The five benchmark packages above remain unchanged; discovery-only candidates are not promoted to benchmark locks or production approvals.

Findings that affected the design

Anthropic's evaluation helpers invoke a specific CLI and create provider-specific command files; these are unsuitable as the portable execution layer. Its archive helper uses an exclusion list without the explicit path/symlink boundary required here. The new helper performs only scaffolding and validation; it does not execute models or package arbitrary source archives. Its evaluation loop uses a held-out score for selecting iterations, so the local rubric reserves an additional untouched acceptance set before making generalization claims. These observations are scoped to the inspected scripts.

Superpowers' brainstorm server tests support that optional UI implementation, not universal design quality. The UI server and provider dispatch rules are not adopted. Its writing-skills testing ideas inform the local evaluator, while the Agent Skills specification governs format conflicts: descriptions explain what and when, metadata is a string mapping, and provider tool allowlists are not a portability prerequisite.

Skillport offers rubric and example material but explicitly marks the evaluator experimental and uses model-family-specific evaluation assumptions. It remains a secondary reference; its labels and examples do not establish production readiness for this collection.

Future evolution

Preserve this baseline until a reviewed change intentionally replaces it. A future comparison can inspect upstream file additions, removals, and modifications against upstreams.lock.json, map them to contribution decisions, and propose improvements with evidence. Do not update the lock solely to silence a changed hash, automatically overwrite local work, or treat a newer commit as a quality verdict.

Agent Skills guidance integration (2026-09-19)

The authoring, design, evaluation, evidence-collection and optimization packages now include originally written guidance influenced by agentskills/agentskills revision 69ef37e9424c0a7ea9dd2293b559e43ec8176379. Reviewed documentation under docs/skill-creation/ is CC-BY-4.0; each consumer carries its own immutable reviewed-file and attribution record. Integration covers sanitized corrections, control proportional to fragility, noninteractive script interfaces, portable evaluation cases and description activation errors. Self-contained routes and untouched final acceptance remain mandatory. No upstream scripts or setup commands are adopted. Structural validation establishes no measured model gain.

Clone this wiki locally