Cursor Harness v0.2.0 is a dependency-free Cursor plugin and benchmark for correctness-first software-engineering workflows. It coordinates scoped implementation, deterministic verification, read-only review, security controls, and evidence-based release decisions.
The project is source-available under the MIT License. It is not currently claimed to be published in Cursor Marketplace or proven by a live authenticated benchmark.
Cursor Harness is a port of the
agentic-os project to Cursor's
plugin format. The skill library under plugin/skills/ is carried over from
agentic-os largely verbatim — most reference files are byte-identical, and the
edits are frontmatter and cross-link adaptations for Cursor. The agents,
commands, Cursor plugin manifest, hooks, validator, benchmark, and release
tooling in this repository are separate work; they share no code with
agentic-os's equivalents. Both projects are MIT-licensed under the same owner.
Objective results are ranked in this order:
- correct trials;
- matched-correct speedup;
- objective quality-contract pass rate;
- non-critical security outcomes;
- resource use.
Correctness and integrity are non-compensable. Installation lifecycle, harness-on correctness non-regression, no off-pass/on-fail fixture regression, the 80% trial and fixture floors, critical-security controls, and telemetry/evaluator/workspace integrity must all pass. Speed, style, or resource savings cannot offset a failed gate. See the evidence policy.
The consumer plugin in plugin/ contains:
-
seventeen agents: `engineer`,
rust-engineer,godot-engineer,phaser-engineer,devops-engineer, andweb3-engineerfor dispatched implementation;data-model-documenterfor cataloging data contracts;code-reviewerandsecurity-reviewerfor read-only review;data-model-verifier,library-investigator,library-reviewer, andadversarial-claims-reviewerfor evidence-only audit;technical-pm,marketer, andgame-design-shaperfor domain shaping; andcapability-probefor capability detection; -
forty-six skills: `prompt-shaping`,
planning-and-task-breakdown,openspec-planning,incremental-implementation,code-review-and-quality, andsecurity-engineeringfor the core shape/plan/implement/review path;data-model-documentation,data-model-verification, andskill-library-reviewfor catalog and library audit method;findings-ledger,session-state, andmemory-extractionfor finding recurrence and durable state;typescript-testing-backend,typescript-testing-frontend,typescript-data-engineering,typescript-analytics, andbrowser-testing-with-devtoolsfor TypeScript and browser work;rust-engineer,godot-engineer, andphaser-engineerfor the Rust and game stacks;deployment-pipelines,release-manager, andsite-reliability-engineeringfor pipelines, release, and runtime ops;library-investigatorandadversarial-claims-reviewerfor auditing this plugin's own surfaces and documents that make formal claims;devops-engineerandweb3-smart-contract-engineeringfor platform and EVM work;game-design-shaper,game-concept-creator,game-systems-designer,game-balancer,game-monetization-strategist,iap-manager, andgame-marketerfor game design and monetization ops;marketing-shaper,content-ops,content-pipeline,conversion-ops,growth-engine,outbound-engine,seo-ops,revenue-intelligence,autoresearch,codebase-cost-estimator,security, andtelemetryfor marketing, growth, measurement, and sanitization; -
eleven commands: `/review-gate`,
/triage-findings,/state,/skill-new,/agent-new,/audit-library,/eval-harness,/openspec-propose,/openspec-apply,/openspec-archive, and/verify-ci-parity; - always-applied factual-correctness, grounding, communication, memory-discipline, orchestration, review-tier, actual-diff, anti-patterns, and briefing rules; and
- nine hook events backed by dependency-free Node scripts: a fail-closed
beforeShellExecutionguard; advisorysessionStart/preCompact/stop(and related) hooks that fail open; plus an opt-in dispatch-gate layer (preToolUse/beforeReadFilefail-closed when enabled — ships disabled; see docs/dispatch-enforcement.md).
The core path is: shape an unclear request, plan an explicit dependency graph, implement in tested increments, run deterministic local verification, dispatch read-only code and security reviews in parallel, then ship only after Tier 0 and evidence-backed Tier 1 findings are addressed. Tier 2 judgment remains advisory and is tracked for recurrence.
The intended v0.2.0 scope is local Cursor Editor and Cursor CLI operation.
Repository contracts verify plugin layout, component metadata, hook behavior,
benchmark isolation, and CLI feature detection. Actual Editor loading and
authenticated CLI outcomes remain manual or unverified. npm run plugin:editor:verify and npm run plugin:cli:verify are the operator-run
scripts that capture evidence for them; neither settles a claim on its own, and
plugin loading verification documents
their options, exit codes, and the limits of what they establish.
Cursor Cloud Agents are excluded: plugin loading, hook enforcement, filesystem semantics, and telemetry parity have not been verified in Cloud. The benchmark does not treat local results as Cloud evidence. See the capability matrix and threat model.
The repository follows Cursor's Marketplace repository and nested plugin manifest layout. Once a reviewed source is available through Cursor, the installation concept is Cursor Settings → Customize → Plugins, select the source, review its contents and permissions, then install. This describes the supported concept; it does not claim that Cursor Harness is currently listed or published.
For local development from a git checkout, a user may manually create a
symlink under ~/.cursor/plugins/local. Run it from the repository root; the
target is the plugin/ subdirectory, not the repository root:
mkdir -p ~/.cursor/plugins/local
ln -s "$PWD/plugin" ~/.cursor/plugins/local/cursor-harnessThat command is wrong for a release archive, where the archive root itself
is the plugin and no plugin/ subdirectory exists — linking "$PWD/plugin"
there produces a dangling symlink. Archive users follow
Installation in the plugin README, which ships
inside the archive and links "$PWD" from the extracted directory instead.
Either way, confirm the link resolved with:
ls ~/.cursor/plugins/local/cursor-harness/.cursor-plugin/plugin.jsonThis repository never creates that symlink or writes ~/.cursor. Remove it
manually with rm ~/.cursor/plugins/local/cursor-harness. Restart Cursor after
local plugin changes if discovery does not refresh.
Experienced Cursor users often already have global agents, rules, and hooks. Symlink install does not replace that tree; it only adds a plugin path. Predict collisions before you rely on a name:
- Agent and rule names may collide. Seventeen plugin agents total; all
except
capability-probemay collide with names under a user's global~/.cursor/agents/. The same question applies tofactual-correctness.mdcamong the rules. UNVERIFIED: this repository has no proven Cursor precedence rule for plugin vs global agents/rules (shadow, reverse, or both listed). Do not assume one. Confirm which definition runs — prefer invokingcapability-probe(unique sentinelcursor-harness-agent-discovered) over trusting a colliding name. - Hooks stack. Plugin hooks add alongside any existing
~/.cursorhooks for the same events; they do not replace them. The plugin registersbeforeShellExecutionwithfailClosed: trueand a 5s timeout — if that hook errors or times out, it can gate every shell command.sessionStartinjectors can therefore run twice (yours plus the plugin's). Dispatch enforcement hooks are registered but disabled by default (enabled: false); enabling is an operator choice — dispatch-enforcement.md. - On-disk symlink ≠ Editor loaded the plugin. The documented
plugins/localsymlink does not writeplugins.json. After install,npm run plugin:editor:verifytypically reportsregisteredInPluginsJson: falseandeditorComponentLoading: not-proven(exit3) until you invokecapability-probe, save the transcript, and re-run with--transcript. Details: plugin loading verification.
If the installed CLI does not expose --plugin-dir, authenticated benchmark
runs can use the project-overlay adapter. It copies agents, rules, and skills
into an isolated trial workspace. It intentionally omits executable hooks
because overlay hook path semantics are not proven safe, so overlay results
are not evidence of live plugin loading or hook enforcement.
Authenticated runs do not accept API keys through environment variables or arguments. They require an absolute path to a protected, pre-authenticated Cursor CLI config-template directory outside every trial workspace. The runner validates and copies that template into a fresh per-trial config home.
Plugins and workspace hooks execute local code with the Cursor process's user permissions. Install only from a reviewed revision, inspect the manifest and hook script, and do not treat Marketplace review as a security boundary. The guard is a default-deny allowlist over safe command forms, not a sandbox, endpoint-protection product, or substitute for backups and least privilege. Report vulnerabilities privately through GitHub's advisory form, not a public issue — see Reporting a vulnerability.
Node.js 22 is required. The lockfile has zero third-party dependencies.
External prerequisite for the planning workflow: the openspec CLI
(@fission-ai/openspec, requires Node >= 20.19) is a global prerequisite like
git or gh — install with npm install -g @fission-ai/openspec. It is never
a package.json dependency here or in consumer repos. Engineering planning
artifacts live under openspec/ in the consuming repository and are validated
with openspec validate <change-id> --strict; see
openspec-planning. The maintainer
checks below do not require it. The CLI collects anonymous usage stats with a
printed notice on every invocation — opt out with export OPENSPEC_TELEMETRY=0.
npm ci --ignore-scripts
npm test
npm run validate
npm run plugin:lifecycle:verify
npm run benchmark:corpus-smoke
npm run probenpm test includes release reproducibility and temporary install/uninstall
coverage. npm run plugin:lifecycle:verify performs clean install,
idempotence/repair, and removal against a temporary Cursor root and emits
machine-readable evidence. npm run validate checks manifests, component
discovery and frontmatter, links, inventory hashes, orchestration and
authenticated-workflow wiring, schemas, and hook safety. The corpus smoke
executes 24 deterministic mock trials without model
calls. npm run probe is network-free but reports machine-specific local
Editor/CLI capability evidence.
See Benchmark methodology for fixture contracts, paired execution, profiles, evidence, scoring, and reproduction. Common commands are:
npm run benchmark:corpus-smoke
npm run plugin:lifecycle:verify
npm run benchmark:smoke:authenticated -- --cursor-config-template /protected/cursor-config
npm run benchmark:release:authenticated -- --cursor-config-template /protected/cursor-config
npm run benchmark:report -- <benchmark-manifest.json> <records.ndjson>
npm run benchmark:export -- --run-root <raw-run> --export-root <sanitized-export>
npm run release:test
npm run release:dry-runAuthenticated profiles make paid model calls and require
--cursor-config-template <absolute-protected-path>. API-key environment and
argument authentication is intentionally unsupported because evaluated tool
subprocess inheritance has not been proven safe. CI uploads only the sanitized
allowlisted export after canary and credential-pattern scanning; raw run roots,
workspaces, and Cursor config homes are never upload paths. smoke-24 is
integration evidence only; release-72 repeats the 12-fixture paired corpus
three times but does not establish generality.
Release packaging validates first, checks all v0.2.0 metadata, and emits a
deterministic .tar.gz, SHA-256 file, and machine-readable release manifest.
It packages only the plugin consumer payload and license; it does not publish,
tag, submit to Marketplace, or create a GitHub release.
- manual Editor/Customize installation and component loading;
- live authenticated smoke or release benchmark outcomes;
- live plugin loading through the locally installed CLI;
- token counts in CLI structured output; and
- subagent parentage, call counts, and concurrency telemetry.
These stay explicitly unavailable until direct evidence is captured. Missing telemetry is never scored as zero. The first and third are the claims plugin loading verification exists to capture evidence for.