Releases: MazenAbbas/groundspec
Release list
groundspec v0.3.0rc2
Release candidate. Prerelease only: not published to PyPI, not a stable release.
Guardrail fixes for weaknesses that real user trials of v0.3.0rc1 exposed. No change to the deterministic core, the Domain Pack SDK, or any pack's rules.
Fixed
- Budgets and other user-set limits are boundaries: an agent must not raise or relax one without the user's yes. A stated limit is checked for feasibility, recorded up front, and never justifies reordering the contract steps.
- Package installs and downloads are explicit external actions that need a yes, with an ordered response to a broken environment (use what is available, ask once with the exact command, or report INCOMPLETE).
- The contract text must be reconciled with what actually happened before evaluating.
groundspec evaluateprints an advisory note for cited legal or regulatory claims whose source is more than 24 months older than its access date, or undated. It does not change the completion state.groundspec creategains--tool-call-budgetand--time-budget-minutes.- A check that could not run is omitted (INCOMPLETE), not recorded false (FAIL).
Verification
- 350 tests passed, 3 skipped (Windows-only limits); ruff and strict mypy clean; CI green on Ubuntu, Windows and macOS with Python 3.11-3.13.
- Three maintainer trials and six independent subagent runs, each re-verified against the produced files: docs/forward-tests-v0.3.0rc2.md.
Known limitations
- Prevention of an unconfirmed install or limit change is model behaviour; detection is deterministic. Six small samples only.
- Not yet exercised live: reconciling stale criterion text after an authorization change, and an agent using the new budget flags.
See CHANGELOG.md. Verify downloads with SHA256SUMS.txt.
groundspec v0.3.0rc1
Release candidate for the Domain Pack SDK. Prerelease only: not published to PyPI, not a stable release.
What's new
- Domain Pack SDK: versioned
pack.tomlmanifests, deterministic discovery (official / project / user), a resolver with dependency closure, version-range and conflict checks, canonical reproduciblegroundspec.lock, a directory-content security scanner, and completion gates that can only downgrade a completion state. - New CLI:
groundspec pack list | inspect | init | validate | resolve | test | lock. - Two new official packs:
product-managementanddata-science-ml(no heavy ML dependencies). software,researchandcontentare wrapped viapack.tomlwith their rule files unchanged; existing contracts and rule IDs keep working.- Meta-Skill routing updated. Natural-language pack selection stays model-dependent; only
pack resolveis deterministic.
Verification
- 330 tests passed, 3 skipped (Windows-only limitations); ruff and strict mypy clean.
- CI green on Ubuntu, Windows and macOS with Python 3.11-3.13.
- Four independent forward tests, each re-verified against the produced files: see docs/forward-tests-v0.3.0rc1.md.
Known limitations
resolve()loads every discoverable pack up front (about 385 ms with five packs, measured on one machine).- The lock file is a reproducibility record, not yet an enforcement mechanism.
- Forward tests: one run per scenario, same model family as the author; not a statistically powered evaluation.
See CHANGELOG.md for details. Verify downloads with SHA256SUMS.txt.
v0.2.0rc3 -- Evidence-integrity fix: claim ledger for material factual claims
Second evidence-integrity regression-fix release. A second independent user test of the exported Meta-Skill (a Riyadh university-student food-delivery PRD) reported an unconditional PASS despite unsupported market/competitor claims with no evidence record, a pilot city silently promoted from illustrative anchor to confirmed scope, a conditional ZATCA VAT rule generalized past its actual conditions, a secondary source treated as official regulatory confirmation, and model-generated thresholds that could be mistaken for sourced benchmarks. Full account and root-cause table: CHANGELOG.md's [0.2.0rc3] entry.
What's fixed
- Contract schema 0.4.0 (additive;
0.1.0-0.3.0untouched):status.claim_ledger, a structured, citable provenance record for material factual claims, with a 7-label evidence taxonomy (MEASURED/PRIMARY_SOURCE_VERIFIED/SECONDARY_SOURCE_SUPPORTED/USER_CONFIRMED/MODEL_EVALUATED/ASSUMPTION/RESEARCH_NEEDED) schema-tied toauthority_leveland citation completeness. A secondary source can never be recorded as primary/official confirmation -- this is now a validation error, not a style note. groundspec.metaskill.completiongainshas_unmapped_material_claims(a material claim with no ledger entry forcesINCOMPLETE) andclaim_ledger_has_disclosed_material_limitations(a material claim resting on disclosed secondary support, model judgment, an assumption, or a bounded research gap forcesPASS_WITH_CAVEATS, never a barePASS).- Meta-Skill content, docs (
architecture.md,schema-reference.md,threat-model.md), and a new regression scenario (tests/regression/food_delivery_riyadh_v0_2_0rc2/) updated to match.
Testing
- 236 tests passing (35 new), ruff/mypy clean.
- Eight independent forward tests, covering all 3 domain packs and all 7 risk overlays this project defines, each independently re-verified against its actual produced files (including personally re-running one scenario's automated test suite, not just reading its reported output). Full writeup and coverage matrix:
docs/forward-tests-v0.2.0rc3.md. - Two negative controls proving the completion-state gates are load-bearing, not decorative.
- Two testing-harness incidents (a stray file write outside the assigned test workspace, both from the same working-directory-reset cause) documented in full rather than omitted; both independently confirmed harmless.
Install
pip install groundspec-0.2.0rc3-py3-none-any.whlNo PyPI publish with this release; install the wheel or sdist directly. SHA256SUMS.txt included for verification.
v0.2.0rc2 -- regression fixes from an independent user test
This release fixes real behavioral defects an independent user test found in v0.2.0rc1's Meta-Skill, and adds hardening found necessary by testing the fix itself. v0.2.0rc1 is unchanged and remains published at its own tag/release.
The defect report
Guided-mode PRD scoping (Saudi university-student food-delivery app) with the explicit instruction "clearly separate verified facts, assumptions, and items requiring external research; do not publish anything or perform any external action":
- Zero clarification questions asked while silently defaulting material product decisions (pilot city, business model, persona, payments, language, monetization).
- "Do not perform any external action" did not stop ten web searches.
- A model self-review was recorded as a verified fact for a criterion admittedly not fully checked.
- "No source exists" was asserted from one bounded search.
- Row/label counts were treated as proof of semantic correctness.
- Model self-review recorded as independently verified.
Completion state: PASSreported despite an unresolved critical risk, weak-evidence 'must' criteria, and the authorization violation.
Full account: tests/regression/food_delivery_v0_2_0rc1/README.md.
What's fixed
- Contract schema 0.3.0 (additive;
0.1.0/0.2.0untouched, still fully supported):status.verified_facts.evidence_labelnarrowed to labels meaning real verification happened (VERIFIED,MEASURED,SOURCE_VERIFIED,USER_CONFIRMED,HUMAN-REVIEWED) -- a self-review can no longer be schema-valid there.status.unverified_claimsgains an optionalevidence_label(MODEL-EVALUATED,ASSUMPTION,RESEARCH_NEEDED,PROPOSED,PENDING_EXTERNAL_VALIDATION,OUT_OF_SCOPE).scope.assumptions.safe_defaultis nowconst: true.status.residual_risksgainsaffects_deliverable_validity(defaulttrue). groundspec.metaskill.completion.derive_completion_stategainsauthorization_boundary_violatedandunresolved_critical_risk_to_validity(both forceFAIL), plushas_deferred_high_value_items(forcesPASS_WITH_CAVEATS) -- the last one found necessary by forward-testing the fix itself (see below).evaluate_acceptance_criterianow flags a 'must' criterion as evidence-inadequate when itsverification_methoddemands real evidence but only a weak label was supplied.- Meta-Skill content rewritten: network access/search/fetch explicitly named as external actions with no exception; a concrete checklist of near-universally-material PRD-scoping dimensions; exact field-name reference; verification-method evidence-strength distinctions; domain-pack-fit guidance; Audit-mode hard-constraint scope;
contract-onlyvsplan-and-executeclarification. - A maintained regression scenario (
tests/regression/food_delivery_v0_2_0rc1/) converts the exact reported test into fixtures and deterministic assertions. - A real pre-existing inconsistency in
examples/metaskill/01-food-delivery-ksa-university-prd(silently defaulted geographic scope, contradicting this same release's own new checklist) was found and fixed.
Independent forward testing
Four fresh, isolated subagents (no memory of the engineering work) exercised the corrected Skill end to end: the food-delivery scenario with all external actions prohibited, the same with read-only research explicitly permitted, a small unambiguous software task, and an audit task with an embedded prompt-injection attempt. All four behaved correctly on the original defects -- zero unauthorized network access, correct source-vs-snippet evidence discipline (including catching a real cross-source contradiction), the injection fully disclosed and not complied with, proportionate ceremony on the trivial task. Every claim was independently re-verified against each agent's actual produced files, not accepted on its word; two of the four runs' own evidence, re-evaluated after the has_deferred_high_value_items fix landed, concretely flip from PASS to PASS_WITH_CAVEATS on the same real data. Full writeup: docs/forward-tests-v0.2.0rc2.md.
Verified locally and in CI
Full test suite (201 passed, 1 skipped, documented reason), ruff, mypy --strict, wheel+sdist build, twine check, forbidden-artifact scan, clean install of the exact tagged wheel (pip check, doctor, both skill export/skill validate targets, create/validate/audit) -- all green, including in GitHub Actions across Ubuntu/Windows/macOS x Python 3.11/3.12/3.13 on this exact tag (CI run). Landed via PR #2, merged only after its own full CI matrix was green.
What still relies on model judgment (honestly)
Mode/intent detection, exactly which items are material enough to ask about beyond the documented checklist, whether a source actually supports a claim, whether a self-review is trustworthy. This release makes the dishonest version of these judgment calls mechanically harder to produce -- it does not make the underlying judgment calls infallible.
Known limitations (not yet done)
Not published to PyPI. Five live Meta-Skill runs across two releases is not a statistically powered evaluation. The 30-scenario deterministic-engine corpus and its 3-arm baseline comparison remain PENDING. External user validation remains PENDING. Only three domain packs ship. The Meta-Skill's export safety checks a symlink at the destination and its immediate parent, not the full ancestor chain.
Checksums
See attached SHA256SUMS.txt.
v0.2.0rc1 -- the Groundspec Meta-Skill
Second public release candidate. Major feature: the Groundspec Meta-Skill -- a reusable, natural-language front end (groundspec skill export --target {claude-code,codex}) that turns "use Groundspec to..." into a validated Task Contract, guided execution, and a mechanically-derived completion state, without ever hand-writing TOML or reading the schema.
Fully backward compatible with v0.1.0rc1. The existing deterministic CLI and per-task compiled Skills (groundspec compile) are unchanged; every v0.1.0rc1 behavior this release doesn't explicitly extend behaves identically (verified by dedicated backward-compatibility tests).
Install
pip install groundspec-0.2.0rc1-py3-none-any.whl
groundspec doctor
groundspec skill export --target claude-code --output . # or --target codex(Not yet published to PyPI -- install the attached wheel directly, or clone the repo and pip install -e ".[dev]".)
What's new
groundspec skill export/groundspec skill validate: one canonical, vendor-neutral Meta-Skill source rendered into Claude Code and Codex formats, byte-identical except frontmatter. Atomic writes, overwrite-safe (--forcerequired), symlink-safe, deterministic, returns a file manifest with per-file SHA-256.- Contract schema 0.2.0 (additive only: a new
high_valueclassification forscope.open_questions). Schema0.1.0is untouched and fully supported side by side --SUPPORTED_CONTRACT_VERSIONSlists both. groundspec.metaskill.completion.derive_completion_state:PASS/PASS_WITH_CAVEATS/FAIL/INCOMPLETE/BLOCKED, computed from real structured evidence -- never from confident-sounding text alone. Wired intogroundspec evaluateas a backward-compatible, additive extension.- Four Meta-Skill example scenarios (
examples/metaskill/), including one genuine live dry run: a fresh, isolated subagent completed a real CSV-export task end to end using only the exported Skill and the real CLI against a real fixture, independently re-verified afterward. It found and led to fixing four real gaps in the Skill's own instructions -- seeexamples/metaskill/03-software-csv-export/dry-run-report.md. - New Meta-Skill-specific threat-model section (prompt injection at the Skill boundary, recursive-invocation and clarification-loop guards, export safety, structural validation).
- Two stale pre-publication claims left over from
v0.1.0rc1(README/CONTRIBUTING implying CI hadn't run / the repo wasn't public) corrected.
Verified locally and in CI
Full test suite (159 passed, 1 skipped -- documented environmental reason: symlink creation needs elevated privileges on Windows), ruff, mypy --strict, wheel+sdist build, twine check, a forbidden-artifact scan, and a clean install of the exact tagged wheel into a fresh virtualenv (pip check, doctor, skill export/skill validate for both targets, and create/validate/audit backward-compat) -- all green, including in GitHub Actions across Ubuntu/Windows/macOS x Python 3.11/3.12/3.13 on this exact tag (CI run). Changes landed via PR #1, merged only after its own full CI matrix was green.
Known limitations (not yet done)
- Not published to PyPI.
- The Meta-Skill has one genuine live-run evaluation (one scenario) and three constructed examples -- not a full evaluation suite. The 30-scenario deterministic-engine corpus and its 3-arm baseline comparison protocol are built but the live-model comparison itself has not been run -- see
docs/evaluation-methodology.md, markedPENDING. - External user validation has not happened -- see
docs/user-validation-protocol.md, markedPENDING. - Only three domain packs ship (software, research, content) -- by design, not an oversight.
- The Meta-Skill's export safety checks a symlink at the destination and its immediate parent, not the full ancestor chain -- see
docs/threat-model.md.
Checksums
See attached SHA256SUMS.txt.
v0.1.0rc1 -- first public release candidate
First public release candidate. This is a release candidate, not a stable release: known limitations are listed below and in the README.
Install
pip install groundspec-0.1.0rc1-py3-none-any.whl
groundspec doctor(Not yet published to PyPI -- install the attached wheel directly, or clone the repo and pip install -e ".[dev]".)
What's in this release
- Versioned Task Contract JSON Schema (
contract_schema_version: "0.1.0"), strictly validated (unknown keys rejected, no type coercion). - Deterministic, dependency-light canonical JSON/TOML serialization with a tested round-trip.
- A rule engine with a closed, non-executable condition language, and ten built-in rule packs (core invariants, seven risk overlays, three domain packs: software/research/content) resolved through a deterministic four-layer precedence order.
- A budget model and a soft-loss scoring rubric that can never let a soft score override a failed hard constraint.
- The
groundspecCLI:init,create,validate,compile(--target generic|claude-code|codex),audit,evaluate,pack validate,example,doctor. - Three compilation adapters (Claude Code Skill, OpenAI Codex Skill, generic prompt) sharing one canonical rendering, verified to agree on all load-bearing content.
- Ten worked example scenarios and a 30-scenario evaluation corpus, both exercised by an automated test suite (99 tests, 0 skipped).
Verified locally and in CI
Full test suite, ruff, mypy --strict, wheel+sdist build, twine check, a forbidden-artifact scan, and a clean install of both the wheel and sdist into fresh virtualenvs -- all green, including in GitHub Actions across Ubuntu/Windows/macOS x Python 3.11/3.12/3.13 (see the CI run for this tag).
Known limitations (not yet done)
- Not published to PyPI.
- The 30-scenario evaluation corpus and its metrics are built and unit-tested, but the actual 3-arm baseline model comparison has not been run against a live model -- see
docs/evaluation-methodology.md. Marked PENDING, not fabricated. - External user validation (non-technical users, students, a PM, a developer, a researcher, a marketer) has not happened -- see
docs/user-validation-protocol.md. Marked PENDING. - Only three domain packs ship (software, research, content) -- by design, not an oversight; see the README's "Product boundaries."
- Two threat-model gaps are documented as real and currently unmitigated rather than papered over: symlink escape into a configured rule-pack search directory, and terminal-control-character injection into CLI output (both low-severity, local-CLI-only). See
docs/threat-model.md.
Checksums
See attached SHA256SUMS.txt.