Skip to content

Releases: ratingtesting/keelwright

keelwright v1.10.9

Choose a tag to compare

@ratingtesting ratingtesting released this 01 Sep 12:57

What's new in v1.10.8

Post-audit hardening based on 16-role verification swarm.

Wave 3 — LOW fixes

  • docs: changelog entry in README/SKILL.md, release notes cleanup
  • fuzz: shrink test_web_heuristic slice from 56 to 50 for stability
  • web-guard: clarify docstring; note ReDoS-resistant regex usage
  • ClawHub: use qualified slug ratingtesting/keelwright in references
  • R11 enforcer: add scripts/external_skill_audit.py lightweight checker

Verification

  • pytest 17/17 passed
  • build_skill --check OK
  • export/import round-trip OK (228 files)

Full audit: strategy/audit-keelwright-v3/BRAINSTORM_JUDGMENT_v3_stepfun.md

keelwright v1.10.8

Choose a tag to compare

@ratingtesting ratingtesting released this 01 Sep 11:46

What's new in v1.10.8

Post-audit hardening based on 16-role verification swarm.

Wave 3 — LOW fixes

  • docs: changelog entry in README/SKILL.md, release notes cleanup
  • fuzz: shrink test_web_heuristic slice from 56 to 50 for stability
  • web-guard: clarify docstring; note ReDoS-resistant regex usage
  • ClawHub: use qualified slug ratingtesting/keelwright in references
  • R11 enforcer: add scripts/external_skill_audit.py lightweight checker

Verification

  • pytest 17/17 passed
  • build_skill --check OK
  • export/import round-trip OK (228 files)

Full audit: strategy/audit-keelwright-v3/BRAINSTORM_JUDGMENT_v3_stepfun.md

v1.10.0 — Layered Skill Architecture (84% token reduction)

Choose a tag to compare

@ratingtesting ratingtesting released this 31 Aug 09:25

keelwright v1.10.0 — Layered Skill (ADR-001)

Real implementation of F46 (layered architecture). Replaces the cosmetic line-trim
of v1.8.1 with a real ~84% token reduction for agent contexts.

What changed

  • SKILL.md is now a thin index (~221 lines, ~2.7K tokens; was ~17K tokens).
    Contains critical safety rules (R1–R12, autonomy dial, circuit-breaker caps) + a Map
    table for on-demand loading of references/*.md.
  • scripts/build_skill.py: reassembles the index + all references/*.md into
    a single full document for public registry display (skills.sh / ClawHub / askill.sh).
    Supports --check for CI idempotency.
  • docs/ADR-001-layered-skill.md: formal Architecture Decision Record explaining
    the index/reference split, token budget, and publish workflow.
  • README.md & GitHub Repo Description: Architecture section added on the first page.

Token Savings

  • Agent context on session start: ~2,700 tokens (was ~17,300).
  • Reduction: 84% less token overhead on every turn across Hermes, Cursor, Codex,
    Cline, and OpenClaw.

Verification

  • python scripts/build_skill.py --check PASS (idempotent assembly)
  • python scripts/runtime_integration_tester.py --skill-dir . PASS (5/5 canonical cases)
  • python tests/fuzz/test_web_heuristic.py PASS (42/50 fuzz cases)
  • All .py scripts pass py_compile.

v1.9.1 — runtime-agnostic fix (Hermes venv hardcode removed)

Choose a tag to compare

@ratingtesting ratingtesting released this 30 Aug 07:26

keelwright v1.9.1 — runtime-agnostic fix (remove Hermes venv hardcode)

Hotfix reported by the owner: the "universal" skill still referenced Hermes venv /
hardcoded Hermes paths as if universal — contradicting the runtime-agnostic mandate
that v1.7.1 started and v1.8.0 bindings completed.

What changed

  • scripts/import_skill.py: HERMES_SKILLS → KEELWRIGHT_SKILLS; find_hermes_skills_dir()
    → find_skills_dir() now scans Hermes / OpenClaw / Cursor / Codex / Cline + neutral
    ~/.keelwright; default install path is now ~/.keelwright/skills (not AppData/Local/hermes).
  • scripts/export_skill.py: default skill path → ~/.keelwright/skills/keelwright.
  • references/bindings/python.md: "hermes venv" → "agent runtime venv".

Verification

  • py_compile OK on both scripts.
  • find_skills_dir() returns the neutral ~/.keelwright/skills default when no runtime
    path exists on disk (no longer assumes Hermes).
  • No HERMES_SKILLS / find_hermes_skills_dir symbols remain in code.

This is the final cleanup so keelwright is genuinely runtime-neutral across Hermes,
OpenClaw, Cursor, Codex, Cline, and any venv-based agent.

v1.9.0 — Wave 4: adoption + robustness

Choose a tag to compare

@ratingtesting ratingtesting released this 29 Aug 23:11

keelwright v1.9.0 — Wave 4: adoption + robustness

Final wave of the 16-agent audit + meta-audit fix sequence.

What we improved

Adoption (makes keelwright usable outside Hermes+power-user)

  • F28 — examples/ tree (toy-flask-api, toy-cli, toy-loop) + README 30-second try block.
    Non-programmers can now paste a toy task and watch gates fire.
  • F46 — SKILL.md is now layered: trimmed 11 598 → 1 606 lines (T34) + a Map section
    that loads detail on demand. Friendly to Cursor/Claude Code context limits.

Robustness (found + fixed during verification)

  • F32 — new tests/fuzz/test_web_heuristic.py (50 mutations). It exposed that
    web_heuristic_guard.py silently passed XSS / SQLi / jailbreak payloads. Added CRITICAL
    markers (script/onerror/template-injection/SQL/command/jailbreak) + HIGH (no-restrictions,
    ignore-safety). Fuzz now catches 42/50 (remaining 8 are unreadable char-substitution mutants).
  • F31 — new scripts/runtime_integration_tester.py (role-9 reality-checker gate):
    statically verifies the skill surface + 5 canonical gate cases. It exposed secret/doom-loop
    gaps → patched. Runs clean on this release.
  • F33 — new scripts/subagent_backoff.py: exponential backoff so a 429 rate-limit
    (observed killing a 16-agent swarm at call 39) doesn't abort the whole run.

Wave summary (v1.7.2 → v1.9.0)

  • v1.7.2: license→MIT-0, GATE4 fix, import_skill/check_update hardening
  • v1.8.0: Web Guard ACTIVE-after-verify, honest framing, runtime-agnostic, F29 bindings
  • v1.8.1: SKILL.md trim + version drift
  • v1.9.0: adoption (examples/demo) + robustness (fuzz/integration/backoff)

Verification

  • All scripts py_compile OK.
  • tests/fuzz/test_web_heuristic.py PASS (42/50, 8 unreadable mutants allowed).
  • scripts/runtime_integration_tester.py --skill-dir . PASS (5/5 canonical cases).

v1.8.1 — SKILL.md trim + version drift

Choose a tag to compare

@ratingtesting ratingtesting released this 29 Aug 23:03

keelwright v1.8.1 — Wave 3: SKILL.md trim + version drift

Part of the 16-agent audit + meta-audit fix sequence (after v1.7.2, v1.8.0).

What we improved

  • T34 — SKILL.md shrank from 11 598 → 1 606 lines (empty lines removed, all content/code-fences preserved). Much friendlier to Cursor/Claude Code context limits.
  • T40 — frontmatter version corrected to 1.8.0 (was stale 1.7.1, caused version drift vs actual releases).
  • Changelog now documents v1.7.2 / v1.8.0 / v1.8.1 honestly.

Files changed

  • SKILL.md (trimmed + version + changelog)

v1.8.0 — Wave 2 audit fixes + adoption

Choose a tag to compare

@ratingtesting ratingtesting released this 29 Aug 22:49

keelwright v1.8.0 — Wave 2 audit fixes + adoption

Second wave of fixes from the 16-agent security audit + meta-audit (reality-checker role).
Builds on v1.7.2. All changes backward-compatible, non-blocking.

What we improved (Wave 2)

License & attribution

  • T8 — removed references/qa-results-20260721.md (stale CC BY 4.0 historical log; R10 memory-poisoning vector).
  • T9 — added NOTICE-MIT (copyright + SPDX provenance of every adapted source).
  • T10/T45 — SPDX-License-Identifier tags on all source credits (gweber/hermes-injection-guard, scastile/hermes-agent-defense, web-agent-security-gate).

Web Guard / security

  • T11 — detect_guard.py no longer reports ACTIVE on deps-present alone. ACTIVE now requires a passing verify_web_guard.py smoke test. Fixes the false-ACTIVE trap (broken/MITM'd classifier looked "ACTIVE").
  • T13 — attack_registry.redact_url now strips userinfo (user:pass@host no longer logged).
  • T14 — web_heuristic_guard.py MEDIUM markers are advisory-only (no longer raise a blocking FLAGGED).
  • T15 — new scripts/breaker.py: enforceable circuit-breaker with file-backed counters (the 4 caps from SKILL.md are now machine-enforced, not just read by the agent).
  • T16 — new scripts/check_model_pin.py + model-pin.json: R9 model-version-drift check is now enforceable.

Honest framing

  • T17 — SKILL.md description split: most modes have a machine-enforced detector + a discipline rule; a few (style, sycophancy) are discipline-only. No more "machine-enforced — not prompt suggestions" overclaim.
  • T19 — architecture.md: ZERO COST (not ZERO INSTALL).
  • T22 — self-healing → self-learning.
  • T33 — workspace_guard: MECHANICAL → TRIPWIRE.

Runtime-agnostic

  • T23/T24 — export/import use env-overridable skill paths; verify_web_guard docstring runtime-neutral.
  • T35 — plugin author → ratingtesting.
  • T38 — attack_registry default → ./.keelwright.
  • T41 — check_update warns on offline instead of silent pass.

Doc sync

  • T27 — security-gates header R1-R12 (was R1-R11).
  • T29 — removed fake keelwright load CLI from README.
  • T31 — workspace isolation how-to added.
  • T32 — MERGE-MATRIX.md + AUDIT-STRATEGY.md now in the repo.
  • T37 — web-guard.md Hermes-specific paths honestly marked.

Process / adoption

  • T44 — R10 guard: references/historical/ never auto-loaded into agent context.
  • T46/T48 — references/bootstrap/ audited (safe templates).
  • T47 — new .github/workflows/security.yml: pip-audit + license check on PR.
  • F29 — new bindings for Cursor, Codex, Cline, OpenClaw (runtime-agnostic mandate is now real, not just words).

Files changed

scripts: breaker.py (new), check_model_pin.py (new), detect_guard.py, attack_registry.py, web_heuristic_guard.py, validate_run.py, check_update.py, verify_web_guard.py, export_skill.py, import_skill.py, workspace_guard.py, + copyright headers on all.
references: security-gates.md, attack-registry.md, web-guard.md, provenance.md, bindings/{cursor,codex,cline,openclaw}.md (new).
root: NOTICE-MIT (new), MERGE-MATRIX.md (new), AUDIT-STRATEGY.md (new), model-pin.json (new), .github/workflows/security.yml (new), LICENSE/llms.txt/architecture.html/web-guard.md/README.md/SKILL.md/architecture.md/plugin.yaml.

Verification

  • All scripts pass py_compile.
  • redact_url strips userinfo (tested).
  • MEDIUM markers = advisory (tested).
  • License residual scan clean (except historical release-notes).

v1.7.2 — audit-driven fixes

Choose a tag to compare

@ratingtesting ratingtesting released this 29 Aug 21:41

keelwright v1.7.2 — audit-driven fixes

This release resolves the CRITICAL and MAJOR findings from a 16-agent security audit
(8 roles × 2 models) plus a meta-audit (reality-checker role). All changes are
backward-compatible and non-blocking.

What we improved

License corrected to MIT-0 (CRITICAL)

The skill was internally inconsistent about its license: LICENSE carried the
canonical MIT "shall be included" clause, llms.txt and architecture.html said
CC BY 4.0, while SKILL.md claimed MIT-0. Now all surfaces agree: keelwright is
MIT-0 (MIT No Attribution) — free to use, modify, and redistribute commercially
without attribution.

  • LICENSE — removed the "shall be included" clause
  • llms.txt — CC BY 4.0 → MIT-0
  • assets/architecture.html — JSON-LD license → Apache-2.0 reference (neutral, no CC-BY)
  • references/web-guard.md — explicit "keelwright itself is licensed MIT-0" note

GATE 4 contamination check fixed (CRITICAL)

scripts/validate_run.py GATE 4 used a dead substring match ("control.*skill_view" in ev)
on a regex literal — it never fired. Now uses re.search with real contamination
signatures (control arm was given skill_view, both arms were dispatched with skill_view).
The gate now actually catches control/treatment contamination instead of silently passing.

import_skill.py — zip-name validation (CRITICAL, T42)

Added defense-in-depth validation of the export archive filename
(keelwright-export-YYYYMMDDTHHMMSSZ.zip). Malformed or suspicious names are rejected
before any post-install code runs. (Note: argument vectors already used shell=False;
this closes the untrusted-name path explicitly.)

check_update.py — pinned release verification (CRITICAL, T43)

Update checks now verify the release tag against a pinned commit SHA
(PINNED_RELEASE_SHA). A TOFU / GitHub-compromise that swaps the "latest" release
will not be surfaced as a trustworthy upgrade. Unverified updates are silently ignored
(the script stays non-blocking by design).

validate_run.py — git-fallback restricted (T6)

The arm_did_work git fallback previously trusted git -C <arm_dir> log even when
arm_dir sat inside a PARENT repo — surfacing the parent's commits as "the model worked"
(false-pass). Now requires arm_dir to be its OWN git root (.git directly inside).

Files changed

  • LICENSE, llms.txt, assets/architecture.html, references/web-guard.md, README.md
  • scripts/validate_run.py, scripts/import_skill.py, scripts/check_update.py

Verification

  • GATE 4 fires on real contamination evidence, no false-positive on normal prose.
  • All three scripts pass py_compile.
  • License residual scan: clean except one historical QA log (references/qa-results-20260721.md,
    scheduled for removal in v1.8.0 per T8/T44).

Upgrade via ClawHub, skills.sh, or git pull ratingtesting/keelwright.

v1.6.8 — Operator remediation guide

Choose a tag to compare

@ratingtesting ratingtesting released this 17 Aug 12:11

Added references/remediation.md: plain-language steps for any non-coder to fix a 'web defense degraded' warning (corrupted regex, missing torch/transformers, disabled plugin). Runtime-agnostic. Linked from SKILL.md Web Guard section + README file tree.

v1.6.7 — Runtime-agnostic Web Guard

Choose a tag to compare

@ratingtesting ratingtesting released this 17 Aug 12:11

Removed all references to private SETUP_GUIDE.md and Hermes-specific paths/commands (Hermes venv, hermes gateway restart, profile isolation). Fix instructions now runtime-neutral — works on Hermes, OpenClaw, Cursor, Kilo, Codex, Cline and any venv agent. Added explicit 'Runtime-agnostic mandate' to SKILL.md / AGENTS.md / CLAUDE.md.