Releases: Cogensec/caap
Releases · Cogensec/caap
Release list
caap-benchmark 0.1.0
caap 0.1.0 (2026-09-30) implementing the CAAP-200 working taxonomy 2.0.0-draft.1.
First public release of the CAAP-200 benchmark, implementing the CAAP 2.0.0-draft.1 working taxonomy. The entries below record what was built and corrected on the way to this release.
Added
- A release process. Releases are git tags
vX.Y.Zonmain; the release workflow checks that the tag matches the version inpyproject.toml,versions.py, and the README, runs the full checks and the generated-files check, builds the distribution, and publishes a GitHub release with the wheel and sdist, the taxonomy JSON and YAML, the schemas,SHA256SUMS, and notes taken from the changelog section for that version.scripts/release.pyprovidesbump(set the version everywhere and cut the changelog),check(readiness), andnotes; repository validation now fails on version drift between those files;caap --versionreports the package and taxonomy versions from one source. The procedure is indocs/RELEASING.md. - Evidence bundles for public conformance claims.
caap attest createpackages a graded assessment session or an observedcaap runreport intocaap-evidence.zipwith asubmission.json(newevidence-bundleschema) naming the assurance tier and its fixed labels and claim boundary, the taxonomy and benchmark versions, the subject, the scorecard and layers, every evidence file with its SHA-256, a badge text, and its own canonical hash.caap attest verifyis the reference verifier a registry runs: it recomputes every hash, re-grades an assessment from the bundled responses, recomputes an observed scorecard from the bundled results, and fails on undeclared files or altered claim text.docs/EVIDENCE_SUBMISSION.mdspecifies the bundle, the checks, the claim rule, and what the website's submission form, registry, and badge need. Observed reports now record the taxonomy and benchmark versions. - Any LLM can now be assessed without tools or file access.
caap assess promptrenders a session as a self-contained prompt (optionally split with--chunk-size), andcaap assess importturns the model's reply, raw JSON or Markdown with ajsonblock, into validated responses, filling the housekeeping fields a model commonly omits and refusing replies from another session. Achat-assistantprofile selects the 37 patterns a tool-less model can answer, anddocs/prompts/coding-agent-self-assessment.mdis a shareable prompt that has a coding agent with shell access run the whole protocol on itself from a plain checkout, with no package manager required. - Agent-native CAAP-200 assessment protocol (
caap assess init,grade, andmock-respond), specified indocs/ASSESSMENT.md: 200 adapter-free paired-trial cases underassessments/cases/(a benign control and an adversarial condition per pattern, 400 trials), example capability profiles underprofiles/, a hash-bound session manifest, capability-aware scope, grading that treats missing evidence as inconclusive, over-blocking and recovery measures, and four-layer reporting. Four new schemas cover the case, response, manifest, and report. Every record and domain now carries anintegrity_layer. Results are always labeledagent_self_assessmentandself_reported_unsigned. - The 175 scaffold cases now carry pattern-specific benign objectives, adversarial conditions, and untrusted fixtures instead of generic template text.
- CAAP
2.0.0-draft.1registry with 200 stable pattern records across 11 domains. - Attack families, definitions, maturity, implementation status, relationships, mappings, severity, safety metadata, and 200 pattern pages.
- 25 executable v1.0-aligned safe-sentinel cases and 175 disabled scaffolds.
- Python CLI, mock/command/HTTP adapters, oracle engine, five-state results, scoring, evidence hashes, and JSON/HTML/JUnit reports.
- Schemas, tests, CI, safety policy, governance, and contributor workflow.
- The package now bundles the CAAP-200 registry and the 25 executable cases, so an installed
caapcan list, show, validate, and run outside a repository checkout. The generator writes these copies and repository validation checks they match the canonical files.
Changed
- Every one of the 200 pattern records now has a pattern-specific definition stating the mechanism, the trust boundary it crosses, and the unsafe result, replacing the single template sentence. Each record also has its own six-axis severity vector, and the baseline score is derived from the vector by the documented CAAP-200 method (impact and exploitability weighted double). The 25 v1.0 reference scores are unchanged and marked
caap-v1.0-baseline; all other scores arevector-derived. Pattern pages show the severity, and validation enforces vector range, rating consistency, and the reference-score tolerance. - GitHub Actions in the CI, CodeQL, and release workflows are pinned to commit SHAs with the resolved version in a comment, and a Dependabot configuration keeps the pins and Python dependencies current. The pins were then raised to
actions/checkoutv7.0.1,actions/setup-pythonv7.0.0,actions/upload-artifactv7.0.1, andgithub/codeql-actionv4.38.2. - The published JSON Schemas are now enforced. The runner validates cases against the test-case schema, the HTTP and command adapters validate target payloads against the adapter-response schema and turn a malformed payload into
test_error, and repository validation checks the registry against the taxonomy schema. Full validation uses the newschemaextra (jsonschema), installed in CI; without it a structural fallback derives its required fields from the schema, which closes the previous gap where the hand-rolled check required three fewer fields than the schema. The schemas are bundled with the package. - The 25 executable reference cases now carry mechanism-specific fixtures, attack-success oracles, and secure-behavior evidence instead of one shared template. Each declares a
mock_scenariowith the secure and vulnerable event traces for its mechanism; the mock adapter replays them, so the vulnerable mode now fails on the mechanism oracle (for example an unapproved memory write or a replayed nonce being accepted) before the sentinel backstop. A shared event vocabulary incaap_benchmark.eventsmaps event types to telemetry keys, and each case requires the keys its traces produce. Scaffolds are unchanged. - CI now runs
ruff checkas a separate lint job, andmake lintruns it locally. Existing findings were cleared; the generator's one-record-per-line taxonomy tables are exempt from the line-length rule only.
Fixed
- The HTTP adapter no longer follows redirects. Previously a loopback endpoint could answer with a 3xx and have the request body and bearer token re-sent to an arbitrary remote host, bypassing the loopback-only default. A redirect now yields a
test_errorresult naming the refused target. caap runnow honors"enabled": falseand skips disabled cases, reporting the count on stderr. Previously the 175 contributor scaffolds could be executed and reported as passes. A new--include-disabledflag runs them on request.- The generated
caap-200.yamlemitted empty lists and objects as bare keys, which YAML parsers read asnull; they are now written as[]and{}so the YAML registry is equivalent to the canonical JSON. A round-trip test guards this.