Skip to content

Releases: SAY-5/playbook

v6.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 27 Sep 23:47
a40e491

The run store now keeps a procedure's prompt versions, reports, traces, grades, scenario bank, proposals and promotions under one prefix named by its slug, on disk and in S3 alike, as prompts/vN.json, reports/vN.json, traces/vN/, grades/vN/, bank.json, proposals/ and promotions/. That layout change is why this is a major release, and there is no migration from the 5.0.0 layout. In 5.0.0 prompt versions, reports, grades and the scenario bank lived only in the local runs directory and the S3 store held nothing but traces and review records, so the deployed runner could load no prompt but the v1 it created in its own /tmp. With PLAYBOOK_RUN_STORE=s3 all of them now live in the bucket, every grade is stored and its score and verdict are written onto the run's DynamoDB item, and playbook ops reads the store rather than a directory. PromptSpec, VersionReport and ScenarioBank no longer save or load files themselves, the CLI and deploy/lambda/handler.py open the store through the same make_store factory, and CONTRIBUTING.md now states the policy this release follows: a change to the CLI or to the stored artifact layout takes a major version.

deploy/lambda/handler.py runs the promoted version when a message names no prompt_version and refuses a procedure with nothing promoted, so after three receives the message lands in the dead-letter queue instead of running an unreviewed prompt; test_lambda_handler_runs_a_stored_version_from_s3 checks the refusal, an explicit version, the stored grade and the promoted default against LocalStack. make lambda-zip resolves the bundle for x86_64 manylinux and Python 3.12 from uv.lock whatever the host, make lambda-check imports it under linux/amd64 in Docker, a lambda-zip CI job imports the handler from it on ubuntu, and deploy/terraform/main.tf declares architectures = ["x86_64"] for the function instead of relying on the AWS default. The image installs from uv.lock with uv pinned to 0.11.7, serve-fakes gained --host, which the compose fakes service sets to 0.0.0.0 so its published ports reach the servers, and version reads the installed package metadata rather than a hard-coded 0.1.0.

Live mode no longer falls back to the local Jira stand-in and empty Jira and Slack tokens: --live requires JIRA_BASE_URL, JIRA_TOKEN and SLACK_TOKEN unless PLAYBOOK_TOOLS=fake asks for the stand-ins, and every trace records in tool_endpoints whether its tools were real or fake and which Jira and Slack hosts it used. Ingest no longer takes any sentence containing "is" for a decision; it recognises a priority matrix such as "full outage is Highest", and it attaches a decision to the step that names the same channel, a matrix to the step that sets priority and a sentence about posting or paging to the first slack.post step before falling back to keyword overlap. On the sample procedures the #ic-oncall page moves from open-incident to announce, the enterprise bump from escalate to create-ticket, and two narrative sentences stop being decisions. playbook review propose queues a reviewer's own rule for approval, playbook loop reports "all scenarios pass" when the last permitted round gets there and no longer calls a round that newly passes a scenario a plateau, and the regress artifact carries the replayed version's scores.

make demo-review runs the promotion gate, the audit trail, the scenario bank and both regression runs on the output of make demo: v1 is refused because it puts a customer's email or phone number in Slack in two scenarios, v3 is promoted, and a deliberately wrong rule proposed and approved by a reviewer becomes v4, which the regression guard fails with triage-10 and triage-12 broken. Both demo scripts now wait for the offline fakes to answer on /health, retrying the start once, before running the pipeline.

The browser demo reads the sample procedures from ../procedures at build time instead of keeping copies, parses the rubrics and scenario sets into JSON in vite.config.ts, and no longer ships framer-motion or yaml; its JavaScript bundle is 227.58 kB (73.29 kB gzipped) against 444.72 kB (142.04 kB) for 5.0.0. Its fonts are served with the page from web/public/fonts instead of the Fontshare and Google Fonts CDNs, both under the SIL Open Font License 1.1 with the licence text beside the files, and the display face is now the latin subset of Space Grotesk as one variable file with a 300 to 700 weight axis, because the ITF Free Font License of Clash Display, the 5.0.0 face, forbids distributing its files through a repository. Section reveals are CSS transitions switched on by one IntersectionObserver hook, and the counters and the full-run log take their reduced-motion switch from the same web/src/hooks/motion.ts, so prefers-reduced-motion is honoured across the page; percentages round the way Python formats them, checked against the table scripts/pct-cases.py writes; the transcript gained a rules tab listing what the offline stand-in parsed out of each prompt version and the prose it ignored; and the title and hero no longer say the agent grades itself.

Vite moves from 5.4.21 to 7.3.6 and @vitejs/plugin-react from 4.7.0 to 5.2.0, web/package.json declares node ^20.19.0 || >=22.12.0 and web/.nvmrc names 22; npm audit over web/package-lock.json reported esbuild (moderate, GHSA-67mh-4wv8-2f99) and Vite (high, three advisories including GHSA-fx2h-pf6j-xcff) at 5.0.0 and reports 0 vulnerabilities at this commit. A web CI job runs npm ci, the type-check, the 51 self-check assertions, the production bundle and a smoke test that drives Chrome through playwright-core and fails on a console error, a request to another origin or horizontal overflow at 1440 and 390 px. 48 tests pass and 2 skip by design at this commit, with the S3 store and handler tests running against LocalStack in CI.

v5.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 10 Sep 09:06

playbook ops reads a runs directory back and reports what an operator asks between runs: the procedures, their prompt versions, the pass rate history, the forbidden actions still open with the criteria that failed, the pending proposals, the promoted version, and the last run with its duration. Every eval, loop and regress now writes a JSON artifact under runs//artifacts/ holding the version scores, the failing scenarios and the measured cost of that run. Tool-call counts per tool, calls per run, tool latency as mean, p95 and max, and run wall time are measured from the traces, and each trace records its own duration. make demo prints the ops summary at the end, and the README carries the real output of that run. Everything still works offline with no API key.

v4.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 10 Sep 09:00

playbook bank keeps every scenario a procedure has been run on, with its tags and how each prompt version scored it, so a case retired from scenarios.yaml keeps its history and stays in the replay. playbook regress replays the bank against a version and compares each scenario with the last version that scored it. A scenario that used to pass and now fails is a regression and exits the run non-zero, naming what broke. A scenario the bank has never scored is new, so its failure is reported separately and does not trip the guard. Pass rates are reported per tag, and --tag or --failures narrow the replay to one segment or to the cases with a recorded failure.

v3.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 10 Sep 08:53

Corrections derived by the feedback loop are now proposals a reviewer decides on before they enter a prompt version, with approve, edit and reject recorded under the reviewer's name. playbook diff compares two versions step by step and reports the steps added, removed and changed, including retitled and reordered ones. playbook promote signs a graded version off for use and refuses while the version still commits a forbidden action or carries undecided proposals, and the refusal is recorded as well. playbook review audit prints the whole decision trail, corrections and versions together, from the proposals and promotions stored alongside the runs. The unattended path is unchanged: playbook loop still decides as the reviewer auto, and offline mode still needs no API key.

v2.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 08 Sep 23:27

Scenario synthesis and coverage. playbook coverage splits every walkthrough decision point into its branches (each severity, tier or impact level named, both sides of a KB match or customer-facing condition, and an otherwise branch) and reports which scenarios cover each one and how many scenarios sit on each side of every rubric criterion; procedures with uncovered branches or one-sided criteria are flagged, and --strict makes that a failing exit. playbook synthesize derives one scenario per branch (or per uncovered branch) from the closest hand-written template, inferring expected outcomes from scenarios that share the criterion's condition variables and tagging the rest needs-expert. Both sample sets cover all 11 and 8 branches offline; 31 tests pass, 2 skip by design.

v1.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 08 Sep 23:20

First release of Playbook. It ingests an expert's SOP and walkthrough into a cited Procedure, renders a versioned prompt, runs a bounded tool-calling agent over the Messages API against Jira, Slack and a knowledge base, grades every run with the expert's rubric, and turns failures into prompt corrections that are re-run and compared. The Terraform stack (S3, SQS, DynamoDB, Secrets Manager, Lambda runner) validates and applies against LocalStack. Everything runs offline with the deterministic fakes; 26 tests pass, 2 skip by design without LocalStack or an API key.