Releases: AIops-tools/Queue-AIops
Release list
v0.6.0
Fixed
- A cluster node's key count no longer poses as the whole dataset.
INFO keyspaceon a Redis Cluster node answers for that node's slots only, sokeyspace/overviewreportedtotalKeys: 100for a 3-master cluster actually holding 300 — a dataset-wide claim wrong by the shard count, with nothing in the payload to say otherwise. Both now carryscope(nodevsserver),clusterMode, and on a cluster a note explaining that the figure covers one node. Measured on a real 3-master Redis 7.4 cluster (100 / 92 / 108). undo applyreplays against the target the original write ran on. It dispatched the inverse against whatever target the caller named — in practice the config's first entry — while the write's own target sat unused in the undo record. On a multi-target config the inverse therefore ran against the wrong host; it only looks harmless because the resource usually is not there, but two hosts holding the same name and the inverse succeeds on the wrong one, silently. An explicitly named target still wins. Line-wide: all 24 copies had the identical defect. Caught live in container-host-aiops, where a stop recorded against a Podman target replayed against a Portainer one.
v0.5.0
Changed (BREAKING)
- Requires MCP SDK 2.0 (
mcp[cli]>=2.0,<3.0).mcp.server.fastmcpno longer exists in 2.0; the server is now built withMCPServerand reports its package version in the stdio handshake.
Fixed
undo applyworks from the CLI. Every write tool is imported lazily inside its own CLI command, so a CLI-driven undo ran in a process where the inverse tool was never registered and failed with "inverse tool is not registered" — for every write tool. Only the MCP entry point, which imports the whole server, worked. Found while live-verifying against a real cluster.- An undetermined outcome is audited
unknown, notok. The harness only classified a result as undetermined when the payload also carried anerrorkey, so a write that looked successful but had not been confirmed was recorded as a success. as_intno longer round-trips integers through float64, which cannot represent values above 2**53 exactly. A line-wide sweep found only one of six vendored copies had actually been fixed after the original precision bug. The bool guard precedes the int short-circuit becauseboolsubclassesint— otherwiseTruewould be returned unchanged and serialised astruerather than a number.
v0.4.0
Release notes — queue-aiops 0.4.0
Previous release: 0.3.0.
BREAKING — the authorization layer is removed
This tool no longer decides whether a write is permitted. Read-only mode
(<PREFIX>_READ_ONLY), the graduated-approval / approver gate, and the
rules.yaml deny engine are all gone. Whether an operation runs is the
agent's judgement, or the permission of the account you connect it with — point
it at a read-only credential and the write fails at the server, the place that
actually owns the permission.
What the tool guarantees instead is that nothing is silent: every operation,
over MCP and the CLI alike, lands a row in the audit log — there is no
unaudited entry point. Destructive writes still capture their before-state and
record an undo token where a clean inverse exists.
- If you set
<PREFIX>_READ_ONLY=1, it now has no effect and the MCP server
logs a warning at startup. Restrict writes via the connecting account instead. <PREFIX>_AUDIT_APPROVED_BY/<PREFIX>_AUDIT_RATIONALEstill work, but are now
optional audit annotations — recorded on the row when set, never required.- The declared
risk_levelis carried into the audit row as a descriptive tier
(a label, not a gate).
The governance harness is now: audit (MCP+CLI, unbypassable) · runaway/budget
safety guard · undo recording · output sanitize. policy.py is a small
risk-tier classifier; governance/readonly.py is deleted.
v0.3.0
Release notes — queue-aiops 0.3.0
Previous release: 0.2.2.
In this tool
redis_config_setrefuses the parameters that lock this tool out of the server:requirepass,masterauth,bind,protected-mode,maxclients,port,unixsocket,aclfile, and anythingtls-*. Settingrequirepassinvalidated the stored credential, and the recorded undo then needed a connection that could no longer authenticate. The refusal happens before the prior value is read, which also stops the live password being copied intoaudit.dband the undo store.list_clientsno longer returns this tool's own connection, andredis_kill_clientrefuses it — by id or by address. The two sibling tools already filtered themselves out on the read path; Redis did not.
Every tool in the line: previews and undetermined outcomes
This release fixes three harness defects that were silently degrading the audit
trail and the undo store.
A write that loses its response is no longer recorded as a failure. The
harness assumed a sanitized error meant nothing had happened. That assumption is
false in exactly the case that matters most: when a write severs its own
connection, the request has already landed, the response cannot come back, and
the operation was recorded as status=error with no undo token created at
all. Transport-level failures are now audited as status=unknown, the result
says plainly that the operation may have taken effect and should be verified
before retrying, and a write that stashed its before-state has its inverse
recorded anyway — flagged effectVerified: false, which undo_list and
undo_apply both surface. Existing undo.db files are migrated in place; their
rows read as verified, which is accurate, since the old code only ever recorded
on the confirmed path.
A dry-run no longer writes an undo token. Previews were recording inverses
built from a before-state they never had: the undo callback's permissive default
filled the gap with a guess, producing a real, applicable token for an operation
that never happened.
A dry-run no longer demands a named approver. Requiring an approval in order
to ask whether something needs approval inverts what a preview is for. The tier
is still computed and still audited, so the preview can tell you an approver
will be needed; it just no longer refuses to answer. The write itself is gated
exactly as before.
The invariant, now stated: a dry_run may read; it must never write. Guards
run on the preview path, which means a preview can and does report that an
operation would be refused.
Also line-wide
- Truncated text now ends in an ellipsis instead of being cut silently. This
line already treats a silent cut as a defect for lists; it was doing exactly
that to strings. - Error messages are capped at 800 characters, not 300. These messages end
with what to do instead, so the cap was removing the most useful sentence of
every long refusal.
v0.2.2
Release notes — queue-aiops 0.2.2
Previous release: 0.2.1.
Fixed: the remaining float-typed integer quantities
0.2.1 converted the obvious counts, but a regex-driven sweep missed 19 more —
consumers, per-node fdUsed/socketsUsed, queues/connections/channels
object totals, connection connectedAt, slowlog id/startTime/durationUs, and
Redis totalCommands. A live RabbitMQ broker reported "consumers": 0.0.
These are integers now. Genuinely fractional values — consumerUtilisation,
opsPerSec, fragmentationRatio, and the *Rate fields — are unchanged.
Live-verified: RabbitMQ
The entire RabbitMQ command group had never been run against a live broker. It has
now been exercised against RabbitMQ 3.13.7: doctor, every read
(overview, queues, queue, connections, channels, policies, nodes) with
queue depth cross-checked against rabbitmqadmin ground truth, the backlog and
churn analyses, and the full governance loop — set_policy really created a policy
on the broker, priorState recorded that it had not existed, and undo_apply
deleted it, with all three calls audited.
The platform guard was confirmed too: a Redis-only analysis against a RabbitMQ
target fails with a teaching error naming the mismatch.
See docs/VERIFICATION.md. Redis cluster/sentinel topologies
and AUTH/TLS connections remain unverified.
v0.2.1
Release notes — queue-aiops 0.2.1
Previous release: 0.2.0.
Fixed: integer quantities were rendered as floats
Key counts, client counts and byte totals were routed through the float coercion
helper, so a live server reported 202.0 keys and 1.0 connected clients. The
values were arithmetically right but semantically wrong — a count cannot be
fractional, and a reader should not have to wonder whether .0 means the number was
rounded.
These now use a new as_int() helper and come back as integers. Genuinely fractional
values (hitRatePct, usedPctOfMax, opsPerSec, expiresPct) are unchanged. Note
that equality assertions cannot catch this class of bug (202 == 202.0), so the
regression test asserts the type.
If you parse these fields, the JSON shape changes from 202.0 to 202.
Live-verified
This release was exercised end-to-end against a real Redis 7.4.9 server: every
Redis read cross-checked against redis-cli ground truth, all four analyses, and the
full governance loop (real redis_config_set → audit row → undo_apply restoring the
prior value). See docs/VERIFICATION.md — RabbitMQ remains
mock-only and is now the largest gap in this repo.
v0.2.0
Release notes — queue-aiops 0.2.0
Previous release: 0.1.1.
Headline: read-only mode
export QUEUE_READ_ONLY=1With this set the 8 write tools are never registered — an MCP
client lists 20 tools instead of 28. The writes are not hidden
behind a flag and not merely refused on call: they are absent from the session,
so a model cannot invoke one and cannot be argued into one. For a reviewer this
is checkable rather than promised — connect, list the tools, and the writes are
not there.
Enforcement is two layers deep: the @governed_tool harness refuses every
non-read operation (covering the CLI and in-process callers too), and the MCP
server removes write tools from list_tools(). Changing entry point does not
get around it.
BREAKING — return shapes changed
This release changes payloads that callers may be parsing. Both changes exist
to stop a result from misrepresenting itself:
- Absent fields are now
null, not"". A missing value and an empty value
were previously indistinguishable, which invited consumers to invent the
difference. Keys are still always present — only the value may be null. - Anything with a
limitnow returns an envelope —
{"<items>": [...], "returned": N, "limit": L, "truncated": bool}. Truncation is
measured (one extra row is fetched), never inferred from the page happening to
be full. Where a genuine pre-cap total is knowable it is reported astotal;
where it isn't,totalis deliberately omitted rather than echoingreturned.
Also in this release
docs/VERIFICATION.md— what the mock suite actually guarantees, a live
verification checklist, and the criteria for claiming this tool verified.skills/queue-aiops/references/agent-guardrails.md— for driving this tool with a
smaller / local model: which guardrails are now enforced for you, and a
ready-made system prompt for the rest.- Expanded operator playbooks in the skill documentation.
- The advertised tool count now matches what an MCP client actually lists
(it includesundo_list/undo_apply), and a release gate keeps it honest. - The
(preview)label has been dropped. It never meant unreleased; verification
status now lives indocs/VERIFICATION.mdwhere it can be specific.