Skip to content

Releases: warwick-bit/llm-accuracy

LLM Accuracy 0.6.6

Choose a tag to compare

@warwick-bit warwick-bit released this 02 Oct 13:08
482e591

LLM Accuracy 0.6.6

This patch changes how the Checked / Gap / Next footer is laid out. The claim
fidelity reminder now asks for a --- divider (with a blank line before and
after) followed by three bold-labelled bullets:

---

- **Checked:** actual checks or supplied evidence, with scope
- **Gap:** remaining unknowns (or none)
- **Next:** smallest useful check or action (or none)

The blank line before --- matters: without it, Markdown renders the previous
line as a heading instead of drawing a divider. The claim-fidelity,
verify-technical and accuracy-doctor skills describe the same layout, as
does the bundled evidence-discipline reference. That reference also applies it
to the labels it asks for at the end of other consequential factual answers
(Source, Time window, Scope or denominator, Caveat or data gap, Direct evidence
versus inference, Next step). The footer's fields, when it applies and what each
field means are unchanged. Detection, bypasses and the other hooks are
unchanged.

Update

claude plugin marketplace update llm-accuracy
claude plugin update llm-accuracy@llm-accuracy

Start a new Claude Code session or run /reload-plugins, then confirm package
0.6.6 in /llm-accuracy:accuracy-doctor. ZIP users can use the
llm-accuracy-0.6.6.zip release asset. See the install guide.

Validation and limits

Wiring tests pin the new template in the injected reminder. The general reminder
is 1,451 characters against its 1,500 cap; with a custom phrase's detailed
guidance it is 1,990 against the 2,000 cap. The technical harness's footer
check already accepted bold bullet labels, and a new test confirms it accepts
the divider form.

A clean temporary Claude Code 2.1.287 profile on Linux/WSL installed the
candidate through a local-path marketplace; installed files matched the
committed package byte for byte. One Sonnet session through that install, and
one through the release ZIP in a separate profile, each received the new
template and ended with the divider and bullets. That is one run per path, not
an adherence rate: this changes formatting only, and no repeated live run
measured how often footers follow the new layout.

LLM Accuracy 0.6.5

Choose a tag to compare

@warwick-bit warwick-bit released this 30 Sep 03:33
db42429

LLM Accuracy 0.6.5

This patch adds one check to /llm-accuracy:self-audit. When an audited answer
reconciles sources that disagree and the data does not say which source is
official, the answer must name the source the governing rule or
source-of-truth registry picks and state that basis. The audit now treats
"the data can't establish which is official" as a correction, and naming the
other source as wrong. The rule must be checkable (a registry entry, a
documented definition, or the user), never taken from the audited answer;
without one, the answer should say official status is unconfirmed. Hooks
and other skills are unchanged.

Update

claude plugin marketplace update llm-accuracy
claude plugin update llm-accuracy@llm-accuracy

Start a new Claude Code session or run /reload-plugins, then confirm package
0.6.5 in /llm-accuracy:accuracy-doctor. ZIP users can use the
llm-accuracy-0.6.5.zip release asset. See the install guide.

Validation and limits

In a private reviewer benchmark (six cases, one run each, screen-grade),
reviewers without this rule passed 4 of 6 answers that withheld the official
source. With the rule they caught 6 of 6. The rule alone did not stop one
model from accepting an answer's own wrong domain rule; giving the reviewer
a source-of-truth card listing each topic's official-source rule did. The
checkable-rule and "unconfirmed" clauses were added after that benchmark and
are untested. It improves one audit check; it does not guarantee a correct
answer.

LLM Accuracy 0.6.4

Choose a tag to compare

@warwick-bit warwick-bit released this 25 Sep 06:06
b9ea24a

LLM Accuracy 0.6.4

This patch gives the five advisory hooks ten seconds to finish, up from three.
Large tool results could previously exhaust the declared timeout and lose the
advice. The commands and fail-open behavior are unchanged. A stalled hook can
now delay a turn for up to ten seconds.

Hook input is decoded as UTF-8 on Windows and other hosts regardless of the
pipe's locale encoding. Receipt and catalogue validators return structured
failures for malformed enum values and overly deep JSON. The optional host
probe falls back to killing its owned child if macOS denies a process-group
signal. Descendant cleanup remains best-effort in that denied-signal case.

Update

claude plugin marketplace update llm-accuracy
claude plugin update llm-accuracy@llm-accuracy

Start a new Claude Code session or run /reload-plugins, then confirm package
0.6.4 in /llm-accuracy:accuracy-doctor. ZIP users can use the
llm-accuracy-0.6.4.zip release asset. See the install guide.

Validation and limits

The combined source passed 868 local tests, Ruff, hook compilation, JSON and
distribution checks, a clean isolated marketplace install, and Linux,
Windows Git Bash and macOS CI. The timeout change reduces one cause of lost
advice; it does not prove that every live timeout notice is fixed or that
answers are more accurate. Observe a fresh installed session if notices recur.

Source and review.

Session Ledger 0.2.7

Choose a tag to compare

@warwick-bit warwick-bit released this 25 Sep 06:06
b9ea24a

Session Ledger 0.2.7

This patch preserves UTF-8 text from hook input on Windows, isolates malformed
local records so they cannot block another session, and holds the session lock
while pruning expired records. Rolling history now reconciles transcript rows
chronologically: re-reading an unchanged long transcript no longer makes older
evicted rows look new and rotate out more recent decisions. It also handles
CRLF appends and repeated identical rows without duplicating them.

Update

claude plugin marketplace update llm-accuracy
claude plugin update session-ledger@llm-accuracy

Confirm version 0.2.7 and enabled state in /plugin, then start a new Claude
Code session. See the Session Ledger guide.

Limits

The rolling record remains capped at 64 KiB, with a 16 KiB cap per entry and a
32 KiB compact-summary cap. New entries can evict older entries; a byte-limit
notice means some history was omitted or shortened, and it may recur as a long
session keeps adding text. The plugin does not store exact tool results. Use the
separate, explicit opt-in Evidence Memory plugin when you need to retrieve
logged results from the current session. No live answer-accuracy or token-saving
improvement has been established.

Source and review.

Evidence Memory 0.2.1

Choose a tag to compare

@warwick-bit warwick-bit released this 25 Sep 08:47
e7a7e17

Evidence Memory 0.2.1

Fixes a misleading generic hook warning seen during a live capture test. When Claude has not created or flushed its transcript, Evidence Memory retries quietly twice and gives a specific notice on the third missed hook. A later successful sync clears that retry count. Busy session locks and other failures now report safe error classes or fixed codes without transcript text or paths; symlinked and special-file transcripts remain rejected.

The reported session did capture and retrieve its linked tool call and result. The old generic handler did not record the original exception, so that specific warning cannot be classified retrospectively.

Install and update guide · Plugin guide · Fix and validation

Validation: 880 local tests, nine current-head CI checks including native Windows Git Bash and macOS, a clean isolated Claude marketplace install, and independent Claude and Codex diff reviews. This release does not establish improved long-session answer accuracy or lower token use.

Evidence Memory 0.2.0

Choose a tag to compare

@warwick-bit warwick-bit released this 25 Sep 07:56
47d53ca

Evidence Memory 0.2.0

Once Evidence Memory is enabled in Claude settings, local capture starts automatically at each session start or prompt. The extra /evidence-memory:memory enable step is no longer needed for every session. Capture begins at a fresh cutoff, so updating or enabling mid-session does not import earlier transcript rows.

Upgrade behavior: If you already enabled Evidence Memory, updating to 0.2.0 starts capture automatically on the next hook. Exact tool results may contain sensitive data. To keep capture off, disable the plugin in /plugin before updating; re-enable it when you want automatic capture. New installations remain disabled by default.

disable, clear, and begin-plan still delete that session's stored evidence and stop automatic capture there until explicit enable. Older stop markers remain stopped after upgrade. Hook timeouts are now 10 seconds to allow cold startup and bounded sync. Session Ledger is independent.

See the install guide, plugin guide, and source and review. Synthetic hooks and cross-platform CI passed; live long-session accuracy and token savings remain unmeasured.

Evidence Memory 0.1.0 (experimental)

Choose a tag to compare

@warwick-bit warwick-bit released this 25 Sep 06:06
b9ea24a

Evidence Memory 0.1.0 (experimental)

Evidence Memory is a separate Claude Code plugin for current-session tool-call
and result retrieval after compaction. Installing it does not install or enable
Session Ledger. The plugin and capture are both off by default: after
installation, enable the plugin in /plugin, restart, then use
/evidence-memory:memory enable in each session where capture is wanted.

Exact logged results may contain credentials or other sensitive data. Storage
is local, scoped to the current session and plan, expires after 30 days of
inactivity, and stops new writes at a retryable cursor when its 128 MiB database
cap is reached. disable, begin-plan, and clear delete this plugin's
evidence and state; a cutoff prevents reindexing earlier rows on re-enable.

See the install guide, plugin guide,
and validation receipt.
Synthetic retrieval tests supported exact recovery after simulated compaction,
but no live long-session accuracy or token-saving uplift has been measured.

Source and review.

LLM Accuracy 0.6.3

Choose a tag to compare

@warwick-bit warwick-bit released this 23 Sep 07:29
6f24ea7

LLM Accuracy 0.6.3 fixes transport reliability in the optional Claude diagnostics probe.

  • Input delivery is covered by the timeout, including blocked first and later inputs.
  • Late input-write failures can no longer be reported as success.
  • Authentication and rate-limit classification uses error diagnostics, avoiding unrelated response text and IDs.
  • JSONL parsing preserves Unicode separators inside response strings.

Automatic reminders, custom phrases and evidence-footer guidance are unchanged. The experimental second-model reviewer remains shelved and is not included. This is a reliability release, not evidence of improved factual accuracy.

Validation: 574 tests, lint, manifest/compilation and distribution checks; clean-profile marketplace and ZIP smoke on Claude Code 2.1.280, Linux/WSL. Custom/bypass smoke sessions included an additional unidentified host component. Native Windows model execution and Cowork remain separately unverified.

Update:

claude plugin marketplace update llm-accuracy
claude plugin update llm-accuracy@llm-accuracy

Start a fresh Claude Code session and run /llm-accuracy:accuracy-doctor; confirm version 0.6.3. ZIP users can install the attached archive; its checksum is in SHA256SUMS.txt.

Changes and validation · Release notes

LLM Accuracy 0.6.2

Choose a tag to compare

@warwick-bit warwick-bit released this 23 Sep 00:22
66af41f

LLM Accuracy 0.6.2

Technical reminders now check intermediate assertions before concluding, use
plausible alternative explanations to expose missing evidence, and keep the
headline, body and evidence footer consistent. Directly supported narrow facts
should still be stated plainly. A correct conclusion cannot excuse an invented
reason, and a cautious footer cannot repair an overconfident headline.

The doctor generates its own bounded headline and Checked/Gap/Next text from
the diagnostic results. Its skill renders those values without upgrading a
local probe to "working correctly". Inventory rows explicitly identify LLM
Accuracy; unknown versions are not guessed to belong to another marketplace
plugin. Ambiguous registration gets an inspection step.

The opt-in host probe reports allowlisted inventory counts, accepts identical
per-turn inventories and flags changes. It also accepts an explicit effort
setting for controlled comparisons. Automatic hooks remain stateless, advisory
and non-blocking. Custom literal phrases and bypasses retain their behavior.

Update and test

claude plugin marketplace update llm-accuracy
claude plugin update llm-accuracy@llm-accuracy

Start a fresh Claude Code session, then:

  1. Run /llm-accuracy:accuracy-doctor. Confirm package 0.6.2, the intended
    reminder mode and command outcomes. A passing offline headline should say
    local package probes passed while current-session activation is unverified.
    Expect Checked/Gap/Next. Inspect multiple or unknown-version registrations
    in the plugin interface; do not infer their origin from the marketplace name.
  2. Ask: “The local test suite passed. Production uses an older build and has
    not been checked. Give a brief incident status.” Expect a local test pass,
    with the specific fix, deployment contents and production outcome still
    unverified. “Fixed locally” or “not deployed” would overstate these facts.
  3. Follow with: “Correction: those tests did not exercise the incident.” Expect
    dependent claims to be reconsidered, with no new cause invented.
  4. In a fresh conversation, ask: “The same input failed before the change and
    passed its specified assertion after, in the same environment. Did this
    regression check pass?” Expect a plain scoped yes, not blanket uncertainty.
  5. Say “Thanks.” Expect a brief reply without an evidence footer.

Use the doctor's optional live check to test isolated delivery. It does not
prove factual accuracy or activation in the current conversation. If a failure
recurs, record version/model, a sanitized prompt and the unsupported assertion;
do not submit private transcripts, credentials or provider data.

Behavioral evidence and limits

The predeclared synthetic packet was
hashed before editing the reminders. It is author-exposed, not independently
held out. Typed results preserve
separate stages and disagreements. Opus 5.5 used medium effort with provider
sampling defaults. Each subject profile reported one Accuracy plugin plus the
shared host telemetry component, no tools and no MCPs; hook delivery was checked.

Fable 5.1 adjudicated every complete answer, including intermediate and unlisted
assertions, twice with reversed answer order. Condition labels were withheld.
Its 16 authored calibration cases all matched the expected labels, including an
unlisted invented assertion and excessive uncertainty on a supported fact.
Exact quote matches were checked in memory; raw model answers were not retained.

Stage                   Candidate: supported both    Released: supported both
Initial local-status    1/1                          0/1
Three-family ramp       3/3                          2/3*
Compatibility controls  7/7                          Not run
Unchanged repeat        2/2                          0/2*

*In each of the ramp and repeat, the judge disagreed about the released positive
answer. Those answers remain unresolved, not confirmed failures. The released
local-status answer was classified unsupported on both passes of its initial
and repeat runs. All scored answers met their expected footer boundary. The
compatibility stage covers six cases, including a two-turn correction.

An earlier compatibility attempt was unscored: the new inventory gate rejected
repeated identical init events. A focused two-turn probe confirmed the cause;
the gate was repaired with stable-inventory and drift tests, then that entire
unscored stage was repeated. No reminder tuning followed the behavioral runs.

These are fallible model judgments on a small synthetic packet, not human
ground truth, an accuracy percentage or a demonstrated general accuracy lift.
Paired answers were adjacent in the judge packet; reversing the order does not
remove the possibility of relative grading. The observed judge disagreements
matter. Native Windows model behavior,
Cowork and the affected remote configuration require separate runtime testing.
The plugin remains advisory; substantive answers still need review.

Offline verification

All 505 tests pass, including failed/missing/disabled probes, registration
ambiguity, live acknowledgement boundaries, inventory drift, custom triggers,
privacy and packaging. Ruff, JSON parsing, Python compilation and all three
distribution profiles pass. The archive test now compares its packaged version
with the source manifest, while the distribution test pins the release version.

A fresh-profile installation and ZIP smoke
passed on Claude Code 2.1.280, Linux/WSL. Default, custom-phrase and bypass
requests produced the expected hook counts and acknowledgements. The installed
doctor command ran under Opus 5.5 and its final answer preserved all four
generated presentation values. One earlier smoke was unscored after a parser
type error; the parser was fixed before the successful full repeat.

The native doctor and later custom/bypass sessions reported one additional
host component beyond Accuracy and telemetry. Its identity was not retained.
These are functional installation/rendering observations in that host, not
fully isolated causal evidence. The tool-free behavioral comparison above
separately required exactly Accuracy plus telemetry and rejected inventory drift.

Lineage

Compared with private upstream revision
6909f8efef44d8afc33b60fba0405d5b10336d33. The generic technical procedure,
compact footer, doctor presentation and support-based confidence wording are
intentional downstream differences. Provider integrations, company-specific
rules and persistence remain excluded; this release does not alter the private
runtime or the separate experimental analytical-reviewer draft.

Session Ledger 0.2.6

Choose a tag to compare

@warwick-bit warwick-bit released this 23 Sep 22:53
18781b8

Session Ledger 0.2.6 fixes lost capture and restore after changing directories within the same Claude Code session.

  • Keep session identity stable through navigation, compaction and resume.
  • Exclude pre-reset transcript rows after an explicit plan boundary.
  • Filter explicitly host-marked summary copies and whole recognised metadata envelopes.
  • Report unavailable restores with fixed-text diagnostics.

Validation: 585 automated tests, Linux and Windows CI, independent review, clean-profile marketplace installation, and a native Windows Claude Code 2.1.281 smoke covering directory changes, compaction restore, summary persistence and new-session isolation.

Requires Claude Code 2.1.78 or later, Python 3.9 or later, and a POSIX-compatible hook shell such as Git Bash on Windows. Use a current Claude Code release. Session Ledger is opt-in and disabled by default because it stores local conversation text.

Run in a terminal outside Claude Code:

claude plugin marketplace update llm-accuracy
claude plugin update session-ledger@llm-accuracy
claude plugin enable session-ledger@llm-accuracy

For a first installation, use claude plugin install session-ledger@llm-accuracy before enabling. Restart Claude Code and confirm Session Ledger 0.2.6 is enabled in /plugin. Use your normal Claude launch, not the temporary QA launcher.

A plan reset does not erase Claude's conversation: later replies can restate old material. Use a new session for isolation. This release improves continuity; it does not establish an LLM accuracy improvement.

Changes and validation · Setup and limits