Repository navigation
Releases: cavi-ai/mcp-eval
Releases · cavi-ai/mcp-eval
Release list
v0.6.0
Added
- Readiness reports and history record measurement profiles covering evaluator
version, platform, architecture, client capabilities, read-only attestation,
skip rules, timeout, and repetition count. Reports identify structural
assessment separately from semantic correctness, which requires a reviewed oracle. - Corpus targets accept reviewed read-only assessment policies and tool
exclusions. Collection and replay retain these policies and detect profile
drift;--require-deploymentsrequires prepared deployment locks.
Changed
- Readiness score deltas and corpus placement require the same standard and
known, matching measurement profiles. Older reports remain readable without
inferring missing conditions; the standard/2 scoring policy is unchanged. - Finding verification binds resolved stdio executables and regular file
arguments by their bytes, with checks before launch and before recording credit.
Changed launch artifacts restart the passing streak or prevent new credit. - Findings tools distinguish unavailable, corrupt, or incomplete index evidence
from an empty valid index, and report an explicit unavailable-evidence error. - Readiness-history writes reject invalid measurement profiles before creating
or extending a journal. Older profile-less readiness records remain supported. - Call, annotation, and readiness-history writes enforce the same 4 MiB encoded
record limit as journal readers, including JSON escaping and the newline.
Oversized records are rejected before creating or extending a journal. - History indexing reuses content-validated call and annotation checkpoints for
append-safe journals, with atomic full reconstruction when cached inputs or
ordering change.index --rebuildexplicitly reconstructs derived rows. - Readiness history coordinates readers and writers with journal locks and
reuses a private, content-validated cache for complete append records. Changed
inputs rebuild atomically; unavailable caches fall back to a locked scan. - Share export streams checked records into disk staging instead of retaining
whole journals in memory. Source and normalized records are limited to 4 MiB;
staging directories are owner-only on Unix. - Doctor streams every nested JSONL journal under a shared lock, limits physical
lines to 4 MiB, and refuses symlinks and unreadable entries instead of skipping
parts of the redaction sweep.
v0.5.0
Added
- Capability-aware
guidancechecks compare server instructions with tool
catalogs under five client profiles and identify references to tools that
are unavailable to the selected profile without exporting instruction prose. - Finding-based
generateproduces class-specific read-only regression cases.
Semantic findings require a reviewed--expectoracle instead of guessing
intended behavior from captured outcomes. - Bounded read-only
workflowprobes evaluate 2–32 ordered steps in one fresh
session, repeated 1–20 times, and stop at the first failed expectation. - Expectations support nested JSON Pointer assertions through
required_result_pathsandequals_paths, alongside literal top-level keys. - Workflow failure guidance identifies the failed call across CLI, JSON,
Markdown, SARIF, and MCP results, and explains how to check state leakage
and verify the full sequence after a repair. - Corpus documentation covers pinned collection, replay, saved evidence checks,
and CI gating from a repository checkout. It distinguishes private target
and report artifacts from content-minimized capture journals. - Dedicated corpus evidence GitHub Action gates on offline integrity and replay
verdicts, exposes fixed step outputs, and writes a summary without private
report payloads, labels, or paths. - Offline corpus evidence checking validates original input identities, report
bytes, projected scores, and replay verdicts without launching an evaluator
or server. Invalid evidence has a distinct exit status from a failed replay. - Optional corpus replay
--outretains verification JSON and digest-bound
reports in a private directory, refuses existing outputs, and cleans up
incomplete bundles on fatal errors. Failed verdicts retain their evidence. - Versioned corpus verification JSON binds replay outcomes to corpus/target
bytes, evaluator identity, comparison policy, and report digests.
mcpeval schema corpus-verificationprints its published schema. - Optional corpus service-state checks compare expected digests before and after
evaluation. Changed or unavailable state prevents observation credit, and
successful observations retain named state identities for replay. - Optional prepared corpus deployments bind runtime and dependency-tree bytes.
Direct launches bypass package resolution; changed bundles lose observation
credit during collection and replay. - Provenance corpus v3, pinned collection targets, declared prerequisite checks,
evaluator/report hashes, and explicit observed, untested, and errored outcomes.
Deterministic corpus fixtures run in CI; no new public measurements are collected. - Recorder and event UUIDv4 identities for new captures and probe journals.
CLI and MCP annotations accept an event ID without legacy session coordinates. - Shared five-minute and 4096-request budgets for probe, score, verification,
and scaffold operations across transports and reconnects. Exhaustion leaves
unfinished cases errored and readiness unmeasured, without verification credit. - MCP service
ping, tool safety annotations, and output schemas with
conforming structured results alongside existing text. - Four HTTP workers and a sixteen-connection queue keep queries responsive
during one admitted evaluation; overlapping evaluations are refused.
Headers and bodies share a five-second request deadline. - Native MCP
scoreandverify_findingtools behindserve --allow-spawn.
CLI and MCP finding verification share the same execution and durable
evidence service; agent verification cannot authorize mutation.
Changed
- Refresh JSON Schema validation, UUID, and public suffix dependencies.
- Verification state and history live in a durable
lifecycle.db, separate
from the rebuildable index. Existing lifecycle rows are imported once;
finding-ID migrations retain every historical verification. - Verification streaks bind to the selected case, timeout, sandbox declaration,
target configuration, and evaluator executable fingerprint. Changed bindings
restart pass credit; legacy unbound history is retained without contributing
to new streaks. Verification run IDs prevent duplicate credit. - Readiness uses
mcpeval-standard/2. Reliability is scaled by successful-call
rate; consistently failing and untested tools earn no reliability credit.
Coverage counts tools with at least one successful call. Reports include
success, tool-error, RPC-error, rejected-argument, transport-error, and
untested-tool counts. - Synthesized arguments are validated against the full input schema before
tool execution. Output schemas are validated fully on every successful
repeat and by the manifest probe, including types, nested constraints, and
local references. External schema retrieval is disabled. shareomits annotation prose by default;--include-annotation-notes
explicitly includes manually reviewed notes.--forcereplaces the
entire envelope.- The existing corpus retains its historical
mcpeval-standard/1label;
readiness comparisons with the new standard are omitted.
Fixed
- Incomplete tool discovery cannot scaffold a manifest, produce readiness,
execute ordinary probe cases, or add verification credit. Failed later
pages and continuation beyond the 20-page limit are errors; pagination
probes remain available to diagnose the broken listing. - Active evaluation rejects non-object tool results and non-boolean
isError
flags as transport errors, without inventing successful outcomes. Passive
capture retains unknown outcomes without exporting malformed payloads. - False-success verification requires a result assertion or expected error
in an instruction-fidelity case or workflow step. Success-only expectations
and unrelated probes cannot close semantic findings. - Annotated successful calls can promote false-success findings. Finding
identity and verification history remain distinct from ordinary error groups. - Corpus collection and replay reject missing, empty, or incorrectly typed
source/platform metadata and non-string standards under the existing v3
contract. Zero-observation candidates remain valid collection artifacts. - Failure windows cannot cross recorder captures or legacy server boundaries.
Ambiguous legacy sequences produce no windows, and ambiguous annotations
cannot affect finding classification. Identical modern event replays count
once; conflicting identities reject the rebuild and retain the prior index. - Stdio evaluation bounds frames, incoming queues, and retained notifications;
pipe writes time out when a server stops reading. Interleaved messages do not
restart response deadlines, and shutdown cancels and joins both I/O pumps. - Input synthesis uses byte and node budgets before allocating strings,
arrays, or schema-supplied defaults. sharescans the exact exported snapshots, including nested JSONL files,
rejects symlinks and overlapping output paths, and preserves an existing
envelope when preparation fails.- Wrongly typed structured output fails with
output-schema-invalid-result
instead of passing a field-presence check.
Removed
- The GitHub Pages documentation site. Documentation is published at
https://cavi-ai.xyz/docs/mcp-eval/ from each release's docs archive.
v0.4.0
Changed
readinessis an absolute score undermcpeval-standard/1: after the
manifest's cases, a read-only standard battery scores the server's whole
catalog on fixed curves (protocol, catalog, context, error honesty,
reliability, coverage). Surface it cannot test counts against the score.
The manifest is the gate: it alone setspassedand the exit code.- Probe reports are
mcpeval.probe-report/v2(gate,readinesswith
areas and lost checks,readiness_error); diffs are
mcpeval.probe-diff/v2.reportanddiffstill read v1 documents. diffmarks readiness not comparable across standards or against a v1
document;trendscompares scores only within one standard.- The shipped corpus is
mcpeval.readiness-corpus/v2: readiness scores and
area scores frommcpeval scoreundermcpeval-standard/1. Text and
markdown reports (probeandscore) place a readiness score among it
only when the corpus names the report's standard; v1 corpora are not
read.scripts/corpus/collect.shandscripts/corpus/verify.mjsrun
every server with no Kubernetes context and no Docker daemon.
verify.mjsfails when any area other than reliability moves, or
reliability moves by more than 10 (one latency band on every tool), and
runs only on the platform the corpus records. - The
corpus-driftworkflow is removed: re-scoring live public servers is
not a deterministic CI gate. Runscripts/corpus/verify.mjsbefore a
release. - README and installation docs list the Homebrew, npm, and Cargo install
commands, and the quickstart runs the installedmcpevaland
mcpeval-demoinstead of./target/releasepaths.
Added
mcpeval scoreruns the standard battery without a manifest.probe --gate-onlyandcompare --gate-onlyskip the standard battery;
probe --confirm-read-onlyandscore --confirm-read-onlyattest
unannotated tools as read-only for it.mcpeval-demo --broken flaky:read_counterfails every second call.- Readiness catalog area: per tool, a description of at least 40
characters, described and typed input properties, declared
readOnlyHint(anddestructiveHinton writers), and a declared
outputSchema. mcpeval-demo --broken undescribed: 20-character descriptions and
undescribed input properties.- Text and markdown reports print each lost check's hint once per reason.
- Readiness protocol area: unknown methods answer -32601,
pinganswers,
unknown tools are refused,tools/listpages cleanly, declared
resources and prompts list cleanly, and version negotiation holds. mcpeval-demo --broken unknown-method: unknown methods answer a result.
The clean demo answerspingand unknown methods with -32601.- Readiness error-honesty area: a read-only tool called with arguments
that violate its input schema must refuse them with -32602 or an
isErrorresult that says why, and the session must answer the next
request. mcpeval-demorefuses arguments of the wrong type with -32602;
--broken lying-errorsaccepts them.- The standard battery synthesizes arguments from each read-only tool's
input schema, so tools with required parameters are exercised; a
required string with only apatternscorescoverage-unsynthesizable. probe --skip-toolandscore --skip-toolkeep a tool uncalled; a
skipped tool the battery would have called scores 0 for reliability and
error honesty, and any skip scores contention and payload 0, so skipping
never raises the score.
Fixed
probeandinitread everytools/listpage (up to 20), as the
standard battery does: a manifest may name a tool listed after the first
page instead of exiting 2,initscaffolds tools from every page, and
discovery-costandtoken-costmeasure the whole catalog. A page after
the first that fails ends the catalog;paginationreports it.protocol-negotiationno longer fails a server that answers the
supported version with an older version it supports.payload-boundsandcontentiondecline requests a tool sends before
it answers, instead of failing the case;contention's second client
finds a tool listed on a latertools/listpage.mcpeval-demotracks replies to its mid-call sampling and elicitation
requests under every--brokenaspect. Under the other aspects the reply
was answered as a new request and broke the next call.- The npm package installs
mcpeval-demonext tomcpeval. It
previously extracted onlymcpevalfrom the release archive, so the
quickstart's demo server was missing on npm installs. mcpeval shareexits 1 when the redaction sweep flags a file, per the
exit-code contract for gates. It previously exited 3.
v0.3.0
Security
mcpeval servevalidates every request:Hostmust name loopback
(403, or 400 when missing or repeated), anOriginmust be loopback
(403), andContent-Typemust beapplication/json(415). Previously
a web page could POST to the loopback endpoint cross-origin, or through
DNS rebinding, and haverun_probelaunch an arbitrary command.run_probeandscaffoldlaunch the server process an agent names, so
servelists and runs them only with the new--allow-spawnflag.
Agent-loop setups that use them must add the flag.mcpeval shim-httpvalidates every request before forwarding it, with
the same checks asserve:Hostmust name loopback (403, or 400 when
missing or repeated), anOriginmust be loopback (403), and a POST's
Content-Typemust beapplication/json(415). Previously it
forwarded any request, itsOriginincluded, to the upstream server.
Added
mcpeval findings --format jsonandserve's finding tools carry
retryable:truewhen every failure of the group was retryable,
falsewhen none was,nullwhen mixed or unreported.- GitHub Action outputs
passed,readiness,report,exit-code,
diff-exit-code, andsarif. - The GitHub Action renders the report as markdown into the job summary.
- GitHub Action
baselineinput runsmcpeval diff --fail-on-regression
against a committed report (fail-on-change: 'true'adds
--fail-on-change) and appends the diff to the job summary; the step
fails with the diff's exit code when the probe passed. - GitHub Action
sarif: 'true'rendersmcpeval.sarifand uploads it
throughgithub/codeql-action/upload-sarif@v4, also when the probe fails. - GitHub Action
report-pathinput sets where the JSON report is written
(defaultmcpeval.report.json). mcpeval init --tool <NAME>(repeatable) restricts the candidates to the
named tools; a name the catalog lacks, or one init cannot call (annotated
as a writer, required arguments, or unattested), exits 2.mcpeval init --dry-runprints each catalog tool's decision; it calls no
tool, writes no file, and skips the--outputexistence check.mcpeval.probe-report/v1gains optional fields:generator(name and
version),manifest_sha256(SHA-256 of the manifest bytes the run
parsed), and per casetool,hint(the remediation for its reason), and
detail(the declared bound a bound-based failure exceeded, with its
limit and the observed value).- Published JSON Schemas for the report and diff documents:
docs/mcp-eval.probe-report.schema.jsonand
docs/mcp-eval.probe-diff.schema.json, printed bymcpeval schema report
andmcpeval schema diff(mcpeval schemastill prints the manifest
schema). - Manifest
timeout_ms(100 to 600000): how long each request waits for
its response, over stdio and HTTP. completionprobe: for a server declaring thecompletionscapability,
onecompletion/completerequest for the manifest's reference
(ref_type/ref_uri) and argument must answer a well-formed completion
—completion.values, an array of strings — withinmax_values
(1..=100). A structured error naming the argument is
completion-argument-unknown(the completion surface out of sync with
the prompt's own declarations); a malformed envelope is
completion-invalid-request; more thanmax_valuesis
completion-value-flood; a transport failure is
completion-stalled-request. Undeclared support passes trivially.
Scores under the contract category. The demo server gained a
--broken completionpersonality (non-string completion values) and
itswelcomeprompt now declares alanguageargument served by
completion/complete. Manifest JSON schema, remediation hints, and
official docs cover the probe.- Corpus drift check:
node scripts/corpus/verify.mjsre-probes every
observation indata/readiness-corpus.jsonwith the current binary and
exits non-zero on any score that moved; acorpus-driftworkflow runs it
weekly and on pull requests touching the corpus or the probe battery, so
"reproducible by anyone" stays an enforced property rather than a claim.
Agreement tests pin the drift check's launch commands to the collector's
arrays. - Corpus grew from 19 to 33 observations:
data/readiness-corpus.jsonnow
holds 33 public servers collected withscripts/corpus/collect.sh—
fifteen additional credential-free servers across the npm and uvx
ecosystems (airbnb, sqlite, docker, mermaid, terraform, tavily, ollama,
calculator, wikipedia, searxng, git, arxiv), minusmcp-atlassian, which
lists no tools without credentials. Every prior observation reproduced
byte-identically on re-run; every observation now carries catalog
measurements. mcpeval.readiness-corpus/v1gains optional fields: top-level
battery(the probe kinds every observation was scored on; defaults to
discovery-cost,token-cost,pagination,surface-listing; an empty
list is rejected) and per observationtool_countandcatalog_tokens.
scripts/corpus/collect.shkeeps each server's JSON report and writes all
three;scripts/corpus/verify.mjsstill compares scores only. An agreement
test pins the collector's battery to the drift check's manifest.- Catalog placement line in text and markdown reports when the corpus
carriescatalog_tokensand the report has a token-cost measurement:
catalog: 566 tokens over 12 tools, lighter than 23 of 33 observed servers (median 1186 tokens). - Findings carry
class(unstable-error-code,false-success,
blocked-optimal-path,recovers-on-retry,retry-did-not-recover,
recurring-error), a one-linehint, anderr_codes(every distinct
code, sorted).findings --format agentaddsclass=and ahint:
line, andcodes=[…]on the cause line when there are several;md
addsClass,Hint, andError codes;export-issuesadds a
Diagnosissection;serve'slist_findingsaddsclass=and
get_findingadds the hint. mcpeval findingswith no promoted findings printsno promoted findings; run mcpeval promote --threshold 0 to see every issueto
stderr and exits 0.
Changed
mcpeval --helporders subcommands battery-first (init,probe,
report,diff,compare,explain,schema,trends,serve),
then capture (shim,shim-http,index,promote,findings,
generate,verify,export-issues,annotate), then hygiene
(doctor,share), and prints aStart here:block with the
init→probe→shimpath. The top-levelaboutline now names
both the CI probe battery and friction capture, instead of capture
alone.- The GitHub Action installs the release archive pinned by its
distribution/release.jsonfor the runner's platform, verifying the
.sha256companion's digest and file name, the pinned SHA-256, and the
size, instead of runningcargo install mcpeval.versioninstalls
another release, verified against that release's ownrelease.json. A
preinstalledmcpevalis still used as is. - The GitHub Action passes every input to its scripts through environment
variables; no input is interpolated into shell code. - The GitHub Action's
commandinput accepts a JSON array of strings for
arguments that contain spaces; any other value is split on whitespace
without quoting. Setting both or neither ofcommandandurlexits 2. - The GitHub Action writes the rendered markdown report (and the baseline
diff, whenbaselineis set) to amarkdownoutput path next to the
JSON report, instead of only the job summary. - The GitHub Action's
baselineandsarifinputs needmcpeval0.3.0
or later (diffandreport --manifest); the pinned release moves
whendistribution/release.jsonis updated at release time. On an
oldermcpeval, the step now fails fast with a usage error naming the
installed version instead of a bare usage error from the probe. mcpeval initandserve'sscaffoldtool also scaffold
pagination,protocol-negotiation, andsurface-listing(when
resourcesorpromptsare declared); with--confirm-read-only, each
tool that answers{}also getsdegradation-over-n,latency-budget,
andoutput-schema(when declared), plus onecontentionand one
payload-boundscase. The summary line reports the case count per probe
kind instead of the schema-guessability count. Non-identifier tool names
fall back tocase-<kind>-<n>ids (case-schema-guessability-<n>, was
case-schema-<n>).mcpeval initandserve'sscaffoldtool never call or scaffold a
tool annotateddestructiveHint: trueorreadOnlyHint: false, with or
without--confirm-read-only.- Without
--confirm-read-only,initscaffolds per-tool cases for
zero-required tools annotatedreadOnlyHint: true. --confirm-read-onlyadds the unannotated zero-required tools to the
candidates.- The
initsummary line appends the number of tools skipped by
annotations. mcpeval-demoannotatesreadOnlyHint: trueon its read tools,
destructiveHint: trueonbreak_session, andreadOnlyHint: falseon
recover_session; its catalog measures 2262 bytes and 566 estimated
tokens (was 1802 and 451).mcpeval-demo'sslow_readdescription states its 400 ms delay (was
200 ms).mcpeval diffreports a case that fails in both documents for different
reasons aschanged(textCHANGED <old> → <new>, JSON verdict
changed, summary countchanged) instead ofunchanged. The new
--fail-on-changeflag exits non-zero on any changed case;
--fail-on-regressionis unchanged.mcpeval diffrefuses two documents that name different servers.- SARIF results carry a location: the manifest file (relative to the
working directory) and the failing case's line. Each run declares
automationDetails.idmcpeval/<server>/, and th...
v0.2.0
Added
- Four 2025-06-18 spec-surface probes:
protocol-negotiation(a fresh handshake with the supported version is
echoed, an unknown date-shaped version is answered with a date-shaped
non-echoed version, and the claimed version is itself echoable on a
third handshake),
sampling(a tool call under a clientsamplingcapability; the server's
sampling/createMessagesub-requests are answered with a stub sample,
bounded bymax_requests),
elicitation(same shape forelicitation/create, answered with the
manifest'srespondaction), and
resource-subscription(for a server declaringresources.subscribe:
read, subscribe, trigger, and requirenotifications/resources/updated
withinmax_wait_seconds, then unsubscribe cleanly; undeclared support
passes trivially). Negotiation, sampling, and elicitation score under the
contract category; resource-subscription under reliability. The demo
server gained--broken negotiation,--broken sampling,
--broken elicitation, and--broken subscriptionpersonalities and
sampled_read,elicited_read, andpublish_statustools. cancellationprobe: issues a read-only call, cancels it with
notifications/cancelled, and requires the server to acknowledge the
cancellation — silence for the cancelled request id withingrace_seconds
or the structured "Request cancelled" error (-32800) that production
servers return; preflight confirms the tool succeeds uncancelled. Works on
stdio and Streamable HTTP (the call POST runs on its own connection with
grace_secondsas its timeout, so the cancellation lands mid-flight). The
demo server gained a--broken cancellationaspect that drops the
notification, and its slow tool now honors cancellation through a
dedicated stdin reader thread.mcpeval compareaccepts one stdio command after--as an additional
comparison column labeledstdio, alongside--endpointHTTP targets;
two or more targets are required.- State of MCP servers guide: the readiness corpus
(data/readiness-corpus.json) holds 16 public servers collected with
scripts/corpus/collect.sh, and the distribution is published with its
method notes. - npm package
@cavi-ai/mcp-evaland Homebrew formulaFormula/mcpeval.rb,
both pinned to the v0.1.0 release archives and SHA-256 digests recorded in
distribution/release.json; the npm installer verifies the checksum
companion, the pinned archive size, and the SHA-256 before extracting the
binary. - CI: distribution contract check,
npm installend-to-end on Linux and
Windows, and Homebrew audit, install, and test on macOS.
Changed
- npm publication uses trusted publishing (OIDC) with provenance; no npm
token is stored in the repository.
mcp-eval v0.1.0
First tagged release.
- 13 deterministic probes (five headline evaluation dimensions, eight supplemental)
- Readiness score calibrated against a corpus of public MCP servers
- Remediation hints for every failure reason (mcpeval explain)
- Friction capture -> findings lifecycle -> verification closing loop
- Native agent loop over MCP (run_probe, scaffold, findings, trends)
- SARIF 2.1.0 for code-scanning annotations; re-renderable baselines
- Share envelope, session-cost interpretation, self-tested GitHub Action
Install: cargo install mcpeval --locked
Prebuilt binaries with SHA256 checksums are attached below.
What's Changed
- feat: expand resilience probe coverage by @sasan1200 in #5
- feat: add concurrent-client probe by @sasan1200 in #6
- feat: add Streamable HTTP probe transport by @sasan1200 in #7
- feat: add local HTTP capture proxy by @sasan1200 in #8
- feat: generate verification manifests by @sasan1200 in #9
- feat: add official release docs pipeline by @sasan1200 in #10
- chore: ignore local artifacts in open-source repo by @sasan1200 in #11
- feat: add streamable HTTP probe transport by @sasan1200 in #12
- feat: token-cost probe, JSON probe reports, and CI by @sasan1200 in #13
- feat: expand probe battery operations by @sasan1200 in #14
- feat: expand probe battery — init scaffolding, readiness scoring, markdown reports, new probes, compare, trends, issue export by @sasan1200 in #15
- feat: serve findings over MCP, GitHub Action, and CI guide by @sasan1200 in #16
- feat: binary release pipeline, crates.io publish, and release runbook by @sasan1200 in #17
- feat: bundled demo server and stdio interop hardening from dogfooding by @sasan1200 in #18
- fix: release pipeline depends on the correct gate job id by @sasan1200 in #19
- feat: payload-bounds, surface-listing, and output-schema probes by @sasan1200 in #20
- feat: remediation hints and calibrated readiness scores by @sasan1200 in #21
- feat: native agent loop — run_probe and scaffold tools on serve by @sasan1200 in #22
- feat: session-cost translation and the share envelope by @sasan1200 in #23
- feat: downstream landing — SARIF, report re-rendering, agent wiring config by @sasan1200 in #24
Full Changelog: https://github.com/cavi-ai/mcp-eval/commits/v0.1.0