Skip to content

Releases: cavi-ai/mcp-eval

v0.6.0

Choose a tag to compare

@github-actions github-actions released this 09 Oct 00:01
v0.6.0
283d4bd

Added

  • Readiness reports and history record measurement profiles covering evaluator
    version, platform, architecture, client capabilities, read-only attestation,
    skip rules, timeout, and repetition count. Reports identify structural
    assessment separately from semantic correctness, which requires a reviewed oracle.
  • Corpus targets accept reviewed read-only assessment policies and tool
    exclusions. Collection and replay retain these policies and detect profile
    drift; --require-deployments requires prepared deployment locks.

Changed

  • Readiness score deltas and corpus placement require the same standard and
    known, matching measurement profiles. Older reports remain readable without
    inferring missing conditions; the standard/2 scoring policy is unchanged.
  • Finding verification binds resolved stdio executables and regular file
    arguments by their bytes, with checks before launch and before recording credit.
    Changed launch artifacts restart the passing streak or prevent new credit.
  • Findings tools distinguish unavailable, corrupt, or incomplete index evidence
    from an empty valid index, and report an explicit unavailable-evidence error.
  • Readiness-history writes reject invalid measurement profiles before creating
    or extending a journal. Older profile-less readiness records remain supported.
  • Call, annotation, and readiness-history writes enforce the same 4 MiB encoded
    record limit as journal readers, including JSON escaping and the newline.
    Oversized records are rejected before creating or extending a journal.
  • History indexing reuses content-validated call and annotation checkpoints for
    append-safe journals, with atomic full reconstruction when cached inputs or
    ordering change. index --rebuild explicitly reconstructs derived rows.
  • Readiness history coordinates readers and writers with journal locks and
    reuses a private, content-validated cache for complete append records. Changed
    inputs rebuild atomically; unavailable caches fall back to a locked scan.
  • Share export streams checked records into disk staging instead of retaining
    whole journals in memory. Source and normalized records are limited to 4 MiB;
    staging directories are owner-only on Unix.
  • Doctor streams every nested JSONL journal under a shared lock, limits physical
    lines to 4 MiB, and refuses symlinks and unreadable entries instead of skipping
    parts of the redaction sweep.

v0.5.0

Choose a tag to compare

@github-actions github-actions released this 07 Oct 16:11
v0.5.0
7f331d7

Added

  • Capability-aware guidance checks compare server instructions with tool
    catalogs under five client profiles and identify references to tools that
    are unavailable to the selected profile without exporting instruction prose.
  • Finding-based generate produces class-specific read-only regression cases.
    Semantic findings require a reviewed --expect oracle instead of guessing
    intended behavior from captured outcomes.
  • Bounded read-only workflow probes evaluate 2–32 ordered steps in one fresh
    session, repeated 1–20 times, and stop at the first failed expectation.
  • Expectations support nested JSON Pointer assertions through
    required_result_paths and equals_paths, alongside literal top-level keys.
  • Workflow failure guidance identifies the failed call across CLI, JSON,
    Markdown, SARIF, and MCP results, and explains how to check state leakage
    and verify the full sequence after a repair.
  • Corpus documentation covers pinned collection, replay, saved evidence checks,
    and CI gating from a repository checkout. It distinguishes private target
    and report artifacts from content-minimized capture journals.
  • Dedicated corpus evidence GitHub Action gates on offline integrity and replay
    verdicts, exposes fixed step outputs, and writes a summary without private
    report payloads, labels, or paths.
  • Offline corpus evidence checking validates original input identities, report
    bytes, projected scores, and replay verdicts without launching an evaluator
    or server. Invalid evidence has a distinct exit status from a failed replay.
  • Optional corpus replay --out retains verification JSON and digest-bound
    reports in a private directory, refuses existing outputs, and cleans up
    incomplete bundles on fatal errors. Failed verdicts retain their evidence.
  • Versioned corpus verification JSON binds replay outcomes to corpus/target
    bytes, evaluator identity, comparison policy, and report digests.
    mcpeval schema corpus-verification prints its published schema.
  • Optional corpus service-state checks compare expected digests before and after
    evaluation. Changed or unavailable state prevents observation credit, and
    successful observations retain named state identities for replay.
  • Optional prepared corpus deployments bind runtime and dependency-tree bytes.
    Direct launches bypass package resolution; changed bundles lose observation
    credit during collection and replay.
  • Provenance corpus v3, pinned collection targets, declared prerequisite checks,
    evaluator/report hashes, and explicit observed, untested, and errored outcomes.
    Deterministic corpus fixtures run in CI; no new public measurements are collected.
  • Recorder and event UUIDv4 identities for new captures and probe journals.
    CLI and MCP annotations accept an event ID without legacy session coordinates.
  • Shared five-minute and 4096-request budgets for probe, score, verification,
    and scaffold operations across transports and reconnects. Exhaustion leaves
    unfinished cases errored and readiness unmeasured, without verification credit.
  • MCP service ping, tool safety annotations, and output schemas with
    conforming structured results alongside existing text.
  • Four HTTP workers and a sixteen-connection queue keep queries responsive
    during one admitted evaluation; overlapping evaluations are refused.
    Headers and bodies share a five-second request deadline.
  • Native MCP score and verify_finding tools behind serve --allow-spawn.
    CLI and MCP finding verification share the same execution and durable
    evidence service; agent verification cannot authorize mutation.

Changed

  • Refresh JSON Schema validation, UUID, and public suffix dependencies.
  • Verification state and history live in a durable lifecycle.db, separate
    from the rebuildable index. Existing lifecycle rows are imported once;
    finding-ID migrations retain every historical verification.
  • Verification streaks bind to the selected case, timeout, sandbox declaration,
    target configuration, and evaluator executable fingerprint. Changed bindings
    restart pass credit; legacy unbound history is retained without contributing
    to new streaks. Verification run IDs prevent duplicate credit.
  • Readiness uses mcpeval-standard/2. Reliability is scaled by successful-call
    rate; consistently failing and untested tools earn no reliability credit.
    Coverage counts tools with at least one successful call. Reports include
    success, tool-error, RPC-error, rejected-argument, transport-error, and
    untested-tool counts.
  • Synthesized arguments are validated against the full input schema before
    tool execution. Output schemas are validated fully on every successful
    repeat and by the manifest probe, including types, nested constraints, and
    local references. External schema retrieval is disabled.
  • share omits annotation prose by default; --include-annotation-notes
    explicitly includes manually reviewed notes. --force replaces the
    entire envelope.
  • The existing corpus retains its historical mcpeval-standard/1 label;
    readiness comparisons with the new standard are omitted.

Fixed

  • Incomplete tool discovery cannot scaffold a manifest, produce readiness,
    execute ordinary probe cases, or add verification credit. Failed later
    pages and continuation beyond the 20-page limit are errors; pagination
    probes remain available to diagnose the broken listing.
  • Active evaluation rejects non-object tool results and non-boolean isError
    flags as transport errors, without inventing successful outcomes. Passive
    capture retains unknown outcomes without exporting malformed payloads.
  • False-success verification requires a result assertion or expected error
    in an instruction-fidelity case or workflow step. Success-only expectations
    and unrelated probes cannot close semantic findings.
  • Annotated successful calls can promote false-success findings. Finding
    identity and verification history remain distinct from ordinary error groups.
  • Corpus collection and replay reject missing, empty, or incorrectly typed
    source/platform metadata and non-string standards under the existing v3
    contract. Zero-observation candidates remain valid collection artifacts.
  • Failure windows cannot cross recorder captures or legacy server boundaries.
    Ambiguous legacy sequences produce no windows, and ambiguous annotations
    cannot affect finding classification. Identical modern event replays count
    once; conflicting identities reject the rebuild and retain the prior index.
  • Stdio evaluation bounds frames, incoming queues, and retained notifications;
    pipe writes time out when a server stops reading. Interleaved messages do not
    restart response deadlines, and shutdown cancels and joins both I/O pumps.
  • Input synthesis uses byte and node budgets before allocating strings,
    arrays, or schema-supplied defaults.
  • share scans the exact exported snapshots, including nested JSONL files,
    rejects symlinks and overlapping output paths, and preserves an existing
    envelope when preparation fails.
  • Wrongly typed structured output fails with output-schema-invalid-result
    instead of passing a field-presence check.

Removed

v0.4.0

Choose a tag to compare

@github-actions github-actions released this 01 Oct 22:32
v0.4.0
8dafabc

Changed

  • readiness is an absolute score under mcpeval-standard/1: after the
    manifest's cases, a read-only standard battery scores the server's whole
    catalog on fixed curves (protocol, catalog, context, error honesty,
    reliability, coverage). Surface it cannot test counts against the score.
    The manifest is the gate: it alone sets passed and the exit code.
  • Probe reports are mcpeval.probe-report/v2 (gate, readiness with
    areas and lost checks, readiness_error); diffs are
    mcpeval.probe-diff/v2. report and diff still read v1 documents.
  • diff marks readiness not comparable across standards or against a v1
    document; trends compares scores only within one standard.
  • The shipped corpus is mcpeval.readiness-corpus/v2: readiness scores and
    area scores from mcpeval score under mcpeval-standard/1. Text and
    markdown reports (probe and score) place a readiness score among it
    only when the corpus names the report's standard; v1 corpora are not
    read. scripts/corpus/collect.sh and scripts/corpus/verify.mjs run
    every server with no Kubernetes context and no Docker daemon.
    verify.mjs fails when any area other than reliability moves, or
    reliability moves by more than 10 (one latency band on every tool), and
    runs only on the platform the corpus records.
  • The corpus-drift workflow is removed: re-scoring live public servers is
    not a deterministic CI gate. Run scripts/corpus/verify.mjs before a
    release.
  • README and installation docs list the Homebrew, npm, and Cargo install
    commands, and the quickstart runs the installed mcpeval and
    mcpeval-demo instead of ./target/release paths.

Added

  • mcpeval score runs the standard battery without a manifest.
  • probe --gate-only and compare --gate-only skip the standard battery;
    probe --confirm-read-only and score --confirm-read-only attest
    unannotated tools as read-only for it.
  • mcpeval-demo --broken flaky: read_counter fails every second call.
  • Readiness catalog area: per tool, a description of at least 40
    characters, described and typed input properties, declared
    readOnlyHint (and destructiveHint on writers), and a declared
    outputSchema.
  • mcpeval-demo --broken undescribed: 20-character descriptions and
    undescribed input properties.
  • Text and markdown reports print each lost check's hint once per reason.
  • Readiness protocol area: unknown methods answer -32601, ping answers,
    unknown tools are refused, tools/list pages cleanly, declared
    resources and prompts list cleanly, and version negotiation holds.
  • mcpeval-demo --broken unknown-method: unknown methods answer a result.
    The clean demo answers ping and unknown methods with -32601.
  • Readiness error-honesty area: a read-only tool called with arguments
    that violate its input schema must refuse them with -32602 or an
    isError result that says why, and the session must answer the next
    request.
  • mcpeval-demo refuses arguments of the wrong type with -32602;
    --broken lying-errors accepts them.
  • The standard battery synthesizes arguments from each read-only tool's
    input schema, so tools with required parameters are exercised; a
    required string with only a pattern scores coverage-unsynthesizable.
  • probe --skip-tool and score --skip-tool keep a tool uncalled; a
    skipped tool the battery would have called scores 0 for reliability and
    error honesty, and any skip scores contention and payload 0, so skipping
    never raises the score.

Fixed

  • probe and init read every tools/list page (up to 20), as the
    standard battery does: a manifest may name a tool listed after the first
    page instead of exiting 2, init scaffolds tools from every page, and
    discovery-cost and token-cost measure the whole catalog. A page after
    the first that fails ends the catalog; pagination reports it.
  • protocol-negotiation no longer fails a server that answers the
    supported version with an older version it supports.
  • payload-bounds and contention decline requests a tool sends before
    it answers, instead of failing the case; contention's second client
    finds a tool listed on a later tools/list page.
  • mcpeval-demo tracks replies to its mid-call sampling and elicitation
    requests under every --broken aspect. Under the other aspects the reply
    was answered as a new request and broke the next call.
  • The npm package installs mcpeval-demo next to mcpeval. It
    previously extracted only mcpeval from the release archive, so the
    quickstart's demo server was missing on npm installs.
  • mcpeval share exits 1 when the redaction sweep flags a file, per the
    exit-code contract for gates. It previously exited 3.

v0.3.0

Choose a tag to compare

@github-actions github-actions released this 23 Sep 15:13
v0.3.0
806c22f

Security

  • mcpeval serve validates every request: Host must name loopback
    (403, or 400 when missing or repeated), an Origin must be loopback
    (403), and Content-Type must be application/json (415). Previously
    a web page could POST to the loopback endpoint cross-origin, or through
    DNS rebinding, and have run_probe launch an arbitrary command.
  • run_probe and scaffold launch the server process an agent names, so
    serve lists and runs them only with the new --allow-spawn flag.
    Agent-loop setups that use them must add the flag.
  • mcpeval shim-http validates every request before forwarding it, with
    the same checks as serve: Host must name loopback (403, or 400 when
    missing or repeated), an Origin must be loopback (403), and a POST's
    Content-Type must be application/json (415). Previously it
    forwarded any request, its Origin included, to the upstream server.

Added

  • mcpeval findings --format json and serve's finding tools carry
    retryable: true when every failure of the group was retryable,
    false when none was, null when mixed or unreported.
  • GitHub Action outputs passed, readiness, report, exit-code,
    diff-exit-code, and sarif.
  • The GitHub Action renders the report as markdown into the job summary.
  • GitHub Action baseline input runs mcpeval diff --fail-on-regression
    against a committed report (fail-on-change: 'true' adds
    --fail-on-change) and appends the diff to the job summary; the step
    fails with the diff's exit code when the probe passed.
  • GitHub Action sarif: 'true' renders mcpeval.sarif and uploads it
    through github/codeql-action/upload-sarif@v4, also when the probe fails.
  • GitHub Action report-path input sets where the JSON report is written
    (default mcpeval.report.json).
  • mcpeval init --tool <NAME> (repeatable) restricts the candidates to the
    named tools; a name the catalog lacks, or one init cannot call (annotated
    as a writer, required arguments, or unattested), exits 2.
  • mcpeval init --dry-run prints each catalog tool's decision; it calls no
    tool, writes no file, and skips the --output existence check.
  • mcpeval.probe-report/v1 gains optional fields: generator (name and
    version), manifest_sha256 (SHA-256 of the manifest bytes the run
    parsed), and per case tool, hint (the remediation for its reason), and
    detail (the declared bound a bound-based failure exceeded, with its
    limit and the observed value).
  • Published JSON Schemas for the report and diff documents:
    docs/mcp-eval.probe-report.schema.json and
    docs/mcp-eval.probe-diff.schema.json, printed by mcpeval schema report
    and mcpeval schema diff (mcpeval schema still prints the manifest
    schema).
  • Manifest timeout_ms (100 to 600000): how long each request waits for
    its response, over stdio and HTTP.
  • completion probe: for a server declaring the completions capability,
    one completion/complete request for the manifest's reference
    (ref_type/ref_uri) and argument must answer a well-formed completion
    — completion.values, an array of strings — within max_values
    (1..=100). A structured error naming the argument is
    completion-argument-unknown (the completion surface out of sync with
    the prompt's own declarations); a malformed envelope is
    completion-invalid-request; more than max_values is
    completion-value-flood; a transport failure is
    completion-stalled-request. Undeclared support passes trivially.
    Scores under the contract category. The demo server gained a
    --broken completion personality (non-string completion values) and
    its welcome prompt now declares a language argument served by
    completion/complete. Manifest JSON schema, remediation hints, and
    official docs cover the probe.
  • Corpus drift check: node scripts/corpus/verify.mjs re-probes every
    observation in data/readiness-corpus.json with the current binary and
    exits non-zero on any score that moved; a corpus-drift workflow runs it
    weekly and on pull requests touching the corpus or the probe battery, so
    "reproducible by anyone" stays an enforced property rather than a claim.
    Agreement tests pin the drift check's launch commands to the collector's
    arrays.
  • Corpus grew from 19 to 33 observations: data/readiness-corpus.json now
    holds 33 public servers collected with scripts/corpus/collect.sh —
    fifteen additional credential-free servers across the npm and uvx
    ecosystems (airbnb, sqlite, docker, mermaid, terraform, tavily, ollama,
    calculator, wikipedia, searxng, git, arxiv), minus mcp-atlassian, which
    lists no tools without credentials. Every prior observation reproduced
    byte-identically on re-run; every observation now carries catalog
    measurements.
  • mcpeval.readiness-corpus/v1 gains optional fields: top-level
    battery (the probe kinds every observation was scored on; defaults to
    discovery-cost, token-cost, pagination, surface-listing; an empty
    list is rejected) and per observation tool_count and catalog_tokens.
    scripts/corpus/collect.sh keeps each server's JSON report and writes all
    three; scripts/corpus/verify.mjs still compares scores only. An agreement
    test pins the collector's battery to the drift check's manifest.
  • Catalog placement line in text and markdown reports when the corpus
    carries catalog_tokens and the report has a token-cost measurement:
    catalog: 566 tokens over 12 tools, lighter than 23 of 33 observed servers (median 1186 tokens).
  • Findings carry class (unstable-error-code, false-success,
    blocked-optimal-path, recovers-on-retry, retry-did-not-recover,
    recurring-error), a one-line hint, and err_codes (every distinct
    code, sorted). findings --format agent adds class= and a hint:
    line, and codes=[…] on the cause line when there are several; md
    adds Class, Hint, and Error codes; export-issues adds a
    Diagnosis section; serve's list_findings adds class= and
    get_finding adds the hint.
  • mcpeval findings with no promoted findings prints no promoted findings; run mcpeval promote --threshold 0 to see every issue to
    stderr and exits 0.

Changed

  • mcpeval --help orders subcommands battery-first (init, probe,
    report, diff, compare, explain, schema, trends, serve),
    then capture (shim, shim-http, index, promote, findings,
    generate, verify, export-issues, annotate), then hygiene
    (doctor, share), and prints a Start here: block with the
    init → probe → shim path. The top-level about line now names
    both the CI probe battery and friction capture, instead of capture
    alone.
  • The GitHub Action installs the release archive pinned by its
    distribution/release.json for the runner's platform, verifying the
    .sha256 companion's digest and file name, the pinned SHA-256, and the
    size, instead of running cargo install mcpeval. version installs
    another release, verified against that release's own release.json. A
    preinstalled mcpeval is still used as is.
  • The GitHub Action passes every input to its scripts through environment
    variables; no input is interpolated into shell code.
  • The GitHub Action's command input accepts a JSON array of strings for
    arguments that contain spaces; any other value is split on whitespace
    without quoting. Setting both or neither of command and url exits 2.
  • The GitHub Action writes the rendered markdown report (and the baseline
    diff, when baseline is set) to a markdown output path next to the
    JSON report, instead of only the job summary.
  • The GitHub Action's baseline and sarif inputs need mcpeval 0.3.0
    or later (diff and report --manifest); the pinned release moves
    when distribution/release.json is updated at release time. On an
    older mcpeval, the step now fails fast with a usage error naming the
    installed version instead of a bare usage error from the probe.
  • mcpeval init and serve's scaffold tool also scaffold
    pagination, protocol-negotiation, and surface-listing (when
    resources or prompts are declared); with --confirm-read-only, each
    tool that answers {} also gets degradation-over-n, latency-budget,
    and output-schema (when declared), plus one contention and one
    payload-bounds case. The summary line reports the case count per probe
    kind instead of the schema-guessability count. Non-identifier tool names
    fall back to case-<kind>-<n> ids (case-schema-guessability-<n>, was
    case-schema-<n>).
  • mcpeval init and serve's scaffold tool never call or scaffold a
    tool annotated destructiveHint: true or readOnlyHint: false, with or
    without --confirm-read-only.
  • Without --confirm-read-only, init scaffolds per-tool cases for
    zero-required tools annotated readOnlyHint: true.
  • --confirm-read-only adds the unannotated zero-required tools to the
    candidates.
  • The init summary line appends the number of tools skipped by
    annotations.
  • mcpeval-demo annotates readOnlyHint: true on its read tools,
    destructiveHint: true on break_session, and readOnlyHint: false on
    recover_session; its catalog measures 2262 bytes and 566 estimated
    tokens (was 1802 and 451).
  • mcpeval-demo's slow_read description states its 400 ms delay (was
    200 ms).
  • mcpeval diff reports a case that fails in both documents for different
    reasons as changed (text CHANGED <old> → <new>, JSON verdict
    changed, summary count changed) instead of unchanged. The new
    --fail-on-change flag exits non-zero on any changed case;
    --fail-on-regression is unchanged.
  • mcpeval diff refuses two documents that name different servers.
  • SARIF results carry a location: the manifest file (relative to the
    working directory) and the failing case's line. Each run declares
    automationDetails.id mcpeval/<server>/, and th...
Read more

v0.2.0

Choose a tag to compare

@github-actions github-actions released this 10 Sep 18:54
8b028d0

Added

  • Four 2025-06-18 spec-surface probes:
    protocol-negotiation (a fresh handshake with the supported version is
    echoed, an unknown date-shaped version is answered with a date-shaped
    non-echoed version, and the claimed version is itself echoable on a
    third handshake),
    sampling (a tool call under a client sampling capability; the server's
    sampling/createMessage sub-requests are answered with a stub sample,
    bounded by max_requests),
    elicitation (same shape for elicitation/create, answered with the
    manifest's respond action), and
    resource-subscription (for a server declaring resources.subscribe:
    read, subscribe, trigger, and require notifications/resources/updated
    within max_wait_seconds, then unsubscribe cleanly; undeclared support
    passes trivially). Negotiation, sampling, and elicitation score under the
    contract category; resource-subscription under reliability. The demo
    server gained --broken negotiation, --broken sampling,
    --broken elicitation, and --broken subscription personalities and
    sampled_read, elicited_read, and publish_status tools.
  • cancellation probe: issues a read-only call, cancels it with
    notifications/cancelled, and requires the server to acknowledge the
    cancellation — silence for the cancelled request id within grace_seconds
    or the structured "Request cancelled" error (-32800) that production
    servers return; preflight confirms the tool succeeds uncancelled. Works on
    stdio and Streamable HTTP (the call POST runs on its own connection with
    grace_seconds as its timeout, so the cancellation lands mid-flight). The
    demo server gained a --broken cancellation aspect that drops the
    notification, and its slow tool now honors cancellation through a
    dedicated stdin reader thread.
  • mcpeval compare accepts one stdio command after -- as an additional
    comparison column labeled stdio, alongside --endpoint HTTP targets;
    two or more targets are required.
  • State of MCP servers guide: the readiness corpus
    (data/readiness-corpus.json) holds 16 public servers collected with
    scripts/corpus/collect.sh, and the distribution is published with its
    method notes.
  • npm package @cavi-ai/mcp-eval and Homebrew formula Formula/mcpeval.rb,
    both pinned to the v0.1.0 release archives and SHA-256 digests recorded in
    distribution/release.json; the npm installer verifies the checksum
    companion, the pinned archive size, and the SHA-256 before extracting the
    binary.
  • CI: distribution contract check, npm install end-to-end on Linux and
    Windows, and Homebrew audit, install, and test on macOS.

Changed

  • npm publication uses trusted publishing (OIDC) with provenance; no npm
    token is stored in the repository.

mcp-eval v0.1.0

Choose a tag to compare

@sasan1200 sasan1200 released this 30 Aug 19:29
ffc0b2d

First tagged release.

  • 13 deterministic probes (five headline evaluation dimensions, eight supplemental)
  • Readiness score calibrated against a corpus of public MCP servers
  • Remediation hints for every failure reason (mcpeval explain)
  • Friction capture -> findings lifecycle -> verification closing loop
  • Native agent loop over MCP (run_probe, scaffold, findings, trends)
  • SARIF 2.1.0 for code-scanning annotations; re-renderable baselines
  • Share envelope, session-cost interpretation, self-tested GitHub Action

Install: cargo install mcpeval --locked
Prebuilt binaries with SHA256 checksums are attached below.

What's Changed

  • feat: expand resilience probe coverage by @sasan1200 in #5
  • feat: add concurrent-client probe by @sasan1200 in #6
  • feat: add Streamable HTTP probe transport by @sasan1200 in #7
  • feat: add local HTTP capture proxy by @sasan1200 in #8
  • feat: generate verification manifests by @sasan1200 in #9
  • feat: add official release docs pipeline by @sasan1200 in #10
  • chore: ignore local artifacts in open-source repo by @sasan1200 in #11
  • feat: add streamable HTTP probe transport by @sasan1200 in #12
  • feat: token-cost probe, JSON probe reports, and CI by @sasan1200 in #13
  • feat: expand probe battery operations by @sasan1200 in #14
  • feat: expand probe battery — init scaffolding, readiness scoring, markdown reports, new probes, compare, trends, issue export by @sasan1200 in #15
  • feat: serve findings over MCP, GitHub Action, and CI guide by @sasan1200 in #16
  • feat: binary release pipeline, crates.io publish, and release runbook by @sasan1200 in #17
  • feat: bundled demo server and stdio interop hardening from dogfooding by @sasan1200 in #18
  • fix: release pipeline depends on the correct gate job id by @sasan1200 in #19
  • feat: payload-bounds, surface-listing, and output-schema probes by @sasan1200 in #20
  • feat: remediation hints and calibrated readiness scores by @sasan1200 in #21
  • feat: native agent loop — run_probe and scaffold tools on serve by @sasan1200 in #22
  • feat: session-cost translation and the share envelope by @sasan1200 in #23
  • feat: downstream landing — SARIF, report re-rendering, agent wiring config by @sasan1200 in #24

Full Changelog: https://github.com/cavi-ai/mcp-eval/commits/v0.1.0