Skip to content

Releases: llm-measurement/fleetdiff

v0.3.1: Clearer investigation reports

Choose a tag to compare

@github-actions github-actions released this 02 Oct 03:18
v0.3.1
ba9f4a3

v0.3.1 Release Notes

v0.3.1 makes investigation reports easier to scan and brings the release binary
in line with the README examples. The report starts with the main finding,
then shows contributor tables, coverage, and concise notes.

Try It

With the verified binary installed, run this from a repository checkout:

fleetdiff investigate --before examples/sessions/data/before \
  --after examples/sessions/data/after --expected app

The synthetic sessions example opens with:

1 of 8 tracked sessions flagged for review: 90.91% of attributed tokens.
Reported tokens: 400 -> 3300; model attempts: 4 -> 13.
  +1592 tokens from attempt count; +1308 tokens from tokens per attempt.

Token-change contributions display as whole tokens. JSON retains the calculated
precision; token-per-attempt averages and share bounds keep their existing display
precision. Unequal share bounds remain outward-rounded ranges.

Upgrade

Replace the binary after following the verification instructions.
The four supported targets remain Linux and macOS on AMD64 and ARM64.

Existing summary files, flags, accounting, and JSON version 1 are unchanged.
Session flags still require a lower-bound share strictly above the configured
threshold and complete relevant observations. No collector upgrade or window
reset is needed. Text consumers should use --format json for structured data;
the human-readable report layout has changed.

v0.3.0: User and session contributors

Choose a tag to compare

@github-actions github-actions released this 01 Oct 22:28
v0.3.0
811b1f3

v0.3.0 Release Notes

These notes describe v0.3.0. Verify the published archive's provenance and checksums
before installation; the installer defaults to this version.

What Changes

fleetdiff investigate can answer two more questions when the input summaries
contain compatible measurements:

  • Which tracked users account for tokens or model attempts?
  • Which sessions have a high enough share to investigate?

Use collector v0.3.0 with optional topk_keys, or compatible sketchkit summary
files containing top_users, top_sessions, or their _requests variants.
Fleetdiff still reads summary files, not raw traces, and makes no network calls.

Session candidates require a share lower bound strictly above --flag-share
(default 0.25) and complete relevant observations. The share is of attributed
sketch weight, not all application traffic. Request weighting works when usage is
missing; token-weighted flags are withheld when usage is incomplete. A flag is a
reason to investigate, not proof of a loop, a root cause, or wasted money.

fleetdiff investigate --before before/ --after after/ --expected app \
  --flag-share 0.25

The existing demo fixtures show prompt attribution; they do not fabricate session
measurements. Collect two complete windows with session ranking enabled to answer
the session question. See Investigation.

Compatibility

  • Comparison JSON remains version 1. Text now says "Model attempts" rather than
    "Requests"; the existing requests JSON field is unchanged.
  • Missing optional measurements in either window or any input producer yield
    cannot_determine, never zero. Differing extraction contracts remain errors.
  • Raw identifiers are not recovered from hashes. Hashes remain hidden unless
    explicitly requested. Do not interpret pseudonymous summaries as anonymous.
  • Upgrade the reader before enabling the collector's new rankings. No migration
    or rewrite of saved v0.2.0 summary files is required. New rankings cannot be
    reconstructed for old windows. Older readers cannot answer these questions.

The installer now detects the platform and verifies attestations and checksums
before extraction. Error messages name missing CLI flags without echoing values.
Supported platforms remain Linux/macOS on amd64/arm64. Installation and rollback
instructions are in Operations.

Verification

The release workflow checks each exact archive on its native platform, including
threshold boundaries, missing usage, old windows, and hidden metadata. Linux also
has an offline, unprivileged smoke test. Archives, dependency inventory, checksums,
and provenance must all be verified before publishing the draft release.

This is a pre-1.0 release, not an enterprise support promise. It does not enforce
budgets, stop agent runs, reconstruct raw requests, or produce a billing ledger.

fleetdiff v0.2.0

Choose a tag to compare

@github-actions github-actions released this 27 Sep 00:14
v0.2.0
84a621f

What changed in your agent application?

This release adds fleetdiff investigate: a local, read-only report that starts
with one application and also works across compatible, disjoint producers.

  • Split observed token changes into request-count and tokens-per-request contributions.
  • Compare tracked prompt contributors with lower and upper bounds.
  • Show missing usage and incomplete observation coverage instead of guessing.
  • Read versioned usage provenance. Older exports remain comparable, with their
    missing source evidence reported as unknown.

Download the archive for Linux or macOS, on AMD64 or ARM64. Verify its provenance
before extraction using the installation instructions.

./fleetdiff --version
./fleetdiff investigate --before ./before --after ./after --expected app

The single-app example
includes complete, missing-usage, and two-stack cases. compare remains available;
its JSON contract stays at version 1.

All four native archives passed tests for both commands. Assets include checksums,
an SPDX dependency inventory, build metadata, runtime licenses, and signed build
provenance. macOS binaries are not Apple Developer ID signed or notarized.

This is measurement, not a causal explanation, invoice reconciliation, budget
enforcement, or loop detection. Session-level token attribution remains unavailable.
No account or upload is required. No new sketch algorithm is introduced.

v0.1.1

Choose a tag to compare

@github-actions github-actions released this 26 Sep 00:57
v0.1.1
2bccb93

Local, read-only comparison of agent-fleet usage across separately operated systems.

  • Compare request and token totals, missing token usage, distinct estimates, and tracked prompt-weight changes.
  • Try the saved-file demo or the live two-collector research-agent example, including MCP resources and tool errors.
  • Download binaries for Linux and macOS on AMD64 or ARM64. Checksums, an SPDX dependency inventory, and signed build provenance are included.
  • Tagged Go installs now report the module version rather than dev.

See README.md for the walkthrough, CHANGELOG.md for details, and docs/OPERATIONS.md for installation and provenance verification. Summary sharing requires authorization and compatible measurement contracts; this tool does not enforce budgets or assess task quality.