Skip to content

Releases: lizhuojunx86/traceguard

v1.6.0

Choose a tag to compare

@lizhuojunx86 lizhuojunx86 released this 10 Sep 12:37
ebc88f9

[1.6.0] - 2026-09-10

SemVer minor, on four counts: SPEC v1.1 → v1.2 (a new nullable traces
column and a new Span method — §6.2 / §6.3), a new experimental opt-in
extension (traceguard.sources), a new audit finding kind
(capture_unmatched), and a new extra (traceguard[anchors]). No existing
signature moves, the normalize algorithm and the audit algo v1 envelope are
untouched, and the frozen 29-symbol public surface is unchanged.

SPEC v1.2 — where a number came from, and whether anyone else can check it.
Motivation and compatibility analysis:
docs/spec-changes/2026-09-10-source-snapshot-approval-binding.md. The four
invariants cover the model, the prompt and the feature cutoff, and stop at the
edge of the process: a pipeline can use the right model on the right prompt with
the right feature_as_of, be handed a vendor value that was rewritten into
existence weeks later, and pass every check while being wrong. v1.2 adds the
three pieces that were missing — a record of what was retrieved and when
(traceguard.sources), a per-request join key for reconciling self-reported
calls against an out-of-band ledger (provider_response_id), and an export a
recipient can verify without the database, the network, or this package
(evidence-bundle/v1).

Added

  • traceguard.sources — experimental opt-in extension (SPEC v1.2 §6.6):
    point-in-time correctness for retrieved data, the one thing the four
    invariants did not cover. A pipeline could use the right model, the right
    prompt and the right feature_as_of, be fed a vendor value rewritten into
    existence weeks later, and pass every invariant while being wrong. Published
    measurement of exactly that: 41.4% of vendor epsActual values differ
    between first sight and today, 15.3% flip a binary entry decision
    (analysis/eps_revision.py recomputes both offline).

    • New contract-external table source_snapshots (own DeclarativeBase;
      created only by sources.enable(engine), never by make_engine).
      Indexed on trace_id, source_uri, content_hash, retrieved_at — the
      last two are what answer "which traces depended on content the source
      later rewrote?".
    • Span.record_source(snapshot, *, strict) — a NEW method (SPEC §6.3
      minor); no existing Span signature changes. strict is keyword-only
      with no default, matching select_model discipline.
    • SourceVerdict: verified / anachronistic / unverifiable /
      unchecked, following the routing_integrity four-way precedent. When a
      source states no published_at, strict mode refuses ("cannot establish
      that the source existed at feature_as_of") and loose mode records
      unverifiable — being unable to prove existence is not proof of absence,
      and neither is it a pass. unchecked (no feature_as_of at the call
      site) stays distinct from unverifiable (the source would not say).
    • Builders: content_digest (raw bytes, no normalization — whitespace and
      key order are part of what was served), from_http_response (duck-typed;
      imports neither httpx nor requests), from_mcp_result (an MCP result is
      already parsed, so the digest covers §4.4 canonical bytes and
      normalizer_id says so). normalized_hash and normalizer_id are
      required together, <name>@<version> — an unnamed normalizer's digest is
      not comparable, and worse, it looks comparable.
    • No retrieved content is stored — digests and metadata only. Contract
      intent, not a mode; archiving the bytes is the consumer's business, joined
      on content_hash.
    • CLI: python -m traceguard.sources --db URL enable|list (list exits 1
      on any actionable verdict, so it can gate CI).
    • Honest layering, including what a snapshot cannot prove (that the host
      actually fetched the bytes from source_uri; that published_at is
      true): docs/sources.md.
    • Its API surface is deliberately not in the contract-guard job while
      experimental — the fields have to survive real use first (revision
      decision D9; revisit after two minors).
  • SPEC §5 invariant 3 now names retrieved external data in its scope, with
    source_snapshots.published_at as its valid_from. No fifth invariant is
    added — that wording was already a general principle. It does add one new
    refusal condition under invariant 3 (strict mode refuses a source stating
    no valid_from), scoped to traceguard.sources; existing instances are
    untouched, since validate_reference_timing takes a definite valid_from
    and has no "unknown" state. Classified minor under §6.4, with that
    clause's warn-first ramp carried by strict being keyword-only with no
    default: no existing call site can be silently converted into an error,
    because there are none and there is no default to decide for them.

  • SPEC §6.6 registers traceguard.approval as planned, not implemented
    (interface, single-use approval_id, float-free payloads). Implementation
    is gated on a real consumer.

  • traces.provider_response_id (SPEC §3.1 v1.2): nullable, indexed — the
    join key for per-request reconciliation. wrap_anthropic records
    message.id, wrap_openai records response.id (chat and responses).
    Streaming calls leave it NULL: the wrappers do not drain the stream, so there
    is no final message to take an id from, and a synthesized id would reconcile
    as a real call. routing_audit.ingest_claude_code fills it from the
    transcript's API message id; lines that fell back to uuid:<line uuid> for
    identity leave it NULL rather than store a locally-minted value that would
    look like a provider's. Rows ingested earlier are not backfilled — same
    posture as agent_id/session_id in 1.5.0; the id stays in output_parsed.
    Legacy databases get the column and its index on open
    (ensure_trace_columns), and its failure message now names the SPEC revision
    that introduced the specific missing column.

  • traceguard.audit.reconcile_requests() + the request-ledger/v1 format —
    capture-fidelity layer L1.5. L1 compares totals, which can cancel: an
    under-reported call and an over-reported one net out, and the provider usage
    API supplies token volume but no call counts. L1.5 asks what totals cannot —
    is this call present on both sides — by joining provider_response_id
    against an out-of-band request ledger. CLI: reconcile --source requests-json:PATH. Any out-of-band source that can export the documented
    shape works; no specific gateway interface is promised.

  • New finding kind capture_unmatched (WARN) carrying a direction of
    out_of_band_only ("a call the capture layer did not see — bypass or wrapper
    coverage gap") or self_reported_only ("a record the provider side does not
    vouch for — fabrication, duplication, or an incomplete ledger"). A new kind
    is a SemVer minor
    under the 2026-08-27 revision's rule A; FINDING_SEVERITY
    and the frozen table in tests/test_audit_api_surface.py were updated
    deliberately, and ChainFinding gained a defaulted direction field (§6.3,
    additive). The aggregate capture_mismatch path is unchanged — the two kinds
    prove different things, and one kind for both would give them one threshold
    and one response.

    • Three disciplines, each because a false positive teaches people to ignore
      the check (SPEC B3.4): a ledger whose declared window does not cover the
      requested one is refused rather than reporting the uncovered stretch as
      missing; traces with a NULL provider_response_id are counted and excluded,
      never reported as fabricated; and a response_id appearing twice on one
      side is reported as a duplicate, not as something missing.
  • traceguard.sources gained drift (A5): the retrieval-to-retrieval
    content_hash sequence per source_uri, summarised as a rate with its n and
    a Wilson 95% interval. The counting discipline follows
    analysis/eps_revision.py and is enforced by tests: unchecked snapshots are
    not observations, a source retrieved only once is in neither side of the rate,
    and an empty denominator reports None rather than 0.0. CLI: python -m traceguard.sources --db URL drift [--source-uri] [--since] [--json].

    • A source whose digests changed without two observations is reported as
      sources_changed_uncomparable rather than vanishing: it cannot enter the
      numerator without a denominator it has not earned, and silence there
      re-opens the same hole by another door.
  • traceguard.audit.export_bundle() / verify_bundle() and the
    evidence-bundle/v1 format (docs/specs/evidence-bundle.md, JSON Schema
    docs/specs/evidence-bundle-v1.schema.json): a self-contained JSON document —
    selected traces, the chain segment covering them, the chain head, anchors,
    source snapshots, findings — that a recipient verifies with no database, no
    network and no signature verification
    . CLI: bundle --out PATH [--since/--until/--trace-ids] [--hash-only] [--anchor-file PATH] and
    verify-bundle PATH (exit 1 on BREAK). A test asserts verify-bundle never
    opens a database.

    • hash_only is honest about being weaker. The algo v1 envelope covers
      content fields; strip them and entry hashes cannot be recomputed at all. So
      a hash_only verify checks chain LINKAGE and the head against an anchor
      only, always emits a content_not_recomputed (INFO) finding, and its
      summary reads "LINKAGE OK (hash_only) … content was NOT recomputed" where
      full reads "VERIFIED (full)". The two modes never share a word — reading a
      hash_only pass as "content verified" is this format's most consequential
      misreading, and a test pins the wording apart. content_not_recomputed is a
      bundle-level annotation, not an audit finding kind: it is not in
      FINDING_SEVERITY and not bound by the §6.6 kind freeze.
    • Anchors are structure-checked, never verified. An rfc3161 anchor is
      re...
Read more

v1.5.0 — the audit layer becomes contract, and traces get an owner

Choose a tag to compare

@lizhuojunx86 lizhuojunx86 released this 27 Aug 23:21
d485c83

SPEC v1.0 → v1.1, and the audit layer graduates. The trigger was the 2026-08-26 METR/Redwood investigation of the OpenAI–Hugging Face incident: roughly 7% of the evaluated transcripts had spoofed tool calls, and agents researched how to spoof, edit or delete their own transcripts. A hash chain answers "was the stored record changed afterwards?" and not "was the record true?".

SPEC v1.1

  • traces.agent_id / traces.session_id — nullable, indexed identity dimensions for correlating several executors after the fact. span / trace / wrap_anthropic / wrap_openai take them keyword-only; TRACEGUARD_AGENT_ID / TRACEGUARD_SESSION_ID are the fallback. An existing database gets the columns on open (ALTER TABLE ADD COLUMN, nullable, plus the index; well under a second on a 58,769-row copy of the routing_audit store). They take no part in input_hash or the invariants, and they sit outside the audit algo v1 hash envelope — append-only under the guard, not attested by the chain.
  • traceguard.audit is no longer experimental. Its __all__, finding kinds and severities, public function parameters and the three boundary statements in docs/audit.md are frozen by tests/test_audit_api_surface.py in the contract-guard CI job. Both golden test files are byte-identical to 1.4.0.

audit v2

  • Anchor sinks: FileAnchorSink / GitNoteAnchorSink / WebhookAnchorSink, anchor_to() (tries every sink, then raises if any failed), AnchorScheduler. CLI: anchor --sink SPEC [--every SECONDS], verify --anchor-file PATH. The interval is the exposure window; the docs say so.
  • reconcile: self-reported token volume per model and UTC bucket against the provider's usage report (Usage Admin API, or a saved JSON). New finding kind capture_mismatch (WARN), with the direction spelled out: traces > provider means self-reports the provider never served; provider > traces means traffic the SDK never recorded. Tokens only — the Usage API has no request counts. CLI: reconcile --source anthropic-usage|json:PATH --window START,END.
  • Claude Code ingest fills the two columns; the OTel exporter emits them as traceguard.agent_id / traceguard.session_id and as gen_ai.agent.id / session.id.

Zero new dependencies. The 29-symbol public surface is unchanged. Full entry: packages/traceguard/CHANGELOG.md; motivation and compatibility analysis: docs/spec-changes/2026-08-27-audit-v2-correlation-schema.md.

PyPI: https://pypi.org/project/traceguard/1.5.0/ (built and published by publish.yml via Trusted Publishing; artifacts carry provenance attestations).

v1.4.0 — a cache audit that can leave the laptop

Choose a tag to compare

@lizhuojunx86 lizhuojunx86 released this 28 Aug 00:17

Published to PyPI on 2026-08-19 (https://pypi.org/project/traceguard/1.4.0/) from tag v1.4.0; this GitHub Release entry was added on 2026-08-28. The notes below are the [1.4.0] section of packages/traceguard/CHANGELOG.md, verbatim.


Contract-external, all inside traceguard.routing_audit. Additive: sections
1-4 and 3b of the cache audit render byte-for-byte as before, verified against
main on the reference store for all three of --format table|md|csv.

One release, one problem: making a cache audit leave the laptop. 1.3.0
finished the report and left it stranded — the numbers were only ever true of
one machine, and nothing in the output let a second person check whether their
numbers and yours described the same thing. Every item below exists so that two
sets of figures can be put beside each other and the comparison mean something.
They are listed in the order that question gets answered: get the numbers out,
say which traffic they came from, stop them changing afterwards, say when they
were taken, and give them somewhere to go.

--emit-share / --show-share: a cache-audit summary you can hand to
someone else.
Comparing cache behaviour across organisations needs a corpus,
and a corpus needs people willing to send a file. That is a trust problem
rather than a serialisation one, so the export is shaped around the reader
rather than the writer: --show-share prints the exact bytes --emit-share
would write, in full, so there is nothing to find out afterwards. The tool
makes no network calls and has no upload path.

  • Aggregates only. Per-model hit rates and token volumes, gap-bucket counts
    and per-bucket cost, the cap band, both ends of the net-benefit range, the
    cross-model switch rate, and the undecidable count. No prompt text, no
    paths, no session ids, no per-trace timestamps, no free-form strings.

  • The invariant, not a list of things remembered. Every string in the
    payload is a schema constant, a whitelisted model id, one of the two window
    bounds, the installed version, or a decimal money literal. model_id is
    whitelisted against the published price sheet rather than passed through,
    because an arbitrary model id can name an internal gateway or an employer;
    anything else folds into one (unrecognized) row that keeps the counts and
    drops the name. A store poisoned with sentinel strings in prompts, paths,
    session ids, model names and unplanned-for keys is asserted to leak none of
    them, and two further tests exist only to prove that assertion can still fail.

  • An open window is refused, not warned about. Every rate and dollar figure
    scales with how long you looked, so --since/--until (or --benchmark)
    are mandatory for an export. A corpus whose members each measured "all time"
    is not comparable with itself, and nothing downstream can repair that.

  • The band is the citable field. recommended_cap_band names the answer
    and argmax_reference_only names the point estimate, so quoting the argmax
    reads wrong at the call site. Net is emitted at both ends of the
    undecidable-gap assumption (measured and pessimistic), never as one
    number.

  • tool_version is read from installed package metadata, never from a
    literal in this repo. A submission claiming a version it was not produced by
    would corrupt the corpus in a way nobody could detect later, and this repo
    has already had to fix one hand-copied number that went stale.

  • New: cross-model switch rate by gap-length decile, each group carrying
    its own bounds in minutes. On the reference store the rate runs from 0% at
    60-72 minutes of idle to 23.1% past 24 hours, which is the empirical reason
    an optimal keep-alive cap is finite rather than infinite. A single overall
    rate averages that away.

  • corpus.fingerprint, because a closed window is necessary and not
    sufficient.
    The window closes over timestamps; the store keeps growing
    inside it, since ingest walks ~/.claude/projects and a transcript that
    only appears later carries messages timestamped weeks ago. The reference
    --benchmark window went from 432 expired gaps over 168 sessions to 439 over
    174 in under a day without moving by a second, and the argmax net moved from
    $811.30 to $806.82 with it. The fingerprint is a sha256 over one tuple per
    trace the window loaded (session, timestamp, model, prompt and output volume,
    source), sorted so row order does not matter and length-delimited so two
    corpora cannot serialise to the same bytes. Session ids are hashed into it and
    none comes back out: one digest covers the whole set and no per-record digest
    is emitted anywhere. The first version of benchmark/README.md told people to
    pick a window and keep it, full stop; that advice was a necessary condition
    sold as a sufficient one and has been corrected in place.

  • --emit-share refuses to overwrite an existing file and exits 2. A corpus
    entry is an immutable record of one corpus, not a document that gets a new
    version: its numbers can already have been cited from a path that cannot be
    updated, and overwriting would rewrite them while the path kept pointing at
    what looks like the same thing. Entries are named
    NNN-<source>-<first 8 of corpus.fingerprint>.json, so a re-run over grown
    traffic is a new entry sitting beside the old one rather than replacing it.

  • generated_at / settling_days: when you pulled it, not just what you
    pulled.
    corpus.fingerprint shows that two submissions are different
    corpora and cannot show why. settling_days — export time minus
    window.until — can: a file exported the day a window shut reports a tail
    still filling in, one exported a month later reports a tail that has probably
    stopped. Proposed by Boris Dzhingarov, from the Google Search Console
    version of the same bug, where the last days of any date range are
    provisional and quietly revise themselves after export so a weekly report
    never reconciles with the previous one. Neither field feeds the fingerprint:
    mixing a clock into it would make every re-run of unchanged traffic look like
    fresh traffic. settling_days may be negative and is not validated away —
    --until <date> closes at 23:59:59.999999, so a 09:00 export legitimately
    sits ten hours before its own bound, and a rule that refused it would pass at
    23:59 and fail at 09:00, which is the opposite of reproducible.

benchmark/ — schema specification, submission criteria, and the first
entry. It holds one file and says plainly that under 20 submissions any
cross-organisation number in it is an anecdote.

v1.3.0 — cap sweep, measured model switch, band-not-argmax

Choose a tag to compare

@lizhuojunx86 lizhuojunx86 released this 28 Aug 00:17

Published to PyPI on 2026-08-18 (https://pypi.org/project/traceguard/1.3.0/) from tag v1.3.0; this GitHub Release entry was added on 2026-08-28. The notes below are the [1.3.0] section of packages/traceguard/CHANGELOG.md, verbatim.


Contract-external, all inside traceguard.routing_audit. The frozen public
surface, the SPEC MUSTs and every existing signature are untouched.

Read this as one conclusion overturned three times, not as a feature list.
1.2.0 shipped a keep-alive counterfactual that answered "should you ping to
hold the cache open?" with a single verdict over a single number. A reader
recomputed it and the answer changed; recomputing it twice more changed it
twice more. Each correction was larger than the thing it corrected, which is
why the report no longer prints a single number at all.

1 — The aggregate verdict was averaging two populations with opposite signs.
1.2.0 said NOT WORTH IT and stopped. Splitting the money by gap bucket shows
that verdict is the sum of a win and a loss, not a finding: 1–4h gaps pay for
themselves ($82.65 of pings against $896.02–$900.03 of avoidable rewrites)
while >4h gaps drown them ($1,962.07 against $1,036.71–$1,041.43). A verdict
averaged over buckets that behave differently is not a decision — which was the
failure the 1.2.0 write-up was itself about, one level down. Section 3 now also
prices a capped policy, one you could actually run: ping until some threshold
of idle, then give up, paying for the pings burned on gaps that outlive the cap
and banking only the ones it bridges.

2 — The cap that made the split look good was hand-picked. That threshold
was 4h because 4h is the 1-4h / >4h bucket boundary; the code comment said
so. New section 3b costs every cap from 1h to 12h in 15-minute steps, plus an
uncapped policy competing on equal terms, and takes the argmax: 10h, not 4h
($811.30 net against $569.10 at 4h). In the same pass, _rewrite_cost — an
upper bound with nothing underneath it — got a floor, so every verdict now
compares against an interval instead of one side of one. Verdicts are
three-state accordingly: below the floor WORTH IT, above the ceiling NOT WORTH
IT, and UNDECIDED in between, a band the two-state version silently scored
as a win. session_gaps and audit take a cap argument; the default solves
it.

That floor moves this corpus by only 0.4% ($1,941.46 → $1,932.73), and that is
not reassurance: a post-gap write averages ~350,000 tokens against a ~1,500-
token session baseline, so the interval is narrow because the two populations
differ by two orders of magnitude, not because either bound is tight. The
report says so where it prints the number.

3 — The argmax that replaced 4h was a point estimate, and a single
measurement moved it further than its own lead.
Section 3 had been listing
"caches are model-scoped, so a mid-session model switch makes the preceding
pings worthless" as a stated approximation — while both model_ids sat in
the store the whole time. Measured: 18 of 378 decidable expired gaps (4.8%)
came back on a different model
, and the rate climbs with idle time, 1.8% in
1-4h against 7.1% in >4h. Those gaps bank nothing while still costing what
they cost, so their savings are deducted before the argmax is taken. The
direction is the whole point — a longer cap collects a larger share of exactly
the gaps this removes, so omitting it had not added noise, it had pushed the
cap systematically long.

The deduction moved 10h by $18.76, while 10h leads the runner-up by
$7.63. A correction bigger than the gap between first and second place is
enough to reorder them, and more remain unquantified — the 54 gaps with a NULL
model_id, the prompt volume frozen at the pre-gap message, the ping cadence
held at 55m and never swept.

So the output is a band and an interval, not a number. Section 3b closes
with the one line meant to be quoted:

RECOMMENDED CAP: 9h..12h (cadence 55m). Within this band the choice costs under
10% of the optimum, which is less than the size of corrections still outstanding.

The argmax stays in the table, marked "not for citation". Two ranges are
reported because they answer different questions and one footnote had been
claiming both: the sign-stable range (net > 0) is 1h15m..12h, 44 grid
points, and says only that capping is the right shape of policy — net inside it
spans $102.13..$811.30, 8x, so it emphatically does not say the caps are
interchangeable. The argmax neighbourhood (net within k of the maximum,
k default 0.10, --peak-band-tolerance) is 9h..12h, 13 points, and is the
one that does. Where a range ends at a grid edge it is flagged censored, and
the flag carries the marginal evidence rather than a shrug: nothing above the
argmax recovers to it, drift out to 12h is -$48.83, and the single observation
beyond the grid is a further -$950.86 — a peak past 12h is unsupported, not
excluded.

Because the cross-model deduction only removes gaps proven to have switched,
the headline is the optimistic end of a range, so the sweep runs the other end
too: treating every undecidable gap as cross-model, the argmax stays at 10h and
its net falls to $663.69. The truth is between the two runs and the report
declines to say where.

Also

  • --benchmark pins a frozen reporting window (2026-05-30..2026-08-16)
    and refuses to combine with --since/--until. The store is appended to
    continuously and the expired-gap count drifted 429 → 432 across one afternoon
    of editing; copying the DB aside fixes one comparison, a closed window fixes
    every future run. Every number quoted above and in the README comes from it.
  • Section 2 gained per-bucket money, the rewrite bracket and a measured model
    switch column reading 3 of 166 (1.8%), 23 unknown — the unknowns are
    deliberately outside the denominator, which the earlier shorthand hid.
  • Sections 1 and 4 are byte-identical to 1.2.0 throughout all of the above.
  • Money is still never guessed: unpriced models, unpriced speed tiers and
    gaps with no comparable model_id are counted and left out of the totals.

Fixed

  • python -m traceguard.routing_audit.ingest_claude_code silently did
    nothing.
    The module holds the parser while the CLI lives in ingest, and
    it had no __main__ guard, so the documented command imported it and exited
    0 without a word — indistinguishable from a backfill that worked. The
    package README documented exactly that command, and without --write besides.
    The module now delegates to the real CLI, the README shows
    python -m traceguard.routing_audit.ingest --write, and a test runs --help
    through every documented entry point.

v1.2.0 — prompt-cache audit + wrapper metering fix

Choose a tag to compare

@lizhuojunx86 lizhuojunx86 released this 15 Aug 17:30
ec7b1a7

[1.2.0] - 2026-08-16

Minor: a read-only prompt-cache efficiency audit under traceguard.routing_audit,
and a metering correction in wrap_anthropic. The contract is untouched — the
frozen 29-symbol surface, all SPEC MUSTs, the normalize algorithm and every
existing signature are unchanged, and no dependency was added. routing_audit
is contract-external and tokens_in is a nullable non-MUST field whose value
convention the SPEC does not pin, so neither change reaches the major bar
(SPEC §6).

Read the wrap_anthropic entry first if you already have a store of traces.
The same API call now records a larger tokens_in than it did in ≤1.1.1, and
old traces are not migrated: on cache-heavy traffic the upgrade boundary is a
break in the series, not a step in it.

Fixed

  • wrap_anthropic severely under-recorded tokens_in on cached traffic.
    This is a recorded-metrics semantics change: the same API call now writes
    a larger tokens_in than it did in ≤1.1.1. Traces written by earlier
    versions are not migrated and are not comparable to new ones for any
    cache-heavy workload — re-derive from output_parsed.usage where it exists,
    or treat the boundary as a break in the series.

    The Messages API reports three mutually exclusive input counts:
    input_tokens covers only the uncached prefix, while cache_read_input_tokens
    and cache_creation_input_tokens cover the rest. The wrapper recorded
    input_tokens alone, so any prompt served from cache was counted as a
    fraction of its real size. Agent-shaped traffic is the worst case: in this
    repo's own routing-audit corpus, opus-4-8 shows 5.1M bare input tokens
    against 5,318M cache-read — an under-count of roughly three orders of
    magnitude, which propagates into any per-token rate, cost estimate or
    routing decision computed from the affected traces.

    tokens_in is now the sum of the three, i.e. full prompt volume. That is
    the convention traceguard.routing_audit.ingest_claude_code and
    routing_audit.rerun have always used, and the wrapper was the one writer
    that disagreed. tokens_out is unchanged.

    The streaming branch is deliberately untouched: usage is not available until
    the caller drains the stream, so it still records parse_status="partial"
    with no tokens rather than a false zero.

    wrap_openai was checked and is correct as-is — OpenAI's convention is
    the opposite. usage.prompt_tokens (Responses: input_tokens) already
    includes the cached prefix, which prompt_tokens_details.cached_tokens
    reports as a subset; summing there would double-count. No metering change on
    that side.

Added

  • traceguard.routing_audit.cache_audit — a read-only prompt-cache
    efficiency audit.
    python -m traceguard.routing_audit.cache_audit --db sqlite:///traces_routing_audit.db, with --format table|md|csv and an
    optional --since / --until window. It opens the store with SQLite
    mode=ro and never writes; like routing_audit.rerun it emits only
    aggregates, token counts and money — no prompt or answer text.

    It answers "your prompt cache hit rate is low" in four sections: per-model
    token-weighted hit rate with input-side list cost against a no-cache
    counterfactual; the distribution of gaps between consecutive requests inside
    a session_id (<5m / 5m–1h / 1–4h / >4h) with an upper bound on
    cache-expiry rewrite cost; a keep-alive-ping counterfactual (one ping per 55
    minutes across every >1h gap, billed at the 0.1× read multiplier) with an
    explicit worth-it / not-worth-it verdict; and the non-claude_code_session
    traffic, checked against each model's minimum cacheable prefix so a
    structurally-uncacheable 0% is not mistaken for a misconfiguration.

    It does not ingest — point it at a store ingest_claude_code already
    filled. No new price table either: every figure comes from
    routing_audit.pricing (price_for, so Sonnet 5's two price eras resolve by
    invoked_at, and cache_creation_split, so both the flat store shape and
    the nested transcript shape reconcile). A model with no list price, or a
    speed tier with no published price, keeps its token counts and reports n/a
    money rather than a guess.

    The one new constant is MIN_CACHEABLE_TOKENS (Opus 5 / Fable 5 512, Opus
    4.8 / Sonnet 5 1,024, Opus 4.7 2,048, Haiku 4.5 4,096), read from the
    Anthropic prompt-caching reference on 2026-08-16. It is not monotonic
    across generations
    — 4.7 needs twice its successor's prefix — so a model
    absent from it reports unknown instead of being interpolated from a
    neighbour.

  • Both wrappers now record the per-kind usage split under
    output_parsed["usage"], so cost can be recomputed from the store instead of
    only at write time.

    For wrap_anthropic the keys are flat and match the block written by
    routing_audit.ingest_claude_code exactly — input_tokens,
    output_tokens, cache_read_input_tokens, cache_creation_input_tokens,
    cache_creation_5m, cache_creation_1h, service_tier, speed — which
    makes a wrapper-produced trace directly consumable by
    routing_audit.pricing.compute_cost_usd (its cache_creation_split reads
    this flat shape), including the 2× one-hour cache-write multiplier. The API
    reports the TTL split nested under usage.cache_creation; the wrapper
    flattens it on write, as ingest does.

    For wrap_openai the keys are OpenAI's own (prompt_tokens /
    completion_tokens / cached_tokens, and input_tokens /
    output_tokens / cached_tokens on the Responses API) — deliberately not
    remapped onto the Anthropic names, since the pricing table covers only
    claude-* models and there is no reader for a cross-provider key convention.

    No public API changed and no cost is computed inside either wrapper: pricing
    stays a routing_audit concern (contract-external), per the layering rule.

v1.1.1

Choose a tag to compare

@lizhuojunx86 lizhuojunx86 released this 05 Aug 04:59
3b331d7

Patch release: a concurrency fix in the audit evidence layer.

Contract untouched — the frozen 29-symbol public surface, all SPEC MUSTs,
the normalize algorithm and the chain hash algorithm (v1) are unchanged, and no
dependency was added.

Fixed

traceguard.audit.enable() could raise
OperationalError: table audit_chain_entries already exists when several
processes reached the first enable() on the same database at once.

ensure_audit_tables handled the TOCTOU race with a single retry, assuming the
loser re-runs create_all and finds everything present. That misses the case
that actually bites: create_all walks several tables and the walk is not
atomic, so two racers can lose to each other on different tables in turn — A
creates t1 while B fails on t1; B retries and reaches t2 just as A gets
there. The second collision escapes an except-block wrapping only one retry.

It now retries in a bounded loop and treats "every audit table is present" as
success regardless of which racer created which table — the postcondition the
function actually promises.

The window is narrow: the existing two-thread regression test passes 30/30
locally and still surfaced the failure once on a slower CI runner. A new
higher-concurrency test (8 racers × 12 rounds) reds reliably against the
single-retry implementation; the original test is unchanged.

PR: #32

Note

First release published through PyPI Trusted Publishing (OIDC, no long-lived
token), added in #31 — the artifacts carry a provenance attestation.

v1.1.0 — audit evidence layer

Choose a tag to compare

@lizhuojunx86 lizhuojunx86 released this 17 Jul 00:01
704a3a3

Post-1.0 first cut of the evidence layer: a new opt-in, experimental submodule traceguard.audit (SemVer minor, purely additive, zero new dependencies).

Three honestly-layered capabilities

  1. ORM-layer append-only guard (anti-mistake) — blocks accidental ORM UPDATE/DELETE and session.execute(update(Trace)...)/delete(Trace) bulk DML against traces. cost_usd-only updates pass (the legal reprice path). Core SQL, raw drivers, legacy bulk APIs and dialect upserts bypass it — documented, not claimed otherwise.
  2. Row hash chain (tamper-evident, not tamper-proof) — every ORM-inserted trace chains sha256(prev_hash || canonical(entry metadata + content)); cost_usd is excluded from the envelope and evidenced via chained cost events. verify_chain() runs two passes with a BREAK/WARN/GAP taxonomy.
  3. Exportable anchor — export_anchor() emits the chain head digest for storage outside the DB. Without an external anchor, full-chain rewrite / tail truncation are undetectable (no MAC in v1).

Guarantees

  • Purely additive: frozen 29-symbol surface, all SPEC MUSTs, the normalize algorithm and every signature untouched; contract-guard green.
  • Import has zero side effects; activation is explicit (enable()/attach()), scoped to attached engines.
  • Fail-open by default (SPEC §4.1), SAVEPOINT-isolated, bounded retries on head races; opt-in strict via enable(strict=True) / TRACEGUARD_AUDIT_STRICT=1.
  • Canonical serialization (algo v1) frozen by golden tests.

Quality

  • 73 new tests (full suite 353), hardened by a two-round adversarial multi-agent review (21 confirmed findings all fixed).
  • New docs/audit.md with the honest three-tier threat model.

PyPI: https://pypi.org/project/traceguard/1.1.0/ · PR: #29

🤖 Generated with Claude Code

v1.0.0 — the freeze-flip

Choose a tag to compare

@lizhuojunx86 lizhuojunx86 released this 12 Jul 10:41
8651a69

The freeze-flip. Zero functional changes — no new symbols, no signature changes, no behavioural changes; every 0.9.0 test passes unmodified.

This release turns the contract that 0.8.0–0.9.0 already enforced — SPEC §3–5 MUSTs including invariant 4, fail-open instrumentation, the 29-symbol curated surface guarded by the required contract-guard CI job — into a formal SemVer commitment: from 1.0.0, breaking any public signature or MUST item requires a major version bump (SPEC §6).

Soak evidence behind the flip

  • Two external consumers writing real traces since 0.9.0: huadian (via the guardian bridge) and quant_alpha_v2 (via manual spans — catching 206 real look-ahead violations in a 2016–2026 backtest scan).
  • A two-week quiet period (2026-06-29 → 2026-07-12) in which the frozen surface needed zero fixes: all interim commits were confined to the contract-external routing_audit module.

Changed

  • Development Status classifier: 3 - Alpha → 5 - Production/Stable.
  • SPEC status flipped from Draft v0.3 to v1.0 (TRACEGUARD_SPEC.md / docs/SPEC.md); the SPEC version now tracks the package major. The Chinese original remains authoritative.
  • ROADMAP marks the 1.0 definition as achieved; INTEGRATING.md no longer says the SDK is unreleased.

Note: this is the first wheel that ships traceguard.routing_audit. It is contract-external and experimental (off the frozen surface, like exporters.otel / contamination / loop) — importing it implies no SemVer promise.

Full details: CHANGELOG.md

v0.9.0

Choose a tag to compare

@lizhuojunx86 lizhuojunx86 released this 29 Jun 14:56
403e677

The release that turned on traceguard's first external adoption (Phase 0 acceptance #7): huadian (126 traces via the guardian bridge) and quant_alpha_v2 (252 traces via manual spans, 206 look-ahead violations caught in a 2016–2026 backtest scan). No breaking changes — public surface only grows (SemVer minor); Development Status stays Alpha (the 1.0 freeze-flip remains gated on soak).

Added

  • feature_as_of on wrap_openai / wrap_anthropic — a datetime, a per-call callable (replay many points in time without touching the create() call site), or None. Makes wrapper traces checkable by the look-ahead invariants (SPEC §3) — "tracing" → "look-ahead protection". Fail-open on a raising callable or naive datetime.
  • resolve_feature_as_of (new public symbol) — the wrappers' point-in-time resolution, for consumers instrumenting by hand (own Tracer.span, e.g. a no-SDK / bare-httpx client).
  • traceguard.bridges.guardian.write_trace_from_guardian — opt-in, off-the-frozen-surface bridge from a pipeline-guardian checkpoint to a trace. Duck-types guardian (never imports it), fail-open, carries feature_as_of. Adopt traceguard with ~5 lines without changing a pinned guardian dependency.
  • register_model(if_exists="error" | "ignore") — idempotent re-registration.
  • examples/manual_span.py — bare-client recipe + the sync-CM / async-body pattern.

Changed

  • Required contract-guard CI job (frozen surface + golden hashes + invariants) and branch-protected main — the freeze is now enforced by mechanism. Build excludes * 2.py sync-conflict copies.

Upgrade

pip install -U traceguard / uv add traceguard==0.9.0. Drop-in — all 0.2.0–0.8.1 signatures unchanged. Public surface is now 29 symbols (+resolve_feature_as_of).

PRs #14–#19.

v0.8.1

Choose a tag to compare

@lizhuojunx86 lizhuojunx86 released this 28 Jun 11:09
ff54378

Patch release: one adoption-blocking bugfix in the SDK wrappers. No API changes — public __all__ (28 symbols) and all signatures unchanged (SemVer patch). Development Status classifier stays Alpha; the 1.0 freeze-flip remains gated on soak.

Fixed

  • Wrapped clients are now transparent to copy.deepcopy / copy.copy. A client returned by wrap_openai / wrap_anthropic previously raised RecursionError when copied (and TypeError on the engine-backed tracer once past it): the delegating __getattr__ forwarded the __setstate__ / __reduce_ex__ dunders the copy/pickle protocol probes on a half-constructed instance to a not-yet-set delegate attribute, recursing forever. Frameworks such as LangChain / LlamaIndex deep-copy LLM clients, so a wrapped client crashed where the raw client would not. Delegation is now factored into a private _DelegatingWrapper mixin that raises AttributeError for private/dunder lookups (so copy/pickle falls back cleanly) and implements __deepcopy__ sharing the process-level Tracer by reference while deep-copying the wrapped client. The wrapper no longer adds a copy-time failure the underlying client didn't already have.

Upgrade

pip install -U traceguard / uv add traceguard==0.8.1. Drop-in — no code changes required.

PR #11 (fix) + #12 (release).

🤖 Generated with Claude Code