Repository navigation
Releases: lizhuojunx86/traceguard
Release list
v1.6.0
[1.6.0] - 2026-09-10
SemVer minor, on four counts: SPEC v1.1 → v1.2 (a new nullable traces
column and a new Span method — §6.2 / §6.3), a new experimental opt-in
extension (traceguard.sources), a new audit finding kind
(capture_unmatched), and a new extra (traceguard[anchors]). No existing
signature moves, the normalize algorithm and the audit algo v1 envelope are
untouched, and the frozen 29-symbol public surface is unchanged.
SPEC v1.2 — where a number came from, and whether anyone else can check it.
Motivation and compatibility analysis:
docs/spec-changes/2026-09-10-source-snapshot-approval-binding.md. The four
invariants cover the model, the prompt and the feature cutoff, and stop at the
edge of the process: a pipeline can use the right model on the right prompt with
the right feature_as_of, be handed a vendor value that was rewritten into
existence weeks later, and pass every check while being wrong. v1.2 adds the
three pieces that were missing — a record of what was retrieved and when
(traceguard.sources), a per-request join key for reconciling self-reported
calls against an out-of-band ledger (provider_response_id), and an export a
recipient can verify without the database, the network, or this package
(evidence-bundle/v1).
Added
-
traceguard.sources— experimental opt-in extension (SPEC v1.2 §6.6):
point-in-time correctness for retrieved data, the one thing the four
invariants did not cover. A pipeline could use the right model, the right
prompt and the rightfeature_as_of, be fed a vendor value rewritten into
existence weeks later, and pass every invariant while being wrong. Published
measurement of exactly that: 41.4% of vendorepsActualvalues differ
between first sight and today, 15.3% flip a binary entry decision
(analysis/eps_revision.pyrecomputes both offline).- New contract-external table
source_snapshots(ownDeclarativeBase;
created only bysources.enable(engine), never bymake_engine).
Indexed ontrace_id,source_uri,content_hash,retrieved_at— the
last two are what answer "which traces depended on content the source
later rewrote?". Span.record_source(snapshot, *, strict)— a NEW method (SPEC §6.3
minor); no existingSpansignature changes.strictis keyword-only
with no default, matchingselect_modeldiscipline.SourceVerdict:verified/anachronistic/unverifiable/
unchecked, following therouting_integrityfour-way precedent. When a
source states nopublished_at, strict mode refuses ("cannot establish
that the source existed at feature_as_of") and loose mode records
unverifiable— being unable to prove existence is not proof of absence,
and neither is it a pass.unchecked(nofeature_as_ofat the call
site) stays distinct fromunverifiable(the source would not say).- Builders:
content_digest(raw bytes, no normalization — whitespace and
key order are part of what was served),from_http_response(duck-typed;
imports neither httpx nor requests),from_mcp_result(an MCP result is
already parsed, so the digest covers §4.4 canonical bytes and
normalizer_idsays so).normalized_hashandnormalizer_idare
required together,<name>@<version>— an unnamed normalizer's digest is
not comparable, and worse, it looks comparable. - No retrieved content is stored — digests and metadata only. Contract
intent, not a mode; archiving the bytes is the consumer's business, joined
oncontent_hash. - CLI:
python -m traceguard.sources --db URL enable|list(listexits 1
on any actionable verdict, so it can gate CI). - Honest layering, including what a snapshot cannot prove (that the host
actually fetched the bytes fromsource_uri; thatpublished_atis
true):docs/sources.md. - Its API surface is deliberately not in the contract-guard job while
experimental — the fields have to survive real use first (revision
decision D9; revisit after two minors).
- New contract-external table
-
SPEC §5 invariant 3 now names retrieved external data in its scope, with
source_snapshots.published_atas itsvalid_from. No fifth invariant is
added — that wording was already a general principle. It does add one new
refusal condition under invariant 3 (strict mode refuses a source stating
novalid_from), scoped totraceguard.sources; existing instances are
untouched, sincevalidate_reference_timingtakes a definitevalid_from
and has no "unknown" state. Classified minor under §6.4, with that
clause's warn-first ramp carried bystrictbeing keyword-only with no
default: no existing call site can be silently converted into an error,
because there are none and there is no default to decide for them. -
SPEC §6.6 registers
traceguard.approvalas planned, not implemented
(interface, single-useapproval_id, float-free payloads). Implementation
is gated on a real consumer. -
traces.provider_response_id(SPEC §3.1 v1.2): nullable, indexed — the
join key for per-request reconciliation.wrap_anthropicrecords
message.id,wrap_openairecordsresponse.id(chat and responses).
Streaming calls leave it NULL: the wrappers do not drain the stream, so there
is no final message to take an id from, and a synthesized id would reconcile
as a real call.routing_audit.ingest_claude_codefills it from the
transcript's API message id; lines that fell back touuid:<line uuid>for
identity leave it NULL rather than store a locally-minted value that would
look like a provider's. Rows ingested earlier are not backfilled — same
posture asagent_id/session_idin 1.5.0; the id stays inoutput_parsed.
Legacy databases get the column and its index on open
(ensure_trace_columns), and its failure message now names the SPEC revision
that introduced the specific missing column. -
traceguard.audit.reconcile_requests()+ therequest-ledger/v1format —
capture-fidelity layer L1.5. L1 compares totals, which can cancel: an
under-reported call and an over-reported one net out, and the provider usage
API supplies token volume but no call counts. L1.5 asks what totals cannot —
is this call present on both sides — by joiningprovider_response_id
against an out-of-band request ledger. CLI:reconcile --source requests-json:PATH. Any out-of-band source that can export the documented
shape works; no specific gateway interface is promised. -
New finding kind
capture_unmatched(WARN) carrying adirectionof
out_of_band_only("a call the capture layer did not see — bypass or wrapper
coverage gap") orself_reported_only("a record the provider side does not
vouch for — fabrication, duplication, or an incomplete ledger"). A new kind
is a SemVer minor under the 2026-08-27 revision's rule A;FINDING_SEVERITY
and the frozen table intests/test_audit_api_surface.pywere updated
deliberately, andChainFindinggained a defaulteddirectionfield (§6.3,
additive). The aggregatecapture_mismatchpath is unchanged — the two kinds
prove different things, and one kind for both would give them one threshold
and one response.- Three disciplines, each because a false positive teaches people to ignore
the check (SPEC B3.4): a ledger whose declaredwindowdoes not cover the
requested one is refused rather than reporting the uncovered stretch as
missing; traces with a NULLprovider_response_idare counted and excluded,
never reported as fabricated; and aresponse_idappearing twice on one
side is reported as a duplicate, not as something missing.
- Three disciplines, each because a false positive teaches people to ignore
-
traceguard.sourcesgaineddrift(A5): the retrieval-to-retrieval
content_hashsequence persource_uri, summarised as a rate with its n and
a Wilson 95% interval. The counting discipline follows
analysis/eps_revision.pyand is enforced by tests:uncheckedsnapshots are
not observations, a source retrieved only once is in neither side of the rate,
and an empty denominator reportsNonerather than 0.0. CLI:python -m traceguard.sources --db URL drift [--source-uri] [--since] [--json].- A source whose digests changed without two observations is reported as
sources_changed_uncomparablerather than vanishing: it cannot enter the
numerator without a denominator it has not earned, and silence there
re-opens the same hole by another door.
- A source whose digests changed without two observations is reported as
-
traceguard.audit.export_bundle()/verify_bundle()and the
evidence-bundle/v1format (docs/specs/evidence-bundle.md, JSON Schema
docs/specs/evidence-bundle-v1.schema.json): a self-contained JSON document —
selected traces, the chain segment covering them, the chain head, anchors,
source snapshots, findings — that a recipient verifies with no database, no
network and no signature verification. CLI:bundle --out PATH [--since/--until/--trace-ids] [--hash-only] [--anchor-file PATH]and
verify-bundle PATH(exit 1 on BREAK). A test assertsverify-bundlenever
opens a database.hash_onlyis honest about being weaker. The algo v1 envelope covers
content fields; strip them and entry hashes cannot be recomputed at all. So
ahash_onlyverify checks chain LINKAGE and the head against an anchor
only, always emits acontent_not_recomputed(INFO) finding, and its
summary reads "LINKAGE OK (hash_only) … content was NOT recomputed" where
fullreads "VERIFIED (full)". The two modes never share a word — reading a
hash_onlypass as "content verified" is this format's most consequential
misreading, and a test pins the wording apart.content_not_recomputedis a
bundle-level annotation, not an audit finding kind: it is not in
FINDING_SEVERITYand not bound by the §6.6 kind freeze.- Anchors are structure-checked, never verified. An
rfc3161anchor is
re...
v1.5.0 — the audit layer becomes contract, and traces get an owner
SPEC v1.0 → v1.1, and the audit layer graduates. The trigger was the 2026-08-26 METR/Redwood investigation of the OpenAI–Hugging Face incident: roughly 7% of the evaluated transcripts had spoofed tool calls, and agents researched how to spoof, edit or delete their own transcripts. A hash chain answers "was the stored record changed afterwards?" and not "was the record true?".
SPEC v1.1
traces.agent_id/traces.session_id— nullable, indexed identity dimensions for correlating several executors after the fact.span/trace/wrap_anthropic/wrap_openaitake them keyword-only;TRACEGUARD_AGENT_ID/TRACEGUARD_SESSION_IDare the fallback. An existing database gets the columns on open (ALTER TABLE ADD COLUMN, nullable, plus the index; well under a second on a 58,769-row copy of the routing_audit store). They take no part ininput_hashor the invariants, and they sit outside the audit algo v1 hash envelope — append-only under the guard, not attested by the chain.traceguard.auditis no longer experimental. Its__all__, finding kinds and severities, public function parameters and the three boundary statements in docs/audit.md are frozen bytests/test_audit_api_surface.pyin the contract-guard CI job. Both golden test files are byte-identical to 1.4.0.
audit v2
- Anchor sinks:
FileAnchorSink/GitNoteAnchorSink/WebhookAnchorSink,anchor_to()(tries every sink, then raises if any failed),AnchorScheduler. CLI:anchor --sink SPEC [--every SECONDS],verify --anchor-file PATH. The interval is the exposure window; the docs say so. reconcile: self-reported token volume per model and UTC bucket against the provider's usage report (Usage Admin API, or a saved JSON). New finding kindcapture_mismatch(WARN), with the direction spelled out: traces > provider means self-reports the provider never served; provider > traces means traffic the SDK never recorded. Tokens only — the Usage API has no request counts. CLI:reconcile --source anthropic-usage|json:PATH --window START,END.- Claude Code ingest fills the two columns; the OTel exporter emits them as
traceguard.agent_id/traceguard.session_idand asgen_ai.agent.id/session.id.
Zero new dependencies. The 29-symbol public surface is unchanged. Full entry: packages/traceguard/CHANGELOG.md; motivation and compatibility analysis: docs/spec-changes/2026-08-27-audit-v2-correlation-schema.md.
PyPI: https://pypi.org/project/traceguard/1.5.0/ (built and published by publish.yml via Trusted Publishing; artifacts carry provenance attestations).
v1.4.0 — a cache audit that can leave the laptop
Published to PyPI on 2026-08-19 (https://pypi.org/project/traceguard/1.4.0/) from tag v1.4.0; this GitHub Release entry was added on 2026-08-28. The notes below are the [1.4.0] section of packages/traceguard/CHANGELOG.md, verbatim.
Contract-external, all inside traceguard.routing_audit. Additive: sections
1-4 and 3b of the cache audit render byte-for-byte as before, verified against
main on the reference store for all three of --format table|md|csv.
One release, one problem: making a cache audit leave the laptop. 1.3.0
finished the report and left it stranded — the numbers were only ever true of
one machine, and nothing in the output let a second person check whether their
numbers and yours described the same thing. Every item below exists so that two
sets of figures can be put beside each other and the comparison mean something.
They are listed in the order that question gets answered: get the numbers out,
say which traffic they came from, stop them changing afterwards, say when they
were taken, and give them somewhere to go.
--emit-share / --show-share: a cache-audit summary you can hand to
someone else. Comparing cache behaviour across organisations needs a corpus,
and a corpus needs people willing to send a file. That is a trust problem
rather than a serialisation one, so the export is shaped around the reader
rather than the writer: --show-share prints the exact bytes --emit-share
would write, in full, so there is nothing to find out afterwards. The tool
makes no network calls and has no upload path.
-
Aggregates only. Per-model hit rates and token volumes, gap-bucket counts
and per-bucket cost, the cap band, both ends of the net-benefit range, the
cross-model switch rate, and the undecidable count. No prompt text, no
paths, no session ids, no per-trace timestamps, no free-form strings. -
The invariant, not a list of things remembered. Every string in the
payload is a schema constant, a whitelisted model id, one of the two window
bounds, the installed version, or a decimal money literal.model_idis
whitelisted against the published price sheet rather than passed through,
because an arbitrary model id can name an internal gateway or an employer;
anything else folds into one(unrecognized)row that keeps the counts and
drops the name. A store poisoned with sentinel strings in prompts, paths,
session ids, model names and unplanned-for keys is asserted to leak none of
them, and two further tests exist only to prove that assertion can still fail. -
An open window is refused, not warned about. Every rate and dollar figure
scales with how long you looked, so--since/--until(or--benchmark)
are mandatory for an export. A corpus whose members each measured "all time"
is not comparable with itself, and nothing downstream can repair that. -
The band is the citable field.
recommended_cap_bandnames the answer
andargmax_reference_onlynames the point estimate, so quoting the argmax
reads wrong at the call site. Net is emitted at both ends of the
undecidable-gap assumption (measuredandpessimistic), never as one
number. -
tool_versionis read from installed package metadata, never from a
literal in this repo. A submission claiming a version it was not produced by
would corrupt the corpus in a way nobody could detect later, and this repo
has already had to fix one hand-copied number that went stale. -
New: cross-model switch rate by gap-length decile, each group carrying
its own bounds in minutes. On the reference store the rate runs from 0% at
60-72 minutes of idle to 23.1% past 24 hours, which is the empirical reason
an optimal keep-alive cap is finite rather than infinite. A single overall
rate averages that away. -
corpus.fingerprint, because a closed window is necessary and not
sufficient. The window closes over timestamps; the store keeps growing
inside it, sinceingestwalks~/.claude/projectsand a transcript that
only appears later carries messages timestamped weeks ago. The reference
--benchmarkwindow went from 432 expired gaps over 168 sessions to 439 over
174 in under a day without moving by a second, and the argmax net moved from
$811.30 to $806.82 with it. The fingerprint is a sha256 over one tuple per
trace the window loaded (session, timestamp, model, prompt and output volume,
source), sorted so row order does not matter and length-delimited so two
corpora cannot serialise to the same bytes. Session ids are hashed into it and
none comes back out: one digest covers the whole set and no per-record digest
is emitted anywhere. The first version ofbenchmark/README.mdtold people to
pick a window and keep it, full stop; that advice was a necessary condition
sold as a sufficient one and has been corrected in place. -
--emit-sharerefuses to overwrite an existing file and exits 2. A corpus
entry is an immutable record of one corpus, not a document that gets a new
version: its numbers can already have been cited from a path that cannot be
updated, and overwriting would rewrite them while the path kept pointing at
what looks like the same thing. Entries are named
NNN-<source>-<first 8 of corpus.fingerprint>.json, so a re-run over grown
traffic is a new entry sitting beside the old one rather than replacing it. -
generated_at/settling_days: when you pulled it, not just what you
pulled.corpus.fingerprintshows that two submissions are different
corpora and cannot show why.settling_days— export time minus
window.until— can: a file exported the day a window shut reports a tail
still filling in, one exported a month later reports a tail that has probably
stopped. Proposed by Boris Dzhingarov, from the Google Search Console
version of the same bug, where the last days of any date range are
provisional and quietly revise themselves after export so a weekly report
never reconciles with the previous one. Neither field feeds the fingerprint:
mixing a clock into it would make every re-run of unchanged traffic look like
fresh traffic.settling_daysmay be negative and is not validated away —
--until <date>closes at 23:59:59.999999, so a 09:00 export legitimately
sits ten hours before its own bound, and a rule that refused it would pass at
23:59 and fail at 09:00, which is the opposite of reproducible.
benchmark/ — schema specification, submission criteria, and the first
entry. It holds one file and says plainly that under 20 submissions any
cross-organisation number in it is an anecdote.
v1.3.0 — cap sweep, measured model switch, band-not-argmax
Published to PyPI on 2026-08-18 (https://pypi.org/project/traceguard/1.3.0/) from tag v1.3.0; this GitHub Release entry was added on 2026-08-28. The notes below are the [1.3.0] section of packages/traceguard/CHANGELOG.md, verbatim.
Contract-external, all inside traceguard.routing_audit. The frozen public
surface, the SPEC MUSTs and every existing signature are untouched.
Read this as one conclusion overturned three times, not as a feature list.
1.2.0 shipped a keep-alive counterfactual that answered "should you ping to
hold the cache open?" with a single verdict over a single number. A reader
recomputed it and the answer changed; recomputing it twice more changed it
twice more. Each correction was larger than the thing it corrected, which is
why the report no longer prints a single number at all.
1 — The aggregate verdict was averaging two populations with opposite signs.
1.2.0 said NOT WORTH IT and stopped. Splitting the money by gap bucket shows
that verdict is the sum of a win and a loss, not a finding: 1–4h gaps pay for
themselves ($82.65 of pings against $896.02–$900.03 of avoidable rewrites)
while >4h gaps drown them ($1,962.07 against $1,036.71–$1,041.43). A verdict
averaged over buckets that behave differently is not a decision — which was the
failure the 1.2.0 write-up was itself about, one level down. Section 3 now also
prices a capped policy, one you could actually run: ping until some threshold
of idle, then give up, paying for the pings burned on gaps that outlive the cap
and banking only the ones it bridges.
2 — The cap that made the split look good was hand-picked. That threshold
was 4h because 4h is the 1-4h / >4h bucket boundary; the code comment said
so. New section 3b costs every cap from 1h to 12h in 15-minute steps, plus an
uncapped policy competing on equal terms, and takes the argmax: 10h, not 4h
($811.30 net against $569.10 at 4h). In the same pass, _rewrite_cost — an
upper bound with nothing underneath it — got a floor, so every verdict now
compares against an interval instead of one side of one. Verdicts are
three-state accordingly: below the floor WORTH IT, above the ceiling NOT WORTH
IT, and UNDECIDED in between, a band the two-state version silently scored
as a win. session_gaps and audit take a cap argument; the default solves
it.
That floor moves this corpus by only 0.4% ($1,941.46 → $1,932.73), and that is
not reassurance: a post-gap write averages ~350,000 tokens against a ~1,500-
token session baseline, so the interval is narrow because the two populations
differ by two orders of magnitude, not because either bound is tight. The
report says so where it prints the number.
3 — The argmax that replaced 4h was a point estimate, and a single
measurement moved it further than its own lead. Section 3 had been listing
"caches are model-scoped, so a mid-session model switch makes the preceding
pings worthless" as a stated approximation — while both model_ids sat in
the store the whole time. Measured: 18 of 378 decidable expired gaps (4.8%)
came back on a different model, and the rate climbs with idle time, 1.8% in
1-4h against 7.1% in >4h. Those gaps bank nothing while still costing what
they cost, so their savings are deducted before the argmax is taken. The
direction is the whole point — a longer cap collects a larger share of exactly
the gaps this removes, so omitting it had not added noise, it had pushed the
cap systematically long.
The deduction moved 10h by $18.76, while 10h leads the runner-up by
$7.63. A correction bigger than the gap between first and second place is
enough to reorder them, and more remain unquantified — the 54 gaps with a NULL
model_id, the prompt volume frozen at the pre-gap message, the ping cadence
held at 55m and never swept.
So the output is a band and an interval, not a number. Section 3b closes
with the one line meant to be quoted:
RECOMMENDED CAP: 9h..12h (cadence 55m). Within this band the choice costs under
10% of the optimum, which is less than the size of corrections still outstanding.
The argmax stays in the table, marked "not for citation". Two ranges are
reported because they answer different questions and one footnote had been
claiming both: the sign-stable range (net > 0) is 1h15m..12h, 44 grid
points, and says only that capping is the right shape of policy — net inside it
spans $102.13..$811.30, 8x, so it emphatically does not say the caps are
interchangeable. The argmax neighbourhood (net within k of the maximum,
k default 0.10, --peak-band-tolerance) is 9h..12h, 13 points, and is the
one that does. Where a range ends at a grid edge it is flagged censored, and
the flag carries the marginal evidence rather than a shrug: nothing above the
argmax recovers to it, drift out to 12h is -$48.83, and the single observation
beyond the grid is a further -$950.86 — a peak past 12h is unsupported, not
excluded.
Because the cross-model deduction only removes gaps proven to have switched,
the headline is the optimistic end of a range, so the sweep runs the other end
too: treating every undecidable gap as cross-model, the argmax stays at 10h and
its net falls to $663.69. The truth is between the two runs and the report
declines to say where.
Also
--benchmarkpins a frozen reporting window (2026-05-30..2026-08-16)
and refuses to combine with--since/--until. The store is appended to
continuously and the expired-gap count drifted 429 → 432 across one afternoon
of editing; copying the DB aside fixes one comparison, a closed window fixes
every future run. Every number quoted above and in the README comes from it.- Section 2 gained per-bucket money, the rewrite bracket and a measured model
switch column reading3 of 166 (1.8%), 23 unknown— the unknowns are
deliberately outside the denominator, which the earlier shorthand hid. - Sections 1 and 4 are byte-identical to 1.2.0 throughout all of the above.
- Money is still never guessed: unpriced models, unpriced speed tiers and
gaps with no comparablemodel_idare counted and left out of the totals.
Fixed
python -m traceguard.routing_audit.ingest_claude_codesilently did
nothing. The module holds the parser while the CLI lives iningest, and
it had no__main__guard, so the documented command imported it and exited
0 without a word — indistinguishable from a backfill that worked. The
package README documented exactly that command, and without--writebesides.
The module now delegates to the real CLI, the README shows
python -m traceguard.routing_audit.ingest --write, and a test runs--help
through every documented entry point.
v1.2.0 — prompt-cache audit + wrapper metering fix
[1.2.0] - 2026-08-16
Minor: a read-only prompt-cache efficiency audit under traceguard.routing_audit,
and a metering correction in wrap_anthropic. The contract is untouched — the
frozen 29-symbol surface, all SPEC MUSTs, the normalize algorithm and every
existing signature are unchanged, and no dependency was added. routing_audit
is contract-external and tokens_in is a nullable non-MUST field whose value
convention the SPEC does not pin, so neither change reaches the major bar
(SPEC §6).
Read the wrap_anthropic entry first if you already have a store of traces.
The same API call now records a larger tokens_in than it did in ≤1.1.1, and
old traces are not migrated: on cache-heavy traffic the upgrade boundary is a
break in the series, not a step in it.
Fixed
-
wrap_anthropicseverely under-recordedtokens_inon cached traffic.
This is a recorded-metrics semantics change: the same API call now writes
a largertokens_inthan it did in ≤1.1.1. Traces written by earlier
versions are not migrated and are not comparable to new ones for any
cache-heavy workload — re-derive fromoutput_parsed.usagewhere it exists,
or treat the boundary as a break in the series.The Messages API reports three mutually exclusive input counts:
input_tokenscovers only the uncached prefix, whilecache_read_input_tokens
andcache_creation_input_tokenscover the rest. The wrapper recorded
input_tokensalone, so any prompt served from cache was counted as a
fraction of its real size. Agent-shaped traffic is the worst case: in this
repo's own routing-audit corpus,opus-4-8shows 5.1M bare input tokens
against 5,318M cache-read — an under-count of roughly three orders of
magnitude, which propagates into any per-token rate, cost estimate or
routing decision computed from the affected traces.tokens_inis now the sum of the three, i.e. full prompt volume. That is
the conventiontraceguard.routing_audit.ingest_claude_codeand
routing_audit.rerunhave always used, and the wrapper was the one writer
that disagreed.tokens_outis unchanged.The streaming branch is deliberately untouched: usage is not available until
the caller drains the stream, so it still recordsparse_status="partial"
with no tokens rather than a false zero.wrap_openaiwas checked and is correct as-is — OpenAI's convention is
the opposite.usage.prompt_tokens(Responses:input_tokens) already
includes the cached prefix, whichprompt_tokens_details.cached_tokens
reports as a subset; summing there would double-count. No metering change on
that side.
Added
-
traceguard.routing_audit.cache_audit— a read-only prompt-cache
efficiency audit.python -m traceguard.routing_audit.cache_audit --db sqlite:///traces_routing_audit.db, with--format table|md|csvand an
optional--since/--untilwindow. It opens the store with SQLite
mode=roand never writes; likerouting_audit.rerunit emits only
aggregates, token counts and money — no prompt or answer text.It answers "your prompt cache hit rate is low" in four sections: per-model
token-weighted hit rate with input-side list cost against a no-cache
counterfactual; the distribution of gaps between consecutive requests inside
asession_id(<5m / 5m–1h / 1–4h / >4h) with an upper bound on
cache-expiry rewrite cost; a keep-alive-ping counterfactual (one ping per 55
minutes across every >1h gap, billed at the 0.1× read multiplier) with an
explicit worth-it / not-worth-it verdict; and the non-claude_code_session
traffic, checked against each model's minimum cacheable prefix so a
structurally-uncacheable 0% is not mistaken for a misconfiguration.It does not ingest — point it at a store
ingest_claude_codealready
filled. No new price table either: every figure comes from
routing_audit.pricing(price_for, so Sonnet 5's two price eras resolve by
invoked_at, andcache_creation_split, so both the flat store shape and
the nested transcript shape reconcile). A model with no list price, or a
speed tier with no published price, keeps its token counts and reportsn/a
money rather than a guess.The one new constant is
MIN_CACHEABLE_TOKENS(Opus 5 / Fable 5 512, Opus
4.8 / Sonnet 5 1,024, Opus 4.7 2,048, Haiku 4.5 4,096), read from the
Anthropic prompt-caching reference on 2026-08-16. It is not monotonic
across generations — 4.7 needs twice its successor's prefix — so a model
absent from it reportsunknowninstead of being interpolated from a
neighbour. -
Both wrappers now record the per-kind usage split under
output_parsed["usage"], so cost can be recomputed from the store instead of
only at write time.For
wrap_anthropicthe keys are flat and match the block written by
routing_audit.ingest_claude_codeexactly —input_tokens,
output_tokens,cache_read_input_tokens,cache_creation_input_tokens,
cache_creation_5m,cache_creation_1h,service_tier,speed— which
makes a wrapper-produced trace directly consumable by
routing_audit.pricing.compute_cost_usd(itscache_creation_splitreads
this flat shape), including the 2× one-hour cache-write multiplier. The API
reports the TTL split nested underusage.cache_creation; the wrapper
flattens it on write, as ingest does.For
wrap_openaithe keys are OpenAI's own (prompt_tokens/
completion_tokens/cached_tokens, andinput_tokens/
output_tokens/cached_tokenson the Responses API) — deliberately not
remapped onto the Anthropic names, since the pricing table covers only
claude-*models and there is no reader for a cross-provider key convention.No public API changed and no cost is computed inside either wrapper: pricing
stays arouting_auditconcern (contract-external), per the layering rule.
v1.1.1
Patch release: a concurrency fix in the audit evidence layer.
Contract untouched — the frozen 29-symbol public surface, all SPEC MUSTs,
the normalize algorithm and the chain hash algorithm (v1) are unchanged, and no
dependency was added.
Fixed
traceguard.audit.enable() could raise
OperationalError: table audit_chain_entries already exists when several
processes reached the first enable() on the same database at once.
ensure_audit_tables handled the TOCTOU race with a single retry, assuming the
loser re-runs create_all and finds everything present. That misses the case
that actually bites: create_all walks several tables and the walk is not
atomic, so two racers can lose to each other on different tables in turn — A
creates t1 while B fails on t1; B retries and reaches t2 just as A gets
there. The second collision escapes an except-block wrapping only one retry.
It now retries in a bounded loop and treats "every audit table is present" as
success regardless of which racer created which table — the postcondition the
function actually promises.
The window is narrow: the existing two-thread regression test passes 30/30
locally and still surfaced the failure once on a slower CI runner. A new
higher-concurrency test (8 racers × 12 rounds) reds reliably against the
single-retry implementation; the original test is unchanged.
PR: #32
Note
First release published through PyPI Trusted Publishing (OIDC, no long-lived
token), added in #31 — the artifacts carry a provenance attestation.
v1.1.0 — audit evidence layer
Post-1.0 first cut of the evidence layer: a new opt-in, experimental submodule traceguard.audit (SemVer minor, purely additive, zero new dependencies).
Three honestly-layered capabilities
- ORM-layer append-only guard (anti-mistake) — blocks accidental ORM
UPDATE/DELETEandsession.execute(update(Trace)...)/delete(Trace)bulk DML againsttraces.cost_usd-only updates pass (the legal reprice path). Core SQL, raw drivers, legacy bulk APIs and dialect upserts bypass it — documented, not claimed otherwise. - Row hash chain (tamper-evident, not tamper-proof) — every ORM-inserted trace chains
sha256(prev_hash || canonical(entry metadata + content));cost_usdis excluded from the envelope and evidenced via chained cost events.verify_chain()runs two passes with a BREAK/WARN/GAP taxonomy. - Exportable anchor —
export_anchor()emits the chain head digest for storage outside the DB. Without an external anchor, full-chain rewrite / tail truncation are undetectable (no MAC in v1).
Guarantees
- Purely additive: frozen 29-symbol surface, all SPEC MUSTs, the normalize algorithm and every signature untouched;
contract-guardgreen. - Import has zero side effects; activation is explicit (
enable()/attach()), scoped to attached engines. - Fail-open by default (SPEC §4.1), SAVEPOINT-isolated, bounded retries on head races; opt-in strict via
enable(strict=True)/TRACEGUARD_AUDIT_STRICT=1. - Canonical serialization (algo v1) frozen by golden tests.
Quality
- 73 new tests (full suite 353), hardened by a two-round adversarial multi-agent review (21 confirmed findings all fixed).
- New
docs/audit.mdwith the honest three-tier threat model.
PyPI: https://pypi.org/project/traceguard/1.1.0/ · PR: #29
🤖 Generated with Claude Code
v1.0.0 — the freeze-flip
The freeze-flip. Zero functional changes — no new symbols, no signature changes, no behavioural changes; every 0.9.0 test passes unmodified.
This release turns the contract that 0.8.0–0.9.0 already enforced — SPEC §3–5 MUSTs including invariant 4, fail-open instrumentation, the 29-symbol curated surface guarded by the required contract-guard CI job — into a formal SemVer commitment: from 1.0.0, breaking any public signature or MUST item requires a major version bump (SPEC §6).
Soak evidence behind the flip
- Two external consumers writing real traces since 0.9.0: huadian (via the guardian bridge) and quant_alpha_v2 (via manual spans — catching 206 real look-ahead violations in a 2016–2026 backtest scan).
- A two-week quiet period (2026-06-29 → 2026-07-12) in which the frozen surface needed zero fixes: all interim commits were confined to the contract-external
routing_auditmodule.
Changed
- Development Status classifier:
3 - Alpha→5 - Production/Stable. - SPEC status flipped from Draft v0.3 to v1.0 (
TRACEGUARD_SPEC.md/docs/SPEC.md); the SPEC version now tracks the package major. The Chinese original remains authoritative. - ROADMAP marks the 1.0 definition as achieved;
INTEGRATING.mdno longer says the SDK is unreleased.
Note: this is the first wheel that ships
traceguard.routing_audit. It is contract-external and experimental (off the frozen surface, likeexporters.otel/contamination/loop) — importing it implies no SemVer promise.
Full details: CHANGELOG.md
v0.9.0
The release that turned on traceguard's first external adoption (Phase 0 acceptance #7): huadian (126 traces via the guardian bridge) and quant_alpha_v2 (252 traces via manual spans, 206 look-ahead violations caught in a 2016–2026 backtest scan). No breaking changes — public surface only grows (SemVer minor); Development Status stays Alpha (the 1.0 freeze-flip remains gated on soak).
Added
feature_as_ofonwrap_openai/wrap_anthropic— adatetime, a per-call callable (replay many points in time without touching thecreate()call site), orNone. Makes wrapper traces checkable by the look-ahead invariants (SPEC §3) — "tracing" → "look-ahead protection". Fail-open on a raising callable or naive datetime.resolve_feature_as_of(new public symbol) — the wrappers' point-in-time resolution, for consumers instrumenting by hand (ownTracer.span, e.g. a no-SDK / bare-httpx client).traceguard.bridges.guardian.write_trace_from_guardian— opt-in, off-the-frozen-surface bridge from apipeline-guardiancheckpoint to a trace. Duck-types guardian (never imports it), fail-open, carriesfeature_as_of. Adopt traceguard with ~5 lines without changing a pinned guardian dependency.register_model(if_exists="error" | "ignore")— idempotent re-registration.examples/manual_span.py— bare-client recipe + the sync-CM / async-body pattern.
Changed
- Required
contract-guardCI job (frozen surface + golden hashes + invariants) and branch-protectedmain— the freeze is now enforced by mechanism. Build excludes* 2.pysync-conflict copies.
Upgrade
pip install -U traceguard / uv add traceguard==0.9.0. Drop-in — all 0.2.0–0.8.1 signatures unchanged. Public surface is now 29 symbols (+resolve_feature_as_of).
v0.8.1
Patch release: one adoption-blocking bugfix in the SDK wrappers. No API changes — public __all__ (28 symbols) and all signatures unchanged (SemVer patch). Development Status classifier stays Alpha; the 1.0 freeze-flip remains gated on soak.
Fixed
- Wrapped clients are now transparent to
copy.deepcopy/copy.copy. A client returned bywrap_openai/wrap_anthropicpreviously raisedRecursionErrorwhen copied (andTypeErroron the engine-backed tracer once past it): the delegating__getattr__forwarded the__setstate__/__reduce_ex__dunders the copy/pickle protocol probes on a half-constructed instance to a not-yet-set delegate attribute, recursing forever. Frameworks such as LangChain / LlamaIndex deep-copy LLM clients, so a wrapped client crashed where the raw client would not. Delegation is now factored into a private_DelegatingWrappermixin that raisesAttributeErrorfor private/dunder lookups (so copy/pickle falls back cleanly) and implements__deepcopy__sharing the process-levelTracerby reference while deep-copying the wrapped client. The wrapper no longer adds a copy-time failure the underlying client didn't already have.
Upgrade
pip install -U traceguard / uv add traceguard==0.8.1. Drop-in — no code changes required.
🤖 Generated with Claude Code