Repository navigation
Releases: JuliusBrussee/caveman
Release list
Caveman 3.2.0
Caveman 3.2.0 brings the release sweep and subsequent native-install fixes to the installer, plugins, skills, and local runtime.
Cache-safe compression and runtime
- Keep previous compression decisions stable across a conversation, including decisions to send raw bytes. Cache-prefix checks record caveman-caused changes in local request telemetry and
caveman status. - Pass through rate-limit responses with
Retry-After; pin accepted raw retries to their conversation lineage. - Build the proxy and engine with tree-sitter support for TypeScript, JavaScript, Python, Rust, Java, C, and C++ in the six-platform binary matrix. The container remains a pure-Go build.
- Route bare
/v1/modelsby request protocol, support CCR on filesystems without hard links, and compress supported JUnit, pytest, and Jest reports. - Route Pi's ChatGPT OAuth traffic through the subscription proxy, including zstd request decoding. Tool-result compression requires a recoverable handle.
Native installs, hooks, and skills
- Native Claude and Codex installs include the three voice skills. Disabling native routing keeps those skills available and disables every configured Claude profile.
- Subagents inherit their session's mode. Headless
claude -pand Agent SDK sessions start with the voice off. - Record session stats at SessionEnd, with project-scoped session lookup.
- Add or extend always-on integration for Codex, Cursor, Copilot CLI, Antigravity CLI, and Grok Build; add CodeBuddy Code installation.
- Install OpenCode cavecrew agents as subagents and warn when an active provider bypasses the proxy.
- Move the standalone installer from
bin/toinstaller/for enterprise marketplace compatibility; improve Windows command-shim handling and agent detection. - Refine voice rules for diff-first answers, questions, comments, and hookless status reporting.
Compression, evaluation, and extension
- Add standalone OpenCode and OpenAI-compatible providers to
caveman-compress, provider-specific model selection, and protectednocompressregions. - Record evaluation usage, add a deterministic fidelity suite, support language-specific prompt sets, and retain partial results after interrupted runs. Existing published style measurements remain labeled with their original rules and methods.
- Browser extension 1.3.0 improves rich-editor snapshots and send-control discovery. Browser-store publication is separate from this GitHub release.
Companion releases
- Signed binaries and container:
bin-v2.1.0. - npm CLI:
@caveman-ai/cli@2.1.0. - npm Pi extension:
@caveman-ai/pi@0.3.0, including the Apache-2.0 license transition and LICENSE/NOTICE files from the unpublished 0.2.0 version. - npm TypeScript middleware:
@caveman-ai/middleware@1.0.1, supporting tested Anthropic SDK versions through 0.131.0.
All four companion releases are published through the repository's signed-tag workflows. All 36 runtime binaries download anonymously; their checksum manifest and release identity are signed. The container signature is verified against the exact release workflow identity. npm packages carry GitHub Actions provenance. Platform integration tests, local cache checks, and release artifacts do not establish customer savings or production outcomes.
Published CLI smoke passed on macOS ARM, Linux x64 and Windows x64: clean npm installation, signed companion setup, wrapper endpoint injection in record mode, compression with byte-exact retrieval, and isolated npm uninstall preserving unrelated configuration.
Full changes: v3.1.0...v3.2.0
@caveman-ai/middleware 1.0.1
- Version gate:
@anthropic-ai/sdk0.129.0, 0.130.0 and 0.131.0 were tested
and are now inside the supported range (>=0.124.0 <0.132). Before this,
those releases passed content through with anunsupported_versionwarning.
Verify this release
- Registry provenance (trusted publishing from this workflow): https://www.npmjs.com/package/@caveman-ai/middleware/v/1.0.1#provenance
- CycloneDX SBOM of the published dependency graph: attached
.cdx.json - Built from annotated, GitHub-verified tag
middleware-ts-v1.0.1onmain
bin-v2.1.0
Signed Caveman runtime companions for macOS, Linux, and Windows. The binaries are licensed under the Apache License 2.0 (LICENSE, NOTICE); see LICENSING.md and the attached third-party notices. The ghcr.io caveman-proxy image is still built without cgo, so its code compressor handles Go source only; the caveman-proxy and caveman-engine binaries here also compress TypeScript, JavaScript, Python, Rust, Java and C/C++.
bin-v2.0.2
Signed Caveman runtime companions for macOS, Linux, and Windows. The binaries are licensed under the Apache License 2.0 (LICENSE, NOTICE); see LICENSING.md and the attached third-party notices.
Caveman 3.1.0
Three skills replace six intensity levels, the reviewed fix stack lands, and the repo maintenance that piled up since 3.0.0 comes with it. 106 commits over main.
The skill, rethought for current models
/caveman is the voice. /ultracave strips the grammar. /megacave speaks classical Chinese. Each is a self-contained skills/<id>/SKILL.md in the same shape: thesis, persistence, why, numbered rules with bad/good pairs, when to break the rules, pre-send check. Simplified Technical English (ASD-STE100) is the floor under all three instead of a competing register. The tool-call silence ban becomes one status line per phase (#1127, #1154), and the host harness outranks the skill on when to speak (#1125). /caveman lite|full|ultra|wenyan* keep working as aliases; stored sessions migrate on read.
Hooks: mode ids are now caveman, ultracave, megacave (+ one-shots, off); legacy values map on every read path; the SessionStart hook injects the active skill whole; per-turn reinforcement is derived from the skill's thesis line; statusline renders [CAVEMAN], [ULTRACAVE], [MEGACAVE]; an outdated statusline copy pinned to a plugin-cache path now triggers the repair nudge. The always-on rule copies (IDE rules, opencode AGENTS.md, OpenClaw, extension, cavecrew agents, Codex/Gemini stubs, Codex plugin mirror) are regenerated or hand-synced with a drift test. Design record: docs/technical/three-skills-refactor.md.
Evals were rerun on claude-opus-5-5 with host settings isolated: vs a plain Answer concisely. control, median output reduction is caveman 3%, ultracave 35%, megacave 9% (n=10, single run, length only). README and HONEST-NUMBERS cite exactly that. Owner decision: the default voice adds little over "be concise" on this model; ultracave carries the savings.
Fix stack (PRs #1163 #1157 #1158 #1159 #1160 #1161 #1162, merged as one)
Proxy session-marker correlation and silent CCR failures (#1105 #1110 #1149 #1113); engine JSON elision, non-finite cells, shrink sentinel (#1151 #1095); CLI dangling hooks, shrink hook permissions, Codex TOML, Remote Control, opencode OAuth, watermark locks (#1137 #1138 #1133 #1134 #1101 #1119 #1116 #1132); statusline nudge path (#1147); Codex SessionStart honors defaultMode (#185 #1139); caveman-compress streaming/ceiling/pairing (#1131 #1107). #1164 (Pi recovery proof, Pi 1.0.0) is not here: it must wait for a caveman-mcp binary release.
Maintenance
maintenance/2026-09-20:/caveman status,defaultMode: "manual".- Contributor PRs merged: #1078 (C++ parse gate), #1076 (GDScript elision), #1156 (syslog severities), #1108 (delete dead
proxy/routing), #1084 (Grok Build install, closes #547), #1062 (eval snapshot contract). - Daily-triage remnants cherry-picked: Python encoding fixes, merge scope check, the conformance pin-bump chain (#978 #1096 #1097 #1098 #1099 #1100 #1109 #1118).
- Dependabot: #1144 #1143 #1141 #1140 merged. #1142 (TypeScript 7, @types/node 26) reverted: its own CI was red.
- Pins bumped to v3.1.0 everywhere the shim-security test checks, plus the Gemini and Codex plugin manifests.
Caveman 3.0.0
Caveman 3.0.0
Everything is Apache-2.0 now
The whole repo, including engine, proxy, browse, MCP server, shrink, the cavemem Go core and shared platform, is licensed Apache-2.0 from 3.0.0 onward. There's no hosted-service restriction and no Change Date, and you don't need a commercial license. Fork it, embed it, run it for your company. LICENSING.md has the details.
Earlier releases keep the license they shipped with. Code that outside contributors wrote under MIT stays available under MIT (LICENSE-MIT).
Caveman Learn runs itself
- Autopilot. After a session ends, Learn re-checks your history in the background at low priority, at most every 6 hours. It never makes a session wait, and
caveman learn autopilot offstops it. - Memory doctor. It finds
MEMORY.mdnotes Claude Code never loads (Claude Code only reads the first 200 lines or 25KB),@importspointing at missing files, the same rule loaded from two files, and stale paths inCLAUDE.md. - Trends. Week-over-week numbers from your own sessions. They use the median session, and you get "not enough data" when there are fewer than 5 sessions.
- Built-in model price list, now including the current Claude and GPT lineups. Numbers no longer depend on which folder you run from.
- Every fix still asks first, gets re-counted, and is undone if it didn't make each message smaller.
Middleware 1.0 and SDK 1.2
@caveman-ai/middlewareandcaveman-middlewareare stable at 1.0.0 and need SDK 1.2.0.- Adapters fail open: if Caveman has a problem, your request goes through untouched and you get a loud warning.
- The proxy adds tenant identity, a Postgres HA store, TLS/mTLS and HA deploy manifests (
deploy/). @caveman-ai/contracts2.0.0 publishes the protocol 1.1 schemas and OpenAPI.
Telemetry
The CLI and its agent hooks send usage stats with a random install ID and your IP address. This is on by default. Turn it off with caveman telemetry off or DO_NOT_TRACK=1, and CI never sends. The skill by itself sends nothing. See SECURITY.md.
Versions
| Artifact | Version |
|---|---|
Runtime binaries + ghcr.io/juliusbrussee/caveman-proxy |
bin-v2.0.0 |
@caveman-ai/cli |
2.0.0 |
@caveman-ai/sdk / caveman-sdk |
1.2.0 |
@caveman-ai/middleware / caveman-middleware |
1.0.0 |
@caveman-ai/contracts |
2.0.0 |
@caveman-ai/pi |
0.2.0 |
Upgrade
npx -y github:JuliusBrussee/cavemanBreaking: the SDK license changes from MIT to Apache-2.0, and middleware 1.0.0 requires SDK >= 1.2.0.
@caveman-ai/sdk 1.2.0
- Breaking (license): relicensed from MIT to Apache-2.0, along with the rest of
the repository in Caveman 3.0.0. Releases before this one keep the MIT license. onDecisionandonDiagnosticaccept async sinks: a rejected promise is
swallowed like a throw, asonReportalready was, instead of an unhandled
rejection that terminates Node.- Callers that only waited on another call's shared capabilities fetch no longer
record its failure in the breaker: one refused connect at cold start is one
failure, not one per waiter. A call its host aborts records nothing (it used to
record a success, so a cancelled half-open probe closed the breaker) and frees
the probe slot. NewCircuitBreaker.release(). - A runtime token is trimmed of surrounding whitespace; any other character
outside printable ASCII isinvalid_configuration, as in Python. - An
https://orsocksHTTPS_PROXY/HTTP_PROXYisinvalid_configuration
with the unsupported scheme named byready()andpreflight(), instead of
runtime_unavailableon every call.MiddlewareErrortakes an optional detail. retrieve()(and so the recovery binding) throws
MiddlewareError('invalid_request')for arguments that are not an object with
a stringhandle, never aTypeError.- The
version_unverifiedwarn-once line readsCaveman middleware is running on an unverified framework version, because that call proceeds. - CommonJS
node16consumers can useimport sdk = require(...)and
import * aswith values: the.d.ctsshims re-export values, not only types. @caveman-ai/sdk/middlewareis stable: the@experimentalmarker is gone,
and it follows semver, because@caveman-ai/middleware1.0 depends on it.- A candidate whose
sourceIdis not a scope token is skipped as
unsupported, as in Python. - Exporter: cost is always exported as
doubleValue, a whole-dollar amount
too, as in Python. - Middleware transport honors
HTTPS_PROXY/HTTP_PROXY/NO_PROXYwithout
NODE_USE_ENV_PROXY, never proxies loopback, and adds acaoption. - Plan validation is total: any malformed plan is
invalid_plan, never
adapter_error. - The breaker, Retry-After and the capabilities TTL use a monotonic clock.
Retrieve and delete have their own concurrency budget. Server-advertised
deadlines are capped (5 s optimize, 30 s retrieve). - Stricter endpoint parsing; BOM-prefixed responses are rejected; a 429 with a
bad UTF-8 body is stillcapacity. preflight()reportsunknown_capabilitywhen no transform is usable.
counts.skippedmeans sent minus replaced. A stale-revision refresh runs
inline through the breaker.- New
unsupported_providerandunsupported_requestreason codes. - Exporter: user-supplied cache-creation attributes pass through again and
are no longer clamped. Cost is also emitted ascaveman.usage.cost_usd
(gen_ai.usage.cost_usdis deprecated). - New read-only
MiddlewareRuntime.strict. decline()accepts any catalogReasonCodeand an optional adapter id,
which the warn-once log line then names.typesVersionsand.d.ctsshims: TypeScript projects using node10 or
node16 CommonJS resolution can import both entry points without
skipLibCheck.@caveman-ai/sdk/middlewareimplements middleware protocol 1.1:- capabilities are parsed tolerantly, cached for 300 s and refreshed
single-flight - plans survive a policy-revision change
- errors are handled as the protocol specifies, including
Retry-After - a new circuit breaker opens after 5 consecutive or 10 of 20 failures and
counts deadlines - per-candidate budgets replace the whole-call bypass
- capabilities are parsed tolerantly, cached for 300 s and refreshed
- Security: replacements must be
exact_ccr, carry the recovery marker and
handle, and be strictly shorter in UTF-8 bytes.recovery: "none"output is
never applied. - Fixed a per-request memory leak (
AbortSignal.any). The runtime token no
longer appears inJSON.stringifyorutil.inspect. - Local data errors no longer count as runtime outages, and 4xx responses no
longer clear cached capabilities. - New options:
maxConcurrency,allowInsecureTransport,onDecision,
tracer,meter. Endpoint path prefixes work, anddeadlineMsdefaults to
the runtime's advertised value. - The constructor never throws: endpoint and option errors warn once and show
up inready()/preflight().decline()no longer throws in strict mode. - Breaking (experimental subpath):
recovery()returnsnullfor an invalid scope.deleteSession()returns aSessionDeleteResult.validateCapabilitiesis removed.- Scopes are normalized: values that aren't valid tokens are hashed to
h-….
- Requests send
Caveman-Middleware-Features,Caveman-Middleware-Clientand
W3C trace context. - Node floor lowered to
>=22.12. CommonJSrequire()works through
require(esm). - OTel exporter:
cacheCreationTokensis emitted as
gen_ai.usage.cache_creation.input_tokens. - Release process: each release gets a GitHub Release with these notes and a
CycloneDX SBOM of its dependency graph.
Verify this release
- Registry provenance (trusted publishing from this workflow): https://www.npmjs.com/package/@caveman-ai/sdk/v/1.2.0#provenance
- CycloneDX SBOM of the published dependency graph: attached
.cdx.json - Built from annotated, GitHub-verified tag
sdk-ts-v1.2.0onmain
caveman-sdk 1.2.0
- Breaking (license): relicensed from MIT to Apache-2.0, along with the rest of
the repository in Caveman 3.0.0. Releases before this one keep the MIT license. - Callers that only waited on another call's shared capabilities fetch no longer
record its failure in the breaker: one refused connect at cold start is one
failure, not one per waiter. An interrupted half-open probe records nothing
and frees the probe slot. NewCircuitBreaker.release(). observe()/observe_background()normalize the receipt scope, as TypeScript
does, and drop a receipt with an invalid scope. Asyncobserve()has its own
receipt worker (one worker, sixteen queued), so receipts never take optimize's
slots.- A runtime token is stripped of surrounding whitespace (a mounted secret file's
trailing newline); a token with any other character outside printable ASCII
isinvalid_configurationinstead of aruntime_unavailableon every call. - An
https://orsocksproxy URL isinvalid_configuration;ready()and
preflight()name the unsupported scheme.MiddlewareErrortakes an optional
detail. endpointis the runtime origin (scheme://host[:port]), as in TypeScript,
and""when the endpoint was refused; it never echoes userinfo.delete_session()validates the response like TypeScript:schema_version
must be 1,originals_deletedis true only for a literaltrue, and
deletedappears only with four safe non-negative integer counts.- The recovery binding passes only
handle/offset/limit/queryto
retrieve(); non-object arguments or a missing handle raise
MiddlewareError("invalid_request"), which adapters turn into a tool error. - Coroutine sinks (
on_report,on_decision,on_diagnostic) run on the
caller's event loop instead of leaking a "never awaited" warning; their
failures never reach the call. - The
version_unverifiedwarn-once line readsCaveman middleware is running on an unverified framework version, because that call proceeds. - The transport bounds the whole exchange (connect, proxy tunnel, TLS, headers
and body) by the deadline, and never reuses connections the server closed
while idle. aclose()never blocks on a stuck worker.- The first capabilities fetch is single-flight. An unusable capabilities view
is refreshed once. Server-advertised deadlines are capped. - JSON is parsed strictly (no NaN; integer-valued floats become ints), and
endpoint parsing is stricter. - Warn-once rules, the decision event's
runtime_build, andadapter_error
for an unserializable manifest now match TypeScript. - New
unsupported_providerandunsupported_requestreason codes. - Exporter: cost is also emitted as
caveman.usage.cost_usd. - The warn-once line now reads
Caveman middleware passed content through unchanged: adapter=… reason=…, the same as TypeScript. Log filters that
match the oldCaveman middleware decision:prefix need updating. - Strict
ready()/preflight()now report the firstdecline(). decline()accepts any catalog reason and an optional adapter id, which the
warn-once line names. It never raises in strict mode.- Python floor lowered from 3.13 to 3.11.
caveman_cloud.middlewareimplements middleware protocol 1.1 and is stable:
it follows semver, becausecaveman-middleware1.0 depends on it. Requests carryCaveman-Middleware-Features,
Caveman-Middleware-Client, and, when atraceris configured,traceparent/tracestatefor the SDK's own span
(never the application's ambient context, as in TypeScript).- Capabilities are parsed tolerantly and cached for 300 s with single-flight
refresh. A new policy revision or transform version no longer rejects a plan. - Only
exact_ccrreplacements that carry the recovery marker and handle and are
shorter in UTF-8 bytes are applied.recovery: "none"is never applied, which
closes a text-injection path. - New circuit breaker: opens after 5 consecutive or 10 of 20 failures, stays
open 30 s, then allows one probe. Deadlines count; local and 4xx errors don't.
Ill-formed Unicode is reported asunsupported_shape.Retry-Afteris honored. - The default deadline comes from capabilities (500 ms before the first fetch).
Retrieve has its own 5 s deadline and its own pool. - Async calls respect their deadline even while queued. Pools are rebuilt after
fork().aclose()resolves in-flight calls asclosedinstead of raising
CancelledError. - Transport: keep-alive reuse; DNS, connect and TLS inside the deadline;
HTTPS_PROXY/HTTP_PROXY/NO_PROXY;ssl_context; a pluggabletransport;
endpoint path prefixes;allow_insecure_transport;max_concurrency. - Per-candidate budgets replace the whole-call bypass at 256 candidates or 4096
manifest items. - Endpoint problems no longer raise at construction.
recovery()returnsNone
for an invalid scope. Scopes are normalized, so emails and spaces are hashed.
delete_session()returns the result. - New: warn-once logging on
caveman.middleware,on_decisionevents
(includingno_candidate), opt-in OTeltracer/meter, andas_sync(),
ensure_sync(),ensure_async(). - The credential is kept out of
vars(),repr()and pickling. OTelExporter.record_span(cache_creation_tokens=)emits
gen_ai.usage.cache_creation.input_tokens.gen_ai.usage.cost_usdis
deprecated; it stays through 1.x.- Release process: each release gets a GitHub Release with these notes and a
CycloneDX SBOM of its dependency graph.
Verify this release
- Registry provenance (trusted publishing from this workflow): https://pypi.org/project/caveman-sdk/1.2.0/#files
- CycloneDX SBOM of the published dependency graph: attached
.cdx.json - Built from annotated, GitHub-verified tag
sdk-python-v1.2.0onmain
@caveman-ai/middleware 1.0.0
- First stable release. Requires
@caveman-ai/sdk1.2.0 or later. - Breaking (license): relicensed from MIT to Apache-2.0, along with the rest of
the repository in Caveman 3.0.0. Releases before this one keep the MIT license. - Release process: prereleases no longer take the
latestdist-tag, each
release gets a GitHub Release with these notes and a CycloneDX SBOM, and the
published dependency graph is audited before publish. - Version gate:
- A framework version outside the tested range, or a prerelease, now
passes content through with oneunsupported_versionwarning instead of
silently doing nothing.acceptFrameworkVersionoverrides it. - Bundled deploys run with a one-time
version_unverifiednotice. - Nothing throws at wrap time; strict mode raises from
ready().
- A framework version outside the tested range, or a prerelease, now
- The framework version is read from the application's installed copy.
Framework peers are declared optional and unranged (*), so Yarn PnP can
resolve them without a plainnpm installfailing with ERESOLVE. require()works, and TypeScript resolves with node10, node16, nodenext
and bundler../compatibilityexposes atier. Importing under
workerd/edge-lightthrows a clear unsupported-runtime error.enginesis
>=22.12.- Every adapter accepts a per-request scope function. Emails and free-text ids
are normalized, and a missingthread_idno longer fails the call. - Adapter exceptions and changed SDK internals pass through as
adapter_error. Every pass-through reason is logged once. - Large histories and earlier images no longer skip the whole call.
manifestBytesandwireBytesare configurable. - A recovery tool-name clash reports
recovery_name_conflict; the OpenAI tools
helper no longer throws. OpenAI Responses turns pass through unless
store:false. The Mastra oversize latch is per thread, not per process. - Entry points that can only record say so at construction and report
recovery_unbound. - LangChain compressed copies no longer embed the original in
lc_kwargs.
Minified bundles keep compressing. - Tiers:
ai-sdk,langchain,openaiandanthropicare certified; the
others are experimental. @anthropic-ai/sdkrange widened to<0.129. Tested up to ai 7.0.114,
openai 7.23.0, @google/genai 2.24.0, langchain 1.5.12, @langchain/core 1.2.12,
strands 1.19.0, mastra 1.70.0 and MCP 1.30.1.- In strict mode, adapter exceptions now raise
MiddlewareError('adapter_error')instead of passing through. - Decline warnings name the adapter instead of
adapter=-. - New
@caveman-ai/middleware/langchain-modelsubpath (withCavemanModel,
CavemanChatModel,scopeFromConfig). It needs only@langchain/core;
/langchainstill exports everything. .d.ctsshims type-check under node16 CommonJS withoutskipLibCheck.CavemanDocumentCompressoris also exported from/langchain-model, so it
needs only@langchain/core.- A
caveman_retrievethe runtime refuses (an unknown or expired handle, or
the runtime being down) now returns{"error":"<code>"}, or an MCP
isErrorresult, instead of crashing the native tool loop. So do
arguments that are not an object with a stringhandle(null, a list):
{"error":"invalid_request"}. - The
fetchoption of the OpenAI and Anthropic wrappers is optional and
defaults to the client's own. Embeddings, files and models calls are no
longer reported as skipped. - Wrapping a client, agent or model twice runs one Caveman layer instead of
turning compression off (or throwing, in LangChain and Strands). - The version gate reads the framework copy the adapter actually runs (openai
and anthropic: the client's own version), whatever the working directory.
It warns when the app resolves a different copy.@ai-sdk/provideris no
longer gated. - LangChain tool errors are marked
status:'error'and never compressed. withCavemanis idempotent, and its bundle stays mutable; only the recovery
tool is frozen.CavemanChatModel.profilereturns the inner model's
profile.- ai-sdk retries reuse one logical call id and one optimization.
- A Mastra thread that overflows the scan budget compresses again on its next
turn. - Recovery context no longer leaks into calls on other clients. The
recovery_unboundhint fires on use. Strict mode raisesadapter_error
from synchronous hooks. - The wire now carries the real package version and the
@langchain/core
version. Every reported reason is a spec §8 catalog code.
Verify this release
- Registry provenance (trusted publishing from this workflow): https://www.npmjs.com/package/@caveman-ai/middleware/v/1.0.0#provenance
- CycloneDX SBOM of the published dependency graph: attached
.cdx.json - Built from annotated, GitHub-verified tag
middleware-ts-v1.0.0onmain
caveman-middleware 1.0.0
- First stable release. Requires
caveman-sdk1.2.0 or later
(caveman-sdk>=1.2,<2): 1.1.0 lacks APIs the adapters import. caveman_middleware.__version__is read from the installed distribution.- A refused
caveman_retrieve(an unknown or expired handle, or a
404/410/503) returns{"error": code}through each framework's tool-error
result and warns once. The run carries on, and cancellation still
propagates. - Recovery arguments that are not an object with a string
handle(None,
a list) return{"error": "invalid_request"}in the OpenAI, Anthropic and
MCP adapters. - Wrapping a client, agent or model twice runs one Caveman layer instead of
stacking middleware and turning compression off. - A wrapped
AnthropicBedrockclient keepsaws_profile, so calls are signed
with the caller's AWS identity. - LiteLLM registers one process-wide callback however many instances are
live, so the host's own callbacks are never dropped. - Importing an adapter whose framework is out of range names the installed
version and the supported range. - Every reported reason is a spec §8 catalog code. Calls that were never LLM
calls (embeddings, token counting, non-LLM ASGI routes) report nothing. Adapter.versionis the installed package version.- Strands 1.43 recovery no longer fails on a missing
cancel_signal. - CI tests the installed wheel, adds a Python 3.14 lane and classifier, and
watchestests/middleware-e2e/**. recovery_name_conflictnow reacheson_diagnostic, and strictready()
raises it.- Version-gate warnings name the adapter through
decline(), on both sync
and async runtimes. - Every adapter fails open: outside strict mode, an adapter or runtime error
sends the original request and records a reason code (adapter_errorfor
adapter bugs). Strict mode raisesMiddlewareError. Every pass-through
reason is logged once oncaveman.middleware. - Version gate: a framework outside its tested range is skipped with
unsupported_versionand one warning.accept_framework_version=True
overrides it, and an unreadable version runs withversion_unverified. New
preflight()andready()for startup checks;COMPATIBILITYlists tiers
and ranges. - Certified tier: langchain, openai, anthropic, litellm. Everything else is
experimental. - openai 2.x (httpx) and 3.x (httpx2) are both supported, so importing the
adapter no longer fails withNo module named 'httpx2'. Stored Responses
calls pass through withprovider_state_retainedunlessstore=Falseor
allow_stored_responses=True. - Anthropic Bedrock and Vertex clients are supported.
- An existing
caveman_retrievetool no longer raises: recovery is disabled
for that registration andrecovery_name_conflictis logged once. - Free-form thread and session ids are hashed into valid scopes. A missing
scope passes through instead of failing the agent call. - Adapters accept a sync or async runtime on either code path.
- Long histories are budgeted (
manifest_bytes, 2 MiB), and bytes and images
are hashed instead of skipping the call. ASGI reportspayload_budgetfor
oversize bodies and no longer requires fastapi or starlette. - Google: wrapping returns clones and never mutates or closes your client, and
unwrap_google()returns the original. - LiteLLM: bounded in-flight state, sync Router support for OpenAI and
Anthropic,unsupported_providerfor other providers. - CrewAI registers its hooks once per process. AutoGen lists workbench tools
once per turn. - Accepted ranges now match the lowest versions actually tested. Python 3.11–3.13.
Thepydantic-aiextra usespydantic-ai-slim. - CI: floor and latest lanes with constraints files, and a nightly canary that
opens an issue on failure. - Breaking (license): relicensed from MIT to Apache-2.0, along with the rest of
the repository in Caveman 3.0.0. Releases before this one keep the MIT license. - Release process: each release gets a GitHub Release with these notes and a
CycloneDX SBOM of its dependency graph.
Verify this release
- Registry provenance (trusted publishing from this workflow): https://pypi.org/project/caveman-middleware/1.0.0/#files
- CycloneDX SBOM of the published dependency graph: attached
.cdx.json - Built from annotated, GitHub-verified tag
middleware-python-v1.0.0onmain