Skip to content

Releases: JuliusBrussee/caveman

Caveman 3.2.0

Choose a tag to compare

@JuliusBrussee JuliusBrussee released this 08 Oct 05:42
v3.2.0
e20f07e

Caveman 3.2.0 brings the release sweep and subsequent native-install fixes to the installer, plugins, skills, and local runtime.

Cache-safe compression and runtime

  • Keep previous compression decisions stable across a conversation, including decisions to send raw bytes. Cache-prefix checks record caveman-caused changes in local request telemetry and caveman status.
  • Pass through rate-limit responses with Retry-After; pin accepted raw retries to their conversation lineage.
  • Build the proxy and engine with tree-sitter support for TypeScript, JavaScript, Python, Rust, Java, C, and C++ in the six-platform binary matrix. The container remains a pure-Go build.
  • Route bare /v1/models by request protocol, support CCR on filesystems without hard links, and compress supported JUnit, pytest, and Jest reports.
  • Route Pi's ChatGPT OAuth traffic through the subscription proxy, including zstd request decoding. Tool-result compression requires a recoverable handle.

Native installs, hooks, and skills

  • Native Claude and Codex installs include the three voice skills. Disabling native routing keeps those skills available and disables every configured Claude profile.
  • Subagents inherit their session's mode. Headless claude -p and Agent SDK sessions start with the voice off.
  • Record session stats at SessionEnd, with project-scoped session lookup.
  • Add or extend always-on integration for Codex, Cursor, Copilot CLI, Antigravity CLI, and Grok Build; add CodeBuddy Code installation.
  • Install OpenCode cavecrew agents as subagents and warn when an active provider bypasses the proxy.
  • Move the standalone installer from bin/ to installer/ for enterprise marketplace compatibility; improve Windows command-shim handling and agent detection.
  • Refine voice rules for diff-first answers, questions, comments, and hookless status reporting.

Compression, evaluation, and extension

  • Add standalone OpenCode and OpenAI-compatible providers to caveman-compress, provider-specific model selection, and protected nocompress regions.
  • Record evaluation usage, add a deterministic fidelity suite, support language-specific prompt sets, and retain partial results after interrupted runs. Existing published style measurements remain labeled with their original rules and methods.
  • Browser extension 1.3.0 improves rich-editor snapshots and send-control discovery. Browser-store publication is separate from this GitHub release.

Companion releases

All four companion releases are published through the repository's signed-tag workflows. All 36 runtime binaries download anonymously; their checksum manifest and release identity are signed. The container signature is verified against the exact release workflow identity. npm packages carry GitHub Actions provenance. Platform integration tests, local cache checks, and release artifacts do not establish customer savings or production outcomes.

Published CLI smoke passed on macOS ARM, Linux x64 and Windows x64: clean npm installation, signed companion setup, wrapper endpoint injection in record mode, compression with byte-exact retrieval, and isolated npm uninstall preserving unrelated configuration.

Full changes: v3.1.0...v3.2.0

@caveman-ai/middleware 1.0.1

Choose a tag to compare

@github-actions github-actions released this 08 Oct 05:16
middleware-ts-v1.0.1
e20f07e
  • Version gate: @anthropic-ai/sdk 0.129.0, 0.130.0 and 0.131.0 were tested
    and are now inside the supported range (>=0.124.0 <0.132). Before this,
    those releases passed content through with an unsupported_version warning.

Verify this release

bin-v2.1.0

Choose a tag to compare

@github-actions github-actions released this 08 Oct 05:26
bin-v2.1.0
e20f07e

Signed Caveman runtime companions for macOS, Linux, and Windows. The binaries are licensed under the Apache License 2.0 (LICENSE, NOTICE); see LICENSING.md and the attached third-party notices. The ghcr.io caveman-proxy image is still built without cgo, so its code compressor handles Go source only; the caveman-proxy and caveman-engine binaries here also compress TypeScript, JavaScript, Python, Rust, Java and C/C++.

bin-v2.0.2

Choose a tag to compare

@github-actions github-actions released this 04 Oct 01:50
bin-v2.0.2
aeb45e2

Signed Caveman runtime companions for macOS, Linux, and Windows. The binaries are licensed under the Apache License 2.0 (LICENSE, NOTICE); see LICENSING.md and the attached third-party notices.

Caveman 3.1.0

Choose a tag to compare

@JuliusBrussee JuliusBrussee released this 04 Oct 04:30

Three skills replace six intensity levels, the reviewed fix stack lands, and the repo maintenance that piled up since 3.0.0 comes with it. 106 commits over main.

The skill, rethought for current models

/caveman is the voice. /ultracave strips the grammar. /megacave speaks classical Chinese. Each is a self-contained skills/<id>/SKILL.md in the same shape: thesis, persistence, why, numbered rules with bad/good pairs, when to break the rules, pre-send check. Simplified Technical English (ASD-STE100) is the floor under all three instead of a competing register. The tool-call silence ban becomes one status line per phase (#1127, #1154), and the host harness outranks the skill on when to speak (#1125). /caveman lite|full|ultra|wenyan* keep working as aliases; stored sessions migrate on read.

Hooks: mode ids are now caveman, ultracave, megacave (+ one-shots, off); legacy values map on every read path; the SessionStart hook injects the active skill whole; per-turn reinforcement is derived from the skill's thesis line; statusline renders [CAVEMAN], [ULTRACAVE], [MEGACAVE]; an outdated statusline copy pinned to a plugin-cache path now triggers the repair nudge. The always-on rule copies (IDE rules, opencode AGENTS.md, OpenClaw, extension, cavecrew agents, Codex/Gemini stubs, Codex plugin mirror) are regenerated or hand-synced with a drift test. Design record: docs/technical/three-skills-refactor.md.

Evals were rerun on claude-opus-5-5 with host settings isolated: vs a plain Answer concisely. control, median output reduction is caveman 3%, ultracave 35%, megacave 9% (n=10, single run, length only). README and HONEST-NUMBERS cite exactly that. Owner decision: the default voice adds little over "be concise" on this model; ultracave carries the savings.

Fix stack (PRs #1163 #1157 #1158 #1159 #1160 #1161 #1162, merged as one)

Proxy session-marker correlation and silent CCR failures (#1105 #1110 #1149 #1113); engine JSON elision, non-finite cells, shrink sentinel (#1151 #1095); CLI dangling hooks, shrink hook permissions, Codex TOML, Remote Control, opencode OAuth, watermark locks (#1137 #1138 #1133 #1134 #1101 #1119 #1116 #1132); statusline nudge path (#1147); Codex SessionStart honors defaultMode (#185 #1139); caveman-compress streaming/ceiling/pairing (#1131 #1107). #1164 (Pi recovery proof, Pi 1.0.0) is not here: it must wait for a caveman-mcp binary release.

Maintenance

  • maintenance/2026-09-20: /caveman status, defaultMode: "manual".
  • Contributor PRs merged: #1078 (C++ parse gate), #1076 (GDScript elision), #1156 (syslog severities), #1108 (delete dead proxy/routing), #1084 (Grok Build install, closes #547), #1062 (eval snapshot contract).
  • Daily-triage remnants cherry-picked: Python encoding fixes, merge scope check, the conformance pin-bump chain (#978 #1096 #1097 #1098 #1099 #1100 #1109 #1118).
  • Dependabot: #1144 #1143 #1141 #1140 merged. #1142 (TypeScript 7, @types/node 26) reverted: its own CI was red.
  • Pins bumped to v3.1.0 everywhere the shim-security test checks, plus the Gemini and Codex plugin manifests.

Caveman 3.0.0

Choose a tag to compare

@JuliusBrussee JuliusBrussee released this 30 Sep 19:25
v3.0.0
b33a395

Caveman 3.0.0

Everything is Apache-2.0 now

The whole repo, including engine, proxy, browse, MCP server, shrink, the cavemem Go core and shared platform, is licensed Apache-2.0 from 3.0.0 onward. There's no hosted-service restriction and no Change Date, and you don't need a commercial license. Fork it, embed it, run it for your company. LICENSING.md has the details.

Earlier releases keep the license they shipped with. Code that outside contributors wrote under MIT stays available under MIT (LICENSE-MIT).

Caveman Learn runs itself

  • Autopilot. After a session ends, Learn re-checks your history in the background at low priority, at most every 6 hours. It never makes a session wait, and caveman learn autopilot off stops it.
  • Memory doctor. It finds MEMORY.md notes Claude Code never loads (Claude Code only reads the first 200 lines or 25KB), @imports pointing at missing files, the same rule loaded from two files, and stale paths in CLAUDE.md.
  • Trends. Week-over-week numbers from your own sessions. They use the median session, and you get "not enough data" when there are fewer than 5 sessions.
  • Built-in model price list, now including the current Claude and GPT lineups. Numbers no longer depend on which folder you run from.
  • Every fix still asks first, gets re-counted, and is undone if it didn't make each message smaller.

Middleware 1.0 and SDK 1.2

  • @caveman-ai/middleware and caveman-middleware are stable at 1.0.0 and need SDK 1.2.0.
  • Adapters fail open: if Caveman has a problem, your request goes through untouched and you get a loud warning.
  • The proxy adds tenant identity, a Postgres HA store, TLS/mTLS and HA deploy manifests (deploy/).
  • @caveman-ai/contracts 2.0.0 publishes the protocol 1.1 schemas and OpenAPI.

Telemetry

The CLI and its agent hooks send usage stats with a random install ID and your IP address. This is on by default. Turn it off with caveman telemetry off or DO_NOT_TRACK=1, and CI never sends. The skill by itself sends nothing. See SECURITY.md.

Versions

Artifact Version
Runtime binaries + ghcr.io/juliusbrussee/caveman-proxy bin-v2.0.0
@caveman-ai/cli 2.0.0
@caveman-ai/sdk / caveman-sdk 1.2.0
@caveman-ai/middleware / caveman-middleware 1.0.0
@caveman-ai/contracts 2.0.0
@caveman-ai/pi 0.2.0

Upgrade

npx -y github:JuliusBrussee/caveman

Breaking: the SDK license changes from MIT to Apache-2.0, and middleware 1.0.0 requires SDK >= 1.2.0.

@caveman-ai/sdk 1.2.0

Choose a tag to compare

@github-actions github-actions released this 30 Sep 22:22
sdk-ts-v1.2.0
b33a395
  • Breaking (license): relicensed from MIT to Apache-2.0, along with the rest of
    the repository in Caveman 3.0.0. Releases before this one keep the MIT license.
  • onDecision and onDiagnostic accept async sinks: a rejected promise is
    swallowed like a throw, as onReport already was, instead of an unhandled
    rejection that terminates Node.
  • Callers that only waited on another call's shared capabilities fetch no longer
    record its failure in the breaker: one refused connect at cold start is one
    failure, not one per waiter. A call its host aborts records nothing (it used to
    record a success, so a cancelled half-open probe closed the breaker) and frees
    the probe slot. New CircuitBreaker.release().
  • A runtime token is trimmed of surrounding whitespace; any other character
    outside printable ASCII is invalid_configuration, as in Python.
  • An https:// or socks HTTPS_PROXY/HTTP_PROXY is invalid_configuration
    with the unsupported scheme named by ready() and preflight(), instead of
    runtime_unavailable on every call. MiddlewareError takes an optional detail.
  • retrieve() (and so the recovery binding) throws
    MiddlewareError('invalid_request') for arguments that are not an object with
    a string handle, never a TypeError.
  • The version_unverified warn-once line reads Caveman middleware is running on an unverified framework version, because that call proceeds.
  • CommonJS node16 consumers can use import sdk = require(...) and
    import * as with values: the .d.cts shims re-export values, not only types.
  • @caveman-ai/sdk/middleware is stable: the @experimental marker is gone,
    and it follows semver, because @caveman-ai/middleware 1.0 depends on it.
  • A candidate whose sourceId is not a scope token is skipped as
    unsupported, as in Python.
  • Exporter: cost is always exported as doubleValue, a whole-dollar amount
    too, as in Python.
  • Middleware transport honors HTTPS_PROXY/HTTP_PROXY/NO_PROXY without
    NODE_USE_ENV_PROXY, never proxies loopback, and adds a ca option.
  • Plan validation is total: any malformed plan is invalid_plan, never
    adapter_error.
  • The breaker, Retry-After and the capabilities TTL use a monotonic clock.
    Retrieve and delete have their own concurrency budget. Server-advertised
    deadlines are capped (5 s optimize, 30 s retrieve).
  • Stricter endpoint parsing; BOM-prefixed responses are rejected; a 429 with a
    bad UTF-8 body is still capacity.
  • preflight() reports unknown_capability when no transform is usable.
    counts.skipped means sent minus replaced. A stale-revision refresh runs
    inline through the breaker.
  • New unsupported_provider and unsupported_request reason codes.
  • Exporter: user-supplied cache-creation attributes pass through again and
    are no longer clamped. Cost is also emitted as caveman.usage.cost_usd
    (gen_ai.usage.cost_usd is deprecated).
  • New read-only MiddlewareRuntime.strict.
  • decline() accepts any catalog ReasonCode and an optional adapter id,
    which the warn-once log line then names.
  • typesVersions and .d.cts shims: TypeScript projects using node10 or
    node16 CommonJS resolution can import both entry points without
    skipLibCheck.
  • @caveman-ai/sdk/middleware implements middleware protocol 1.1:
    • capabilities are parsed tolerantly, cached for 300 s and refreshed
      single-flight
    • plans survive a policy-revision change
    • errors are handled as the protocol specifies, including Retry-After
    • a new circuit breaker opens after 5 consecutive or 10 of 20 failures and
      counts deadlines
    • per-candidate budgets replace the whole-call bypass
  • Security: replacements must be exact_ccr, carry the recovery marker and
    handle, and be strictly shorter in UTF-8 bytes. recovery: "none" output is
    never applied.
  • Fixed a per-request memory leak (AbortSignal.any). The runtime token no
    longer appears in JSON.stringify or util.inspect.
  • Local data errors no longer count as runtime outages, and 4xx responses no
    longer clear cached capabilities.
  • New options: maxConcurrency, allowInsecureTransport, onDecision,
    tracer, meter. Endpoint path prefixes work, and deadlineMs defaults to
    the runtime's advertised value.
  • The constructor never throws: endpoint and option errors warn once and show
    up in ready()/preflight(). decline() no longer throws in strict mode.
  • Breaking (experimental subpath):
    • recovery() returns null for an invalid scope.
    • deleteSession() returns a SessionDeleteResult.
    • validateCapabilities is removed.
    • Scopes are normalized: values that aren't valid tokens are hashed to
      h-….
  • Requests send Caveman-Middleware-Features, Caveman-Middleware-Client and
    W3C trace context.
  • Node floor lowered to >=22.12. CommonJS require() works through
    require(esm).
  • OTel exporter: cacheCreationTokens is emitted as
    gen_ai.usage.cache_creation.input_tokens.
  • Release process: each release gets a GitHub Release with these notes and a
    CycloneDX SBOM of its dependency graph.

Verify this release

caveman-sdk 1.2.0

Choose a tag to compare

@github-actions github-actions released this 30 Sep 22:27
sdk-python-v1.2.0
92c295f
  • Breaking (license): relicensed from MIT to Apache-2.0, along with the rest of
    the repository in Caveman 3.0.0. Releases before this one keep the MIT license.
  • Callers that only waited on another call's shared capabilities fetch no longer
    record its failure in the breaker: one refused connect at cold start is one
    failure, not one per waiter. An interrupted half-open probe records nothing
    and frees the probe slot. New CircuitBreaker.release().
  • observe()/observe_background() normalize the receipt scope, as TypeScript
    does, and drop a receipt with an invalid scope. Async observe() has its own
    receipt worker (one worker, sixteen queued), so receipts never take optimize's
    slots.
  • A runtime token is stripped of surrounding whitespace (a mounted secret file's
    trailing newline); a token with any other character outside printable ASCII
    is invalid_configuration instead of a runtime_unavailable on every call.
  • An https:// or socks proxy URL is invalid_configuration; ready() and
    preflight() name the unsupported scheme. MiddlewareError takes an optional
    detail.
  • endpoint is the runtime origin (scheme://host[:port]), as in TypeScript,
    and "" when the endpoint was refused; it never echoes userinfo.
  • delete_session() validates the response like TypeScript: schema_version
    must be 1, originals_deleted is true only for a literal true, and
    deleted appears only with four safe non-negative integer counts.
  • The recovery binding passes only handle/offset/limit/query to
    retrieve(); non-object arguments or a missing handle raise
    MiddlewareError("invalid_request"), which adapters turn into a tool error.
  • Coroutine sinks (on_report, on_decision, on_diagnostic) run on the
    caller's event loop instead of leaking a "never awaited" warning; their
    failures never reach the call.
  • The version_unverified warn-once line reads Caveman middleware is running on an unverified framework version, because that call proceeds.
  • The transport bounds the whole exchange (connect, proxy tunnel, TLS, headers
    and body) by the deadline, and never reuses connections the server closed
    while idle.
  • aclose() never blocks on a stuck worker.
  • The first capabilities fetch is single-flight. An unusable capabilities view
    is refreshed once. Server-advertised deadlines are capped.
  • JSON is parsed strictly (no NaN; integer-valued floats become ints), and
    endpoint parsing is stricter.
  • Warn-once rules, the decision event's runtime_build, and adapter_error
    for an unserializable manifest now match TypeScript.
  • New unsupported_provider and unsupported_request reason codes.
  • Exporter: cost is also emitted as caveman.usage.cost_usd.
  • The warn-once line now reads Caveman middleware passed content through unchanged: adapter=… reason=…, the same as TypeScript. Log filters that
    match the old Caveman middleware decision: prefix need updating.
  • Strict ready()/preflight() now report the first decline().
  • decline() accepts any catalog reason and an optional adapter id, which the
    warn-once line names. It never raises in strict mode.
  • Python floor lowered from 3.13 to 3.11.
  • caveman_cloud.middleware implements middleware protocol 1.1 and is stable:
    it follows semver, because caveman-middleware 1.0 depends on it. Requests carry Caveman-Middleware-Features,
    Caveman-Middleware-Client, and, when a tracer is configured, traceparent/tracestate for the SDK's own span
    (never the application's ambient context, as in TypeScript).
  • Capabilities are parsed tolerantly and cached for 300 s with single-flight
    refresh. A new policy revision or transform version no longer rejects a plan.
  • Only exact_ccr replacements that carry the recovery marker and handle and are
    shorter in UTF-8 bytes are applied. recovery: "none" is never applied, which
    closes a text-injection path.
  • New circuit breaker: opens after 5 consecutive or 10 of 20 failures, stays
    open 30 s, then allows one probe. Deadlines count; local and 4xx errors don't.
    Ill-formed Unicode is reported as unsupported_shape. Retry-After is honored.
  • The default deadline comes from capabilities (500 ms before the first fetch).
    Retrieve has its own 5 s deadline and its own pool.
  • Async calls respect their deadline even while queued. Pools are rebuilt after
    fork(). aclose() resolves in-flight calls as closed instead of raising
    CancelledError.
  • Transport: keep-alive reuse; DNS, connect and TLS inside the deadline;
    HTTPS_PROXY/HTTP_PROXY/NO_PROXY; ssl_context; a pluggable transport;
    endpoint path prefixes; allow_insecure_transport; max_concurrency.
  • Per-candidate budgets replace the whole-call bypass at 256 candidates or 4096
    manifest items.
  • Endpoint problems no longer raise at construction. recovery() returns None
    for an invalid scope. Scopes are normalized, so emails and spaces are hashed.
    delete_session() returns the result.
  • New: warn-once logging on caveman.middleware, on_decision events
    (including no_candidate), opt-in OTel tracer/meter, and as_sync(),
    ensure_sync(), ensure_async().
  • The credential is kept out of vars(), repr() and pickling.
  • OTelExporter.record_span(cache_creation_tokens=) emits
    gen_ai.usage.cache_creation.input_tokens. gen_ai.usage.cost_usd is
    deprecated; it stays through 1.x.
  • Release process: each release gets a GitHub Release with these notes and a
    CycloneDX SBOM of its dependency graph.

Verify this release

  • Registry provenance (trusted publishing from this workflow): https://pypi.org/project/caveman-sdk/1.2.0/#files
  • CycloneDX SBOM of the published dependency graph: attached .cdx.json
  • Built from annotated, GitHub-verified tag sdk-python-v1.2.0 on main

@caveman-ai/middleware 1.0.0

Choose a tag to compare

@github-actions github-actions released this 30 Sep 22:30
middleware-ts-v1.0.0
b33a395
  • First stable release. Requires @caveman-ai/sdk 1.2.0 or later.
  • Breaking (license): relicensed from MIT to Apache-2.0, along with the rest of
    the repository in Caveman 3.0.0. Releases before this one keep the MIT license.
  • Release process: prereleases no longer take the latest dist-tag, each
    release gets a GitHub Release with these notes and a CycloneDX SBOM, and the
    published dependency graph is audited before publish.
  • Version gate:
    • A framework version outside the tested range, or a prerelease, now
      passes content through with one unsupported_version warning instead of
      silently doing nothing. acceptFrameworkVersion overrides it.
    • Bundled deploys run with a one-time version_unverified notice.
    • Nothing throws at wrap time; strict mode raises from ready().
  • The framework version is read from the application's installed copy.
    Framework peers are declared optional and unranged (*), so Yarn PnP can
    resolve them without a plain npm install failing with ERESOLVE.
  • require() works, and TypeScript resolves with node10, node16, nodenext
    and bundler. ./compatibility exposes a tier. Importing under
    workerd/edge-light throws a clear unsupported-runtime error. engines is
    >=22.12.
  • Every adapter accepts a per-request scope function. Emails and free-text ids
    are normalized, and a missing thread_id no longer fails the call.
  • Adapter exceptions and changed SDK internals pass through as
    adapter_error. Every pass-through reason is logged once.
  • Large histories and earlier images no longer skip the whole call.
    manifestBytes and wireBytes are configurable.
  • A recovery tool-name clash reports recovery_name_conflict; the OpenAI tools
    helper no longer throws. OpenAI Responses turns pass through unless
    store:false. The Mastra oversize latch is per thread, not per process.
  • Entry points that can only record say so at construction and report
    recovery_unbound.
  • LangChain compressed copies no longer embed the original in lc_kwargs.
    Minified bundles keep compressing.
  • Tiers: ai-sdk, langchain, openai and anthropic are certified; the
    others are experimental.
  • @anthropic-ai/sdk range widened to <0.129. Tested up to ai 7.0.114,
    openai 7.23.0, @google/genai 2.24.0, langchain 1.5.12, @langchain/core 1.2.12,
    strands 1.19.0, mastra 1.70.0 and MCP 1.30.1.
  • In strict mode, adapter exceptions now raise
    MiddlewareError('adapter_error') instead of passing through.
  • Decline warnings name the adapter instead of adapter=-.
  • New @caveman-ai/middleware/langchain-model subpath (withCavemanModel,
    CavemanChatModel, scopeFromConfig). It needs only @langchain/core;
    /langchain still exports everything.
  • .d.cts shims type-check under node16 CommonJS without skipLibCheck.
  • CavemanDocumentCompressor is also exported from /langchain-model, so it
    needs only @langchain/core.
  • A caveman_retrieve the runtime refuses (an unknown or expired handle, or
    the runtime being down) now returns {"error":"<code>"}, or an MCP
    isError result, instead of crashing the native tool loop. So do
    arguments that are not an object with a string handle (null, a list):
    {"error":"invalid_request"}.
  • The fetch option of the OpenAI and Anthropic wrappers is optional and
    defaults to the client's own. Embeddings, files and models calls are no
    longer reported as skipped.
  • Wrapping a client, agent or model twice runs one Caveman layer instead of
    turning compression off (or throwing, in LangChain and Strands).
  • The version gate reads the framework copy the adapter actually runs (openai
    and anthropic: the client's own version), whatever the working directory.
    It warns when the app resolves a different copy. @ai-sdk/provider is no
    longer gated.
  • LangChain tool errors are marked status:'error' and never compressed.
  • withCaveman is idempotent, and its bundle stays mutable; only the recovery
    tool is frozen. CavemanChatModel.profile returns the inner model's
    profile.
  • ai-sdk retries reuse one logical call id and one optimization.
  • A Mastra thread that overflows the scan budget compresses again on its next
    turn.
  • Recovery context no longer leaks into calls on other clients. The
    recovery_unbound hint fires on use. Strict mode raises adapter_error
    from synchronous hooks.
  • The wire now carries the real package version and the @langchain/core
    version. Every reported reason is a spec §8 catalog code.

Verify this release

caveman-middleware 1.0.0

Choose a tag to compare

@github-actions github-actions released this 30 Sep 22:34
middleware-python-v1.0.0
b33a395
  • First stable release. Requires caveman-sdk 1.2.0 or later
    (caveman-sdk>=1.2,<2): 1.1.0 lacks APIs the adapters import.
  • caveman_middleware.__version__ is read from the installed distribution.
  • A refused caveman_retrieve (an unknown or expired handle, or a
    404/410/503) returns {"error": code} through each framework's tool-error
    result and warns once. The run carries on, and cancellation still
    propagates.
  • Recovery arguments that are not an object with a string handle (None,
    a list) return {"error": "invalid_request"} in the OpenAI, Anthropic and
    MCP adapters.
  • Wrapping a client, agent or model twice runs one Caveman layer instead of
    stacking middleware and turning compression off.
  • A wrapped AnthropicBedrock client keeps aws_profile, so calls are signed
    with the caller's AWS identity.
  • LiteLLM registers one process-wide callback however many instances are
    live, so the host's own callbacks are never dropped.
  • Importing an adapter whose framework is out of range names the installed
    version and the supported range.
  • Every reported reason is a spec §8 catalog code. Calls that were never LLM
    calls (embeddings, token counting, non-LLM ASGI routes) report nothing.
  • Adapter.version is the installed package version.
  • Strands 1.43 recovery no longer fails on a missing cancel_signal.
  • CI tests the installed wheel, adds a Python 3.14 lane and classifier, and
    watches tests/middleware-e2e/**.
  • recovery_name_conflict now reaches on_diagnostic, and strict ready()
    raises it.
  • Version-gate warnings name the adapter through decline(), on both sync
    and async runtimes.
  • Every adapter fails open: outside strict mode, an adapter or runtime error
    sends the original request and records a reason code (adapter_error for
    adapter bugs). Strict mode raises MiddlewareError. Every pass-through
    reason is logged once on caveman.middleware.
  • Version gate: a framework outside its tested range is skipped with
    unsupported_version and one warning. accept_framework_version=True
    overrides it, and an unreadable version runs with version_unverified. New
    preflight() and ready() for startup checks; COMPATIBILITY lists tiers
    and ranges.
  • Certified tier: langchain, openai, anthropic, litellm. Everything else is
    experimental.
  • openai 2.x (httpx) and 3.x (httpx2) are both supported, so importing the
    adapter no longer fails with No module named 'httpx2'. Stored Responses
    calls pass through with provider_state_retained unless store=False or
    allow_stored_responses=True.
  • Anthropic Bedrock and Vertex clients are supported.
  • An existing caveman_retrieve tool no longer raises: recovery is disabled
    for that registration and recovery_name_conflict is logged once.
  • Free-form thread and session ids are hashed into valid scopes. A missing
    scope passes through instead of failing the agent call.
  • Adapters accept a sync or async runtime on either code path.
  • Long histories are budgeted (manifest_bytes, 2 MiB), and bytes and images
    are hashed instead of skipping the call. ASGI reports payload_budget for
    oversize bodies and no longer requires fastapi or starlette.
  • Google: wrapping returns clones and never mutates or closes your client, and
    unwrap_google() returns the original.
  • LiteLLM: bounded in-flight state, sync Router support for OpenAI and
    Anthropic, unsupported_provider for other providers.
  • CrewAI registers its hooks once per process. AutoGen lists workbench tools
    once per turn.
  • Accepted ranges now match the lowest versions actually tested. Python 3.11–3.13.
    The pydantic-ai extra uses pydantic-ai-slim.
  • CI: floor and latest lanes with constraints files, and a nightly canary that
    opens an issue on failure.
  • Breaking (license): relicensed from MIT to Apache-2.0, along with the rest of
    the repository in Caveman 3.0.0. Releases before this one keep the MIT license.
  • Release process: each release gets a GitHub Release with these notes and a
    CycloneDX SBOM of its dependency graph.

Verify this release