Repository navigation
Highlights
- No More Silent Option Loss: New
ContentResponse.Warningsreports every caller option a provider dropped, clamped or substituted before sending the request - Vendor-Native Reasoning Control: Thinking is sent in each model family's own mechanism—effort level, token budget,
thinking_level,enable_thinking, Anthropicthinkingobject or Ollamathink; combinations a vendor rejects fail locally with typed errors instead of a 400 - Sampling Preserved:
temperature,top_p,top_kand penalties are removed only where the vendor actually rejects them, no longer from every model that reasons - Lossless Reasoning Round-Trips: Signed and encrypted thinking blocks keep their order and signatures across turns on Anthropic and Bedrock, fixing replay 400s in agent loops
- New Model Generations: GPT-6 (Astra, Sol, Luna), grok-4.7, Qwen 3.5–3.8, GLM 5.x, Kimi K2.6/K3, MiniMax M3, DeepSeek V4, Gemma 4, Claude Mythos Preview, gpt-oss-safeguard on Bedrock
- Security: Fixed SQL injection through pgvector metadata filters; JSON Schemas can no longer read local files via
$ref - Breaking Changes: API removals and new typed errors—read the section below before upgrading
Breaking Changes
This release deliberately does not keep requests byte-identical: the wire changes only where the vendor rejected the old request or where a silently lost option now reaches the vendor. There is no global switch back; on the OpenAI provider llms.WithExtraBody can still put any field back into the request.
API Changes
tools/perplexity: removedModelSonarReasoningandModelR11776(retired by the vendor)—useModelSonarReasoningProllms/googleai/vertex: removed ten names duplicated fromgoogleai(RoleUser,ResponseMIMETypeJson,ErrNoContentInResponse, …)—switch to thegoogleaipackage;vertex.Vertexnow embeds*googleai.GoogleAI, sovertex.Vertex{CallbacksHandler: h}literals no longer compiletesting/llmtest:TestLLMtakes options—declareWithoutStreaming()/WithoutToolCalls()for providers without them;MockLLM.GenerateContentStreamis gone (usellms.WithStreamingFunc)anthropic.ErrModelRefusalis now an alias ofllms.ErrModelRefusal;reasoning.ClaudeSupportsEffortWithBudgettakes a provider;reasoning.OffWiregained four values and was renumbered—never persist its numeric valueContentChoice,ContentResponse,ContentReasoningandCallOptionsgained fields—unkeyed struct literals no longer compile
Behavior Changes
- Typed errors before the request where the vendor would answer 400 or silently ignore the option:
reasoning.ErrReasoningOffUnsupported—WithReasoningDisabled()on models that cannot stop thinking (base GPT-5, gpt-6-astra, DeepSeek R1, grok-4.7, Kimi K3, GLM 5.3 via Mistral, Fable/Mythos and gpt-oss on Bedrock, …)reasoning.ErrEffortWithTools—an explicit effort combined with function tools on gpt-5.4–5.6, gpt-6-sol and gpt-6-lunareasoning.ErrThinkingBudgetUnsupported,reasoning.ErrChatToolsUnsupported,anthropic.ErrAssistantPrefillUnsupported,anthropic.ErrForcedToolUseWithThinkingllms.ErrStructuredOutputUnsupported—WithStructuredOutputfor DeepSeek, GLM and MiniMax M via the OpenAI provider, Ollama Cloud, and Bedrock models whose AWS model card lacks structured output (see the opt-in fallbacks below)
- Google AI / Vertex: default output limit raised from 2048 to 16384 tokens;
WithGRPCClient/WithGRPCConnreturnErrOptionNotHonored;vertex.Newneeds a project and location (options orGOOGLE_CLOUD_PROJECT/GOOGLE_CLOUD_LOCATION) or returnsErrMissingCloudTarget; VertexStopReasonuses the vendor spelling (MAX_TOKENS,STOP) - HuggingFace: moved to
router.huggingface.coand/v1/chat/completions(the old host no longer resolves);top_k,min_lengthandrepetition_penaltyare no longer sent; useWithMaxTokensinstead ofWithMaxLength; default modelmeta-llama/Llama-3.1-8B-Instruct - Mistral: default model
open-mistral-7b→ministral-8b-latest - Bedrock: legacy InvokeModel counters in
GenerationInfoareint(wereint32); Nova reasoning applies only toamazon.nova-2-lite - pgvector: metadata filter keys must be plain identifiers, otherwise
ErrInvalidFilterKey - Structured output: schemas with an external
$refno longer compile - Extra body: moved out of
CallOptions.Metadata—read it withllms.ExtraBody(opts)
New Features
Option Warnings
ContentResponse.Warnings lists each caller option that did not reach the vendor as asked—option, model, requested and sent value, reason and kind (drop, clamp, substitute). Every provider fills it, and a warning appears only where a value really changed.
Reasoning & Sampling per Model Family
- Claude behind OpenAI-compatible gateways (bare names or LiteLLM
anthropic/,bedrock/,vertex_ai/routes) receives Anthropic'sthinkingobject for budget, adaptive and off; public OpenAI-compatible hosts keep the standard wire - Qwen on DashScope:
reasoning_effortfor 3.8, a working disable viaenable_thinkingfor 3.5–3.7, andthinking_budgetas its own field—also for GLM, Kimi and DeepSeek served by DashScope - GLM, Kimi, MiniMax, DeepSeek: thinking is switched off the way each vendor documents; their reasoning (and Qwen's) is replayed on every assistant turn
- Gemini: Gemini 3 is disabled through the level scale,
MINIMALreaches the generations that accept it, Gemma 4 takesHIGH/MINIMAL, Gemini 2.5 budgets are clamped into the documented range - Bedrock: only families with their own mechanism get a thinking config (Claude, Nova 2 Lite, grok, gpt-oss); DeepSeek, GLM, Kimi, MiniMax and Qwen no longer receive the Anthropic object
- Ollama: sends
think(gpt-oss gets its three levels) and surfaces the server's native thinking - Claude: thinking budgets are raised to the vendor minimum of 1024; a caller's
top_p≥ 0.95 survives thinking ReasoningSupportForreports the mechanism (budget, adaptive or both) and no longer guesses effort tiers for unclassified models
Reasoning Blocks
ContentReasoning.Blocks keeps every reasoning block—readable text with its signature, or encrypted data—in the order the vendor produced it, and replays it unchanged. Content and Signature work as before; use HasContent() to tell a model that actually reasoned from one that only signed its answer (Gemini 3 signs every answer).
Truncation, Tool Choice & Output Limit
ContentChoice.Truncatednormalizes thelength,max_tokensandmodel_lengthstop reasons across providers;llms.WithFailOnTruncation()turns truncation into a typed error that still carries the partial answer- Tool choice is classified once (
llms.ClassifyToolChoice) and reaches the wire in every spelling on OpenAI, Anthropic, Bedrock Converse and Google AI; Mistral accepts all vendor values - The output-limit field follows the route:
max_tokensfor grok, Qwen and DeepSeek,max_completion_tokensfor OpenAI
New Options & Provider APIs
- Call options:
WithVerbosity,WithLogProbs/WithTopLogProbs,WithInferenceSpeed(Anthropic fast mode),WithFailOnTruncation, and provider-neutralllms.WithExtraBody(merged into nested objects on the OpenAI provider, reported inWarningselsewhere) - Anthropic:
WithFederation—Workload Identity Federation exchanges an OIDC assertion for short-lived tokens instead of a static API key;ListModelsreturns per-model capabilities - Google AI:
ListModels; Vertex AI now runs on the Google GenAI SDK backend through the same implementation - OpenAI: vLLM and other self-hosted backends work without an API key—only known public providers require one (
RequiresAPIKey) - Structured output fallbacks:
openai.WithStructuredOutputFallback()(DeepSeek, Z.ai GLM, MiniMax M) andollama.WithCloudStructuredOutputFallback()(Ollama Cloud) put the schema into the prompt and validate the answer locally instead of refusing - pgvector:
WithMetadataIndexescreates indexes overcmetadatakeys (with optional exclusions), so filtered searches stop scanning the whole table - Embeddings:
mistral.WithEmbeddingModel/WithEmbeddingHTTPClient,voyageai.WithBaseURL,embeddings.CheckEmbeddings
Security Fixes
- pgvector SQL injection: metadata filters and the collection name were interpolated into SQL, so a filter value containing a quote could rewrite the query (match every row or append a statement). They are now bound as query arguments—upgrade if user input can reach a filter
- Local file access via JSON Schema: the schema compiler resolved an external
$refthrough a file loader, letting a schema passed toWithStructuredOutputread local files and probe paths; external references are now refused
Bug Fixes
Anthropic
- Replayed turns keep every thinking block with its text and signature—fixes
400 … thinking.thinking: Field requiredin agent chains - All tool results of a multi-tool message are sent (previously only the first);
top_kreaches the API; tool-call arguments keep full numeric precision - Reasoning tokens are reported; a failing stream returns the text already received; the legacy completions stream no longer panics
Bedrock
- Converse carries the caller's tool choice (was pinned to
auto),top_k, images as images, one message per caller turn, and keeps parallel streamed tool calls apart - Each family gets the reasoning shape it accepts: DeepSeek R1 no longer fails with
ValidationException, gpt-oss and gpt-oss-safeguard receivereasoning_effort, Nova reasoning works on legacy InvokeModel - Encrypted reasoning (
redactedContent) is read and replayed on both APIs; streamed token counters and tool-call integer precision are fixed
OpenAI-Compatible Providers
- Agent calls with function tools on gpt-5.4+ no longer fail with 400 because of
reasoning_effort; DeepSeek no longer loses the output limit - Streamed responses report the same usage as non-streamed ones; an abandoned stream no longer leaks a goroutine
- Error messages from gateways with a non-OpenAI error shape (e.g. xAI) are preserved instead of a bare status code
- Mistral thinking chunks are read and replayed
Google AI / Vertex
- Answers are no longer cut at 2048 tokens by default;
seedand both penalties reach the generation config - Blocked prompts and exhausted token budgets are reported as distinct errors; Gemini 3 function calls without a signature are signed; streamed tool calls without an ID get one
- A failing stream keeps the partial answer
Other Providers
- Ollama: partial answers kept on any failure;
-cloudmodels on a local server are treated as Ollama Cloud; tool-call arguments decoded exactly - Mistral: tool choice,
top_p, JSON mode and seed reach the wire; embeddings honor the caller's model, endpoint, HTTP client and retries; streaming keeps finish reason and usage - Embeddings: empty or short responses return errors instead of panicking or returning a partial slice; Jina keeps the requested batch size
- pgvector: a failed or cancelled init no longer leaks or closes the caller's connection
Performance
- OpenAI and Ollama streaming build text with
strings.Builderand parse each chunk once—4000 chunks of a ~100 KB answer drop from 20.8 ms / 212 MB to 0.057 ms / 0.5 MB allocated
Deprecations
llms.GetModelContextSizeandllms.CalculateMaxTokens—their table stops at GPT-4o and returns 2048 for every newer modelgoogleai.WithRest()and(*vertex.Vertex).Close()are no-ops
Dependencies Updates
- Removed
cloud.google.com/go/vertexai—Vertex AI now goes through the Google GenAI SDK; no new dependencies
Testing & CI
- The CI test step could never fail (output piped into
teewithoutpipefail)—fixed - Gateway response headers are scrubbed from all recordings; Bedrock cassettes replay without an AWS environment; the conformance suite's streaming subtest now actually runs
Contributors
- @sirozha - Option warnings, vendor-native reasoning and sampling rules across all providers, reasoning block round-trips, truncation and tool-choice parity, Anthropic federation/fast mode/model listing, Vertex migration to the GenAI SDK, HuggingFace router, pgvector SQL injection fix, live verification against vendor APIs
- @asdek - Structured output fallbacks for OpenAI-compatible and Ollama Cloud models, pgvector metadata indexes, GPT-6 generation support, API-key-less OpenAI-compatible backends, test infrastructure
Full Changelog: v0.1.14-update.7...v0.1.14-update.8