Skip to content

v0.1.14-update.8

Latest

Choose a tag to compare

@asdek asdek released this 29 Sep 06:32
5d59c61

Highlights

  • No More Silent Option Loss: New ContentResponse.Warnings reports every caller option a provider dropped, clamped or substituted before sending the request
  • Vendor-Native Reasoning Control: Thinking is sent in each model family's own mechanism—effort level, token budget, thinking_level, enable_thinking, Anthropic thinking object or Ollama think; combinations a vendor rejects fail locally with typed errors instead of a 400
  • Sampling Preserved: temperature, top_p, top_k and penalties are removed only where the vendor actually rejects them, no longer from every model that reasons
  • Lossless Reasoning Round-Trips: Signed and encrypted thinking blocks keep their order and signatures across turns on Anthropic and Bedrock, fixing replay 400s in agent loops
  • New Model Generations: GPT-6 (Astra, Sol, Luna), grok-4.7, Qwen 3.5–3.8, GLM 5.x, Kimi K2.6/K3, MiniMax M3, DeepSeek V4, Gemma 4, Claude Mythos Preview, gpt-oss-safeguard on Bedrock
  • Security: Fixed SQL injection through pgvector metadata filters; JSON Schemas can no longer read local files via $ref
  • Breaking Changes: API removals and new typed errors—read the section below before upgrading

Breaking Changes

This release deliberately does not keep requests byte-identical: the wire changes only where the vendor rejected the old request or where a silently lost option now reaches the vendor. There is no global switch back; on the OpenAI provider llms.WithExtraBody can still put any field back into the request.

API Changes

  • tools/perplexity: removed ModelSonarReasoning and ModelR11776 (retired by the vendor)—use ModelSonarReasoningPro
  • llms/googleai/vertex: removed ten names duplicated from googleai (RoleUser, ResponseMIMETypeJson, ErrNoContentInResponse, …)—switch to the googleai package; vertex.Vertex now embeds *googleai.GoogleAI, so vertex.Vertex{CallbacksHandler: h} literals no longer compile
  • testing/llmtest: TestLLM takes options—declare WithoutStreaming() / WithoutToolCalls() for providers without them; MockLLM.GenerateContentStream is gone (use llms.WithStreamingFunc)
  • anthropic.ErrModelRefusal is now an alias of llms.ErrModelRefusal; reasoning.ClaudeSupportsEffortWithBudget takes a provider; reasoning.OffWire gained four values and was renumbered—never persist its numeric value
  • ContentChoice, ContentResponse, ContentReasoning and CallOptions gained fields—unkeyed struct literals no longer compile

Behavior Changes

  • Typed errors before the request where the vendor would answer 400 or silently ignore the option:
    • reasoning.ErrReasoningOffUnsupported—WithReasoningDisabled() on models that cannot stop thinking (base GPT-5, gpt-6-astra, DeepSeek R1, grok-4.7, Kimi K3, GLM 5.3 via Mistral, Fable/Mythos and gpt-oss on Bedrock, …)
    • reasoning.ErrEffortWithTools—an explicit effort combined with function tools on gpt-5.4–5.6, gpt-6-sol and gpt-6-luna
    • reasoning.ErrThinkingBudgetUnsupported, reasoning.ErrChatToolsUnsupported, anthropic.ErrAssistantPrefillUnsupported, anthropic.ErrForcedToolUseWithThinking
    • llms.ErrStructuredOutputUnsupported—WithStructuredOutput for DeepSeek, GLM and MiniMax M via the OpenAI provider, Ollama Cloud, and Bedrock models whose AWS model card lacks structured output (see the opt-in fallbacks below)
  • Google AI / Vertex: default output limit raised from 2048 to 16384 tokens; WithGRPCClient/WithGRPCConn return ErrOptionNotHonored; vertex.New needs a project and location (options or GOOGLE_CLOUD_PROJECT/GOOGLE_CLOUD_LOCATION) or returns ErrMissingCloudTarget; Vertex StopReason uses the vendor spelling (MAX_TOKENS, STOP)
  • HuggingFace: moved to router.huggingface.co and /v1/chat/completions (the old host no longer resolves); top_k, min_length and repetition_penalty are no longer sent; use WithMaxTokens instead of WithMaxLength; default model meta-llama/Llama-3.1-8B-Instruct
  • Mistral: default model open-mistral-7b → ministral-8b-latest
  • Bedrock: legacy InvokeModel counters in GenerationInfo are int (were int32); Nova reasoning applies only to amazon.nova-2-lite
  • pgvector: metadata filter keys must be plain identifiers, otherwise ErrInvalidFilterKey
  • Structured output: schemas with an external $ref no longer compile
  • Extra body: moved out of CallOptions.Metadata—read it with llms.ExtraBody(opts)

New Features

Option Warnings

ContentResponse.Warnings lists each caller option that did not reach the vendor as asked—option, model, requested and sent value, reason and kind (drop, clamp, substitute). Every provider fills it, and a warning appears only where a value really changed.

Reasoning & Sampling per Model Family

  • Claude behind OpenAI-compatible gateways (bare names or LiteLLM anthropic/, bedrock/, vertex_ai/ routes) receives Anthropic's thinking object for budget, adaptive and off; public OpenAI-compatible hosts keep the standard wire
  • Qwen on DashScope: reasoning_effort for 3.8, a working disable via enable_thinking for 3.5–3.7, and thinking_budget as its own field—also for GLM, Kimi and DeepSeek served by DashScope
  • GLM, Kimi, MiniMax, DeepSeek: thinking is switched off the way each vendor documents; their reasoning (and Qwen's) is replayed on every assistant turn
  • Gemini: Gemini 3 is disabled through the level scale, MINIMAL reaches the generations that accept it, Gemma 4 takes HIGH/MINIMAL, Gemini 2.5 budgets are clamped into the documented range
  • Bedrock: only families with their own mechanism get a thinking config (Claude, Nova 2 Lite, grok, gpt-oss); DeepSeek, GLM, Kimi, MiniMax and Qwen no longer receive the Anthropic object
  • Ollama: sends think (gpt-oss gets its three levels) and surfaces the server's native thinking
  • Claude: thinking budgets are raised to the vendor minimum of 1024; a caller's top_p ≥ 0.95 survives thinking
  • ReasoningSupportFor reports the mechanism (budget, adaptive or both) and no longer guesses effort tiers for unclassified models

Reasoning Blocks

ContentReasoning.Blocks keeps every reasoning block—readable text with its signature, or encrypted data—in the order the vendor produced it, and replays it unchanged. Content and Signature work as before; use HasContent() to tell a model that actually reasoned from one that only signed its answer (Gemini 3 signs every answer).

Truncation, Tool Choice & Output Limit

  • ContentChoice.Truncated normalizes the length, max_tokens and model_length stop reasons across providers; llms.WithFailOnTruncation() turns truncation into a typed error that still carries the partial answer
  • Tool choice is classified once (llms.ClassifyToolChoice) and reaches the wire in every spelling on OpenAI, Anthropic, Bedrock Converse and Google AI; Mistral accepts all vendor values
  • The output-limit field follows the route: max_tokens for grok, Qwen and DeepSeek, max_completion_tokens for OpenAI

New Options & Provider APIs

  • Call options: WithVerbosity, WithLogProbs/WithTopLogProbs, WithInferenceSpeed (Anthropic fast mode), WithFailOnTruncation, and provider-neutral llms.WithExtraBody (merged into nested objects on the OpenAI provider, reported in Warnings elsewhere)
  • Anthropic: WithFederation—Workload Identity Federation exchanges an OIDC assertion for short-lived tokens instead of a static API key; ListModels returns per-model capabilities
  • Google AI: ListModels; Vertex AI now runs on the Google GenAI SDK backend through the same implementation
  • OpenAI: vLLM and other self-hosted backends work without an API key—only known public providers require one (RequiresAPIKey)
  • Structured output fallbacks: openai.WithStructuredOutputFallback() (DeepSeek, Z.ai GLM, MiniMax M) and ollama.WithCloudStructuredOutputFallback() (Ollama Cloud) put the schema into the prompt and validate the answer locally instead of refusing
  • pgvector: WithMetadataIndexes creates indexes over cmetadata keys (with optional exclusions), so filtered searches stop scanning the whole table
  • Embeddings: mistral.WithEmbeddingModel/WithEmbeddingHTTPClient, voyageai.WithBaseURL, embeddings.CheckEmbeddings

Security Fixes

  • pgvector SQL injection: metadata filters and the collection name were interpolated into SQL, so a filter value containing a quote could rewrite the query (match every row or append a statement). They are now bound as query arguments—upgrade if user input can reach a filter
  • Local file access via JSON Schema: the schema compiler resolved an external $ref through a file loader, letting a schema passed to WithStructuredOutput read local files and probe paths; external references are now refused

Bug Fixes

Anthropic

  • Replayed turns keep every thinking block with its text and signature—fixes 400 … thinking.thinking: Field required in agent chains
  • All tool results of a multi-tool message are sent (previously only the first); top_k reaches the API; tool-call arguments keep full numeric precision
  • Reasoning tokens are reported; a failing stream returns the text already received; the legacy completions stream no longer panics

Bedrock

  • Converse carries the caller's tool choice (was pinned to auto), top_k, images as images, one message per caller turn, and keeps parallel streamed tool calls apart
  • Each family gets the reasoning shape it accepts: DeepSeek R1 no longer fails with ValidationException, gpt-oss and gpt-oss-safeguard receive reasoning_effort, Nova reasoning works on legacy InvokeModel
  • Encrypted reasoning (redactedContent) is read and replayed on both APIs; streamed token counters and tool-call integer precision are fixed

OpenAI-Compatible Providers

  • Agent calls with function tools on gpt-5.4+ no longer fail with 400 because of reasoning_effort; DeepSeek no longer loses the output limit
  • Streamed responses report the same usage as non-streamed ones; an abandoned stream no longer leaks a goroutine
  • Error messages from gateways with a non-OpenAI error shape (e.g. xAI) are preserved instead of a bare status code
  • Mistral thinking chunks are read and replayed

Google AI / Vertex

  • Answers are no longer cut at 2048 tokens by default; seed and both penalties reach the generation config
  • Blocked prompts and exhausted token budgets are reported as distinct errors; Gemini 3 function calls without a signature are signed; streamed tool calls without an ID get one
  • A failing stream keeps the partial answer

Other Providers

  • Ollama: partial answers kept on any failure; -cloud models on a local server are treated as Ollama Cloud; tool-call arguments decoded exactly
  • Mistral: tool choice, top_p, JSON mode and seed reach the wire; embeddings honor the caller's model, endpoint, HTTP client and retries; streaming keeps finish reason and usage
  • Embeddings: empty or short responses return errors instead of panicking or returning a partial slice; Jina keeps the requested batch size
  • pgvector: a failed or cancelled init no longer leaks or closes the caller's connection

Performance

  • OpenAI and Ollama streaming build text with strings.Builder and parse each chunk once—4000 chunks of a ~100 KB answer drop from 20.8 ms / 212 MB to 0.057 ms / 0.5 MB allocated

Deprecations

  • llms.GetModelContextSize and llms.CalculateMaxTokens—their table stops at GPT-4o and returns 2048 for every newer model
  • googleai.WithRest() and (*vertex.Vertex).Close() are no-ops

Dependencies Updates

  • Removed cloud.google.com/go/vertexai—Vertex AI now goes through the Google GenAI SDK; no new dependencies

Testing & CI

  • The CI test step could never fail (output piped into tee without pipefail)—fixed
  • Gateway response headers are scrubbed from all recordings; Bedrock cassettes replay without an AWS environment; the conformance suite's streaming subtest now actually runs

Contributors

  • @sirozha - Option warnings, vendor-native reasoning and sampling rules across all providers, reasoning block round-trips, truncation and tool-choice parity, Anthropic federation/fast mode/model listing, Vertex migration to the GenAI SDK, HuggingFace router, pgvector SQL injection fix, live verification against vendor APIs
  • @asdek - Structured output fallbacks for OpenAI-compatible and Ollama Cloud models, pgvector metadata indexes, GPT-6 generation support, API-key-less OpenAI-compatible backends, test infrastructure

Full Changelog: v0.1.14-update.7...v0.1.14-update.8