Skip to content

rig-v0.40.0

Choose a tag to compare

@github-actions github-actions released this 11 Jul 00:50
2f37dfc

Added

Fixed

Other

Contributors

Changed

  • (agent) [breaking] max_turns and default_max_turns now bound the exact total number of model calls, including the initial call, tool continuations, and retries. A budget of 0 makes no model call, while 1 permits only the initial call. Unconfigured tool-then-answer flows now need an explicit total budget of 2. To preserve the former maximum allowance of an explicit old budget n, account for the old effective n + 2 calls; otherwise, set the intended literal total.

  • (tool) [breaking] flatten Tool / ToolDyn metadata: tool authors now implement description() and parameters() directly, and Tool::definition(prompt) / ToolDyn::definition(prompt) are removed. ToolDefinition remains a provider/request artifact generated from registered tools, with Tool::NAME / Tool::name() / ToolDyn::name() as the single source of truth for advertised and dispatched tool names.

  • (providers) [breaking] migrate llamafile onto the shared GenericCompletionModel<Ext> / GenericEmbeddingModel<Ext> path, deleting its hand-rolled completion model, request types, message flattening, and streaming profile. llamafile::CompletionModel / llamafile::EmbeddingModel are now type aliases for the generic models; the provider-specific StreamingCompletionResponse type is replaced by the shared OpenAI one. Requests now serialize messages in the shared OpenAI shape (single-text user content still flattens to a string; system/multi-part content is sent as a content-part array, which llama.cpp-family servers accept).

  • (openai) [breaking] new OpenAICompatibleProvider trait (mirroring AnthropicCompatibleProvider) is now required by GenericCompletionModel's Ext parameter; it carries the telemetry provider name (so minimax/zai/xiaomimimo spans stop reporting as "openai") and an EMITS_COMPLETE_SINGLE_CHUNK_TOOL_CALLS flag for llama.cpp-style streaming tool calls.

  • (providers) [breaking] migrate the remaining OpenAI-chat-compatible providers onto GenericCompletionModel<Ext> β€” groq, deepseek, mistral, together, moonshot (OpenAI side), perplexity, hyperbolic, mira, azure, and huggingface all lose their hand-rolled CompletionModel structs, request types, and TryFrom<message::Message> conversions; CompletionModel in each module is now a type alias for the generic model. Provider wire dialects live in OpenAICompatibleProvider hooks: an associated Response type, completion_path (Azure deployment URLs, /v1-prefixed routes), prepare_request (Groq native-tool folding, Moonshot required tool-choice coercion, HuggingFace Fireworks model ids, Perplexity/Mira tool stripping), finalize_request_body (DeepSeek string content + thinking-gated tool choice, Mistral "any" tool choice + prefix field + reasoning stripping, Mira raw-message flattening), and SUPPORTS_RESPONSE_FORMAT / STREAM_INCLUDE_USAGE consts. Provider-specific StreamingCompletionResponse types are replaced by the shared OpenAI one.

  • (openai) [breaking] ToolChoice gains a Function { name } variant serializing OpenAI's {"type":"function","function":{"name":...}} form, so message::ToolChoice::Specific with one function is now supported instead of erroring; CompletionRequest fields are now public; OpenAIRequestParams gains a supports_response_format field.

  • (openai) the shared TryFrom<message::ToolResult> conversion now prefers call_id over id for tool_call_id (matching provider-issued call ids); the shared streaming delta accepts reasoning as an alias for reasoning_content (Groq), and the deprecated function_call finish reason maps to tool-call handling.

  • (providers) behavior notes from the migration: max_tokens is now forwarded by deepseek, together, hyperbolic, and azure (previously silently dropped); together's streaming request uses standard stream/stream_options instead of stream_tokens, and a rig-level ToolChoice::Required now serializes as required instead of erroring; perplexity's non-streaming endpoint drops its stray /v1 prefix (matching its streaming path and the real API); mira's preamble is sent as a system message instead of user; response_format derived from output_schema is deferred while tools are pending a result (groq/mistral/azure previously applied it unconditionally); groq's streaming usage no longer falls back to the legacy x_groq.usage envelope.

  • (openrouter) [breaking] de-fork OpenRouter's parallel message model (issue #2035 phase 4): openrouter::{Message, UserContent, ImageUrl} are now re-exports of the shared OpenAI types, and the fork's FileContent/VideoUrlContent are replaced by shared FileData/VideoUrl. To support this, the shared OpenAI types gain OpenRouter's optional extensions β€” UserContent::Video, ImageUrl.detail becomes Option<ImageDetail> (OpenAI still sends "detail":"auto"), and Message::Assistant gains a skip-when-empty reasoning_details field, an inbound-only images field (never serialized back into requests), plus a deserialize-only role: "model" alias. ReasoningDetails/ResponseImage move into the openai module (re-exported from openrouter). OpenRouter-specific message conversion now goes through openrouter::messages_from_rig_message; TryInto<Vec<openrouter::Message>> resolves to the plain shared conversion. OpenRouter keeps its own request/response/streaming layer (provider preferences, cost accounting, reasoning-details grouping, generated-image extraction) as a documented exception.

  • (openai) the UserContent audio part now serializes its tag as input_audio (matching OpenAI's actual API); audio is still accepted when deserializing.

  • (openai) StreamingCompletionResponse is now generic over the provider's streaming usage payload (StreamingCompletionResponse<U = Usage>, selected via OpenAICompatibleProvider::StreamingUsage), so Mistral's cached-token fallbacks and DeepSeek's cache hit/miss counters survive streaming instead of being narrowed to OpenAI's usage shape.

  • (providers) pre-migration request filtering is preserved where provider support is unverified: hyperbolic still drops tools/tool_choice/output_schema with warnings, and perplexity flattens text-only message content back to plain strings (mixed multimodal content is passed through for sonar models). llamafile keeps the current mapping of output_schema to a json_schema response format (as on the shared path since the llamafile migration; modern llama.cpp servers support it).

  • (llamafile) the chat cassettes are now recorded against an actual llama.cpp llama-server, confirming the shared OpenAI wire shape (content-part arrays, tool calls, tool results) against the real llamafile-family server rather than an OpenAI-compatible proxy.

  • (openai) the assistant tool-call echo now serializes call_id (falling back to id) so it stays consistent with the tool-result side when history recorded via the Responses API is replayed through chat completions; streaming delta content tolerates content-part arrays (Mistral reasoning models) instead of dropping the chunk.

  • (providers) review fixes: mira and perplexity no longer send stream_options (their APIs never received it pre-migration); moonshot rejects a specific forced tool client-side again; openrouter serializes plain assistant reasoning under its documented reasoning key; azure telemetry spans report azure.openai again; mira usage math saturates instead of overflowing; perplexity strips tool-exchange remnants from shared histories.

  • (providers) second review round: openrouter tool-result messages prefer the provider-issued call_id (matching the assistant echo side); Azure's deployment URL stays pinned to the model the handle was created with (a per-request model override only changes the body, as pre-migration); shared streaming spans record gen_ai.system_instructions again; providers without tool support (perplexity, mira, and now hyperbolic) sanitize tool-exchange remnants from shared histories via one shared helper that also preserves strict role alternation (tool-call-only assistant turns are dropped and consecutive assistant turns merged); openrouter's dead pre-migration ToolChoice type is removed, and ToolChoice::Specific with multiple function names now errors client-side for openrouter (the old fork serialized a non-standard array).

  • (moonshot) [breaking] reasoning-only assistant history turns are no longer preserved: the shared conversion drops assistant messages with neither text nor tool calls. Reasoning attached to text or tool-call turns still round-trips via reasoning_content.

  • (providers) [breaking] responses with empty assistant content and no tool calls now surface the shared path's "empty response" error for hyperbolic, perplexity, and huggingface (previously they returned an empty text completion).

  • (providers) [breaking] additional removed public items: the raw response types of perplexity, hyperbolic, and huggingface (each module keeps a CompletionResponse alias to the shared OpenAI payload; Message/Choice/Usage/Delta/Role companions are gone), together::ToolChoice/ToolChoiceFunctionKind, moonshot::ToolChoice, groq::send_compatible_streaming_request and deepseek::send_compatible_streaming_request (use openai::send_compatible_streaming_request), and openrouter's UserContent builder helpers (image_url, file_base64, video_url, ...) β€” construct the shared openai content variants directly.

  • (openai) [breaking] sending rig Video user content to providers on the shared conversion now serializes a video_url content part (an OpenRouter/gateway extension) instead of returning a client-side conversion error; providers without video support will reject it server-side.

  • (providers) [breaking] telemetry: migrated providers' streaming spans are now named chat with gen_ai.operation.name = "chat" (previously chat_streaming). GenAI message-content span fields (gen_ai.input.messages / gen_ai.output.messages) are intentionally left empty instead of recording serialized request/response messages, preserving the privacy/cardinality behavior from #2065; the public SpanCombinator::record_model_output helper is removed. gen_ai.request.model reports the per-request model override when one applies.

  • (providers) third review round: history sanitization treats refusal parts as text when flattening and, for alternation-strict perplexity, merges consecutive same-role turns (dropping a tool exchange could previously leave user/user adjacency its API rejects); streaming no longer overwrites caller-supplied stream_options; openrouter's encrypted reasoning details now correlate with the wire tool-call id and its non-streaming usage uses the reported completion_tokens (no underflow); base64 videos with unrecognized MIME types round-trip as data-URI URLs instead of failing conversion.

  • (providers) fourth review round β€” agent structured output: GenericCompletionModel no longer claims native structured output composes with tools for every provider; it now follows SUPPORTS_RESPONSE_FORMAT. Agents with tools plus an output schema on deepseek/together/moonshot/huggingface/hyperbolic/perplexity/mira fall back to tool-mode schema enforcement as their pre-migration models did (the migration had silently dropped the schema entirely); groq/mistral/azure now compose natively like openai.

  • (openai) new OpenAICompatibleProvider::SUPPORTS_TOOLS const (default true): perplexity, hyperbolic, and mira set it false and tools/tool_choice are dropped with a warning during request conversion β€” before tool-choice validation, so a multi-name ToolChoice::Specific no longer errors client-side on providers that ignored it pre-migration.

  • (openai) streaming robustness: include_usage is inserted into caller-supplied stream_options instead of being skipped (or clobbering the caller's keys, the pre-migration behavior); a delta carrying both reasoning_content and reasoning no longer fails as a serde duplicate-field error that dropped the whole chunk; streaming tool-call index defaults to 0 when omitted (Mistral marks it optional); CompletionResponse.object/created are defaulted on deserialization for gateways that omit them (HuggingFace router sub-providers).

  • (openrouter) non-streaming usage falls back to total - prompt (saturating) when the gateway omits completion_tokens; streaming spans follow the shared telemetry behavior of leaving GenAI message-content fields empty.

  • (openai) [breaking] ToolChoice is now #[non_exhaustive]; GenericCompletionModel's strict_tools/tool_result_array_content fields are private (use the with_* builder methods) and the redundant with_model constructor is removed (use new).

Removed

  • (derive) [breaking] remove the unused public rig_derive::ProviderClient derive macro and its deluxe dependency; Embed and rig_tool are unchanged, and no replacement is provided.
  • (core) [breaking] remove unused Extractor::{get_inner, into_inner} and the always-failing TryFrom<String> for Nothing; no direct replacements are provided.
  • (core) [breaking] remove the unused public streaming::stream_completion_to_stdout helper; use the high-level agent::stream_to_stdout helper instead.
  • (core) [breaking] remove the unused public AudioGeneration<M>, ImageGeneration<M>, and Transcription<M> wrapper traits; use the corresponding AudioGenerationModel, ImageGenerationModel, and TranscriptionModel APIs and request builders directly.
  • (core) [breaking] remove the unused evals module (Eval trait, judge metrics, and builders) along with the experimental feature flag that gated it
  • (anthropic) [breaking] remove the unused public providers::anthropic::decoders module; Anthropic streaming uses the shared SSE machinery.
  • (providers) [breaking] remove the Galadriel provider integration (providers::galadriel), including its client, model constants, environment-variable support, and ignored live tests.