You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
llmdialect/bridge.Model carries a model's figures, so a model list
can tell a client its window and its prices: ContextWindow, MaxOutputTokens, InputModalities, and Pricing, a bridge.ModelPricing of decimal strings per 1,000,000 tokens
(Currency, Input, Output, CachedInput, CacheWrite); a price
quoted per another count is the caller's to convert. ModelList and ModelEntry write them after the members each wire's clients read:
OpenAI's as context_window, max_output_tokens, input_modalities
and pricing (cached_input, cache_write); Anthropic's as its own max_input_tokens and max_tokens beside the same input_modalities
and pricing; Google's as its own inputTokenLimit and outputTokenLimit, with no modalities or prices; the lux wire's as
the OpenAI entry, since the lux dialect's JSON is snake case. Every
pricing member writes per: 1000000. A zero figure or an empty price
is left out, so an entry without figures renders byte for byte as
before.
Changed
llmdialect/openairesp: the Responses frontend keeps a reasoning
model's reasoning across turns, so a Responses client of the gateway
no longer loses it on the way to an OpenAI Responses upstream. include: ["reasoning.encrypted_content"] sets ir.Request.ReasoningReplay
instead of recording include loss. An input item of type reasoning
that carries encrypted_content becomes an opaque block of dialect openai-responses and kind reasoning in the assistant turn where it
stands, its JSON in the form ir.Opaque documents, so the backend
replays it as the client sent it and the other backends drop it with ir.LossOpaque (opaque) instead of reasoning. A reasoning item
without encrypted_content is still dropped with ir.LossReasoningItems: only the upstream's store could resolve it,
and this surface stores nothing. EncodeResponse writes such an opaque
block back as the output item it was, byte for byte, where it stands,
and a thinking block right before it no longer gets a reasoning item
of its own, since the item carries its summary. The stream does the
same: the thinking block's response.output_item.added and summary
deltas open the item, its response.output_item.done carries the
item as it came, and response.completed lists it; a reasoning item
without a thinking block before it gets an output_item.added with its
id and an empty summary. The output_item.done of a thinking block is
written when the next event arrives, not at its block_stop. The added frame names the item by a synthetic id, since the upstream's
id arrives only with the item; take the item from the done frame.
Opaque blocks of other dialects or kinds are still left out.