Skip to content

v0.89.0

Choose a tag to compare

@github-actions github-actions released this 27 Sep 14:33
· 27 commits to main since this release

Added

  • llmdialect/bridge.Model carries a model's figures, so a model list
    can tell a client its window and its prices: ContextWindow,
    MaxOutputTokens, InputModalities, and Pricing, a
    bridge.ModelPricing of decimal strings per 1,000,000 tokens
    (Currency, Input, Output, CachedInput, CacheWrite); a price
    quoted per another count is the caller's to convert. ModelList and
    ModelEntry write them after the members each wire's clients read:
    OpenAI's as context_window, max_output_tokens, input_modalities
    and pricing (cached_input, cache_write); Anthropic's as its own
    max_input_tokens and max_tokens beside the same input_modalities
    and pricing; Google's as its own inputTokenLimit and
    outputTokenLimit, with no modalities or prices; the lux wire's as
    the OpenAI entry, since the lux dialect's JSON is snake case. Every
    pricing member writes per: 1000000. A zero figure or an empty price
    is left out, so an entry without figures renders byte for byte as
    before.

Changed

  • llmdialect/openairesp: the Responses frontend keeps a reasoning
    model's reasoning across turns, so a Responses client of the gateway
    no longer loses it on the way to an OpenAI Responses upstream. include: ["reasoning.encrypted_content"] sets ir.Request.ReasoningReplay
    instead of recording include loss. An input item of type reasoning
    that carries encrypted_content becomes an opaque block of dialect
    openai-responses and kind reasoning in the assistant turn where it
    stands, its JSON in the form ir.Opaque documents, so the backend
    replays it as the client sent it and the other backends drop it with
    ir.LossOpaque (opaque) instead of reasoning. A reasoning item
    without encrypted_content is still dropped with
    ir.LossReasoningItems: only the upstream's store could resolve it,
    and this surface stores nothing. EncodeResponse writes such an opaque
    block back as the output item it was, byte for byte, where it stands,
    and a thinking block right before it no longer gets a reasoning item
    of its own, since the item carries its summary. The stream does the
    same: the thinking block's response.output_item.added and summary
    deltas open the item, its response.output_item.done carries the
    item as it came, and response.completed lists it; a reasoning item
    without a thinking block before it gets an output_item.added with its
    id and an empty summary. The output_item.done of a thinking block is
    written when the next event arrives, not at its block_stop. The
    added frame names the item by a synthetic id, since the upstream's
    id arrives only with the item; take the item from the done frame.
    Opaque blocks of other dialects or kinds are still left out.