Skip to content

v0.88.0

Choose a tag to compare

@github-actions github-actions released this 27 Sep 08:55
· 33 commits to main since this release

Added

  • llmdialect/ir.BlockOpaque, a content block that carries one
    provider item verbatim for replay to the dialect that produced it:
    Block.Opaque holds an ir.Opaque with the producing Dialect, the
    item's Kind within it, and its JSON as Raw. A backend of that
    dialect writes Raw unchanged where the block stands; the other
    backends drop the block and record ir.LossOpaque (opaque), never an
    error. The lux dialect carries it both ways as {"type":"opaque", "opaque":{"dialect","kind","raw"}}, with raw as the JSON value
    itself, so a harness that stores sessions as lux JSON keeps the item
    byte for byte; luxsdk.Opaque and luxsdk.BlockOpaque re-export the
    vocabulary. Raw is kept in the form encoding/json writes (compact,
    with <, > and & escaped), which every later marshal leaves as it
    is; the lux decoder brings a body written by another encoder into that
    form. In a stream the block is one block_start whose header
    carries the payload, then its block_stop. The Anthropic, Chat and
    Responses frontends leave it out of the responses and streams they
    write, and the Anthropic stream shifts later content indices down past
    it so they stay dense. ir.PrefixCacheKeys hashes the block's
    dialect, kind and raw JSON after the fields every block contributes;
    keys of requests without opaque blocks are unchanged.
  • llmdialect/ir.Request.ReasoningReplay asks for the model's reasoning
    in a form the next request of the conversation can carry back, which
    a reasoning model served over OpenAI Responses needs to keep its
    reasoning across turns. The openairesp backend then writes include: ["reasoning.encrypted_content"] (beside the logprobs member when that
    is asked too) and store: false, and keeps each reasoning item that
    comes back with encrypted_content as an opaque block of kind
    reasoning, from the response body and from the stream's
    response.output_item.done frame alike, after the summary's thinking
    block. Encoding a request replays such a block as its input item, and
    the thinking block right before it travels inside the item instead of
    being reported as thinking loss. A reasoning item without
    encrypted_content is not kept, so a caller that does not ask sees
    the same blocks and streams as before. The lux dialect carries the
    ask as reasoning_replay; the anthropic backend needs no ask, since
    its thinking blocks carry their signatures; the openaichat backend
    records ir.LossReasoningReplay (reasoning_replay).