Skip to content

v0.3.4

Choose a tag to compare

@yangglive yangglive released this 14 Sep 10:15

What's New

Long Conversations

  • Compact long conversations safely before oversized model requests and between complete tool exchanges, while retaining the current task, recent inputs, the latest tool results, attachments, and provider continuation state.
  • Enforce model-aware input budgets that reserve output space and estimation headroom, and preserve generated replies that end at a provider length limit while rejecting incomplete compact summaries.
  • Keep compaction progress, savings, failures, cancellation, and refresh recovery stable inside the transcript without duplicating streamed tool output or creating empty assistant rows.

Providers and Models

  • Import structured model candidates from OpenRouter, BuzzHive, Anthropic, and Gemini, including display names, context windows, output limits, and explicitly reported image, audio, and tool capabilities.
  • Prefer endpoint metadata, fill only missing fields from exact model presets, and leave saved provider profiles unchanged until a candidate is explicitly added.
  • Refresh the OpenRouter presets with its free router and six commonly used paid models, with model-specific capabilities and limits.

Integration and Reliability

  • Attribute model traffic to Pudding across OpenAI Chat Completions, OpenAI Responses, Anthropic, and Gemini-compatible gateways using the standard OpenRouter application headers.
  • Preserve canonical history during compaction with atomic snapshot and active-turn ownership checks, monotonic lifecycle events, and no database migration.
  • Clarify the supported file-search context range and guide larger-context reads to precise file slices.