You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Compact long conversations safely before oversized model requests and between complete tool exchanges, while retaining the current task, recent inputs, the latest tool results, attachments, and provider continuation state.
Enforce model-aware input budgets that reserve output space and estimation headroom, and preserve generated replies that end at a provider length limit while rejecting incomplete compact summaries.
Keep compaction progress, savings, failures, cancellation, and refresh recovery stable inside the transcript without duplicating streamed tool output or creating empty assistant rows.
Providers and Models
Import structured model candidates from OpenRouter, BuzzHive, Anthropic, and Gemini, including display names, context windows, output limits, and explicitly reported image, audio, and tool capabilities.
Prefer endpoint metadata, fill only missing fields from exact model presets, and leave saved provider profiles unchanged until a candidate is explicitly added.
Refresh the OpenRouter presets with its free router and six commonly used paid models, with model-specific capabilities and limits.
Integration and Reliability
Attribute model traffic to Pudding across OpenAI Chat Completions, OpenAI Responses, Anthropic, and Gemini-compatible gateways using the standard OpenRouter application headers.
Preserve canonical history during compaction with atomic snapshot and active-turn ownership checks, monotonic lifecycle events, and no database migration.
Clarify the supported file-search context range and guide larger-context reads to precise file slices.