Skip to content
This repository was archived by the owner on Jul 15, 2026. It is now read-only.

v0.0.4 - Token Optimization & Prompt Caching

Latest

Choose a tag to compare

@jorgepb96 jorgepb96 released this 09 Apr 07:48
· 7 commits to main since this release

What's New

Automatic Conversation Compaction

When conversation history exceeds the token limit, evicted messages are now automatically summarized into a compact context summary instead of being silently discarded. This preserves key context across long conversations.

Changes:

  • Engine auto-summarizes evicted messages using the same provider/model
  • Re-summarization triggers only when ≥10 new messages are evicted (avoids per-turn cost)
  • Shows a visible "🔄 Compacting conversation..." notification to the user
  • Summary is persisted to disk alongside conversation history

Prompt Caching for OpenRouter (Anthropic models)

Anthropic models routed through OpenRouter now benefit from prompt caching, matching the native Anthropic provider behavior.

Changes:

  • Detects Anthropic models on OpenRouter (anthropic/* or claude)
  • Applies cache_control: { type: "ephemeral" } on system prompt and prefix messages
  • Applied to chat(), chatWithTools(), and chatStream() methods
  • Graceful fallback: non-Anthropic models on OpenRouter are unaffected

Prompt Caching Summary (all providers)

Provider Caching Type
Anthropic ✅ Explicit cache_control (already implemented)
OpenAI ✅ Automatic (prefix-based, ≥1024 tokens)
Google Gemini ✅ Implicit automatic
Groq ✅ Automatic
OpenRouter ✅ cache_control passthrough for Anthropic models

UI: Uninstall All Skills Button

Added a convenience button to uninstall all built-in skills at once when all are already installed.

Full Changelog

v0.0.3...v0.0.4