You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
CostLedger — structured cost tracking across attempts and escalations; PlanResult.ledger records every LLM call with reason, model, tokens, cost, and success/failure
Strict budget ceiling — router.route() now raises ValueError when no model fits the budget (controlled via budget_strict parameter; default True)
Decomposition visibility — Plan now exposes decomposition_source ("llm" or "fallback") and decomposition_error so callers know when task decomposition fell back
TTL on failed models — proxy's failed-model set now expires entries after 5 minutes (configurable) and caps at 50 entries, preventing unbounded growth
Shared HTTP client — providers reuse the proxy's httpx.AsyncClient instead of creating a new client per request, reducing connection overhead
Ledger table in CLI — tokenwise plan --execute now prints a Rich cost breakdown table with wasted-cost summary
Step-level capabilities — Step.required_capabilities explicitly tracks what each step needs, used during escalation filtering
Structured error codes — StepResult.http_status_code captures the HTTP status from provider errors for reliable error classification
Changed
Escalation ordering — executor and proxy now escalate to stronger tiers first (FLAGSHIP → MID) instead of trying budget tier first; fallback candidates are filtered by the full set of required capabilities
Error classification — split HTTP codes into unusable (402, 403, 404) vs transient (500, 502, 503, 504); _is_model_error checks the integer status code instead of brittle string matching
Router — budget error path now picks the cheapest model by estimate_cost() (input + output) instead of input_price alone
Router — budget is strict by default; planner uses budget_strict=False for its own internal routing
Package name — PyPI distribution renamed to tokenwise-llm (import name tokenwise unchanged)
Fixed
Proxy failed_models set no longer grows without bound across the server lifetime
Removed HTTP 400 from retryable/fallback codes (400 is a request schema error, not a model outage)