Skip to content

v0.3.0

Choose a tag to compare

@itsarbit itsarbit released this 21 Feb 07:21
· 40 commits to master since this release

What's New in v0.3.0

Added

  • CostLedger — structured cost tracking across attempts and escalations; PlanResult.ledger records every LLM call with reason, model, tokens, cost, and success/failure
  • Strict budget ceiling — router.route() now raises ValueError when no model fits the budget (controlled via budget_strict parameter; default True)
  • Decomposition visibility — Plan now exposes decomposition_source ("llm" or "fallback") and decomposition_error so callers know when task decomposition fell back
  • TTL on failed models — proxy's failed-model set now expires entries after 5 minutes (configurable) and caps at 50 entries, preventing unbounded growth
  • Shared HTTP client — providers reuse the proxy's httpx.AsyncClient instead of creating a new client per request, reducing connection overhead
  • Ledger table in CLI — tokenwise plan --execute now prints a Rich cost breakdown table with wasted-cost summary
  • Step-level capabilities — Step.required_capabilities explicitly tracks what each step needs, used during escalation filtering
  • Structured error codes — StepResult.http_status_code captures the HTTP status from provider errors for reliable error classification

Changed

  • Escalation ordering — executor and proxy now escalate to stronger tiers first (FLAGSHIP → MID) instead of trying budget tier first; fallback candidates are filtered by the full set of required capabilities
  • Error classification — split HTTP codes into unusable (402, 403, 404) vs transient (500, 502, 503, 504); _is_model_error checks the integer status code instead of brittle string matching
  • Router — budget error path now picks the cheapest model by estimate_cost() (input + output) instead of input_price alone
  • Router — budget is strict by default; planner uses budget_strict=False for its own internal routing
  • Package name — PyPI distribution renamed to tokenwise-llm (import name tokenwise unchanged)

Fixed

  • Proxy failed_models set no longer grows without bound across the server lifetime
  • Removed HTTP 400 from retryable/fallback codes (400 is a request schema error, not a model outage)