Putting a dollar cap on one ADK session, using LiteLlm #7386
domondi1
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
max_llm_callslimits how many calls an ADK session makes, but not what they cost, and sub-agents or a failing tool can still run up a bill on one request. Here's one way to give each session a hard dollar budget.ADK's
LiteLlmmodel takesapi_baseandextra_headers, so the session's id and budget can go to a small gateway that runs in the same process:Each call reserves its worst-case cost before it's sent, and a call that doesn't fit is refused with a 402 before it reaches the provider. In ADK 2.x that arrives as an event with
error_codeset rather than an exception, so stop reading events at that one (the runner raises if you keep iterating). Sub-agents share the budget if they use the same session's model.inferrail work <session id>shows what it cost.Caveats: only Chat Completions models through
LiteLlm, not native Gemini; setmax_tokens. Tested withgoogle-adk2.10.0 and 2.11.0,litellm1.103.2.Full setup: guide. I maintain Inferrail (open source, Apache-2.0), so take the recommendation with that in mind. Curious how others bound the cost of a single run here.
All reactions