A hard dollar limit per job for IChatClient (OpenAI), with a DelegatingHandler and AsyncLocal #7802
domondi1
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
With
UseFunctionInvocation(), oneGetResponseAsynccan turn into several model calls, and a job that fans out into parallel requests has no single place that stops it at a dollar amount. Usage only comes back after each call, so a cap you check in your own code is always a batch behind.Here's a pattern that needs nothing new from Microsoft.Extensions.AI: an
AsyncLocaljob scope, aDelegatingHandlerthat stamps the job's id and budget on every HTTP request the OpenAI client sends, and an OpenAI-compatible gateway that enforces the budget before forwarding. The handler half is useful on its own too, for example to send per-tenant or per-job attribution headers to any proxy.Run the gateway next to the app (it uses your provider key):
Then wrap the job:
and check what it cost afterwards:
Every call with the same job id draws from one budget. Each call reserves its worst-case cost before it's sent, so parallel calls (and parallel tool loops) in one job can't overspend together. A call that doesn't fit is refused with HTTP 402 before it reaches OpenAI, which surfaces as
ClientResultExceptionwithStatus == 402. The default OpenAI retry policy doesn't retry 402, and the two headers aren't forwarded upstream.The same
OpenAIClientworks unchanged for Agent Framework (openai.GetChatClient(model).AsIChatClient().AsAIAgent(...)) and Semantic Kernel (builder.AddOpenAIChatCompletion(model, openai), where the refusal shows up asHttpOperationException).What I tested: .NET 10, Microsoft.Extensions.AI.OpenAI 10.10.1, Microsoft.Agents.AI.OpenAI 1.23.0, Semantic Kernel 1.80.1, inferrail 0.4.12, against a local stub upstream: six parallel tool-calling requests on one job id. With a tight budget, the calls that didn't fit were refused before reaching the upstream and the job's recorded total stayed under the cap. I haven't run it from .NET against the real OpenAI API yet, so reports are welcome.
Caveats:
AsyncLocalflows throughawaitandTask.WhenAll, but not into work you hand to a separate queue or hosted service. Begin the scope where the job actually runs.gpt-4o-mini,gpt-4.1-miniandgpt-4.1are built in;inferrail modelslists the rest and you can add your own prices).Setup details: guide. I maintain Inferrail (open source, Apache-2.0), so take the suggestion with that in mind. If there's a more idiomatic M.E.AI place for per-request headers than the transport, I'd like to hear it.
All reactions