Resource-aware scheduling for local LLMs across multiple agents? #10971
Replies: 1 comment 2 replies
|
@Parikalp-Bhardwaj Thanks for checking in on this — your distinction from #7951 is right. That issue is about choosing a model; this is about coordinating access to a shared inference endpoint.
I would not start with priority queues or model-affinity scheduling. Instrumentation plus an opt-in concurrency cap is a much smaller experiment, and the resulting data can show whether priority or model-switch avoidance is actually needed. There is also a directly related RFC, #10970, covering shared admission and resource bounds. Its current review is asking for a daemon-owned first slice across turn origins, fair queues, a separate tool-execution ceiling, and bounded reserved interactive capacity. Your endpoint-specific proposal adds a useful dimension: different agents can compete for the same inference server even when their own turn budgets are separate. Could you add your proposal to #10970, outlining how to identify shared endpoints, which provider calls the limit would cover, and what to measure? That would help establish how the provider-call limit fits the broader admission design before you start implementation. |
Uh oh!
There was an error while loading. Please reload this page.
I've been exploring ZeroClaw and was curious about how it handles resource usage when multiple agents share a local inference provider like Ollama or llama.cpp.
I came across Discussion #8671, where someone ran into model switching because summarization was using a different local model.
It got me thinking about what happens when multiple agents and background tasks are running on a machine with limited RAM/VRAM.
For example:
Depending on the provider's configuration and available memory, repeated model switching could add noticeable latency, especially on smaller machines.
I was wondering whether ZeroClaw could benefit from an optional lightweight scheduling policy for local inference requests.
A few things that came to mind:
I'm not suggesting that ZeroClaw should manage GPU memory directly. Ollama and llama.cpp already handle model loading and execution.
I'm more curious about whether coordinating requests at the ZeroClaw runtime level would help when multiple agents are sharing limited local resources.
I also noticed #7951, which discusses effort-based local/cloud model routing. This seems slightly different to me: that issue is about choosing a model based on the task, while this is more about coordinating when local model calls run.
A few questions for the maintainers:
I'm interested in understanding how this fits with ZeroClaw's lightweight design. If there's a useful gap here, I'd be happy to explore it further and work on a small implementation.
All reactions