Question — idle-based auto-stop for the kubernetes managed sandbox provider( or microsanboxes?) #8016
|
We're evaluating self-hosted Omnigent (omnigent-server-kubernetes, v0.14.0) as a team tool, with sandbox.provider: kubernetes for managed-sandbox sessions and auth enabled for multiple users. Realistically, once a team is using it, people will open sessions and rarely bother explicitly deleting them — so we need something to reclaim compute from idle sessions automatically, without losing conversation history. We found that kubernetes is a can_resume/resume_stopped provider (unlike e.g. Modal), and resume_managed_host() will relaunch a fresh Job under the same host id when a message arrives for a session whose host has gone offline. But nothing marks a healthy, connected, simply-unused session's host as offline — the heartbeat keeps going regardless of activity, so ManagedSandboxReaper's offline-age cutoff never starts, and the session (and its Job) just runs forever unless a human calls DELETE /v1/sessions/{id} (which also deletes the conversation, so it's not a good fit for "reclaim idle compute"). Is there a supported/recommended way to voluntarily mark an idle-but-healthy kubernetes-provider host as stopped/offline (so it becomes eligible for the same relaunch-on-resume path), without deleting the session? Or is directly deleting the underlying Job ourselves (bypassing the API) the intended way operators are expected to do this today? |
Replies: 1 comment 1 reply
|
There's a directly related proposal: #3225 describes an opt-in Your distinction between inactivity and offline age checks out for the reaper's normal active-generation path. I ran v0.14.0's extracted I'd link your multi-user deployment requirements to that PR and ask about the intended operator-facing stop/resume interface. The proposal is relevant to preserving conversations, but it doesn't establish that manually deleting a Job is the supported equivalent. |
There's a directly related proposal: #3225 describes an opt-in
sandbox.kubernetes.idle_timeout_spolicy that retains the session and host binding. It's still open and unmerged, so I wouldn't treat its proposed setting as released configuration.Your distinction between inactivity and offline age checks out for the reaper's normal active-generation path. I ran v0.14.0's extracted
_is_reap_candidatewith its realhost_is_livehelper: a fresh online heartbeat returned false; an offline host older than the cutoff returned true. The boundary and pending-cleanup controls passed too. This was a predicate test, not Kubernetes teardown or resume validation.I'd link your multi-user deployment requ…