-
Notifications
You must be signed in to change notification settings - Fork 3
Engine Apis
src/engine/apis/ implements the engine's provider API layer. It registers two Clio-specific API providers with the engine registry: openai-completions (for all OpenAI-compatible servers including LM Studio, llama.cpp router, LiteLLM gateway, and vLLM) and ollama-native (for Ollama's native /api/chat endpoint). Both providers wrap Pi SDK's built-in streaming APIs with Clio-specific logic: residency management for local servers, thinking payload mutations, sampling overrides, response metadata observation, sentinel stripping, and degraded-inference watching.
The area also contains a shared capacity-aware model-residency reconciler (residency.ts) that decides load and evict for local runtimes. Each local runtime provides an adapter that plugs into this shared policy, so co-residency math, protection tiers, TTL dedup, and cross-process locking all come from one place.
src/engine/apis/index.ts exports registerClioApiProviders(), which calls registerEngineApiProvider (from src/engine/api-registry.ts) twice: once for openAICompletionsApiProvider with source "clio", once for ollamaNativeApiProvider with source "clio". The registry's wrappedProvider verifies that model.api matches the registered API before dispatching. The engine's dispatch path (engineStream / engineStreamSimple in api-registry.ts) resolves the registered provider for a given api string and calls its stream or streamSimple.
The provider's stream and streamSimple both call streamCompletions, which builds a pipeline of stream wrappers over Pi's openAICompletionsApi(). The pipeline order (outermost to innermost) is:
-
guardMalformedToolCalls— fails the stream when a tool call arrives with empty arguments andstopReason !== "length"(the latter case is Pi's length-truncation salvage, which Clio does not treat as malformed). -
withReasoningTokenEstimate— addsusage.reasoningTokenswhen the provider did not report reasoning usage. -
stripSentinelsFromStream— removes tokenizer special-token sentinels (e.g.,<tool_call>,) from text deltas. Skipped for diffusion providers. -
stripNeverReasoningFromStream— strips thinking events and thinking content fromdone/errormessages when the model's thinking mechanism is"never". -
filterGemmaChannelStream— splits reasoning and text channels for Gemma models that use channel markers. -
withLiteLLMRouteFailureAdvice— enriches error messages with LiteLLM routing context. -
withResponseModelIdCapture— wraps thefetchto observe response model ID, backend timings, gateway routing headers, and diffusion frames; annotatesdoneanderrorevents. -
withLocalResidency— ensures local residency (LM Studio, llama.cpp, Ollama via gateway) and wraps the stream in the degraded-inference watchdog. -
start— the actual Pi SDK stream (piOpenAICompletions.stream).
Arguments forwarded: normalizeContext(preserveInterruptedReasoning(effectiveContext, requestModel, resolved)) and withRemainingContextBudget(requestModel, effectiveContext, withSamplingOverrides(requestModel, capturedOptions, resolved)).
Thinking payload mutations are applied by composeThinkingOnPayload, which calls applyThinkingPayload and applyLmStudioPayload after Pi applies its own sampling. The mechanism branches:
-
"none": strips thinking request fields; restoresenable_thinkingetc. frommodel.samplingParamsfor Inception P1. -
"effort-levels": writesreasoning_effortwhen the family resolved one; off also carrieschat_template_kwargs.enable_thinking=falsefor strict templates. -
"budget-tokens": writesthinking: { type: "enabled", budget_tokens }when the family declaresthinkingFormat: "anthropic-extended"and the runtime is not vLLM. -
"on-off": writeschat_template_kwargs.enable_thinking. -
"always-on": does not touch the payload.
LM Studio-specific: applyLmStudioPayload deletes chat_template_kwargs (LM Studio does not use it), sets ttl and draft_model from the load profile, and maps reasoning levels through lmStudioReasoningEffort.
LiteLLM-specific: withLiteLLMRequestOptions sets x-litellm-* headers (tags, session-id, timeout, stream-timeout, num-retries) and forces maxRetries: 0 so the OpenAI SDK does not retry below Clio's visible recovery loop.
llama.cpp-specific: applyLlamaCppPromptCachePayload sets cache_prompt based on cacheRetention.
The provider's stream and streamSimple both call runStream, which:
- Builds request headers from
model.headersandoptions.headers. - Wraps the stream in
createDegradedInferenceStream(the watchdog). - Calls
reconcileOllamaResidencyto decide whether to pin the model withkeep_alive: -1. - Builds the request via
buildRequest, which translates messages to Ollama format and applies thinking and sampling options. - Streams via
streamOllamaChatfrom./ollama-http.js. - Emits
text_start/text_delta/text_endevents for text,thinking_start/thinking_delta/thinking_endfor reasoning, andtoolcall_start/toolcall_delta/toolcall_endfor tool calls. - Tracks ownership: when
pinis true and the first response arrives, the model is recorded inownedModelsByTargetandmarkClioLoadedis called. AreportClioModelLoadforwards the load to the orchestrator in a worker.
Message translation: buildMessages calls translateMessage for each message, which handles user (text + images), assistant (text + tool calls + thinking), and toolResult (text with tool_name) roles.
Thinking: ollamaThinkValue maps the applied thinking mechanism to Ollama's think field: "always-on" returns undefined (Ollama owns it), "none" returns the thinking-active boolean, "effort-levels" returns the effort level or true.
Ownership and release on exit: releaseClioLoadedOllamaModels is registered via registerExitRelease. It reads the resident list first, then unloads only models that are both in ownedModelsByTarget and isClioLoaded. A scope narrows the release (used by one-shot probes). The release is bounded by EXIT_RELEASE_MS (2000ms) via an AbortController.
reconcileResidency(adapter) is the single decision point for load and evict. It:
- Checks the TTL cache (60s) — a clean reconcile within the window returns the cached decision.
- For
"router"strategy, acquires the cross-process lock before reading capacity. - Calls
adapter.listResident()to get the resident set. - Calls
adapter.capacity()(router strategy) andadapter.keepModelTags()to gather facts. - Classifies residents:
loadedByClio(from theclioLoadedregistry),protection("tag"forpinned:true/role:scout,"config"for operator-config references,"worker"for adopted worker loads), androle(from config). - Calls
decideResidency(facts)(pure) to produce aResidencyPlan:- Stress notices: context length exceeding model max, CPU-split residents, over-capacity.
-
Observe-only:
facts.managed === false(user-managed lifecycle). - Already resident: co-residents stay; reports once per TTL.
-
JIT strategy (LM Studio): attempts co-residency first;
fallbackEvictcarries ranked candidates for retry after capacity rejection. - Scheduler strategy (Ollama): releases only Clio's own unprotected stragglers.
-
Router strategy (llama.cpp): computes slots needed; declines if unknown capacity with existing residents; evicts only Clio-owned candidates (unprotected first, then config-protected only if
keepTagProtectedis false).
- If the plan would evict, calls
adapter.assertLoadable()first — aResidencyPreconditionErrordeclines the reconcile before any mutation. - Emits notices, then performs mutations under the lock: unload each evicted model, then load the keep model.
- On load failure,
restoreEvictedModelsputs back the evicted models viaadapter.reloadEvicted. - Marks the keep model as Clio-loaded and caches the decision.
Protection tiers: "tag" (server operator pins), "config" (operator's Clio configuration references), "worker" (adopted worker loads). No Clio profile silently evicts another profile's model.
Notice system: declareRuntimeNoticeProducer registers a producer with specific kinds; deliverNotice emits to the active sink (shared bus or stderr). The reconciler declares kinds "will-not-fit", "about-to-evict", "swap", "co-resident", "stress".
ensureLlamaCppResidency(input) builds a ResidencyAdapter with strategy: "router":
-
listResident: fetches/v1/modelsand filters to resident states (loaded,loading,sleeping). -
capacity: readsmax_instancesfrom/props. -
assertLoadable: checks the snapshot for the keep model; throwsResidencyPreconditionErrorif absent. -
load: POSTs to/models/load, then polls/modelsuntil the model isloadedorsleeping(sleeping is a resident state — a router with--sleep-idle-secondsparks idle models while keeping the slot). -
unload: POSTs to/models/unload. -
reloadEvicted: forced reload with its own timeout (recovery must survive request cancellation). - After loading,
restoreDisplacedPinnedreloads any tag-pinned residents the load displaced.
listLlamaCppResidentModels is exported for the degraded-inference watchdog.
ensureLmStudioResidency(model, options) handles LM Studio's just-in-time loading:
- For a managed target with an explicit load profile, acquires the residency lock and calls
ensureLmStudioResidencyUnlocked. - Reads the catalog via
listLmStudioModels. - Resolves the instance via
resolveLmStudioInstance. - If the instance drifted from the load profile (
lmStudioLoadDrift), unloads it and reloads — unless another Clio process holds a lease. - Releases Clio's earlier loads on this server (from
clioOwnership) before loading. - Calls
reconcileResidencywithstrategy: "jit"andttlMs: 0(fresh catalog each turn). - On capacity rejection, tries
plan.fallbackEvictcandidates and restores them on retry failure. - Returns a
LmStudioResidencywithwireModelIdandrelease(the lease).
ensureGatewayLmStudioResidency handles LiteLLM gateway routes: reads the deployment from /v1/model/info, builds a control model, and calls ensureLmStudioResidency on the LM Studio server. The request itself still goes through the gateway.
Cross-process ownership via a state file per LM Studio server:
-
recordClioLoad: adds a load record to the state file. -
leaseClioModel: adds a lease (pid, birth token, host, timestamp) and returns a release function. The lease is checked withleaseHeld(pid alive and birth token matches, or host mismatch or synthetic lease within 6h age). -
clioOwnership: reads loads and live leases. - All operations are best-effort: an unreadable file reads as empty, and a lock failure runs unlocked.
createDegradedInferenceStream(options) wraps an event stream to watch for slow token generation:
- Starts a watchdog on the
startevent (model load is not judged as slow generation). - Counts tokens from
text_delta,thinking_delta, andtoolcall_deltaevents. - After
DEGRADED_GRACE_MS(30s), checks the rate every 5s. If belowDEGRADED_FLOOR_TOKENS_PER_SECOND(2), emits a"degraded"notice once with the resident model summary. -
runningDegradedInferenceWatchdogs()returns the count of active timers (every finished turn returns it to zero).
The watchdog never cancels a turn; it observes and reports.
Defines EngineApiProvider<TApi, TOptions> with api, stream, and streamSimple methods.
Defines ResidentModelInfo (modelId, aliasIds, sizeVramBytes, sizeBytes, tags) and residentMatchesKeep (matches on modelId or any aliasIds).
pickSamplingProfile selects the quirks-based profile for thinking or instruct mode. samplingParamsFromProfile maps the profile to wire parameters (top_p, top_k, min_p, repeat_penalty/repetition_penalty, presence_penalty, frequency_penalty). Run-scoped overrides from core/run-overrides.ts merge on top.
remainingContextMaxTokens computes the output token budget: min(requested, model.maxTokens, contextWindow - inputTokens - CONTEXT_BUDGET_SAFETY_TOKENS). The requested value comes from (in order): tool-turn limit, globalDefaultMaxOutputTokens (set at session start), model.maxTokens, or DEFAULT_MAX_OUTPUT_TOKENS.
sequenceDiagram
participant ChatLoop as Chat Loop
participant Registry as api-registry
participant Provider as openAICompletionsApiProvider
participant Residency as reconcileResidency
participant LMStudio as ensureLmStudioResidency
participant Ownership as lmstudio-ownership
participant Pi as Pi SDK
ChatLoop->>Registry: engineStream(model, context, options)
Registry->>Provider: stream(model, context, options)
Provider->>Residency: reconcileResidency(adapter)
Residency->>Residency: decideResidency(facts) [pure]
Residency->>Residency: unload evicted, load keep model
Residency-->>Provider: ResidencyPlan
Provider->>LMStudio: ensureLmStudioResidency(model, options)
LMStudio->>Ownership: clioOwnership(serverKey)
Ownership-->>LMStudio: { loads, leased }
LMStudio->>LMStudio: unload drifted / Clio-owned instances
LMStudio->>LMStudio: reconcileResidency (jit)
LMStudio->>Ownership: recordClioLoad(serverKey, instanceId, modelKey)
LMStudio->>Ownership: leaseClioModel(serverKey, modelKey)
Ownership-->>LMStudio: release function
LMStudio-->>Provider: { wireModelId, release }
Provider->>Pi: openAICompletionsApi().stream(model, context, options)
Pi-->>Provider: AssistantMessageEventStream
Provider-->>ChatLoop: wrapped stream
ChatLoop->>ChatLoop: ... stream ends ...
Provider->>LMStudio: release() [lease dropped]
Provider->>Residency: markClioLoaded / forgetReleasedModel
For the Ollama path, runStream calls reconcileOllamaResidency (which builds a scheduler-strategy adapter and calls reconcileResidency), then streamOllamaChat with keep_alive: -1 when pinning.
Boundary: Pi SDK imports. Only src/engine/** imports @earendil-works/pi-*. The API providers receive Model, Context, and StreamOptions shapes from the Pi SDK.
Boundary: Residency protection. The reconciler's decideResidency enforces:
- Residents tagged
pinned:trueorrole:scoutare never evicted. - Residents referenced by the operator's Clio configuration (via
setProtectedModelsProvider) carry their role and are never evicted while an unprotected candidate exists. - Worker-loaded models (adopted via
adoptWorkerLoadedModel) carry"worker"protection. - The
keepTagProtectedflag prevents a pinned keep model from evicting a config-protected resident (one-way swap that the configured role could never undo).
Boundary: Cross-process locking. Mutations are serialized through withResidencyLock (state-dir lock file) for the "router" strategy. LM Studio uses withResidencyLock for load-profile enforcement.
Boundary: Lease protection. A model with a live lease (another Clio process streaming on it) is never unloaded. leaseHeld checks pid liveness and birth token.
Lifecycle: Release on exit. Each runtime registers via registerExitRelease. releaseClioLoadedModelsOnExit runs all releasers in parallel, each bounded by EXIT_RELEASE_MS (2000ms). Failures are swallowed (best-effort).
Lifecycle: Scoped release. releaseModelsLoadedDuring snapshots the clioLoaded registry before the task, then releases only models that became Clio-loaded during the task. Used by one-shot probes.
New API provider: Add a module in src/engine/apis/ that exports an EngineApiProvider object, import it in src/engine/apis/index.ts, and call registerEngineApiProvider(provider, "clio") in registerClioApiProviders.
New residency strategy: The reconciler's ResidencyStrategy type ("router", "jit", "scheduler") is the extension point. A new strategy would be added to decideResidency's branching, and a new adapter would implement the ResidencyAdapter interface.
New runtime adapter: Implement the ResidencyAdapter interface (targetKey, listResident, unload, load, capacity, assertLoadable, reloadEvicted, withLock) and call reconcileResidency. The llama.cpp router adapter in llamacpp-residency.ts and the LM Studio adapter in lmstudio.ts are reference implementations.
New notice kind: Add the kind to the RuntimeNoticeKind union and declare a producer via declareRuntimeNoticeProducer. A contract test checks that every kind is registered by some producer.
Demonstrates LM Studio load profile enforcement:
-
Settings validation:
validateSettingsacceptsparallel,speculativeDraftMaxTokens, and per-model overrides; rejectsparallel: 0and unknownmtpkeys. -
Per-model override:
effectiveLmStudioLoadmerges the target profile with the model-specific override on top. -
Drift detection:
lmStudioLoadDriftnames only the fields the instance reports and the profile sets. - Direct LM Studio target loads with profile: A turn triggers one load with the profile wire keys; a second turn reuses the matching instance (no reload).
- Reload on drift: An instance loaded with GUI defaults is unloaded and reloaded with the profile.
-
User-managed target is never touched:
lifecycle: "user-managed"produces no loads or unloads. -
Gateway route loads LM Studio model: A LiteLLM gateway route with
model_info.runtime: "lm-studio"triggers a load on the LM Studio server; the request itself still goes through the gateway; the gateway key never reaches LM Studio. - Cross-process ownership: A load releases a model an earlier Clio process loaded on the same server; back-to-back turns on one model stay warm; a model another client loaded is never released; a model a live Clio process is streaming on stays and is released once that process is gone; a drifted instance is not reloaded under a live Clio request; a stream holds its lease while the request runs and drops it after.
Demonstrates Ollama residency and ownership:
- Preserves operator models, unloads Clio-pinned on model switch: An operator-loaded model is never evicted; Clio's own pinned model is unloaded on a model switch.
- Does not claim a failed load: A failed chat does not record ownership; a subsequent operator load is not treated as Clio's.
- Does not transfer ownership between servers: Ownership is per-server.
- Forgets successful unloads, retains after failed unloads: A failed unload retains ownership; a successful unload forgets it.
- Releases on exit only Clio-loaded models: After a model switch, exit release unloads only Clio's current model; an operator model is untouched.
- Bounds release on exit: A hung server does not block exit past the timeout.
-
Reports worker loads once:
reportClioModelLoadreports a cold model once; a warm model is never reported. - Releases worker-reported models at exit: An adopted worker load is released at exit; operator models are untouched.
- Refuses worker reports for wrong target/model/node: Security checks on the adoption path.
- Never evicts a worker-reported model mid-session: A chat turn does not evict a worker-loaded model.
-
Scoped release for probes:
releaseModelsLoadedDuringwithreleaseScopeForreleases only the probed model.
-
The
openAICompletionsApiProviderpipeline is order-sensitive. The wrapper chain instreamCompletionsis built outermost-to-innermost. Moving a wrapper changes which events it sees (e.g.,withResponseModelIdCapturemust be outsidewithLocalResidencyto observe the raw fetch;stripSentinelsFromStreammust be insidestripNeverReasoningFromStreamto see the unstripped stream). -
The residency TTL cache (60s) can mask changes. A clean reconcile within the window returns the cached decision without re-listing. The LM Studio JIT adapter passes
ttlMs: 0to force a fresh catalog each turn. A new adapter must decide whether it needs the TTL fast path. -
sleepingis a resident state for the llama.cpp router. A router with--sleep-idle-secondsparks idle models while keeping the slot and weights. Reading it as unloaded costs twice: capacity math sees a free slot that does not exist, and the load path re-requests a model the router is already running. - The LM Studio ownership file is an optimization, not a precondition. An unreadable file reads as empty, and a lock failure runs unlocked. The system degrades to co-residency rather than failing a turn.
-
The degraded-inference watchdog uses a monotonic clock. A stepped wall clock would fake a collapse on a forward step and hide one on a backward step. The default is
performance.now. -
Ollama's
keep_alive: -1pins the model forever. The only release path isunloadOllamaModel(which fireskeep_alive: 0) orreleaseClioLoadedOllamaModelsat exit. A model Clio pinned but never released holds its weights indefinitely on a shared server. -
The
exactOptionalPropertyTypesconvention is enforced. Optional fields are passed with...(x !== undefined ? { x } : {}), neverx: undefined. TheLlamaCppResidencyInputandWatchDegradedInferenceOptionsinterfaces rely on this. -
The
piOpenAICompletionsconst is a module-level singleton (openAICompletionsApi()). It is not per-request; the stream wrappers are what make each call distinct.
Source and generation metadata
title: "Engine apis"
summary: "The engine's provider API layer: two registered providers (OpenAI-completions and Ollama-native), a shared capacity-aware residency reconciler for local runtimes, a degraded-inference watchdog, and per-runtime adapters for LM Studio, llama.cpp router, and Ollama."
sources:
- "src/engine/apis/index.ts"
- "src/engine/apis/openai-completions.ts"
- "src/engine/apis/ollama-native.ts"
- "src/engine/apis/residency.ts"
- "src/engine/apis/lmstudio.ts"
- "src/engine/apis/llamacpp-residency.ts"
- "src/engine/apis/degraded-inference.ts"
- "src/engine/apis/lmstudio-ownership.ts"
tests:
- "tests/contracts/lmstudio-load-profile.test.ts"
- "tests/extended/ollama-residency.test.ts"
invariants:
- "Only `src/engine/**` imports `@earendil-works/pi-*`; the API providers receive `Model` and `Context` shapes from the Pi SDK."
- "The residency reconciler evicts only Clio-attributed loads; protected residents (`pinned:true`, `role:scout` tags or operator config references) are never evicted while an unprotected candidate exists."
- "A model loaded by one Clio process is never released by another while the first holds a lease, preventing a process from pulling a model out from under another's request."
- "The degraded-inference watchdog observes and reports but never cancels a turn; a slow model is still a working model."Clio Coder · Repository · Website · Documentation
Wiki v0.1 · Developing implementation reference · Source snapshot: 657dce13d. Authored architecture documents define the product contracts.
- Clio Coder GUI Client
- apps / clio-coder-gui
- Apps clio coder gui server
- Apps clio coder gui tests
- apps
- Architecture
- Command-line surfaces
- Core
- Domains agents
- Config Domain
- Context Domain
- Dispatch domain
- Domains evidence
- Domains extensions
- Domains gateway
- domains
- Domains interop
- Domains lifecycle
- Domains memory
- Middleware Domain
- Domains mux
- Domains observability
- Domains plugins
- Prompt Compiler
- Domains providers
- Domains quota
- Domains resources
- Domains safety
- Domains scheduling
- Domains session
- Vendored Tool Registry and Resolution
- Engine
- Engine acp
- Engine apis
- engine
- Entry point
- Interactive
- interactive
- Interactive overlays
- Interactive renderers
- clio-coder wiki
- Scripts
- Contract tests
- Tests extended
- tests
- Tools
- Tools data
- tools
- Tools verify
- Worker runtime