-
Notifications
You must be signed in to change notification settings - Fork 0
Providers and Agents
English | 中文 | 日本語 | 한국어 | Español | Português | Русский
lib/llm-client.js is the protocol layer. All four streams share one openSseStream() skeleton (fetch / error classification / budget negotiation / abort / SSE framing / finalization); each protocol keeps only buildRequest + pure event parsers. SSE data: payload extraction is single-sourced (sseDataPayload, multi-data-line join per spec):
| apiStyle | Endpoint | Notes |
|---|---|---|
chat |
/v1/chat/completions |
also parses DeepSeek-style reasoning_content
|
responses |
/v1/responses |
native input_* multimodal spelling |
anthropic |
/v1/messages |
explicit cache_control on two breakpoints (system + last message) — Anthropic has no implicit prefix caching |
runs |
Hermes /v1/runs
|
agent protocol: approvals / clarifications / tools / thinking events |
thinking: 'inline' | 'omit' (default omit) controls whether reasoning text rides ONE <thinking> block inline in the delta stream; only the main chat and detail thread use inline. The runs reasoning.available field is an answer replay, not thinking — the echo guard must drop it (ADR-0004).
The per-provider request shape is rebuilt four times in the main chat (initial / overflow rebuild / continuation / timestamp rewrite), unified as createTurnRequest (prepare / rebuildFrom / continueWith / rewriteWith) — chat-handler keeps the WHEN, the ladder owns the HOW. The detail thread deliberately does not use it (isolated-session semantics, ADR-0007).
- Default
max_tokens32768; servers that REJECT an over-cap budget with a 400 get onerenegotiateOutputCapretry with the cap parsed from the error text. -
finish_reason === 'length'→ ONE silent continuation pass (anti-repetition instruction, deltas swallowed); still truncated → DONE carriesoutputTruncated, the panel toasts + shows a one-click "continue" button. There is deliberately NO manual max_tokens field — "a knob that requires per-model knowledge is a design bug".
| Agent | Channel | Sessions |
|---|---|---|
| Hermes | /v1/runs |
server-side (sessionId in storage); current-turn parts use canonical text/image_url, conversation_history is strings-only (strict-layer 422, verified 2026-09-24) |
| OpenCode |
opencode serve HTTP |
server-side; random port by default → recommend fixed --port
|
| OpenSquilla | local gateway ws://…/ws
|
one gateway session per conversation; >60K-char attaches upload as a page-context.md document; origin allowlist — see Security Model |
| Agent Bridge |
@xiaohuzai/agent-bridge local daemon |
adapts codex/claude/pi/gemini behind one HTTP protocol; subscription logins work as model sources |
The shared layer lib/agent-turn.js: an agent turn sends ONLY the user's current turn + the trailing page-context run (the transcript lives server-side; history is never resent). Images go through pickTurnImages (≤8 images / ≤3MiB URL budget; attach-time and send-time gates mirror each other). Switching to an agent provider mid-conversation offers 「带上当前对话继续」= a one-shot backfill (plain-text transcript, 200K-char tail cap).
For agent providers the transcript already lives server-side (agent turns send only the current turn) — "let the agent's own UI take over" needs discoverability, not data movement. Two pieces:
-
Automatic naming: after a successful Hermes turn, one-shot
PATCH /api/sessions/{id}(upstream-original API; the server sanitizes titles and rejects exact conflicts) sets the title tobrowsa:+ the first line of the current turn's user text (48-char cap). Stamp-once semantics (hermesSessionTitled_<provider>, same lifecycle as the session id): success AND 4xx stamp (a title conflict must not become a per-turn retry loop); only transport failures (status 0) stay unstamped for the next successful turn. The PATCH times out at 10s and its await races a 2.5s cap so DONE is never hung. Channels today: Hermes (PATCH /api/sessions/{id}) and bridge/codex (the daemon'sPOST /threads/{id}/title→ app-serverthread/name/set, live-verified on codex 0.149.1 — the name persists into codex's state DB andcodex resumeresolves by name or id; Picker-visibility verdicts for the four bridge agents (source-verified 2026-10-01): codex ✓ (picker predicatehas_user_event = 1 AND title <> ''— real turns satisfy it; global state DB, no cwd scoping; named via the daemon'sPOST /threads/{id}/title→ app-serverthread/name/set). claude ✓ since the bridge'stranscriptFix: 'claude'rewrites the self-stampedentrypoint: sdk-*toclion disk after every settled turn (agent-bridge #56, Mac-confirmed) — no client rename channel (claude-agent-acp auto-titles), resume works from the picker or by ID. pi ✓ (ACP sessions land in pi's own cwd-scoped store the picker reads; scope toggle + named filter; pi's nativeset_session_nameRPC writes asession_infoname entry — pi-acp only exposes it as the in-turn/namecommand, not yet wired to the bridge). gemini ✓ (the list excludes onlykind:"subagent"; bridge sessions arekind:"main"; store~/.gemini/tmp/<cwd>/chats/, cwd-scoped; no naming channel — the title comes from the first user message). opencode/squilla unverified. The detail thread's dedicated sessions are deliberately NOT titled.) -
Session-ID display:
storage.getAgentSessionInfosingle-owns the per-kind session-key shapes (bridge keyed per endpoint viaactiveModel || baseUrl); the sessions drawer renders an "Agent session" line (short id + copy) above the list, hidden for LLM providers / when no session exists.
Agent tool approvals and clarification prompts relay through lib/handlers/approval-relay.js (main chat keyed by tabId, detail thread by subId). The pending-entry shape IS the dispatch interface. On the UI side, turn-chrome.js is the turn chrome shared by both surfaces (wait indicator, tool progress, approval/clarify cards, usage chip).
- The dropdown lists only CONFIGURED providers, reachable-first (stable sort); zero configured → one disabled placeholder.
- If the stored activeProvider isn't configured, the first configured one is auto-selected AND persisted (state repair, not preference); the first reachable ping also auto-switches (once).
- Multi-model providers: comma-separated model IDs; one dropdown entry per model;
resolveChatModelhonors activeModel only while it still belongs to that provider. - Both groups (LLM / agent) render as tabbed cards (agent group collapsed by default; tab order [bridge, opencode, hermes] is test-pinned).
Every CHAT turn's system prompt = user systemPrompt + reply-language line + CAPABILITY_HINTS + CHOICE_REQUEST_HINT; it must stay a byte-stable prefix (per-turn keyword gating breaks the KV prompt cache from position 0). Every remaining entry paid for itself with a real rendering bug — don't "tighten" a why-clause without reading its AGENTS.md history. The fix for a missed format is a better hint, not runtime detection.
Source of truth: AGENTS.md "Provider API styles", chat/subchat-handler, provider-resolver, per-agent sections; ADR-0004 / 0007 / 0010. Synced 2026-10-01 (incl. cross-entry handoff).
English
- Home
- Architecture
- Rendering Pipeline
- Storage Model
- Providers and Agents
- ASR and Video Analysis
- Security Model
- Design Decisions
- Contributing
中文
相关 / Related
日本語
한국어
Español
- Inicio
- Arquitectura
- Pipeline de renderizado
- Modelo de almacenamiento
- Proveedores y agentes
- ASR y análisis de vídeo
- Modelo de seguridad
- Decisiones de diseño
- Contribuir
Português
- Início
- Arquitetura
- Pipeline de renderização
- Modelo de armazenamento
- Provedores e agentes
- ASR e análise de vídeo
- Modelo de segurança
- Decisões de design
- Contribuindo
Русский