Skip to content

Providers and Agents

Hermes Agent edited this page Oct 1, 2026 · 7 revisions

Providers and Agents

English | 中文 | 日本語 | 한국어 | Español | Português | Русский

Four wire protocols (apiStyle)

lib/llm-client.js is the protocol layer. All four streams share one openSseStream() skeleton (fetch / error classification / budget negotiation / abort / SSE framing / finalization); each protocol keeps only buildRequest + pure event parsers. SSE data: payload extraction is single-sourced (sseDataPayload, multi-data-line join per spec):

apiStyle Endpoint Notes
chat /v1/chat/completions also parses DeepSeek-style reasoning_content
responses /v1/responses native input_* multimodal spelling
anthropic /v1/messages explicit cache_control on two breakpoints (system + last message) — Anthropic has no implicit prefix caching
runs Hermes /v1/runs agent protocol: approvals / clarifications / tools / thinking events

thinking: 'inline' | 'omit' (default omit) controls whether reasoning text rides ONE <thinking> block inline in the delta stream; only the main chat and detail thread use inline. The runs reasoning.available field is an answer replay, not thinking — the echo guard must drop it (ADR-0004).

The turn-request ladder

The per-provider request shape is rebuilt four times in the main chat (initial / overflow rebuild / continuation / timestamp rewrite), unified as createTurnRequest (prepare / rebuildFrom / continueWith / rewriteWith) — chat-handler keeps the WHEN, the ladder owns the HOW. The detail thread deliberately does not use it (isolated-session semantics, ADR-0007).

Output budget & truncation

  • Default max_tokens 32768; servers that REJECT an over-cap budget with a 400 get one renegotiateOutputCap retry with the cap parsed from the error text.
  • finish_reason === 'length' → ONE silent continuation pass (anti-repetition instruction, deltas swallowed); still truncated → DONE carries outputTruncated, the panel toasts + shows a one-click "continue" button. There is deliberately NO manual max_tokens field — "a knob that requires per-model knowledge is a design bug".

The four agents

Agent Channel Sessions
Hermes /v1/runs server-side (sessionId in storage); current-turn parts use canonical text/image_url, conversation_history is strings-only (strict-layer 422, verified 2026-09-24)
OpenCode opencode serve HTTP server-side; random port by default → recommend fixed --port
OpenSquilla local gateway ws://…/ws one gateway session per conversation; >60K-char attaches upload as a page-context.md document; origin allowlist — see Security Model
Agent Bridge @xiaohuzai/agent-bridge local daemon adapts codex/claude/pi/gemini behind one HTTP protocol; subscription logins work as model sources

The shared layer lib/agent-turn.js: an agent turn sends ONLY the user's current turn + the trailing page-context run (the transcript lives server-side; history is never resent). Images go through pickTurnImages (≤8 images / ≤3MiB URL budget; attach-time and send-time gates mirror each other). Switching to an agent provider mid-conversation offers 「带上当前对话继续」= a one-shot backfill (plain-text transcript, 200K-char tail cap).

Cross-entry handoff: agent-side session naming & ID (2026-10-01)

For agent providers the transcript already lives server-side (agent turns send only the current turn) — "let the agent's own UI take over" needs discoverability, not data movement. Two pieces:

  • Automatic naming: after a successful Hermes turn, one-shot PATCH /api/sessions/{id} (upstream-original API; the server sanitizes titles and rejects exact conflicts) sets the title to browsa: + the first line of the current turn's user text (48-char cap). Stamp-once semantics (hermesSessionTitled_<provider>, same lifecycle as the session id): success AND 4xx stamp (a title conflict must not become a per-turn retry loop); only transport failures (status 0) stay unstamped for the next successful turn. The PATCH times out at 10s and its await races a 2.5s cap so DONE is never hung. Channels today: Hermes (PATCH /api/sessions/{id}) and bridge/codex (the daemon's POST /threads/{id}/title → app-server thread/name/set, live-verified on codex 0.149.1 — the name persists into codex's state DB and codex resume resolves by name or id; Picker-visibility verdicts for the four bridge agents (source-verified 2026-10-01): codex ✓ (picker predicate has_user_event = 1 AND title <> '' — real turns satisfy it; global state DB, no cwd scoping; named via the daemon's POST /threads/{id}/title → app-server thread/name/set). claude ✓ since the bridge's transcriptFix: 'claude' rewrites the self-stamped entrypoint: sdk-* to cli on disk after every settled turn (agent-bridge #56, Mac-confirmed) — no client rename channel (claude-agent-acp auto-titles), resume works from the picker or by ID. pi ✓ (ACP sessions land in pi's own cwd-scoped store the picker reads; scope toggle + named filter; pi's native set_session_name RPC writes a session_info name entry — pi-acp only exposes it as the in-turn /name command, not yet wired to the bridge). gemini ✓ (the list excludes only kind:"subagent"; bridge sessions are kind:"main"; store ~/.gemini/tmp/<cwd>/chats/, cwd-scoped; no naming channel — the title comes from the first user message). opencode/squilla unverified. The detail thread's dedicated sessions are deliberately NOT titled.)
  • Session-ID display: storage.getAgentSessionInfo single-owns the per-kind session-key shapes (bridge keyed per endpoint via activeModel || baseUrl); the sessions drawer renders an "Agent session" line (short id + copy) above the list, hidden for LLM providers / when no session exists.

Approval / clarify relay

Agent tool approvals and clarification prompts relay through lib/handlers/approval-relay.js (main chat keyed by tabId, detail thread by subId). The pending-entry shape IS the dispatch interface. On the UI side, turn-chrome.js is the turn chrome shared by both surfaces (wait indicator, tool progress, approval/clarify cards, usage chip).

Provider selection UI (settled rules)

  • The dropdown lists only CONFIGURED providers, reachable-first (stable sort); zero configured → one disabled placeholder.
  • If the stored activeProvider isn't configured, the first configured one is auto-selected AND persisted (state repair, not preference); the first reachable ping also auto-switches (once).
  • Multi-model providers: comma-separated model IDs; one dropdown entry per model; resolveChatModel honors activeModel only while it still belongs to that provider.
  • Both groups (LLM / agent) render as tabbed cards (agent group collapsed by default; tab order [bridge, opencode, hermes] is test-pinned).

CAPABILITY_HINTS economics (ADR-0010 — do not re-propose shrinking)

Every CHAT turn's system prompt = user systemPrompt + reply-language line + CAPABILITY_HINTS + CHOICE_REQUEST_HINT; it must stay a byte-stable prefix (per-turn keyword gating breaks the KV prompt cache from position 0). Every remaining entry paid for itself with a real rendering bug — don't "tighten" a why-clause without reading its AGENTS.md history. The fix for a missed format is a better hint, not runtime detection.


Source of truth: AGENTS.md "Provider API styles", chat/subchat-handler, provider-resolver, per-agent sections; ADR-0004 / 0007 / 0010. Synced 2026-10-01 (incl. cross-entry handoff).

Clone this wiki locally