·
4 commits
to main
since this release
v0.1.0-alpha.216 - Haiku LLM fallback for intent classification
Completes the two-stage classifier the architecture has documented
since alpha.182. The Stage 2 Haiku LLM fallback was stubbed-but-
unbuilt all this time; ships today.
Architecture (now actually two-stage):
Stage 1: heuristic (classifyIntent)
- Keyword + structural rules
- ~0ms, deterministic
- Covers high-confidence cases (~90% of sends)
Stage 2: Haiku LLM (classifyIntentLlm)
- POST /v1/intent/classify on the edge worker
- ~200-400ms, ~$0.0003/call
- Fires when heuristic returns confidence < 0.7
- Authoritative — overrides heuristic
User-facing impact: short ambiguous follow-ups ("continue", "retry",
"yes", "go", "actually make it red") now route correctly. The
heuristic's keyword lists missed every "continue"-flavored word
because none of them are agent verbs OR chat phrases — they fell
through to baseline-bias, which for an agentic thread is "agent"
(good) but for a fresh thread is "chat" (bad). Haiku reads the
thread context and resolves correctly either way.
Implementation:
- New lib function classifyIntentLlm() POSTs to the edge endpoint
with prompt + threadIsAgentic + last 3 message snippets.
- ChatSurface.send() runs heuristic first; if confidence < 0.7
AND we have an auth key, awaits the LLM result (with 4s
timeout). LLM result overrides heuristic.
- Failure modes (network, 5xx, parse): classifyIntentLlm returns
null; ChatSurface falls back to heuristic. No turn-blocking.
Cost projection: assuming 10% of sends fire stage 2, the per-turn
effective classifier cost is ~$0.00003. Trivial compared to the
~$0.05+ a mis-classified agent turn costs in container spinup +
wasted work.
Note: the legacy heuristic STAYS as the fast path. We don't add
~300ms latency to the 90% of sends that are unambiguous. Only the
hard cases pay.
Pairs with server-side commit landing in qlaud_router today that
adds the /v1/intent/classify endpoint.