-
Notifications
You must be signed in to change notification settings - Fork 0
Writing a Transport
A transport is an async generator that yields string chunks. That is the entire contract.
async function* myTransport(prompt, { config }) {
yield 'Hello';
yield ', world';
}Three transports ship: echo (repeats what you said, for building UI before a model
exists), local (a stub for a browser-resident model), and remote (a stub for a
server). The stubs are stubs on purpose — see
Why nothing here calls a model.
Because the interface is uniform, no component knows where its text came from. That is what makes offline-first a configuration rather than a rewrite.
async function* hosted(prompt, { config }) {
const response = await fetch(config.endpoint, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({ model: config.model, prompt, stream: true })
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
for (;;) {
const { done, value } = await reader.read();
if (done) break;
yield decoder.decode(value, { stream: true });
}
}Never put an API key in page JavaScript.
config.endpointshould point at your own server, which holds the credential and talks to the provider. Anything in the page is readable by anyone who opens devtools.
Most providers stream SSE rather than raw text, so unwrap it:
async function* sse(prompt, { config }) {
const response = await fetch(config.endpoint, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({ prompt, stream: true })
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = '';
for (;;) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
// Lines arrive split across chunks; keep the incomplete tail.
const lines = buffer.split('\n');
buffer = lines.pop();
for (const line of lines) {
if (!line.startsWith('data:')) continue;
const data = line.slice(5).trim();
if (data === '[DONE]') return;
const text = JSON.parse(data).delta?.text;
if (text) yield text;
}
}
}This is the case the package exists to make possible. With WebLLM, the shape is the same:
import * as webllm from '@mlc-ai/web-llm';
let engine;
async function* local(prompt, { config }) {
engine ??= await webllm.CreateMLCEngine(config.model || 'Llama-3.2-1B-Instruct-q4f32_1-MLC', {
initProgressCallback: ({ text }) => {
// The first load downloads a model measured in hundreds of megabytes.
// Saying nothing for two minutes reads as a broken page.
services.conversation.update(pendingId, { text });
}
});
const stream = await engine.chat.completions.create({
messages: [{ role: 'user', content: prompt }],
stream: true
});
for await (const part of stream) {
const text = part.choices[0]?.delta?.content;
if (text) yield text;
}
}Two things to handle that a hosted model does not have:
First load is slow and large. Report progress into the pending record, or into a
system turn. Cache it — WebLLM uses the Cache API, so a second visit is fast.
Not every device can run it. Check WebGPU and fall back rather than failing:
const canRunLocally = 'gpu' in navigator && await navigator.gpu?.requestAdapter();
services.config.transport = canRunLocally ? 'local' : 'remote';A model in the browser has no per-token cost. The conversation also never leaves the device, which is a privacy property you cannot buy from a hosted provider at any price.
| Situation | With a hosted model | With a local one |
|---|---|---|
| Long-tail support | Cost per conversation can exceed the deflection's value | Zero marginal cost |
| Kiosk / in-store | Unmetered usage by passers-by | Unmetered costs nothing |
| Classroom | Per-seat spend nobody budgeted | One download, any number of sessions |
| Internal tooling | Hard to justify a line item | No line item |
| Field application | Needs connectivity | Works offline |
| Sensitive intake | Conversation leaves the device | It does not |
The honest caveat: a 1–3B model in a browser is not a frontier model. It is good at classification, extraction, rephrasing, triage and scripted guidance — which is most of what a product chat surface actually does. Route the hard questions out; keep the volume in.
The shipped TRANSPORTS map is module-private, so compose at the daemon instead:
import { Daemon } from '@machfivetechchicago/machvive-chat-syncopation-ai/services';
class MyDaemon extends Daemon {
async send(text) {
if (this.busy) return null;
// …or override the transport lookup for your own names
return super.send(text);
}
}In practice the simplest route is to drive the conversation yourself and skip the daemon entirely — it is a convenience, not a requirement:
const { conversation } = services;
async function ask(text) {
conversation.add({ role: 'user', text });
const reply = conversation.add({ role: 'assistant', text: '', status: 'pending' });
try {
for await (const chunk of myTransport(text, { config: services.config })) {
conversation.append(reply.id, chunk);
}
conversation.update(reply.id, { status: 'complete' });
} catch (error) {
// Make the failure visible. A silent drop leaves the user staring at a
// prompt that appears to have done nothing.
conversation.update(reply.id, { status: 'error', meta: { 'mcs:error': error.message } });
}
}Every component renders this correctly, because this is exactly what the daemon does.
import { META } from '@machfivetechchicago/machvive-chat-syncopation-ai';
conversation.update(reply.id, {
status: 'complete',
meta: {
[META.MODEL]: 'claude-opus-5',
[META.TOKENS]: usage.output_tokens,
[META.LATENCY]: Date.now() - startedAt,
[META.SOURCE]: 'remote'
}
});META.SOURCE is worth setting honestly — 'local' means no token cost and no data
egress, and that is a distinction a dashboard will want later.
Three reasons, in order of how much they matter:
- Stack agnosticism is the point. The moment this package has an opinion about a provider, it has an opinion about your architecture, your billing and your data residency. A chat bubble should not.
- Credentials belong on a server. A transport bundled here would invite someone to pass an API key to a web component.
- The interesting transport is the local one, and that is a large dependency with its own release cadence and device requirements. Making it a seam rather than a dependency keeps this package installable in seconds.
The suite enforces this: a test fails if any module gains a fetch,
XMLHttpRequest, WebSocket or an external URL. A claim of absence needs a test
asserting that absence, or it goes false one feature later with nothing failing.
Next: Theming.