Skip to content

Writing a Transport

neodigm edited this page Oct 4, 2026 · 1 revision

Writing a Transport

A transport is an async generator that yields string chunks. That is the entire contract.

async function* myTransport(prompt, { config }) {
  yield 'Hello';
  yield ', world';
}

Three transports ship: echo (repeats what you said, for building UI before a model exists), local (a stub for a browser-resident model), and remote (a stub for a server). The stubs are stubs on purpose — see Why nothing here calls a model.

Because the interface is uniform, no component knows where its text came from. That is what makes offline-first a configuration rather than a rewrite.

A hosted model

async function* hosted(prompt, { config }) {
  const response = await fetch(config.endpoint, {
    method: 'POST',
    headers: { 'content-type': 'application/json' },
    body: JSON.stringify({ model: config.model, prompt, stream: true })
  });

  const reader = response.body.getReader();
  const decoder = new TextDecoder();
  for (;;) {
    const { done, value } = await reader.read();
    if (done) break;
    yield decoder.decode(value, { stream: true });
  }
}

Never put an API key in page JavaScript. config.endpoint should point at your own server, which holds the credential and talks to the provider. Anything in the page is readable by anyone who opens devtools.

Server-sent events

Most providers stream SSE rather than raw text, so unwrap it:

async function* sse(prompt, { config }) {
  const response = await fetch(config.endpoint, {
    method: 'POST',
    headers: { 'content-type': 'application/json' },
    body: JSON.stringify({ prompt, stream: true })
  });

  const reader = response.body.getReader();
  const decoder = new TextDecoder();
  let buffer = '';

  for (;;) {
    const { done, value } = await reader.read();
    if (done) break;
    buffer += decoder.decode(value, { stream: true });

    // Lines arrive split across chunks; keep the incomplete tail.
    const lines = buffer.split('\n');
    buffer = lines.pop();

    for (const line of lines) {
      if (!line.startsWith('data:')) continue;
      const data = line.slice(5).trim();
      if (data === '[DONE]') return;
      const text = JSON.parse(data).delta?.text;
      if (text) yield text;
    }
  }
}

A model in the user agent

This is the case the package exists to make possible. With WebLLM, the shape is the same:

import * as webllm from '@mlc-ai/web-llm';

let engine;

async function* local(prompt, { config }) {
  engine ??= await webllm.CreateMLCEngine(config.model || 'Llama-3.2-1B-Instruct-q4f32_1-MLC', {
    initProgressCallback: ({ text }) => {
      // The first load downloads a model measured in hundreds of megabytes.
      // Saying nothing for two minutes reads as a broken page.
      services.conversation.update(pendingId, { text });
    }
  });

  const stream = await engine.chat.completions.create({
    messages: [{ role: 'user', content: prompt }],
    stream: true
  });

  for await (const part of stream) {
    const text = part.choices[0]?.delta?.content;
    if (text) yield text;
  }
}

Two things to handle that a hosted model does not have:

First load is slow and large. Report progress into the pending record, or into a system turn. Cache it — WebLLM uses the Cache API, so a second visit is fast.

Not every device can run it. Check WebGPU and fall back rather than failing:

const canRunLocally = 'gpu' in navigator && await navigator.gpu?.requestAdapter();
services.config.transport = canRunLocally ? 'local' : 'remote';

Why this matters commercially

A model in the browser has no per-token cost. The conversation also never leaves the device, which is a privacy property you cannot buy from a hosted provider at any price.

Situation With a hosted model With a local one
Long-tail support Cost per conversation can exceed the deflection's value Zero marginal cost
Kiosk / in-store Unmetered usage by passers-by Unmetered costs nothing
Classroom Per-seat spend nobody budgeted One download, any number of sessions
Internal tooling Hard to justify a line item No line item
Field application Needs connectivity Works offline
Sensitive intake Conversation leaves the device It does not

The honest caveat: a 1–3B model in a browser is not a frontier model. It is good at classification, extraction, rephrasing, triage and scripted guidance — which is most of what a product chat surface actually does. Route the hard questions out; keep the volume in.

Registering your transport

The shipped TRANSPORTS map is module-private, so compose at the daemon instead:

import { Daemon } from '@machfivetechchicago/machvive-chat-syncopation-ai/services';

class MyDaemon extends Daemon {
  async send(text) {
    if (this.busy) return null;
    // …or override the transport lookup for your own names
    return super.send(text);
  }
}

In practice the simplest route is to drive the conversation yourself and skip the daemon entirely — it is a convenience, not a requirement:

const { conversation } = services;

async function ask(text) {
  conversation.add({ role: 'user', text });
  const reply = conversation.add({ role: 'assistant', text: '', status: 'pending' });

  try {
    for await (const chunk of myTransport(text, { config: services.config })) {
      conversation.append(reply.id, chunk);
    }
    conversation.update(reply.id, { status: 'complete' });
  } catch (error) {
    // Make the failure visible. A silent drop leaves the user staring at a
    // prompt that appears to have done nothing.
    conversation.update(reply.id, { status: 'error', meta: { 'mcs:error': error.message } });
  }
}

Every component renders this correctly, because this is exactly what the daemon does.

Recording what a turn cost

import { META } from '@machfivetechchicago/machvive-chat-syncopation-ai';

conversation.update(reply.id, {
  status: 'complete',
  meta: {
    [META.MODEL]: 'claude-opus-5',
    [META.TOKENS]: usage.output_tokens,
    [META.LATENCY]: Date.now() - startedAt,
    [META.SOURCE]: 'remote'
  }
});

META.SOURCE is worth setting honestly — 'local' means no token cost and no data egress, and that is a distinction a dashboard will want later.

Why nothing here calls a model

Three reasons, in order of how much they matter:

  1. Stack agnosticism is the point. The moment this package has an opinion about a provider, it has an opinion about your architecture, your billing and your data residency. A chat bubble should not.
  2. Credentials belong on a server. A transport bundled here would invite someone to pass an API key to a web component.
  3. The interesting transport is the local one, and that is a large dependency with its own release cadence and device requirements. Making it a seam rather than a dependency keeps this package installable in seconds.

The suite enforces this: a test fails if any module gains a fetch, XMLHttpRequest, WebSocket or an external URL. A claim of absence needs a test asserting that absence, or it goes false one feature later with nothing failing.

Next: Theming.

Clone this wiki locally