Pooled v1.0: peer-to-peer inference in the browser. One open model, split across the devices you already have. Open pooled.run on a laptop, a desktop and a phone, and each device holds some of the model's layers and runs them on its own GPU, right in the browser tab. Chat with it, or switch to Code and let it build a small web app you can play with in the tab.
No install, no account. Pooled was called SwarmLLM until September 2026; old links redirect.
Demo (2:29, attached below): Qwen 3.6 35B MoE across a Mac and an iPhone. The phone opens the room and chats with the model, then Code mode builds a Tetris game, hits a bug, fixes it on its own, and restyles it on request.
What's in v1
Models
- Qwen 3.6 35B MoE (22.5 GB), Qwen 3.8 27B (17 GB) and Qwen3 1.7B (4 GB), split by layers across the devices in a room.
- Long context: the MoE defaults to 32K tokens (up to 64K), the 27B to 16K (up to 32K), the 1.7B to 8K.
Speed (one GB10, Chrome, tok/s; see the README's Performance section and docs/bench-log.md)
- 35B MoE: about 50 plain, 70-80 speculative on coding prompts. Sampling now runs on the GPU, so 16 bytes come back per token instead of the 1 MB logits vector.
- 27B: 11.1 plain, 20-27 speculative. Qwen3 1.7B: about 61 on a short chat, 33 at 4K context.
- Two devices, 20 ms emulated one-way latency: 35B MoE 12.9 plain / 20.6 speculative.
- Real devices: a Mac and an iPhone ran the 35B at 30-60 tok/s in the demo; the first two-machine room (GB10 + M5 Max over Tailscale) ran it at 22-32.
How a room works
- Each device downloads only its own layers. For every token, only the hidden state crosses the network: a few kilobytes per hop, never the weights.
- The model drafts a few tokens ahead and verifies them in one pass, and every device keeps the attention cache for its own layers, so follow-ups only send the new words through the room.
- Our own WGSL kernels on WebGPU, including the mixture-of-experts routing. Apple GPUs (Metal) are supported.
Code mode
- A coding harness in the browser, running on the room's model: it writes files, serves them on a virtual localhost, reads its own errors and fixes them. Previews scale to fit, on phones too.
- Anyone in the room can drive it, not only the host; the phone gets Agent / Preview / Files tabs.
- Small models are much steadier: tool calls are repaired and constrained while sampling. Qwen3 1.7B completes 7 of 8 real tasks (was 4 of 8).
The room
- A new landing page and room: join with a code, lend memory, pick a model that fits, and watch each device download its own layers.
- Every device has its own screen: Serving, its layers, tok/s and ms per hop.
- Device colours match on every screen, a host that reloads goes straight back into its room, and names can't inject markup.
Known limits
- Chrome or Edge on a laptop or desktop is the tested host; Safari on an iPhone joins and holds a few layers. Firefox and Linux Chromium need WebGPU switched on.
- "localhost" in Code mode is a label: nothing outside the tab can reach the preview.
- Prompt processing is still behind native llama.cpp on the same machine.
Everything is open source (MIT): the kernels, the protocol, the harness, the benchmarks and the roadmap. Contributions welcome.
Full list: CHANGELOG.md.