Skip to content

relay server

Eric Busboom edited this page Sep 13, 2026 · 5 revisions

Relay Server

Serve USB-attached relays over TCP, so a radio bridge is a socket away instead of a walk across the room — and connect to a robot by name.

Relay Server

A relay is only useful to whoever is sitting at the machine it is plugged into. The relay server (mbrelay) fixes that: it runs on the host with the boards and hands them out over TCP. A client connects to the pool port, the daemon binds it to a free relay, and from then on the socket is a transparent byte pipe to that board's serial port.

The design target is a drop-in replacement for opening the serial port directly. No client library, no framing, no handshake — nc, a serial terminal pointed at a TCP port, or a few lines of socket code all work unchanged.

   client                    mbrelay (on the host with the boards)         radio
  ─────────►   TCP port  ►  ┌──────────────────────────────────┐  ►  other relays,
   nc / socket / terminal   │  pool: pick a free relay          │     robots,
  ◄─────────             ◄  │  reset it, verify factory default │  ◄  MakeCode
                            │  then get out of the way          │     micro:bits
                            └──────────────────────────────────┘

What you are promised

When you are bound, you have a board that has just been reset and verified at factory defaults: channel 0, group 10, RAW250, power 7, echo off, frag off. The first thing you read is the board's announcement banner, exactly as a direct serial open would give you.

When you disconnect, the server resets the board and restores those defaults before anyone else can have it.

You do not choose which board you get, but you usually get the same one back: the pool remembers which boards your machine used recently and prefers them. That keeps per-robot work on the same hardware so its logs stay comparable.

It is a preference, not a reservation. If your board is taken you get another — the least recently used one — so nothing blocks and wear stays spread. Read the banner to see which you got.

That last part is why this is a service rather than a socat one-liner. The relay firmware persists its configuration in flash across resets and power-cycles (see the Protocol Reference), so a board someone left on channel 23 with echo on would come back that way for the next person — silently, on the wrong frequency, with no error anywhere. The server normalizes on both release and acquire, so a crash or a power cut cannot leave a dirty board in the pool.

Connecting

# Name the ROBOT. The server works out where it is and which board to use.
mbrelay connect tovez

# You do not have to know the address either: the servers announce themselves.
mbrelay connect

# Or name one. A typed address skips discovery entirely.
mbrelay connect <host>:<port>
mbrelay connect tovez@<host>      # a robot, through a named server

# Anything that speaks bytes works, too.
nc <host> <port>

mbrelay connect tovez is the one to reach for. You name the robot, not a channel, a group, a board, a host or a port: the server looks tovez up in its name registry (below), takes whichever relay is free, tunes it, and drops you into the data plane talking to the robot.

From Python, with no dependency on mbrelay at all:

import socket

s = socket.create_connection((HOST, PORT))
s.setsockopt(socket.IPPROTO_TCP, socket.TCP_NODELAY, 1)

print(s.recv(200))            # DEVICE:RADIOBRIDGE:relay:<name>:<serial>
s.sendall(b"!C 5\n")          # channel 5, same grammar as over serial
s.sendall(b"!GO\n")           # from here every byte is radio payload
s.sendall(b"hello over the radio")

Set TCP_NODELAY. Nagle delays small writes that follow one another closely, which can add tens of milliseconds to a radio round-trip. If you pace your writes a few hundred milliseconds apart you will not notice the difference — but it costs nothing to set, and it matters as soon as you send back-to-back.

Which server?

mbrelay discover lists every relay server advertising itself on your network:

$ mbrelay discover
NAME     HOST            ADDRESS       PORT  VERSION
-------  --------------  ------------  ----  ------------
torture  torture.local   192.168.1.12  8760  0.20260826.9
agony    agony.local     192.168.1.19  8760  0.20260826.9

mbrelay connect with no address does the same lookup for you: one server and it connects straight away, several and it asks which. Add --probe to discover to see which of them is actually answering on its port, as opposed to merely advertising.

Which boards?

mbrelay devices --remote lists the boards on every server it finds, not just the servers — which ones are free, and who is holding the rest:

$ mbrelay devices --remote
HOST     NAME   STATE  ROLE         FIRMWARE      SESSION  UID
-------  -----  -----  -----------  ------------  -------  --------
torture  gozop  busy   RADIOBRIDGE  0.20260913.2  s-804    4f02a351
torture  getez  free   RADIOBRIDGE  0.20260913.2  -        17449eac
torture  guvov  free   RADIOBRIDGE  0.20260913.2  -        52f41cc6
torture  zetog  free   RADIOBRIDGE  0.20260913.2  -        f92f913d

$ mbrelay devices torture       # one server, named; skips discovery

It reads each server's HTTP port (the same one the name registry uses, below), never the pool port, so looking does not take a board away from anyone. A server that does not answer is named on stderr and the others are still listed; one running an mbrelay older than 0.20260913.1 is named as needing an upgrade.

If nothing is listed you are probably on a different network from the servers — discovery is link-local and does not cross a router or most VPNs. Name the address instead (mbrelay connect 192.168.1.12:8760); everything else works the same.

You do not always need !GO

!GO is a one-way door — the only way back to the command plane is to disconnect and reconnect. For a single query that is a heavy way to do it, and the command plane can already reach the radio:

> ping            # send one line over the radio, no !GO needed
< pong            # anything received arrives on a "<" line

HELLO re-requests the board's announcement at any time, which is the quick way to confirm which board you are holding.

Use !GO when you want a transparent byte stream; use > for one-shot request/response, which covers most interactive use. !HELP on the board lists the rest of the grammar.

Three things that will catch you out

You cannot reset the board mid-session. BREAK and DTR have no representation in a TCP stream. This matters because the relay's data plane has no in-band escape: once you send !GO, a reset is the only way back to the command plane. The substitute is to disconnect and reconnect — and since the server resets and re-verifies on every bind, the reconnect lands on a freshly clean board.

Reconnecting instantly may be refused when the pool is full. Releasing a board takes two to three seconds: the server has to close the port, reopen it (which is what resets the board), confirm the command plane, restore the settings and verify them. If your script disconnects and immediately reconnects it normally gets a different board and succeeds — the problem only appears when every other board is taken and it needs its own back. Wait a moment, or ask for the server's acquire_wait_ms to be raised so the connect blocks instead of failing.

There is no flow control. Writing flat out at 115200 will overrun the board's USB receive buffer, because the radio is much slower than the serial link. That is exactly what happens on a direct serial connection too, so the server reproduces it faithfully rather than papering over it. Pace your writes — a 10 ms gap between frames is the usual remedy.

When nothing is free

A byte pipe has no error channel, so the server says why in the relay's own comment syntax and then closes the connection:

$ nc <host> <port>
# ERROR: no relay available (4 devices, 3 in use, 1 being handed back)
$

Any client that already ignores # lines — which is any client written against the relay protocol — is unaffected, and a human gets a readable answer instead of a silent hang-up.

Note the two counts. A board that is in use is held by another client; a board being handed back is mid-reset and belongs to nobody — it will be free in a second or two. Lumping them together made it look as though a colleague had a board when nobody did.

Running one

pip install microbit-relayd
mbrelay serve                  # foreground; SIGTERM drains and cleans up

Day-to-day operation:

mbrelay devices                # what is attached, and what state it is in
mbrelay devices --remote       # the boards on every server on the LAN
mbrelay status --watch         # live sessions
mbrelay kick s-3               # boot a session off a board
mbrelay reset <name> --force   # force one board back to factory defaults
mbrelay disable <name>         # take a board out of the pool
mbrelay events --follow        # stream server events
mbrelay flash --all-relays     # reflash every board
mbrelay config show            # merged config, and where each value came from

mbrelay devices works without the server running, which is what you want when you are trying to work out why the server sees nothing. It then asks every board directly, all at once: its name, role and firmware version come from the board itself — relays via HELLO and !VER?, robots via their own HELLO (device NEZHA2 robot tovez …) and VER. Opening a port resets the board; --no-probe lists USB only.

$ mbrelay devices
NAME   STATE        ROLE         FIRMWARE       PORT                     SESSION  UID
-----  -----------  -----------  -------------  -----------------------  -------  --------
zapig  free         RADIOBRIDGE  <0.20260913.2  /dev/cu.usbmodem2121302  -        07d057b7
vitut  busy         RADIOBRIDGE  -              /dev/cu.usbmodem2121202  -        8939f0a5
tovez  busy         NEZHA2       -              /dev/cu.usbmodem2121102  -        a8fdb5e4
  vitut: not asked: port held by another program; name and role from …/mbdeploy-registry.json
  tovez: not asked: port held by node scripts/dev.mjs (pid 56603); name and role from …/radio-robot-lib/config/robots/devices.json

A board that cannot be asked is still named: a port another program holds shows busy and says which program, and its name and role come from mbdeploy's registry, ./config/devices.json, or any file listed in [devices] registries — the line under the table says which. A relay whose FIRMWARE reads <0.20260913.2 answered but predates !VER?.

Plain mbrelay devices only ever looks at the machine you are on; for boards on a server elsewhere, use mbrelay devices --remote or mbrelay devices <host>.

Exit codes are stable, so scripts can branch on them: 0 ok, 1 error, 2 usage, 3 server not running, 4 device not found, 5 no free device, 6 hardware or flash failure.

On the League fleet the server is installed by Ansible and runs under systemd, so systemctl status mbrelay and journalctl -u mbrelay are the first things to try when something looks wrong.

The name registry — where a robot actually is

A micro:bit's five-letter name derives a (channel, group) all by itself (see the Protocol Reference), which is how mbrelay connect tovez works with no configuration at all. But the mapping spreads 3125 names over 73 channels, so about 43 names share each one, and two robots on one channel collide on the air. When they do you have to move one, and its name then no longer says where it is.

So the derived pair is a default, and the registry is what records the exceptions. Asking it about a name always answers:

mbrelay names                          # everything on record
mbrelay names get tovez                # where tovez is
mbrelay names set tovez 12/4           # move it
mbrelay names clear tovez              # back to its derived address

A name nobody has ever asked about is not an error — its default is computed and recorded on the spot, so mbrelay names really is the list of every robot this relay knows about. Only a malformed name is refused: pipip is a legal address nobody happens to be on, while robot1 has no address at all.

Over HTTP, on registry.port (8761 by default), because the people who need this are usually not on the relay host — building a robot's config, or running a channel survey across the fleet:

GET    /names            every association, with conflicts and channel conflicts
GET    /names/<name>     where that robot is, and who it clashes with; creates on a miss
PUT    /names/<name>     {"channel": 12, "group": 4}
DELETE /names/<name>     back to the derived address
GET    /devices          the boards this server holds (what `devices --remote` reads)
GET    /status           version and counts

There is no authentication. This is an internal lab service whose entire content is which radio channel a robot sits on, which anyone with an antenna can determine anyway. Do not expose the port beyond your LAN. If a node should keep the registry to itself, set registry.bind = "127.0.0.1".

Three layers answer a lookup, highest first:

source where when
config [registry.names] in the TOML pinned; a survey's output
registry <state.dir>/names.json set through the API or CLI
derived the name itself computed, then recorded

A config pin outranks anything set through the API — that is the point of it, so a restart cannot quietly reinstate a stale learned value — and mbrelay names set refuses to shadow one rather than pretending to succeed. Changing a pin needs a daemon restart.

The registry reports two kinds of clash, at two severities:

clash what it means severity
conflict two robots on one channel and group error — each receives the other's packets and acts on its commands; move one
channel conflict two robots on one channel, different groups warning — the group byte keeps them from hearing each other, but they share one frequency, so their transmissions collide whenever both run

The mapping never gives two names the same pair, so a conflict only ever comes from a move; two names that derive the same channel are a channel conflict. One clash is reported once — robots that share a link are not also flagged for sharing its channel, unless a third robot is on that channel in another group.

Both show up in mbrelay names (a CONFLICT column, then one ERROR or warning line per clash), in GET /names (conflicts and channel_conflicts, plus conflict / channel_conflict on each robot's row), as counts in GET /status, and on the row for a single robot, so mbrelay names set and mbrelay connect <robot> print them as they happen.

Neither is refused: a survey is expected to pass through a clash halfway, an operator moving robots by hand needs to see it rather than be stopped by it, and mbrelay connect into a clash may be exactly how it gets sorted out.

Moving a robot is a two-sided change. A robot derives its own address from its own name at boot, so it has to be reconfigured to match (the deploy-time channel and group constants in pxt-nezha-diffdrive). The registry only tells the relay where to tune; it cannot move a robot.

The relay firmware cannot see the registry. Its own !N <name> tunes to a name's derived default (see the Protocol Reference), which is the wrong place for a robot that was moved — so mbrelay connect <robot> never uses it: it resolves the name here and tells the board a channel and a group with !CG. That also means mbrelay connect <robot> works against every firmware version in the fleet, since !CG is as old as the protocol. If the registry cannot be reached at all, mbrelay connect falls back to the derived address and says so on stderr — never silently.

A registry saved under an older mapping is corrected on startup: rows that were only ever derived are recomputed, so a change of mapping cannot pin robots to their old defaults. Rows someone set explicitly are left alone.

The radio is shared — check your channel

Four boards means four simultaneous users, and they all transmit into the same air. Two clients that pick the same channel will hear each other's robots, and each robot will act on the other's commands. Nothing in the service prevents this: a shared channel is sometimes exactly what you want, and the data plane is a transparent pipe with no place to interpose.

What the server does do is make it visible. The channel each client selects shows up in mbrelay status, and a collision is logged:

channel_collision channel=4 sessions=s-2(10.0.0.9:52344), s-5(10.0.0.14:41022)

If you are driving a robot, check mbrelay status before you start and agree a channel with whoever else is using the pool.

Two robots can also collide, which is a different problem: about 43 names derive each channel, so two robots can land on one channel no matter who is driving them. Agreeing anything cannot fix that, because the address comes from the name. That is what the name registry is for — it warns about the clash, and you move one of them and record where it went.

If a board stops answering

A board can end up stranded in the data plane — for example if the server was killed outright rather than stopped cleanly, so it never got to reset the board it was holding. A stranded board answers HELLO with silence, because in the data plane your text is radio payload rather than a command.

The server recovers this by itself: when a board does not answer, it sends a break condition, which reboots it. That fallback exists because resetting a micro:bit turns out to be platform-specific — closing and reopening the serial port resets the board on macOS but does nothing at all on Linux, which is what the fleet runs. (Neither does a DTR pulse, nor a 1200-baud touch. A break does, every time.)

If a board still will not answer, reflash it: that resets the chip unconditionally. And prefer systemctl stop mbrelay over killing the process — SIGTERM lets the daemon hand every board back cleanly first.

How boards are identified

The key is the DAPLink USB UID, not the device path — /dev/ttyACM* renumbers on every replug, and the UID does not.

A DAPLink UID is laid out as board(4) family(4) hic(8) unique(16) pad(8) hic(8), and both ends are shared by every board carrying the same interface chip. Four micro:bits from the same batch will end in the identical sixteen characters. If you are matching UIDs by eye, use the middle — which is what mbrelay devices prints.

Identifying a board means opening its port and asking HELLO — then !VER? on a relay, or VER on a robot — and opening the port reboots the board. So the server caches what it learns: a board it already knows is never re-probed, a board someone is using is never probed at all, and a board that stays silent backs off to a five-minute retry rather than being rebooted every few seconds. A cached identity from a server too old to have asked for the firmware version is probed once more, so FIRMWARE fills in after an upgrade.

A robot plugged into a relay host answers in its own dialect (device NEZHA2 robot tovez 2314287040) and is named, with its role and firmware, as foreign — known, but never offered to a client.

Only RADIOBRIDGE boards are offered. The older MakeCode RADIORELAY firmware announces itself too, but does not accept !ECHO ON or !MODE, so the server cannot confirm it has restored a known state — and handing out a board it cannot clean up would break the one promise it makes.

On Linux an installed udev rule also creates /dev/microbit/<uid> symlinks, which are stable across replugs.

Reflashing

mbrelay flash reflashes boards in place, driving pyOCD over SWD via mbdeploy. It takes the board out of the pool, kicks any session on it, flashes, and puts it back.

The server itself has no runtime dependency on that toolchain — a host that only serves relays does not need it installed.