Repository navigation
relay server
Serve USB-attached relays over TCP, so a radio bridge is a socket away instead of a walk across the room — and connect to a robot by name.
A relay is only useful to whoever is sitting at the machine it is plugged into.
The relay server (mbrelay) fixes that: it runs on the host with the boards
and hands them out over TCP. A client connects to the pool port, the daemon
binds it to a free relay, and from then on the socket is a transparent byte pipe
to that board's serial port.
The design target is a drop-in replacement for opening the serial port
directly. No client library, no framing, no handshake — nc, a serial terminal
pointed at a TCP port, or a few lines of socket code all work unchanged.
client mbrelay (on the host with the boards) radio
─────────► TCP port ► ┌──────────────────────────────────┐ ► other relays,
nc / socket / terminal │ pool: pick a free relay │ robots,
◄───────── ◄ │ reset it, verify factory default │ ◄ MakeCode
│ then get out of the way │ micro:bits
└──────────────────────────────────┘
When you are bound, you have a board that has just been reset and verified at factory defaults: channel 0, group 10, RAW250, power 7, echo off, frag off. The first thing you read is the board's announcement banner, exactly as a direct serial open would give you.
When you disconnect, the server resets the board and restores those defaults before anyone else can have it.
You do not choose which board you get, but you usually get the same one back: the pool remembers which boards your machine used recently and prefers them. That keeps per-robot work on the same hardware so its logs stay comparable.
It is a preference, not a reservation. If your board is taken you get another — the least recently used one — so nothing blocks and wear stays spread. Read the banner to see which you got.
That last part is why this is a service rather than a socat one-liner. The
relay firmware persists its configuration in flash across resets and
power-cycles (see the Protocol Reference), so a board someone left on
channel 23 with echo on would come back that way for the next person — silently,
on the wrong frequency, with no error anywhere. The server normalizes on both
release and acquire, so a crash or a power cut cannot leave a dirty board in
the pool.
# Name the ROBOT. The server works out where it is and which board to use.
mbrelay connect tovez
# You do not have to know the address either: the servers announce themselves.
mbrelay connect
# Or name one. A typed address skips discovery entirely.
mbrelay connect <host>:<port>
mbrelay connect tovez@<host> # a robot, through a named server
# Anything that speaks bytes works, too.
nc <host> <port>mbrelay connect tovez is the one to reach for. You name the robot, not a
channel, a group, a board, a host or a port: the server looks tovez up in its
name registry (below), takes whichever relay is free, tunes it, and drops
you into the data plane talking to the robot.
From Python, with no dependency on mbrelay at all:
import socket
s = socket.create_connection((HOST, PORT))
s.setsockopt(socket.IPPROTO_TCP, socket.TCP_NODELAY, 1)
print(s.recv(200)) # DEVICE:RADIOBRIDGE:relay:<name>:<serial>
s.sendall(b"!C 5\n") # channel 5, same grammar as over serial
s.sendall(b"!GO\n") # from here every byte is radio payload
s.sendall(b"hello over the radio")Set TCP_NODELAY. Nagle delays small writes that follow one another closely,
which can add tens of milliseconds to a radio round-trip. If you pace your
writes a few hundred milliseconds apart you will not notice the difference —
but it costs nothing to set, and it matters as soon as you send back-to-back.
mbrelay discover lists every relay server advertising itself on your network:
$ mbrelay discover
NAME HOST ADDRESS PORT VERSION
------- -------------- ------------ ---- ------------
torture torture.local 192.168.1.12 8760 0.20260826.9
agony agony.local 192.168.1.19 8760 0.20260826.9
mbrelay connect with no address does the same lookup for you: one server and it
connects straight away, several and it asks which. Add --probe to discover to
see which of them is actually answering on its port, as opposed to merely
advertising.
mbrelay devices --remote lists the boards on every server it finds, not
just the servers — which ones are free, and who is holding the rest:
$ mbrelay devices --remote
HOST NAME STATE ROLE FIRMWARE SESSION UID
------- ----- ----- ----------- ------------ ------- --------
torture gozop busy RADIOBRIDGE 0.20260913.2 s-804 4f02a351
torture getez free RADIOBRIDGE 0.20260913.2 - 17449eac
torture guvov free RADIOBRIDGE 0.20260913.2 - 52f41cc6
torture zetog free RADIOBRIDGE 0.20260913.2 - f92f913d
$ mbrelay devices torture # one server, named; skips discovery
It reads each server's HTTP port (the same one the name registry uses, below),
never the pool port, so looking does not take a board away from anyone. A server
that does not answer is named on stderr and the others are still listed; one
running an mbrelay older than 0.20260913.1 is named as needing an upgrade.
If nothing is listed you are probably on a different network from the servers —
discovery is link-local and does not cross a router or most VPNs. Name the
address instead (mbrelay connect 192.168.1.12:8760); everything else works the
same.
!GO is a one-way door — the only way back to the command plane is to
disconnect and reconnect. For a single query that is a heavy way to do it, and
the command plane can already reach the radio:
> ping # send one line over the radio, no !GO needed
< pong # anything received arrives on a "<" line
HELLO re-requests the board's announcement at any time, which is the quick way
to confirm which board you are holding.
Use !GO when you want a transparent byte stream; use > for one-shot
request/response, which covers most interactive use. !HELP on the board lists
the rest of the grammar.
You cannot reset the board mid-session. BREAK and DTR have no representation
in a TCP stream. This matters because the relay's data plane has no in-band
escape: once you send !GO, a reset is the only way back to the command plane.
The substitute is to disconnect and reconnect — and since the server resets
and re-verifies on every bind, the reconnect lands on a freshly clean board.
Reconnecting instantly may be refused when the pool is full. Releasing a board takes two to three
seconds: the server has to close the port, reopen it (which is what resets the
board), confirm the command plane, restore the settings and verify them. If your script disconnects and immediately reconnects it normally gets a
different board and succeeds — the problem only appears when every other
board is taken and it needs its own back. Wait a moment, or ask for the
server's acquire_wait_ms to be raised so the connect blocks instead of
failing.
There is no flow control. Writing flat out at 115200 will overrun the board's USB receive buffer, because the radio is much slower than the serial link. That is exactly what happens on a direct serial connection too, so the server reproduces it faithfully rather than papering over it. Pace your writes — a 10 ms gap between frames is the usual remedy.
A byte pipe has no error channel, so the server says why in the relay's own comment syntax and then closes the connection:
$ nc <host> <port>
# ERROR: no relay available (4 devices, 3 in use, 1 being handed back)
$
Any client that already ignores # lines — which is any client written against
the relay protocol — is unaffected, and a human gets a readable answer instead of
a silent hang-up.
Note the two counts. A board that is in use is held by another client; a board being handed back is mid-reset and belongs to nobody — it will be free in a second or two. Lumping them together made it look as though a colleague had a board when nobody did.
pip install microbit-relayd
mbrelay serve # foreground; SIGTERM drains and cleans upDay-to-day operation:
mbrelay devices # what is attached, and what state it is in
mbrelay devices --remote # the boards on every server on the LAN
mbrelay status --watch # live sessions
mbrelay kick s-3 # boot a session off a board
mbrelay reset <name> --force # force one board back to factory defaults
mbrelay disable <name> # take a board out of the pool
mbrelay events --follow # stream server events
mbrelay flash --all-relays # reflash every board
mbrelay config show # merged config, and where each value came frommbrelay devices works without the server running, which is what you want
when you are trying to work out why the server sees nothing. It then asks every
board directly, all at once: its name, role and firmware version come from the
board itself — relays via HELLO and !VER?, robots via their own HELLO
(device NEZHA2 robot tovez …) and VER. Opening a port resets the board;
--no-probe lists USB only.
$ mbrelay devices
NAME STATE ROLE FIRMWARE PORT SESSION UID
----- ----------- ----------- ------------- ----------------------- ------- --------
zapig free RADIOBRIDGE <0.20260913.2 /dev/cu.usbmodem2121302 - 07d057b7
vitut busy RADIOBRIDGE - /dev/cu.usbmodem2121202 - 8939f0a5
tovez busy NEZHA2 - /dev/cu.usbmodem2121102 - a8fdb5e4
vitut: not asked: port held by another program; name and role from …/mbdeploy-registry.json
tovez: not asked: port held by node scripts/dev.mjs (pid 56603); name and role from …/radio-robot-lib/config/robots/devices.json
A board that cannot be asked is still named: a port another program holds shows
busy and says which program, and its name and role come from mbdeploy's
registry, ./config/devices.json, or any file listed in [devices] registries —
the line under the table says which. A relay whose FIRMWARE reads
<0.20260913.2 answered but predates !VER?.
Plain mbrelay devices only ever looks at the machine you are on; for boards on a
server elsewhere, use mbrelay devices --remote or mbrelay devices <host>.
Exit codes are stable, so scripts can branch on them: 0 ok, 1 error, 2
usage, 3 server not running, 4 device not found, 5 no free device, 6
hardware or flash failure.
On the League fleet the server is installed by Ansible and runs under systemd, so
systemctl status mbrelay and journalctl -u mbrelay are the first things to
try when something looks wrong.
A micro:bit's five-letter name derives a (channel, group) all by itself (see
the Protocol Reference), which is how mbrelay connect tovez works
with no configuration at all. But the mapping spreads 3125 names over 73
channels, so about 43 names share each one, and two robots on one channel
collide on the air. When they do you have to move one, and its name then no
longer says where it is.
So the derived pair is a default, and the registry is what records the exceptions. Asking it about a name always answers:
mbrelay names # everything on record
mbrelay names get tovez # where tovez is
mbrelay names set tovez 12/4 # move it
mbrelay names clear tovez # back to its derived addressA name nobody has ever asked about is not an error — its default is computed
and recorded on the spot, so mbrelay names really is the list of every robot
this relay knows about. Only a malformed name is refused: pipip is a legal
address nobody happens to be on, while robot1 has no address at all.
Over HTTP, on registry.port (8761 by default), because the people who need
this are usually not on the relay host — building a robot's config, or running
a channel survey across the fleet:
GET /names every association, with conflicts and channel conflicts
GET /names/<name> where that robot is, and who it clashes with; creates on a miss
PUT /names/<name> {"channel": 12, "group": 4}
DELETE /names/<name> back to the derived address
GET /devices the boards this server holds (what `devices --remote` reads)
GET /status version and counts
There is no authentication. This is an internal lab service whose entire
content is which radio channel a robot sits on, which anyone with an antenna can
determine anyway. Do not expose the port beyond your LAN. If a node should keep
the registry to itself, set registry.bind = "127.0.0.1".
Three layers answer a lookup, highest first:
| source | where | when |
|---|---|---|
config |
[registry.names] in the TOML |
pinned; a survey's output |
registry |
<state.dir>/names.json |
set through the API or CLI |
derived |
the name itself | computed, then recorded |
A config pin outranks anything set through the API — that is the point of it, so
a restart cannot quietly reinstate a stale learned value — and mbrelay names set refuses to shadow one rather than pretending to succeed. Changing a pin
needs a daemon restart.
The registry reports two kinds of clash, at two severities:
| clash | what it means | severity |
|---|---|---|
| conflict | two robots on one channel and group | error — each receives the other's packets and acts on its commands; move one |
| channel conflict | two robots on one channel, different groups | warning — the group byte keeps them from hearing each other, but they share one frequency, so their transmissions collide whenever both run |
The mapping never gives two names the same pair, so a conflict only ever comes from a move; two names that derive the same channel are a channel conflict. One clash is reported once — robots that share a link are not also flagged for sharing its channel, unless a third robot is on that channel in another group.
Both show up in mbrelay names (a CONFLICT column, then one ERROR or
warning line per clash), in GET /names (conflicts and channel_conflicts,
plus conflict / channel_conflict on each robot's row), as counts in
GET /status, and on the row for a single robot, so mbrelay names set and
mbrelay connect <robot> print them as they happen.
Neither is refused: a survey is expected to pass through a clash halfway, an
operator moving robots by hand needs to see it rather than be stopped by it, and
mbrelay connect into a clash may be exactly how it gets sorted out.
Moving a robot is a two-sided change. A robot derives its own address from its own name at boot, so it has to be reconfigured to match (the deploy-time channel and group constants in pxt-nezha-diffdrive). The registry only tells the relay where to tune; it cannot move a robot.
The relay firmware cannot see the registry. Its own !N <name> tunes to a name's
derived default (see the Protocol Reference), which is the wrong
place for a robot that was moved — so mbrelay connect <robot> never uses it: it
resolves the name here and tells the board a channel and a group with !CG. That
also means mbrelay connect <robot> works against every firmware version in
the fleet, since !CG is as old as the protocol. If the registry cannot be
reached at all, mbrelay connect falls back to the derived address and says so
on stderr — never silently.
A registry saved under an older mapping is corrected on startup: rows that were only ever derived are recomputed, so a change of mapping cannot pin robots to their old defaults. Rows someone set explicitly are left alone.
Four boards means four simultaneous users, and they all transmit into the same air. Two clients that pick the same channel will hear each other's robots, and each robot will act on the other's commands. Nothing in the service prevents this: a shared channel is sometimes exactly what you want, and the data plane is a transparent pipe with no place to interpose.
What the server does do is make it visible. The channel each client selects shows
up in mbrelay status, and a collision is logged:
channel_collision channel=4 sessions=s-2(10.0.0.9:52344), s-5(10.0.0.14:41022)
If you are driving a robot, check mbrelay status before you start and agree a
channel with whoever else is using the pool.
Two robots can also collide, which is a different problem: about 43 names derive each channel, so two robots can land on one channel no matter who is driving them. Agreeing anything cannot fix that, because the address comes from the name. That is what the name registry is for — it warns about the clash, and you move one of them and record where it went.
A board can end up stranded in the data plane — for example if the server was
killed outright rather than stopped cleanly, so it never got to reset the board
it was holding. A stranded board answers HELLO with silence, because in the
data plane your text is radio payload rather than a command.
The server recovers this by itself: when a board does not answer, it sends a break condition, which reboots it. That fallback exists because resetting a micro:bit turns out to be platform-specific — closing and reopening the serial port resets the board on macOS but does nothing at all on Linux, which is what the fleet runs. (Neither does a DTR pulse, nor a 1200-baud touch. A break does, every time.)
If a board still will not answer, reflash it: that resets the chip
unconditionally. And prefer systemctl stop mbrelay over killing the process —
SIGTERM lets the daemon hand every board back cleanly first.
The key is the DAPLink USB UID, not the device path — /dev/ttyACM*
renumbers on every replug, and the UID does not.
A DAPLink UID is laid out as
board(4) family(4) hic(8) unique(16) pad(8) hic(8), and both ends are shared by every board carrying the same interface chip. Four micro:bits from the same batch will end in the identical sixteen characters. If you are matching UIDs by eye, use the middle — which is whatmbrelay devicesprints.
Identifying a board means opening its port and asking HELLO — then !VER? on a
relay, or VER on a robot — and opening the port reboots the board. So the
server caches what it learns: a board it already knows is never re-probed, a
board someone is using is never probed at all, and a board that stays silent
backs off to a five-minute retry rather than being rebooted every few seconds. A
cached identity from a server too old to have asked for the firmware version is
probed once more, so FIRMWARE fills in after an upgrade.
A robot plugged into a relay host answers in its own dialect
(device NEZHA2 robot tovez 2314287040) and is named, with its role and firmware,
as foreign — known, but never offered to a client.
Only RADIOBRIDGE boards are offered. The older MakeCode RADIORELAY firmware
announces itself too, but does not accept !ECHO ON or !MODE, so the server
cannot confirm it has restored a known state — and handing out a board it cannot
clean up would break the one promise it makes.
On Linux an installed udev rule also creates /dev/microbit/<uid> symlinks, which
are stable across replugs.
mbrelay flash reflashes boards in place, driving pyOCD over SWD via
mbdeploy. It takes the board out of
the pool, kicks any session on it, flashes, and puts it back.
The server itself has no runtime dependency on that toolchain — a host that only serves relays does not need it installed.