Skip to content

troubleshooting

Eric Busboom edited this page Sep 24, 2026 · 1 revision

Troubleshooting

Diagnosis steps for agents: daemon unreachable, port in use, 127.0.1.1, USB permission denied, stale peers.

Troubleshooting

First: a health check

Run on the host in question:

systemctl is-active mbregistry                                   # Linux
sudo launchctl print system/org.jointheleague.mbregistry | grep -E 'state|pid|last exit'   # macOS LaunchDaemon
mbregistry list --json > /dev/null; echo "exit=$?"               # 0 = healthy, 3 = no daemon reachable
journalctl -u mbregistry -n 50                                   # Linux log; macOS: /Library/Logs/mbregistry.log

Exit codes: 0 ok · 1 error · 2 usage · 3 no daemon · 4 no such device · 5 locked · 6 flash failed · 130 interrupted.

Daemon not reachable (exit 3, "is the daemon running?")

  1. Is it running? If it is in a restart loop, the log's last traceback says why. It is usually a port in use (next section).
  2. Where is the client looking? The error names the socket path. Clients try the user's own socket first, then the system one:
    • Linux: $XDG_RUNTIME_DIR/mbregistry/api.sock or ~/.cache/mbregistry/api.sock, then /run/mbregistry/api.sock
    • macOS: ~/Library/Application Support/mbregistry/api.sock, then /var/run/mbregistry/api.sock
    • Windows: \\.\pipe\mbregistry
  3. Stale per-user socket. A client uses the first socket file that exists, even with nothing listening. A leftover per-user socket from a killed per-user daemon hides a healthy system daemon. Check that no per-user daemon runs (pgrep -fl "mbregistry.*run"), delete the stale file, and retry.
  4. Environment override. A leftover MBREGISTRY_SOCKET beats discovery. Check with env | grep MBREGISTRY.
  5. Custom --socket on the daemon. Pass the same --socket to the client.
  6. Socket exists, connection refused. The daemon died uncleanly, or a second daemon replaced its socket. Restart the service; it recreates the socket on start.

Port already in use

The daemon logs OSError: [Errno 98] Address already in use (Linux) or [Errno 48] (macOS), and the service manager keeps restarting it.

sudo ss -ltnp | grep -E ':(7440|744[2-5])\b'                   # Linux
sudo lsof -nP -iTCP -sTCP:LISTEN | grep -E ':(7440|744[2-5])'   # macOS/Linux
pgrep -fl "mbregistry.*run"
  • Usual cause: a second mbregistry. A leftover foreground run, or a per-user daemon next to the system one. Stop the extra one, then restart the real service. The second daemon bound its local socket before it failed on 7440, so it may have replaced the real daemon's socket or left a stale per-user one.
  • If another program owns the port, move mbregistry with --remote-port/--peer-pub-port/--peer-snapshot-port, and do the same on every host. 7444/7445 are fixed. --no-relay-pool frees 7444.

Peers can't reach this host / it advertises 127.0.1.1

Symptom: other hosts list this host's boards, but remote deploy/serial time out or are refused. Or mDNS shows a loopback address.

avahi-browse -rt _mbregistry._tcp      # check the 'address = [...]' for this host
ip route show default                  # empty = no default route
nc -vz <this-host> 7440                # from a peer
  • Old build. Builds up to v0.20260924.2 fall back to resolving the hostname when there is no default route. On Debian that gives 127.0.1.1. Upgrade to a newer build, which uses the first real interface instead.
  • Wrong interface chosen. The daemon skips docker*, br-*, veth*, virbr*, vmnet*, bridge*, utun*, tun*, tap*, loopback and link-local. If the LAN is on a skipped name, or another interface comes first, give the host a default route via the LAN.
  • Firewall. Open 7440/7442/7443 TCP and 5353 UDP. See Networking and peering.

Permission denied on USB (local clients, Linux)

Symptom: mbserial <name>, mbdeploy deploy <name> or mbrelay connect against a board on this host fails with Permission denied: '/dev/ttyACM0', or pyOCD finds no probe or hangs. Boards on peers are not affected, and the root daemon never is.

ls -l /dev/ttyACM*                                   # want group plugdev, crw-rw----
id | grep -o plugdev                                 # want: plugdev
ls /etc/udev/rules.d/99-mbregistry-cmsis-dap.rules   # want: present

Fix:

sudo /opt/mbtools/bin/mbregistry install-service     # (re)writes the udev rule
sudo udevadm control --reload-rules && sudo udevadm trigger
sudo usermod -aG plugdev "$USER"
# log out and back in (new SSH session): group changes don't reach existing sessions

Stopgap: sudo mbserial .... See Installing mbtools.

Stale or missing peers

  • peer unreachable: that host's daemon is down or unreachable. Check it there. The rows recover by themselves when it reconnects.

  • A retired host never disappears. Peers are never deleted. To purge one, do this on each host that still lists it (Linux root paths shown; use the database path for your platform):

    sudo systemctl stop mbregistry
    sudo /opt/mbtools/bin/python - <<'EOF'
    import sqlite3
    db = sqlite3.connect("/var/lib/mbregistry/devices.db")
    db.execute("DELETE FROM device WHERE host = ?", ("OLDHOST",))
    db.execute("DELETE FROM peer WHERE host = ?", ("OLDHOST",))
    db.commit()
    EOF
    sudo systemctl start mbregistry

    The blunt alternative is to move devices.db aside and restart. Local boards are re-probed, and robot names return from peers' snapshots. Check afterwards with mbrelay names list.

  • A peer never shows up. Multicast is blocked (VLANs, Wi-Fi isolation) or ports are firewalled. Add --peer HOST (see Networking and peering), and browse _mbregistry._tcp from both sides.

  • A peer shows up but its boards don't. Its 7443 is blocked, or the auth tokens differ. The log has a peering warning.

  • No peers after a reboot until the daemon is restarted. The daemon started before the network was up. Restart it (systemctl restart mbregistry / launchctl kickstart -k ...).

  • macOS LaunchAgent sees no peers but Terminal does. Allow it under System Settings → Privacy & Security → Local Network, or use the root LaunchDaemon.

Other symptoms

Symptom Cause / fix
Board shows gone Unplugged, or USB re-enumerated. Replug. The daemon polls every 2 s (--interval).
Board shows no-firmware Nothing announced when it was probed. It may be blank, or running firmware that doesn't announce. Flash it.
mbdeploy deploy refuses a board Its role looks like a relay/bridge. Use --force-relay only if you mean it.
Exit 5 (locked) Another session holds the board. mbregistry list shows who (locked by ...).
Remote op fails unauthorized That host runs with --auth-token. Clients can't send it (known gap).
Robot-console finds no relays Check avahi-browse -rt _mbrelay._tcp. Make sure --no-relay-pool isn't set and 7444/7445 are open. See Robot-console compatibility.
Windows service stops right after start (error 1053) Its binPath= lacks --windows-service. Re-run install-service and use the printed sc.exe commands.

Clone this wiki locally