Repository navigation
protocol
The v6 ASCII line grammar, protocol handler, adapter contract, and the sequence-id reliability layer — the wire authority.
Status: built and self-contained. This document is the wire format
and the design — there is no external spec file. §1–§8 are the design as
implemented; §9 is what the implementation actually found: resolved
ambiguities, one deliberate omission, and the gaps this work exposed,
including the 2026-08-20 space/#id grammar migration, the 2026-08-21
addition of debug and RUN (§6.2/§6.3), the 2026-08-21 reliability
layer — mandatory sequence ids plus cumulative ack/nack, replacing the
undelivered 3×-reply-repeat idea outright (§8) — and the 2026-08-22 changes
(§8.9): a decode failure now NAKs instead of acking, ERR_DUPLICATE_ID is
deleted, PING joins the unsequenced exemption set, lastDone/its reason
move from the handler to the Adapter, every ack/nack gains a reason
token, and the six motion-api §9.1 verbs
(WHEELS_X/WHEELS_V/MOVE_X/MOVE_V/GO_TO_R/GO_TO_W, plus STOP's
own now token) are implemented at the wire/handler layer (WHEELS is
renamed WHEELS_V). 2026-08-26 (§8.5): the telemetry-piggybacked
reliability line is deleted outright — ack/nack is emitted only
in direct reply to an inbound line, never periodically, and
gapOutstanding_ left the handler state with it. 2026-08-27 (§8.3,
§9.11): the unsequenced set grows from three verbs to seven — HELP,
ID, VER and STATUS join HELLO/ESTOP/PING — under a stated
rule (a verb is sequenced iff its correctness depends on its position in
the stream); status gains done=/reason=; ID gains a mandatory
name field (§6.4); an unknown GET name now returns err 1 like
SET instead of nothing; and gapOutstanding_ returns, scoped, as the
predicate for a conditional "your last command didn't land" reminder that
five unsequenced verbs carry when — and only when — the stream is
stalled. §8.0's "~5%" loss premise is withdrawn as unsupported.
Three objects, each with one job, arranged so that the middle one is the only thing that ever knows what a wire byte looks like.
bytes from the port bytes to the port
│ ▲
▼ │
┌────────────────────────────────────────────────────────────────────┐
│ ProtocolHandler │
│ feed(data, length) → reassemble lines → parse → dispatch │
│ reply formatting: ack / nack / err / estop / get / thdr / t / … │
│ telemetry emission on a cadence the app drives │
└────────────────────────────────────────────────────────────────────┘
│ calls typed methods ▲ returns typed results
▼ │
┌────────────────────────────────────────────────────────────────────┐
│ Adapter (you implement — one class, all the callable methods) │
│ onWheels(left, right, duration, id) → Result │
│ onSet(name, value, id) → Result onGet(name, out) → bool │
│ onEstop() onStop(id) onPing() → now … │
└────────────────────────────────────────────────────────────────────┘
│ calls
▼
DifferentialDrive, config store, clock — the actual machine
The adapter never writes a wire byte and never parses one. It receives
decoded, typed arguments and returns a typed result. Everything textual lives
in ProtocolHandler. That is what makes the wire format single-sourced, the
adapter unit-testable without a socket, and a second language implementation a
matter of porting one class.
The transport is not an object here at all. feed() takes bytes from wherever
you got them and a Sink takes bytes back out — a serial port, a radio frame,
a UDP datagram, or a test's std::vector.
One grammar. ASCII. No framing layer.
line ::= sp? verb ( sp field )* sp? '\n'
sp ::= ' '+
verb ::= [A-Za-z][A-Za-z0-9_]*
field ::= any bytes except ' ' and '\n'
id ::= '#' [0-9]+ (mandatory, trailing, §2.2/§8)
That is the entire wire format, both directions, every message. No COBS, no
CRC, no length prefix, no binary plane, no protobuf, no generated codec —
readline() is the transport. It replaced an earlier binary/cleartext-mixed
scheme whose complexity (COBS framing, an app-level CRC duplicating one the
radio hardware already computes, a nine-FieldKind protobuf walker, four
different cleartext sub-grammars) bought portability and integrity guarantees
this project does not use: endpoints are rev-locked by policy, and real
losses are whole packets, about which a CRC says nothing.
-
Terminator is
'\n'(0x0A). A lone'\r'before it is stripped (terminal artifact);'\r'appears nowhere else. -
The first space ends the verb. One or more spaces separate fields; a run
of spaces is ONE separator, and leading/trailing whitespace on the line is
ignored.
strtol/strtof's own whitespace-skipping implements most of this for free (§9.4's leading-whitespace finding covers the residue). - A blank or all-whitespace line is ignored silently — a terminal artifact, not an error; it does not count malformed.
- Fields are positional, fixed arity per verb. Wrong arity is a rejection, not a best-effort parse.
-
Max line: 240 bytes including the terminator. Chosen to sit inside a
radio MTU and a safe UDP payload on the transports this format was designed
for, so a message never fragments — a property this library's
feed()relies on for its overflow-discard behavior (§3.1), even though this library itself is transport-agnostic. -
Every wire value is a base-10 ASCII integer, optionally signed, except
config values (§7.2), which are decimal. No exponents, no
NaN, noinf.flagsis the one exception to base-10: lowercase hex, no0x.
Commands (host → robot) UPPERCASE. Replies (robot → host) lowercase. Verb lookup is case-sensitive.
Not cosmetic. On a shared radio channel a robot hears every other robot. A
robot's own debug output being a syntactically valid command to every other
robot on the channel produces a self-sustaining flood
(.claude/rules/hardware-bench-testing.md, radio-robot-elite). Under this
grammar a reply can never parse as a command, so that class is closed
structurally instead of by keeping the channel private.
A verb starting lowercase is dropped silently and does not count malformed — that is another robot's reply, not an error.
Superseded, 2026-08-21. This section used to describe the id as an
optional, host-assigned correlation token with an ack-suppression spelling
(#0) and a per-verb "required/optional" split. That design is gone,
replaced wholesale by §8's reliability layer: every sequenced command now
carries a mandatory, strictly incrementing id, starting at 1, and the id
is the sequence number the ack/nack scheme tracks — it is no longer a
free-standing correlation token a caller could pick arbitrarily.
An id is still spelled #<n> and is still always the last token of
its line — commands and replies alike.
-
Mandatory on every sequenced verb (
GET SET TLM WHEELS_X WHEELS_V MOVE_X MOVE_V GO_TO_R GO_TO_W STOP RUN— see §8.3 for the seven unsequenced verbs,HELLO,ESTOP,PING, and (as of 2026-08-27)HELP,ID,VER,STATUS, none of which ever carry one at all). A sequenced verb arriving with no id is malformed.ID VER STATUS HELPleft this list on 2026-08-27 — see §8.3 for the rule that moved them and §9.11 for why. - The digits are bare and unsigned:
#+5,#-5, and# 5are all malformed — §2's "optionally signed" applies to data fields, not the id, and needs a dedicated digits-only parser rather than the general integer one. This part is unchanged from the old design. - Because it is always the line's last token, it never shifts position regardless of a verb's own field count, and it is recoverable even from a line that otherwise fails to parse — see §8.4 for how malformed-line recovery works under the new scheme.
-
#0is deleted as a special case. It used to mean "no ack wanted, execute silently"; that spelling is incoherent once every message must be sequenced and acked (a suppressed ack is a hole in the stream by definition). Since ids start at 1 and the session'sexpectedNext_counter is always ≥ 1, an inbound#0simply always compares less thanexpectedNext_— it falls into the ordinary stale/retransmit bucket (§8.1) with no special-casing at all: acked against the highest already-accepted id, never executed. There is no way to suppress a reply any more, by design — suppression is incompatible with a scheme that must see every id to detect gaps. -
ERR_DUPLICATE_ID(code 11) is DELETED (2026-08-22). The old design had the adapter detect and reject a reused id. Under the reliability scheme the handler itself enforces strict monotonicity before an id ever reaches the adapter — an id is only ever dispatched when it exactly equalsexpectedNext_, which then advances past it, so the adapter can never be handed the same id twice.Result::kDuplicateIdwas kept declared-but-unreachable through 2026-08-21 (flagged in §9.8); the 2026-08-22 pass removed the enumerator and its wire code entirely — it is structurally unreachable, not merely unused, and there is nothing left to keep it declared for. Keep the duplicate-retransmit re-ack logic (theid < expectedNext_→ re-ack-without-re-executing row of §8.1's table) — that is a different thing, essential, and unaffected by this removal.
The old malformed-line recovery rule ("if the line's last token is a
well-formed nonzero #id, reply err #<id> <code>, with ESTOP the one
exception") is gone. It was designed for a world where the id was an
optional correlation token; under the mandatory-sequence-number design an
id is either present, well-formed, and sequence-checked, or the line simply
cannot be classified at all. See §8.4 for the replacement rule, and §8.3
for ESTOP's own (still exceptional, but differently exceptional) status.
Stakeholder direction: the banner carries colons because it belongs to
a separately specified protocol — the device-announcement format in
microbit-radio-relay/docs/announce.md — and not to the v6 line grammar
described above.
DEVICE:<role>:<common_name>:<device_name>:<serial>
DEVICE:NEZHA2:robot:vevov:1198504156 (a robot)
DEVICE:RADIOBRIDGE:relay:getez:1779042496 (a relay)
Five fields, sentinel first, in the same order for both device classes.
That is the whole point: one discovery parser finds robots and relays
alike. Through 2026-08-26 this library emitted the same five fields
space-delimited and lowercased (device NEZHA2 robot vevov 1198504156),
which carried identical information and was silently unparseable by the
tooling that consumes it — probe_type accepts only the colon form, so
every robot probed came back with role, common_name, device_name
and serial unset. Two spellings of one format is the defect; this
resolves it toward the format that was specified first and elsewhere.
The announced <serial> is the decimal FICR.DEVICEID[1], which is the
same value the device registry already holds, so it cross-checks against
device_id for free.
Consequences, stated because this line is a deliberate foreign body:
- It is exempt from §9.6's colon-to-space migration. That decision governs the v6 line grammar. The banner is not a v6 line, has no verb, no fields in the §2 sense, and never carries an id — the migration never had jurisdiction over it.
- It is exempt from §2.1's case rule, and this one has a real cost — see the flag below.
-
HELLOitself is unchanged: still uppercase, still zero-arity, still unsequenced, still resetsexpectedNext_(§8.3). Only the shape of the line it emits in reply is specified elsewhere.
FLAGGED for the stakeholder — the uppercase sentinel interacts with §2.1. That section's guarantee is that a robot's own output can never parse as a command, which is what makes a shared radio channel safe from self-sustaining floods: a lowercase first token is dropped silently and is explicitly not counted malformed.
DEVICE:NEZHA2:robot:vevov:...is uppercase, so on a shared channel every other robot tokenises it as one unknown token with no trailing#id, fails id resolution, and takes §8.4's item 1/2 path. It does not flood — that path emits no reply at all — but it does incrementmalformedCount()on every listening robot, so a diagnostic counter now counts ordinary neighbour traffic.A lowercase sentinel (
device:NEZHA2:robot:vevov:1198504156) would keep the colon delimiter, all five fields, and the shared parse, while preserving §2.1 exactly. Its cost is one extra accepted spelling inprobe_type— which that code needs regardless, since relays will keep announcingDEVICE:. The uppercase form is what is specified and what is implemented here; the alternative is recorded so the choice is visible rather than inherited.
namespace Protocol {
class Sink { // where finished lines go
public:
virtual ~Sink() = default;
virtual void write(const char* data, size_t length) = 0;
};
class ProtocolHandler {
public:
ProtocolHandler(Adapter& adapter, Sink& sink);
// Feed an arbitrary block from the port. Partial lines are buffered
// across calls; complete lines are parsed and dispatched immediately.
void feed(const char* data, size_t length);
// Unsolicited emissions the app drives, not the wire.
void sendBanner(); // DEVICE:NEZHA2:robot:<name>:<serial> — §2.4
void sendReady(); // ready
void sendDebug(const char* text); // debug <text> -- robot-to-host ONLY
void emitTelemetry(const Snapshot& snapshot); // thdr once, then t per frame
uint32_t malformedCount() const;
};
} // namespace ProtocolThis is the method most likely to be got wrong, because on a bench it is always handed one tidy line and in the field it never is. It must handle:
- a block containing several complete lines,
- a block ending mid-line (buffer the remainder, dispatch on the next feed),
- a block that is only a line fragment,
-
\r\n— a lone\rbefore the terminator is stripped as a terminal artifact;\rappears nowhere else (§2), - a blank or all-whitespace line (ignored silently — §2, NOT malformed),
- a line longer than the 240-byte maximum — discard to the next
\nand count it malformed, rather than overflowing or truncating into a half-line that parses as something valid.
That last one matters more than it looks: a truncated line whose surviving prefix is still a legal verb with legal arity is a command the host never sent. Discard-to-terminator is the only safe recovery.
The entire codec is a tokenizer over runs of ' '. Firmware constraints mean
no dynamic allocation and no std::string: the handler owns a fixed
char[240] line buffer, and tokenizing overwrites separator spaces with \0
to produce field pointers into that same buffer. Numbers come out with
strtol/strtof. The correlation id (§2.2), where a verb carries one, is a
trailing, self-marking #<n> token rather than a positional field — recovered
by a separate backward scan over the raw line, done before the forward
tokenizer mutates it, so the id stays recoverable even past the small
fixed field-token array's own storage cap
(protocol_handler.cpp's tokenizeLine()).
Rewritten 2026-08-22 for the six stakeholder-directed changes. The
biggest structural change: lastDone/its completion reason MOVE here from
a ProtocolHandler field that nothing ever wrote (§8.8), onWheels
renames to onWheelsV, five new motion methods join it (one per
motion-api §9.1 verb), onStop gains an immediate
argument, and kDuplicateId is deleted outright (§2.2).
namespace Protocol {
// Maps 1:1 onto the wire outcome (§8.2): kOk means nothing further is
// emitted beyond the ack dispatch() already sent (no more standalone
// `ok`); every other value means an `err <code> #<id>` follows that same
// ack (id last, §8.6). kDuplicateId is GONE (2026-08-22) -- it was
// already structurally unreachable (§2.2), and there is nothing left to
// keep a deleted enumerator declared for.
enum class Result : uint8_t {
kOk, // → (ack alone; no further reply)
kUnknown, // → err 1 #<id> ERR_UNKNOWN
kBadArg, // → err 2 #<id> ERR_BADARG
kRange, // → err 3 #<id> ERR_RANGE
kFull, // → err 4 #<id> ERR_FULL
kUnimplemented, // → err 6 #<id> ERR_UNIMPLEMENTED
kNotReady, // → err 8 #<id> ERR_NOT_CONFIGURED
kBusy, // → err 10 #<id> ERR_BUSY
};
// The reliability layer's completion-reason vocabulary (§8.8, motion-
// api.md §5.3): the four reasons a motion can finish, plus kNone for
// "nothing has completed yet" (lastDone() == 0's own pairing).
enum class DoneReason : uint8_t {
kNone, // → "none"
kStop, // → "stop" -- the stop condition was met, or stop() ended it
kTimeout, // → "timeout" -- the backstop fired
kEstop, // → "estop" -- a panic stop ended it
kAborted, // → "aborted" -- the caller abandoned it
};
class Adapter {
public:
virtual ~Adapter() = default;
// ---- session ----
virtual void identity(Identity& out) const = 0; // name, serial, version, …
virtual uint32_t now() const = 0; // [ms] for pong
virtual void status(StatusFields& out) const = 0;
// ---- motion: the six verbs (motion-api.md §9.1), plus STOP/ESTOP.
// Angles (rotation, omega) arrive already decoded from the wire's
// milliradian integers into float milliradians -- degrees-at-the-API
// is a LANGUAGE BINDING's conversion, not this seam's.
virtual Result onWheelsV(float left, float right, // [mm/s] [mm/s]
uint32_t duration, // [ms]
uint32_t id) = 0; // renamed from onWheels
virtual Result onWheelsX(float left, float right, // [mm] [mm]
float cruise, // [mm/s]
uint32_t timeout, // [ms]
uint32_t id) = 0;
virtual Result onMoveX(float distance, float rotation, // [mm] [mrad]
float cruise, uint32_t timeout, // [mm/s] [ms]
uint32_t id) = 0;
virtual Result onMoveV(float v_x, float omega, // [mm/s] [mrad/s]
uint32_t duration, uint32_t id) = 0; // [ms]
virtual Result onGoToR(float x, float y, float speed, // [mm] [mm] [mm/s]
float arrive, uint32_t timeout, // [mm] [ms]
uint32_t id) = 0;
virtual Result onGoToW(float x, float y, float speed,
float arrive, uint32_t timeout,
uint32_t id) = 0;
virtual Result onStop(bool immediate, uint32_t id) = 0; // STOP [now]
virtual void onEstop() = 0; // never sequenced, never queued -- the
// handler replies `estop` itself (§8.3),
// this method's own return is still void
// ---- configuration ----
virtual bool onGet(const char* name, float& out) const = 0;
virtual Result onSet(const char* name, float value, uint32_t id) = 0;
virtual size_t fieldCount() const = 0; // for bare GET
virtual const char* fieldName(size_t index) const = 0;
// ---- telemetry ----
virtual Result onTlm(TlmMode mode) = 0;
// ---- the reliability layer's completion channel (§8.8, MOVED here
// 2026-08-22 from a ProtocolHandler field nothing ever wrote) ----
virtual uint32_t lastDone() const = 0;
virtual DoneReason lastDoneReason() const = 0;
// ---- invocation by name (§6.3) ----
virtual Result onRun(const char* name,
const char* const* argv, size_t argc,
char* result, size_t resultCapacity,
bool& hasResult) = 0;
};
} // namespace ProtocolkUnimplemented and kBusy are carried for completeness against the wire's
error-code space even though DiffDriveAdapter does not itself produce
them — kUnimplemented is reserved for a recognized-but-unwired verb,
kBusy for a subsystem refusing because it is mid-motion.
Returning a Result rather than writing a reply is the deliberate choice. It
means the adapter cannot emit a malformed reply, cannot forget to reply, and
cannot invent a reply shape — the handler does all three, once, for every verb.
onEstop() returns void on purpose: ESTOP never carries a sequence id and
is never part of the ack/nack scheme, because it must not queue behind
anything, including an ack (§8.3). The handler itself still replies estop
after calling this — that reply is not this method's own concern; it is
formatted and sent by ProtocolHandler, exactly like every other reply.
onStop's immediate flag carries STOP now's own request (motion-
api.md §3.7/§9.1) — a deceleration CHOICE, not a different verb.
DiffDriveAdapter accepts and ignores it (its neutral() was already
immediate either way, §5.1); an adapter that owns a real ramp is where the
flag first makes a behavioral difference.
completion channel, now Adapter-owned
The handler POLLS these two methods fresh every time it formats an
ack/nack line (§8.1) — no callback, no clock, no cached copy
anywhere in ProtocolHandler. Monotonic contract: a later lastDone()
value implies every earlier id has also completed, since this library's
motion runs one command at a time, in order. An adapter with no
completion event of its own (DiffDriveAdapter, §8.8.1) returns 0/
kNone forever — wire-correct, functionally inert on that adapter
specifically. See §8.8 for the full story, including why this moved off
the handler.
The concrete adapter that closes the loop for testing:
WHEELS_V <left> <right> <duration> #<id>
[mm/s] [mm/s] [ms]
│
▼ scale by countsPerLength [counts/mm] ← the robot's geometry
left, right [counts/s]
│
▼ velocity = (left+right)/2 , twist = (right-left)/2
DifferentialDrive::drive(velocity, twist, lease=duration)
Renamed from WHEELS/onWheels (2026-08-22) — same fields, same
meaning; motion-api §9.2 confirms WHEELS is
wheels_v. DiffDriveAdapter implements only this one motion verb for
real. The other five (WHEELS_X/MOVE_X/MOVE_V/GO_TO_R/GO_TO_W)
all answer Result::kUnknown on DiffDriveAdapter specifically — it has
no planner, which is honest and already documented (this same posture
RUN already takes for an adapter with an empty registration table, §6.3)
— not kUnimplemented, a deliberate choice recorded in §9.8.
Three things fall out of this that are worth stating before anyone writes it:
-
durationislease, unchanged. Both are[ms], both mean "stop if nobody talks to me again within this window", both exist so a dead host cannot mean a runaway. No reinterpretation, no clamping surprise beyond DiffDrive's ownkLeaseMaxand the wire's 5000 ceiling (enforced by the adapter — the handler holds no bounds table). -
countsPerLengthis the only geometry in the whole path, and it lives in the adapter because DiffDrive deliberately has no millimetres in it (diffdrive §1.1). It is a constructor argument, not a config field — it is a property of the robot's gearing/wheel, not a tunable control-law gain, so it is not reachable throughGET/SET. -
twistis the half-differential, CCW-positive. Getting the sign wrong here is the single most repeated bug in this project's history — a robot whose "left" wheel was physically the right one negated every wheel-derived heading while leaving forward motion correct, so nothing surfaced it and it was patched four times downstream. The adapter is the one place that ordering is decided, and it needs a test that would fail if the two wheels were swapped.
There is no queue in DiffDriveAdapter specifically, because it has no
planner — WHEELS_V reaches drive() directly, and the other five motion
verbs answer kUnknown on it (§5). This is a property of that one concrete
adapter, not a limit the wire or the reliability layer impose: an adapter
that DOES own a planner (this library's own tests/protocol/ fake_motion_adapter.h, built to exercise exactly this) queues motion
commands and runs them to completion one at a time, in arrival order — see
§8.8's own monotonic-lastDone contract, which is stated for Adapter in
general and assumes exactly that ordering guarantee. So neither STOP nor
ESTOP "waits its turn behind an active move" on DiffDriveAdapter; that
framing belongs to a full motion-planner robot (see
motion-api §3.7/§9.1, which specifies
stop/stop(immediate=True)/estop as a richer three-way choice
DiffDriveAdapter collapses onto one behavior). What STOP [now] #<id> and
ESTOP actually do, traced to the kernel:
| verb | adapter call | kernel effect |
|---|---|---|
STOP [now] #<id> |
onStop(immediate, id) → drive_.neutral()
|
writes a neutral command to the mailbox; duty zeroes this cycle, immediately — stageStop() is a bare stageDuty(0, 0), no ramp. immediate (the optional now token, motion-api.md §3.7/§9.1) is accepted but has no effect here — see §9.8 |
ESTOP |
onEstop() → drive_.estop()
|
sets a latch that forces the same neutral path from the next cycle regardless of the mailbox's state, and additionally refuses every new motion command (kRefusedEstopped) until estopClear()
|
The two are not distinguished by how fast duty goes to zero — both are
immediate at the kernel level. They are distinguished by persistence and
guarantee: neutral() is an ordinary mailbox write that the very next
drive() call overrides; estop() latches outside the normal command
handshake specifically so it is effective even if that handshake is wedged,
and it blocks new motion until explicitly cleared. onStop() always returns
kOk — neutral() has no refusal path of its own, even pre-begin() or
mid-estop.
ESTOP's wire reply changed, 2026-08-21 (§8.3). It used to never reply
at all; it now always replies the bare word estop, with the kernel call
(onEstop()) executed before that reply is written, so the reply can
never be mistaken for having queued ahead of the actual stop.
This is a materially smaller distinction than a full motion-planner robot's
"planned stop queues behind the active move, ramps down at a decel ceiling"
versus "estop halts now" — that contrast (measured elsewhere at 39.8 cm/5.9 s
versus 2.9 cm/0.10 s) is about a queue, not about deceleration profile, and
it describes a planner this library does not have. Do not cite that
measurement as characterizing DiffDriveAdapter::onStop(); it does not apply
here.
DifferentialDrive::output() already publishes everything a t frame needs
(see diffdrive §4) — positions, velocities, applied duties,
timing, and the health flags. The adapter's telemetry job is a projection, not
a computation: pick the columns for the active TLM mode, convert counts to
the wire's units, and hand the handler an array.
thdr is emitted once on subscribe and names the columns; t carries the
values in that order. The frame is self-describing, so a consumer never
hardcodes a column index.
This is a reduced projection, not a world-frame pose. A full robot's
POSE/FULL telemetry carries x/y/h fused from OTOS and encoder
odometry — neither of which lives in this library. What DiffDrive actually
publishes is per-wheel counts and counts/s, so that is what this adapter
projects: posl/posr [mm] and vell/velr [mm/s ×10], converted
through the one geometry factor the adapter owns; TLM FULL adds
lambda/biasl/biasr/cyc — the kernel's own learned-state and heartbeat
fields. flags is a local bit layout (computeFlags() in
diffdrive_adapter.cpp), not any externally-numbered scheme — this library has
no OTOS/line/colour/planner, so reusing bit numbers that meant those things
elsewhere would actively mislead a reader.
See §10 for the full telemetry-frame specification built on top of this
projection — TLM <mode> subscription semantics, the thdr/t line
grammar and emission rules, this library's exact column tables, and the
TLM HDR header-recovery command. This section states what gets projected;
§10 states the wire contract around it.
Enough to test DiffDrive over the wire, plus the full six-verb motion
surface motion-api §9.1 specifies — the WIRE and
HANDLER side of all six is implemented; DiffDriveAdapter itself only
gives one of them (WHEELS_V) real effect (§5).
Every row marked "sequenced" below requires a mandatory #<id> — see
§8. The reply column shows each verb's own informational reply only;
every sequenced verb ALSO emits the transport-layer ack/nack
described in §8, as a separate line, alongside whatever is shown here.
PING is unsequenced as of 2026-08-22, and HELP/ID/VER/
STATUS joined it on 2026-08-27 — none of the seven ever carries an id
(§8.3). An unsequenced verb emits no ack/nack of its own on a clean
stream, but five of them DO carry a conditional reminder when the stream
is stalled — see §8.3's reminder rule, which is a different thing from
the transport reply a sequenced verb always gets.
| verb | sequenced? | command | own reply | notes |
|---|---|---|---|---|
HELLO |
no | — | DEVICE:NEZHA2:robot:<name>:<serial> |
resets the sequence (§8.3); not a v6 line — see §2.4 |
PING |
no (2026-08-22) | — | pong <now> |
now = robot clock [ms]; maximally forgiving, like ESTOP — see §8.3, §9.8 |
ID |
no (2026-08-27) | — | id <drivetrain> <profile> <version> <name> |
name is the wire's authoritative board identity, MANDATORY — see §6.4; maximally forgiving, like PING
|
VER |
no (2026-08-27) | — | ver <version> |
maximally forgiving, like PING
|
STATUS |
no (2026-08-27) | — | status ready=1 active=0 connL=1 connR=1 otos=0 wedge=0 flags=<hex> tlm=off next=<n> done=<n> reason=<tok> |
k=v, order not guaranteed, unknown keys ignored; done=/reason= are NEW 2026-08-27 (§8.7) and read fresh off the Adapter at format time |
HELP |
no (2026-08-27) | — | help HELLO PING ID VER STATUS HELP GET SET TLM WHEELS_X WHEELS_V MOVE_X MOVE_V GO_TO_R GO_TO_W STOP ESTOP FUNCS RUN |
rest-of-line; generated from the same table dispatch() uses, so it cannot drift |
GET |
yes | [name] #id |
get name value (one field) or one get line per field (bare GET) |
unknown name → err 1 #<id> alongside the ack (2026-08-27) — symmetric with SET; was a silent no-get-line answer through 2026-08-26, see §9.11 |
SET |
yes | name value #id |
— (accepted: none; rejected: err <code> #<id>) |
an in-order ack is the acceptance; a value that fails to PARSE is a decode failure (§8.9), not a rejection |
TLM |
yes | mode #id |
— |
OFF/POSE/FULL/NOW/AUTO/BUFFER decoded; an unrecognized mode token is a decode failure (§8.9); the adapter's own Result never surfaces on the wire |
WHEELS_X |
yes | left right cruise timeout #id |
— | per-wheel commanded DISTANCE, bounded by encoder travel + timeout (motion-api.md §3.1); kUnknown on DiffDriveAdapter (§5) |
WHEELS_V |
yes | left right duration #id |
— |
renamed from WHEELS (2026-08-22) — maps onto drive() with no planner (§5) |
MOVE_X |
yes | distance rotation cruise timeout #id |
— | body displacement + heading (motion-api.md §3.3); kUnknown on DiffDriveAdapter
|
MOVE_V |
yes | v_x omega duration #id |
— | body twist held for duration (motion-api.md §3.4); kUnknown on DiffDriveAdapter
|
GO_TO_R |
yes | x y speed arrive timeout #id |
— | drive to a robot-frame point (motion-api.md §3.5); kUnknown on DiffDriveAdapter
|
GO_TO_W |
yes | x y speed arrive timeout #id |
— | drive to a world-frame point (motion-api.md §3.6); kUnknown on DiffDriveAdapter
|
STOP |
yes | [now] #id |
— (accepted: none; rejected: err <code> #<id>) |
the optional now token (motion-api.md §3.7/§9.1) sits safely before the id since the id is self-marking; see §5.1 |
ESTOP |
no | — | estop |
never sequenced, never nacked, maximally forgiving; see §5.1/§8.3 |
FUNCS |
yes | — | one funcs <name> [<signature>] line per registered function |
enumerates the adapter's RUN registry; zero lines is a valid answer — see §6.5 |
RUN |
yes | function [arg...] #id |
ret <value> #<id> (accepted, function returned a value) / — (accepted, void) / err <code> #<id> (rejected) |
invocation by name; see §6.3 |
| — | — | — | debug <text> |
robot-to-host ONLY, no inbound wire form; see §6.2 |
| — | — | — |
ack <n> <lastDone> <reason> / nack <n> <lastDone> <reason>
|
transport layer; the <reason> token is NEW 2026-08-22 — see §8.8 |
Angles (rotation, omega) are milliradian integers on the wire
(motion-api.md §9.1) — degrees only exist at a language binding, which is
not this library's concern; the handler decodes them with the ordinary
signed-integer field parser, the same as any other field.
WHEELS_X and the four body/positional verbs (MOVE_X/MOVE_V/
GO_TO_R/GO_TO_W) have no prior wire form before 2026-08-22 — they were
motion-api.md's own proposal, now implemented at the wire/handler layer.
SEED/CAL remain deferred: they need OTOS/odometry this library does not
own.
Rewritten 2026-08-21, updated 2026-08-22 — see §8 for the full design.
ok is deleted: acceptance is signaled by the transport-layer ack
alone. done is deleted as a standalone verb: the lastDone/<reason>
pair carried by every ack/nack is the completion channel.
| reply | meaning |
|---|---|
ack <n> <lastDone> <reason> |
transport: the highest in-order id accepted so far arrived correctly (§8.1). <reason> (NEW 2026-08-22, §8.8) describes lastDone — none when lastDone is 0 |
nack <n> <lastDone> <reason> |
transport: n is the next id the robot actually needs — either a numeric gap, OR (2026-08-22, §8.9) an in-order id whose own content was a DECODE FAILURE and so was never accepted at all |
err <code> #<id> |
application: a command's content was rejected — either a MERITS rejection (arrived intact, the adapter's own Result refused it, §4) or the "content" half of a decode failure (§8.9). Field order: the id is always the LAST token, matching every other line in this grammar (§8.6). |
estop |
ESTOP only — confirms the stop executed (§8.3) |
ret <value> #<id> |
RUN only (§6.3) — the invoked function returned a value, emitted IN ADDITION to the ack
|
funcs <name> [<signature>] |
FUNCS only (§6.5) — one line per registered function, emitted IN ADDITION to the ack; zero lines when the registry is empty |
A command is never "just" accepted or "just" rejected in isolation — but which TRANSPORT reply it gets now depends on which of two kinds of rejection it is (§8.9, the central 2026-08-22 change):
-
Decoded fine, refused on merit (e.g. an out-of-range speed):
ack(it arrived, the sequence advances) pluserr <code> #<id>. -
Decode failure (unknown verb, wrong arity, unparseable field — the
line did not arrive intact):
nack <n> <lastDone> <reason>(the sequence does NOT advance —nnames the SAME id that just failed) pluserr <code> #<id>.
| code | name | meaning |
|---|---|---|
| 1 | ERR_UNKNOWN |
no such verb or field name |
| 2 | ERR_BADARG |
malformed/non-finite argument, wrong arity |
| 3 | ERR_RANGE |
declared bound violated |
| 4 | ERR_FULL |
queue full |
| 6 | ERR_UNIMPLEMENTED |
recognized, not wired on this build |
| 8 | ERR_NOT_CONFIGURED |
refused pre-ready
|
| 10 | ERR_BUSY |
subsystem in motion; retry at rest |
Code 11 (ERR_DUPLICATE_ID) is deleted, not merely unreachable, as of
2026-08-22 — see §2.2/§9.8.
Wire shape: debug <free text>, lowercase, a rest-of-line verb exactly
like HELP's own reply — everything after the first space is one field.
ProtocolHandler::sendDebug(const char* text) is the only way this line is
ever emitted; it is an unsolicited emission the application drives
(alongside sendBanner()/sendReady()), never a reply to an inbound
command.
The host never sends it, and there is no inbound wire form to reject.
Because it is lowercase, an inbound debug ... line is dropped by the same
mechanism every other lowercase-led line is (§2.1) — silently, and not
counted malformed. This is the structural fix for the v5 DBG:-flood
incident (.claude/rules/hardware-bench-testing.md in the robot repo this
library was extracted from): under v5 a robot's own debug output was a
syntactically valid command to every other robot on a shared channel, and
the flood was self-sustaining. Under this grammar a debug line can never
parse as a command, closing that class structurally rather than by keeping
the channel private.
Sanitization: strip, don't reject. text is arbitrary and must not be
able to forge a second line — '\n' and '\r' bytes are stripped before
they can reach the sink. The alternative (reject the whole call) was
considered and rejected: sendDebug() is void, with no channel to report
a rejection back through, so discarding the entire message over one bad
byte would lose strictly more information than delivering everything else
in it. The whole line (debug + space + text + terminator) is also
truncated, never overflowed, to fit the 240-byte cap — the same posture
feed() itself takes on an overlong inbound line (§3.1).
sendDebug("") and sendDebug(nullptr) are the same case. Both emit
the bare line debug\n — no trailing space before the terminator. A text
that sanitizes down to nothing (e.g. it was entirely '\n'/'\r' bytes)
collapses onto this same bare shape, rather than leaving a dangling
separator space (debug \n) that no other field-less reply in this file
ever produces — consistent with the grammar's own "an empty token cannot
exist between spaces" rule.
Wire shape: RUN <function> [arg...] #id — id is mandatory as of
2026-08-21 (§8), same as every other sequenced verb.
Division of responsibility — this is the important design decision, and
it must not move. The handler parses and nothing else: it extracts the
function name and the remaining raw argument tokens as const char*
pointers into its own line buffer, and hands them to
Adapter::onRun(name, argv, argc, result, resultCapacity, hasResult). The
handler holds no function table, does no name resolution, and does no type
conversion — the same "the handler holds no tables" property that makes
GET/SET pure delegation (§7).
The adapter owns resolution, type conversion, and invocation. In
MicroPython or JavaScript that is globals()[name] plus argument
introspection — nearly free. In C++ there is no lookup-by-name and no
parameter-type reflection, so a C++ adapter needs an explicit
registration table declaring each function's name, arity, and per-argument
types to implement onRun() at all. Say this plainly to any porter: RUN
is the first verb where the C++ archetype does substantially more work
than the dynamic ports, not less — a porter reading this handler should not
conclude that registration machinery is itself part of the wire contract.
This library's own DiffDriveAdapter registers nothing and answers every
RUN with ERR_UNKNOWN, which is the correct behavior for an adapter with
an empty allowlist, not a stub left unfinished.
The registration table IS the security boundary. Whatever a concrete
adapter registers is invocable by name from the wire by anything that can
talk to the robot, including any other host on a shared radio channel.
Treat the table as an explicit allowlist, not an implementation detail —
a function should be registered because it is meant to be remotely
callable, not because it happened to be convenient to expose. FUNCS
(§6.5) makes that table discoverable, which sharpens this rather than
softening it: registering a name now publishes it to everything that
can hear the channel, rather than merely making it guessable.
Replies — ret is a lowercase reply verb, always carrying the
(now-mandatory) id:
| outcome | reply |
|---|---|
| function returned a value |
ret <value> #<id> — emitted IN ADDITION to the ack (§8.2) |
| function returned nothing (void) | nothing beyond the ack
|
| unknown function |
ack + err 1 #<id> (ERR_UNKNOWN) |
| wrong arity, or an argument that will not convert |
ack + err 2 #<id> (ERR_BADARG) |
RUN with no function name at all |
malformed (§8.4) — if the id itself is present and well-formed but consumes the only field (RUN #7), still ack + err 2 #<id>; if there is no field at all (bare RUN), no reply of any kind |
#0 no longer suppresses anything (2026-08-21). The old "#0 means no
ack wanted, run silently" rule is deleted along with #0 itself (§2.2) —
since ids start at 1, an inbound #0 is always < expectedNext_ and is
therefore always treated as a stale retransmit: the function is never
invoked, and the reply is the ordinary retransmit ack (§8.1), not
silence. There is no longer any way to run a RUN (or any other verb)
without a reply.
The id is unconditionally the line's last token now — no content
inspection needed. Before 2026-08-21, RUN's open arity meant the
handler had to look at whether the last field's content began with '#'
to decide whether an id was even present. That branch is gone: every
sequenced verb's id is mandatory and is stripped from the line by the same
central step (§8.4), before any verb-specific parsing runs at all, so
RUN's own handler never has to make that decision — it only ever sees
the function name and its real arguments, with the id already resolved.
The one genuine expressiveness limit this leaves is unchanged in spirit and
if anything sharper now: a function's own final argument can never begin
with '#', because that token position is unconditionally the id, not
merely "usually" the id. A function needing a literal '#'-led value as
its logically-last argument cannot be called that way at all — it has to
take that value as a non-final argument, or the caller reorders the call.
Two limitations, not implementation gaps:
-
An argument cannot contain a space. The grammar makes a space the
field separator, so string arguments are single-token only. This is a
genuine constraint on what
RUNcan express, not an oversight. -
onRun()is called synchronously fromfeed(), so a slow registered function stalls line processing for as long as it runs. Registered functions must return promptly; anything long-running must be deferred by the calling application.
Sanitization of the return value. The adapter's own result string is
sanitized by the handler exactly like debug's text — '\n'/'\r'
stripped, truncated to fit the line cap — before it reaches the sink. A
concrete onRun() does not need to pre-sanitize its own output; the
handler treats it as untrusted content regardless, the same way it treats
every other free-form string it ever formats onto the wire.
Added 2026-08-27. ID's reply is id <drivetrain> <profile> <version> <name>. Fields 0-2 are unchanged, so the extension is strictly
additive and a parser reading three fields is unaffected — but name
is MANDATORY, not optional.
profile is build provenance: it is baked into the hex at deploy time.
It is not identity, and treating it as identity has already failed in the
field — a fleet-wide incident had every robot on the playfield reporting
the same profile, because that was the profile in the image they were all
flashed from. name is sourced from the board itself
(microbit_friendly_name()), which makes it the wire's only
authoritative answer to "which robot am I talking to."
Why mandatory rather than optional: a host cannot treat a field as
authoritative identity if the wire says it might be absent. It would have
to fall back to profile exactly when name is missing — reproducing the
original incident, more rarely and much harder to spot. An optional
identity field is worse than none, because it looks like a guarantee.
Identity.name was already plumbed for HELLO's banner before this
change; execId() simply did not read it.
Added 2026-09-07 (stakeholder-directed). Wire shape: FUNCS #id,
no data fields. Reply: one funcs <name> [<signature>] line per
registered function, emitted in addition to the ack every sequenced
verb gets.
FUNCS #7
funcs add int,int->int
funcs blink int->void
funcs ping ->int
ack 7 0 none
§6.3 gave RUN a registration table and called it the security
boundary, then left it opaque: a host could only discover a callable
name by guessing it and reading the ERR_UNKNOWN. FUNCS closes that
— it is to RUN what HELP is to the verb table.
The handler still holds no function table. FUNCS walks the
Adapter's own runCount() / runName(i) / runSignature(i), exactly
the way a bare GET walks fieldCount() / fieldName(i) (§7). It
discloses the adapter's registration table; it does not keep one, and
the §6.3 division of responsibility is unchanged.
One line per function, not one rest-of-line reply. HELP's single
help A B C … line was the obvious alternative and is wrong here for
two reasons. HELP lists a table that is fixed at compile time and
known to fit; a registration table's size is the concrete adapter's
business, and one that outgrew the 240-byte line cap (§3.1) would be
silently truncated — a host would read a short allowlist as the whole
allowlist, which is the one failure mode an allowlist must never have.
And a rest-of-line reply has nowhere to put a per-function signature
without inventing an intra-field encoding the grammar does not have.
The ack is the terminator, which is why FUNCS is sequenced. A
variable number of reply lines needs some way for the host to know it
has them all, and this grammar already has exactly one such mechanism
in use: bare GET's dump, which ends when the ack for that id
arrives. FUNCS reuses it rather than inventing a sentinel line or a
count prefix. That is a deliberate departure from the other pure
queries — HELP/ID/VER/STATUS went unsequenced on 2026-08-27
(§8.3) precisely because each answers in exactly one line and so needs
no terminator. FUNCS does not have that property.
The funcs lines themselves carry no #<id>, again following bare
GET's get name value rather than RUN's ret <value> #<id>: the
id-bearing line in this exchange is the ack that closes it.
An empty registry emits nothing at all — just the ack. This is the
wire-visible form of an empty allowlist, not an error and not an
unimplemented stub. It is the correct and expected answer from any
adapter that owns no callable surface, including this library's own
DiffDriveAdapter (§5), whose RUN already answers every name with
ERR_UNKNOWN. A host reads "nothing is callable here" from the absence
of funcs lines between its command and its ack.
The signature is optional and its format is unspecified. It is a
single token — the grammar makes a space the field separator, so it can
be no more than that — and int,int->int is a readable convention, not
a contract. It is advisory text for a human or a tool; nothing in this
protocol parses it, and a porter is free to emit a different convention
or none. An adapter declaring no signature for an entry omits the field
entirely (funcs blink), rather than emitting an empty token, which
this grammar has no spelling for.
Registered names must be single tokens too, for the harder reason
that a name containing a space could never be addressed by RUN in
the first place — the same expressiveness limit §6.3 already documents
for RUN's arguments.
Both strings are adapter-supplied free-form text and are sanitized,
not trusted. '\n'/'\r' are stripped and the line is truncated to
fit the cap, exactly as debug's text (§6.2) and RUN's returned value
(§6.3) are — an adapter cannot forge a second wire line through a
registered name any more than it can through a return value. An entry
whose name sanitizes down to nothing is skipped rather than emitted as
a bare funcs line, since a nameless entry is not addressable.
Arity: FUNCS takes no data fields. FUNCS x #1 is a decode
failure (§8.9) — nack plus err 2, sequence not advanced — the same
answer ID/VER/STATUS/HELP give to a stray field, and it comes
from sharing their decodeNoFields.
Decision (stakeholder, 2026-08-20): neither the handler nor the kernel implements configuration storage. A configuration system may come later, as its own thing; it is not core work and it is not in this library.
What that means concretely:
-
No config field table lives in this library.
GET/SETare pure delegation. The handler parses the line, decodes the value, callsonGet/onSet, and formats the reply. It holds no field table, no bounds, no storage. Which names are valid is entirely the adapter's business, and an unknown SET name is justerr 1 #<id>coming back from the adapter, layered on top of the ack every in-orderSETgets regardless (§8.2). As of 2026-08-27GETbehaves the same way — an unknown GET name also produceserr 1 #<id>alongside its ack. Through 2026-08-26 it produced nothing beyond the ack, an asymmetry withSETthat this file documented without ever justifying; see §9.11. -
Each library carries only the configuration it needs, as its own type.
DiffDrive already has this:
DifferentialDrive::Configplus the fluent setters, holding gains, limits, and the cycle period.DiffDriveAdaptermaps 15 wire names 1:1 ontoConfig's members — the whole field table is a name/member-pointer pair per row, indiffdrive_adapter.cpp. -
maxDuty,fullDutyVelocity,cyclePeriodare hard-coded, not wired.DifferentialDrive::begin()needs values for these to leave its fail-closed default, but they are not tuning gains — stakeholder decision, 2026-08-20: "I don't see that max duty, full duty velocity, and cycle period need to be configurable, so you can just hard code them." They are build-time constants onDiffDriveAdapter(kMaxDuty/kFullDutyVelocity/kCyclePeriod), applied to the kernel at adapter construction, so building aDiffDriveAdapteralone is sufficient forbegin()to succeed with no external arming step. -
The robot's geometry is not in either library.
countsPerLength(§5) belongs to the adapter, because it is a property of a particular robot and neither a wheel control law nor a line parser is.
Stakeholder design, 2026-08-21, replacing §8's old "should WHEELS emit
done?" question outright (the old text is preserved as §9.9's own
changelog entry, not repeated here). Where §2 through §7 describe the wire
grammar and per-verb payloads, this section describes a layer that sits
underneath every one of them: what it means for a command to arrive, as
opposed to merely being well-formed.
Before this change, the protocol had no delivery guarantee at all. §2.1
(pre-2026-08-21) promised a 3×-repeat of every id-carrying reply "on
consecutive cycles" so an outcome would survive packet loss — but nothing in
ProtocolHandler ever implemented it (§9.2, now historical), because doing
so honestly needs a periodic entry point and a notion of time, which the
handler deliberately does not have. Measured loss on the radio link this
protocol targets is real and nothing in this library protected against
it.
The "~5%" figure this paragraph carried through 2026-08-26 is
withdrawn (2026-08-27). It cited .clasi/knowledge/ in the robot repo
this library was extracted from, and no measurement since has come near
it. Three instrumented runs on ch4 that day put per-line delivery at
66.5% (n=200), 75.0% (n=240) and 83.3% (n=60) — 17-33% loss — against a
wired control of 99.5% (n=200) on the same firmware and the same verbs,
which is what establishes the loss as the RF path rather than this stack.
No replacement figure is stated, deliberately. Those runs span 17 points inside about two hours, so the link is not merely worse than the old number, it is unstable, and its true rate is uncharacterised. Substituting "~75%" would repeat the original mistake — quoting a point estimate from a small sample as a settled property of the link. What this design can honestly assume is a lossy, time-varying link an order of magnitude worse than the withdrawn figure. The scheme below works at those rates; its rationale simply should not lean on a number nobody has demonstrated. See §9.11. The 3×-repeat idea is deleted, not deferred — this section is its full replacement, not an addition alongside it.
Every command carries a mandatory sequential id, #<n>, starting at 1.
The host may pipeline freely — it never has to wait for an ack before
sending its next command. The robot acknowledges cumulatively: one ack
covers every earlier id too, which is what lets the scheme survive loss
with no ring, no per-id storage, and no eviction policy — the entire
receiver-side state is a single number.
Handler state, in full (2026-08-22: lastDone_ is GONE from this
list — see §8.8; 2026-08-26: gapOutstanding_ was removed — see §8.5;
2026-08-27: it is BACK, scoped, as a reply predicate — see §8.3's
reminder rule and §9.11):
uint32_t expectedNext_ = 1; // next sequence id expected from the host
bool gapOutstanding_ = false; // a gap or decode-failure stall is opengapOutstanding_ is set on a gap (the table's third row below) and on a
decode-failure stall (§8.9), cleared when the missing id finally arrives
in order, and cleared by HELLO. It cannot be derived from
expectedNext_, which is why it has to be stored: expectedNext_ = 5
alone cannot distinguish "clean, waiting for #5 the host has not sent
yet" from "stalled, #6 was discarded, still waiting for #5."
Its scope is narrow, deliberately: it is a predicate on a REPLY. Its
only reader is §8.3's conditional reminder — whether an inbound
unsequenced line's reply carries a trailing nack. It restores no
periodic emission, no telemetry piggyback, and no beacon of any kind;
§8.5's rule and its invariant are untouched.
Deliberately, there is no tick() and no clock anywhere in this list.
There is no periodic half to the scheme at all (2026-08-26, §8.5):
feed() is the only origin of every ack/nack, each one a direct
reply to an inbound line. Keeping the handler clock-free is load-bearing,
not incidental: it is the property that lets feed() stay a pure
function of its input bytes plus this small, explicit state, with
nothing owed later "on its own."
Every inbound id, for every sequenced verb (§8.3's exemption set is the only carve-out), is classified into exactly one of three cases:
| inbound id | action | reply |
|---|---|---|
== expectedNext_ |
decode the verb's own fields FIRST (§8.9); only if that succeeds, dispatch to the adapter and expectedNext_ = id + 1
|
ack <id> <lastDone> <reason> on a decode success; nack <expectedNext_> <lastDone> <reason> on a decode failure (§8.9) |
< expectedNext_ |
do NOT re-execute — a retransmit whose ack was lost | ack <expectedNext_ - 1> <lastDone> <reason> |
> expectedNext_ |
discard, do NOT execute — a numeric gap | nack <expectedNext_> <lastDone> <reason> |
<lastDone>/<reason> are read fresh off Adapter::lastDone()/
lastDoneReason() (§8.8) every time either reply is formatted — this
table's middle column used to read a handler-owned lastDone_ field
before 2026-08-22.
The middle row is the one easy to get wrong: a resent WHEELS_V (the host
never saw the first ack, so it resends) must not drive the wheels a
second time. The reply for a retransmit echoes the already-accepted
id (expectedNext_ - 1), not the resent one — telling the host "I already
have everything through here," which is exactly what a resend needs to
hear to stop resending.
A gap stalls the stream on purpose: every subsequent command, however
well-formed, is discarded and nacked until the missing id arrives, giving
strict in-order delivery. Because every new command re-triggers the same
nack <expectedNext_> ..., a lost nack self-heals — the host will see
the next one along with the next command it sends. A decode failure on
an in-order id (§8.9) holds the stream exactly the same way — the failed
id itself becomes the thing every subsequent nack keeps asking for, until
a well-formed line finally supplies it.
nack carries next-expected, not "last good id": it tells the host
exactly what to resend with no +1 inference on either side, and it avoids
overloading 0 as both "nothing accepted yet" and "resend from here."
Updated 2026-08-22 — see §8.9 for the change this section's own
"unknown verb" paragraph below is superseded by. ack/nack answers
one question only: did the bytes arrive, in order, INTACT? err
answers a different one: was the content accepted, once it was known to
have arrived? A message can be perfectly in-order, decode fine, and
still be garbage on the merits (WHEELS_V 99999 0 100 #7 decodes cleanly
and is rejected on range by the adapter). So an in-order command the
ADAPTER rejects on merit emits both: the ack (it arrived, decoded,
the sequence advances) and err <code> #<id> (§6.1) — two lines. The
error code is never folded into ack itself; that would conflate a
transport signal with an application one.
This does NOT generalize to a handler-level decode failure any more (2026-08-22, reversing the pre-2026-08-22 text this paragraph used to carry). An unknown verb, a known verb with the wrong field count, or an unparseable field is a DIFFERENT case from a merits rejection — the line did not "arrive fine", so it NACKs (§8.9), not acks. See §8.9 for the full story and the stakeholder's own rationale for keeping the two distinct.
Every reply so far bare ok/err gains a mandatory id, and ok itself
is gone. SET/WHEELS_V/STOP/RUN(void) success now produces
nothing beyond the ack — the ack is the acceptance signal. RUN's
ret is the one exception: a returned value is genuinely new information
the ack alone cannot carry, so it is still emitted, in addition to
the ack, not instead of it (§6.3).
Sequenced: GET SET TLM WHEELS_X WHEELS_V MOVE_X MOVE_V GO_TO_R GO_TO_W STOP RUN.
Unsequenced: ESTOP, HELLO, (2026-08-22) PING, and (2026-08-27)
HELP, ID, VER, STATUS.
Through 2026-08-26 this set was a list of three structural exceptions with no general principle behind it. It now has one, and the list follows from it rather than the other way round:
A verb is sequenced iff its correctness depends on its position in the stream — either because executing it twice changes the robot, or because answering it out of order yields a wrong answer.
Sequencing a verb costs a mandatory id, a silent drop when the id is absent, and a stale re-ack that answers nothing when it is resent. That price buys delivery ordering, which is worth paying only where ordering is part of being correct.
-
State-changing verbs (
SET TLM STOP RUN WHEELS_* MOVE_* GO_TO_*) are sequenced on the first clause: executing one twice moves the robot twice. -
GETis sequenced on the SECOND clause, and is NOT an exception to the rule. It is read-only, but it is order-dependent: aGETracing a pendingSETreturns the pre-SETvalue with nothing marking it stale, which is a silently wrong answer to a config question. The reordering half of the rule covers it exactly. Do not document this as a stakeholder-directed carve-out — it is the rule working, and the weaker framing invites someone to relitigate it. -
ID/VER/HELPanswer session constants — identity, firmware version, and a verb list generated from a compile-time table. Nothing in the config plane can reach them (onSet's value parameter is afloatand every settable field is a float member;Identity's fields areconst char*and are not in that table), so no ordering can change the answer. -
STATUSanswers live physical state plusnext=. Physical state is time-valued, not order-valued — it changes whether or not the host sends anything — so ordering buys nothing. And see §8.7: sequencing it made its own documented resync job impossible.
The stakeholder's own framing was "every message must have an ID number" —
PING's own exemption is a LATER, explicit stakeholder direction
(2026-08-22, verbatim: "ESTOP, ping, and HELLO shouldn't require IDs"),
not this file's own call the way ESTOP/HELLO's exemption originally
was. All three are exempted because the scheme is structurally
unbootstrappable and unsafe without them:
-
HELLOresets the sequence.expectedNext_ = 1— the handler's entire reliability state since 2026-08-26 (§8.5) — then the banner is emitted; this is the session-start resync a host performs on (re)connect. It does NOT reset the Adapter's ownlastDone()/lastDoneReason()any more (2026-08-22, §8.8) — that state moved off this class entirely, and a handler-level reset has no business reaching into the Adapter to clear something it does not own. A verb that resets the sequence cannot itself be inside the sequence without a chicken-and-egg problem (what id would the very firstHELLOcarry, and against what would it be checked?).HELLO's own arity is unchanged (zero fields) — it does not accept a trailing id at all, and aHELLOwith one is wrong arity, same as any other extra field. -
ESTOPis outside the sequence entirely: no id, never sequenced, never nacked. This is safety-critical: if#4goes missing and the stream is stalled waiting for it, anESTOPsent as#5(or with no id at all) must still execute — it cannot be discarded as "out of order," and it cannot be made to wait behind the missing#4the way an ordinary sequenced command would. -
ESTOPis maximally forgiving. ANY line whose verb token isESTOPexecutes the stop, regardless of trailing junk or arity —ESTOP,ESTOP 1 2 3,ESTOP #5,ESTOP #abcall stop the robot. A panic stop must never be refused over a syntax nit. (A verb that isn't the literal tokenESTOP— e.g.ESTOPXYZ, no space — is a different, ordinary verb token under the tokenizer's own rules, not "ESTOP plus junk"; this forgiveness is about content AFTER the verb, not about fuzzy verb matching.) -
ESTOPREPLIES. Stakeholder, verbatim: "Agree about ESTOP, but if it is not acked, it should be acknowledged, with anestopresponse." Executing the stop before writing the reply means a panic stop never queues behind an outbound reply.ESTOP's reply is the bare wordestop, no fields, ever. -
PINGis liveness and must answer even while the stream is stalled on a gap (2026-08-22, the stakeholder's own words) — the same structural reasonESTOPis exempted: a command gated behind a missing id cannot serve as a liveness probe for the very link that id is missing on.PING's own reply (pong <now>) is unchanged — it never carried anack/id of its own even before this change, so the exemption is purely about the SEQUENCE GATING, not the reply shape. -
Whether
PINGshould be maximally forgiving (likeESTOP) or strict zero-arity (likeHELLO) is THIS FILE'S OWN CALL, not spelled out by the stakeholder's direction — resolved forgiving (§9.8), so an old-style host still appending#<id>toPINGout of habit from before this change keeps working unchanged.
HELP/ID/VER/STATUS take PING's posture, not HELLO's: any
line whose verb token is one of them answers, whatever follows it. ID,
ID #1, ID #99, and ID junk are byte-identical.
This is not cosmetic. Per §9.8 item 7 a malformed unsequenced verb gets
no reply at all — there is no ack to anchor an err against. Strict
zero-arity would therefore make ID #1 wrong-arity and answer it with
silence, trading "dropped as stale" for "dropped as malformed": the same
symptom with a new cause, and it would break every host that still appends
an id out of habit.
Stakeholder direction, verbatim: "I don't actually mind if ID, VER, and help also return an ACK/NAK. What I mind is that they require an ACK/NAK … What I don't want is requiring IDs to have a sequence number, or for help to have a sequence number and just be able to issue those any time, but they can also trigger an ACK/NAK."
Sequence gating and reply emission are separable, and only the first was ever objected to. A verb can be issuable with no id and still carry a reliability line back. So:
When a gap or decode-failure stall is outstanding,
PING,HELP,ID,VER, andSTATUSemitnack <expectedNext_> <lastDone> <reason>AFTER their own reply. On a clean stream they emit nothing.
A reminder, not a receipt — it fires only when something is wrong. It is
the full three-field form (§6.1), read fresh off the Adapter at format
time like every other nack; a two-field abbreviation would break any
host parser written to the documented shape.
Excluded, and firmly:
-
ESTOP— its reply is the bare wordestop, no fields, ever, and it must never queue behind an outbound reply. A panic stop does not carry diagnostic freight; the safety rule wins. -
HELLO— it setsexpectedNext_ = 1and clearsgapOutstanding_, so nothing can be outstanding after it. A reminder there would report on state it just erased.
STATUS keeps the reminder even though next=/done=/reason= already
say the same thing in its own payload. Known and accepted redundancy:
"every unsequenced verb except ESTOP and HELLO carries the reminder"
is a rule that fits in one's head, and "…except STATUS, which tells you
the same thing a different way" is not.
-
PING— is it alive? Cheapest, shortest reply. -
STATUS— alive, and where does the sequence stand?next=,done=,reason=,ready=in one line. This is the diagnostic. -
HELLO— start over. It resets the sequence.
HELLO is not a health check. Firing it at a live session sets the
robot's counter to 1 while the host's stands at N; the host's next command
reads as a numeric gap, and since every already-acked id has been retired
from the host's pending buffer there may be nothing left to resend — the
robot wants #1 forever. A probe that manufactures the wedge it was
checking for. Use HELLO only where losing the sequence is the intent.
This set is deliberately narrow — flagged here prominently so the stakeholder can find and overrule it easily: everything else in this library's scope is sequenced, and the rule above is what decides it.
This whole section describes the 2026-08-21 design, replaced by §8.9's
2026-08-22 rewrite. Preserved for the changelog record, the way §8.0/
§9.2 preserve the eras before them. The pre-2026-08-21 malformed-line
recovery rule (§2.3, historical) is gone. In its place, for any line whose
verb is neither ESTOP nor HELLO (§8.3) and does not start lowercase
(§2.1, unchanged):
- No trailing field at all (the line was just the verb) → malformed, no reply. There is nothing to sequence-check.
-
A trailing field is present but is not a well-formed
#[0-9]+(missing#, non-digit content,#+5, digit overflow) → malformed, no reply. Same reasoning: nothing valid to compare againstexpectedNext_. -
A well-formed id is present → classify it via §8.1's three-way
table, using only the id — the verb's own identity and field
content are not even inspected yet:
- out of order (
<or>expectedNext_) →ack/nackper §8.1, and nothing else is examined or executed — not even whether the verb is recognized. A stalled stream costs nothing but the id comparison itself. - in order (
== expectedNext_) → the sequence advances and theackis sent UNCONDITIONALLY; then, and only then, the verb is looked up and its fields validated. (This is exactly the step §8.9 changes: as of 2026-08-22, decoding happens BEFORE the ack/nack decision, not after.) An unrecognized verb, wrong field count, or an unparseable field at this point behaved exactly like an adapter-level rejection (§8.2, its own pre-2026-08-22 text):err <code> #<id>followed theack, usingERR_UNKNOWN(1) for an unrecognized verb orERR_BADARG(2) for anything else.
- out of order (
This was a strictly cleaner story than the id-recovery rule it replaced (no
verb-specific carve-out to remember, since ESTOP was excluded at the top
by verb identity) — but it could not tell a garbled square-tour leg apart
from a merits-rejected one on the wire, which is exactly the gap §8.9
closes.
Stakeholder direction, verbatim: "An ack or a nack is only a
response to a message, not a beacon. I do not want to have my
connections littered with 5 acks or nacks a second." Observed live on
an idle bridge session: < ack 0 0 none arriving at the telemetry
cadence with no command ever sent.
Through 2026-08-25 this section said the opposite — emitTelemetry()
appended the current reliability line (nack <expectedNext_> <lastDone> <reason> if a gap was outstanding, ack <expectedNext_ - 1> <lastDone> <reason> otherwise) to every telemetry frame, on the theory that the
loss-survival argument needed ack/nack arriving regularly, not
only in reply. That theory bought one thing (a quiet host eventually
heard the robot's state anyway) and cost the wire an unsolicited ack
several times a second, forever. Deleted, not moved: no periodic,
beacon, or telemetry-carried ack/nack of any kind exists any more.
The invariant, which is what actually does the work: every emission
originates in feed(), and the handler holds no clock and no periodic
entry point. Periodicity is therefore structurally impossible, not
merely prohibited — checkable by reading the call graph rather than by
trusting a rule. It is stated first because it is stronger than the
sentence that follows, and because it is what makes §8.1's
gapOutstanding_ safe to hold: a predicate on a reply cannot become a
beacon when nothing but inbound bytes can reach an emitter.
The rule itself is one sentence: an ack/nack line is emitted only
in direct response to an inbound line — the three rows of §8.1's table,
§8.9's decode-failure path, and (2026-08-27) §8.3's conditional reminder
on an unsequenced verb — and nothing else in the handler ever emits one.
emitTelemetry() emits thdr/t frames only.
Widened 2026-08-27 from "an inbound sequenced line" to "an inbound line." Through 2026-08-26 this section was read as forbidding any reliability line from an unsequenced verb. It never said that: the objection recorded above is to periodicity, and a line replying to an inbound unsequenced verb is still a reply, not a beacon. An idle connection stays exactly as silent as it already was. See §8.3.
Loss recovery is host-driven, which §8.9 already required the host to be capable of anyway (the host must own its give-up/retry path — this library has no clock and structurally cannot own it):
- A lost
ackheals on the host's own resend: §8.1's< expectedNext_row re-acks without re-executing. - A lost
nackheals because every subsequent inbound line — however well-formed — re-triggersnack <expectedNext_>while the gap stands (§8.1). A gap still stalls the stream on purpose; it re-nacks per inbound line now, not per telemetry frame. - A host that goes quiet after its last command and wants confirmation
reads
STATUS, unsequenced as of 2026-08-27, which reportsnext=,done=andreason=in its own payload (§8.7) — no ack needed, no sequence id consumed, and it answers while the stream is stalled. -
A host with a lost command and nothing further to send probes by
RESENDING ITS OLDEST STILL-PENDING COMMAND, not by putting a fresh
sequenced line on the wire to shake a
nackloose (stakeholder decision, 2026-08-27). §8.1's middle row is built for exactly this: if the command did arrive, it re-acks without re-executing; if it was the lost one, the stream advances. The rejected alternative consumed a sequence id per probe and, on a stalled stream, was itself just one more line the robot discarded. This closes a hazard the change would otherwise have opened: a host's only automatic retransmit trigger is anack, and anackonly ever answers a sequenced line — so unsequencingSTATUSwithout this would have left a quiet host sitting on a lost command indefinitely, indistinguishable from a dead robot.
The no-timer/no-clock property §8.1 insists on is untouched — now
trivially, since there is no periodic half left to schedule. This
deletion also removed gapOutstanding_ from the handler state (§8.1):
its only reader was emitTelemetry()'s ack-vs-nack choice, and §8.1's
table decides every reply from the inbound id alone. Partially reversed
2026-08-27 — the flag is back, but as a predicate on a reply (§8.3's
conditional reminder), never as the trigger for a periodic emission.
Nothing about the deletion of the piggyback itself is undone.
This subsection (as it stood through 2026-08-21) described lastDone_ as
handler-owned state, "plumbed, not wired" in this library because
WHEELS had no completion event of its own. §8.8 below is the full
2026-08-22 replacement: the field moved OFF the handler entirely, onto
Adapter, and the "plumbed but inert" story now applies specifically to
DiffDriveAdapter, not to every adapter this library can host — a
step()-driven test adapter (tests/protocol/fake_motion_adapter.h) makes
it genuinely live for the first time.
§2.2 (pre-2026-08-21) already stated the invariant "an id is always the
LAST token of its line, commands and replies alike" — but replyErr()
did not follow its own rule: it emitted err #<id> <code>, id first, code
last, undocumented as an exception. Nothing broke from this before now
because this library's own robot side never parses its own replies — but
an archetype (§9.4) must not carry an undocumented exception to its own
stated invariant, because a host parser written to that invariant
uniformly would silently mis-parse every err line. Fixed: err <code> #<id>. The bare form (no id — which no longer exists at all now that
every err implies a prior ack for the same mandatory id) is retired
along with it.
status gains a next=<expectedNext_> key (§6) so a reconnecting host
can resync its own tracking without forcing a full HELLO reset.
This did not work through 2026-08-26, and the reason is worth stating
plainly: STATUS was itself sequenced. A host that has lost sequence
tracking cannot choose an id for it. Guess low and it lands in §8.1's
stale row — re-acked against expectedNext_ - 1, with no status line
emitted at all. Guess high and it opens a gap, stalling the very stream
it was sent to diagnose. The one verb whose documented purpose is
recovering from desync was gated behind not being desynced. Unsequencing
it (2026-08-27, §8.3) is what makes this section true rather than
aspirational. As of
2026-08-22 (§8.8), a HELLO reset no longer clears any completion state at
all — lastDone()/lastDoneReason() live on the Adapter and are
untouched by the handler's own reset — so the original motivation for this
key ("useful because a HELLO reset also clears lastDone_") is narrower
than it was, but next= remains useful on its own merits (resync without
re-establishing the session). status now also reports done=<lastDone> and reason=<tok>
(2026-08-27), closing what §9.8 flagged as a gap rather than a
considered omission. Both are read fresh off Adapter::lastDone()/
lastDoneReason() at format time, never cached (§8.8), exactly as
replyAck() reads them.
This was a prerequisite, not a nicety. Since §8.5 deleted the
telemetry piggyback, (lastDone, reason) only ever rides a direct
reply — so a host awaiting a completion has to provoke one, and the host
in this repo did so with a sequenced STATUS, purely to draw the ack
that carries the pair. Unsequencing STATUS removes that ack. Landing
these two keys FIRST is what keeps completion delivery working across the
change; done in the other order it would have broken silently.
Because status is k=v with unknown keys ignored (§6), both keys are
additive and a host written before this change is unaffected.
Stakeholder direction, verbatim: "[lastDone] is currently a handler
counter that nothing ever writes, so it is permanently 0 — wire-correct
and inert. Replace it with a virtual on Adapter that the handler
POLLS whenever it formats an ack/nack."
// Most recently completed motion, for the reliability piggyback. Monotonic:
// a later value implies every earlier one completed (motion runs one at a
// time, in order), which is what makes a dropped completion recoverable.
virtual uint32_t lastDone() const = 0;
virtual DoneReason lastDoneReason() const = 0; // see §8.8.1This removes handler state entirely (ProtocolHandler now carries only
expectedNext_, §8.1 — at the time of this change it also carried
gapOutstanding_, itself deleted 2026-08-26, §8.5), needs no callback
and no clock,
and makes the field genuinely live once an Adapter that actually completes
motions drives it — which nothing in this library did before this change.
replyAck()/replyNack() call adapter_.lastDone()/lastDoneReason()
fresh every time they run; there is no cache.
A real, undecided design fork this file resolves on its own: the
brief settled the VALUE (lastDone) but not whether HELLO's own reset
should reach into the Adapter to clear it too — the pre-2026-08-22 text
had HELLO reset lastDone_ = 0 as part of the same call, back when it was
handler state. Resolved: no. HELLO's reset now only touches
expectedNext_ (and, until its 2026-08-26 removal, gapOutstanding_,
§8.5) — state the HANDLER owns. Reaching
across the seam to clear something the ADAPTER now owns would reintroduce
exactly the kind of handler-into-adapter coupling this whole move was
meant to avoid, and there is a real argument that a completed motion
SHOULD survive a reconnect (the host may be re-establishing a session
after a dropped link, not asking the robot to forget what it just did). An
adapter that wants HELLO to also clear its own completion state is free to
do so from wherever it observes HELLO itself.
8.8.1 DoneReason — the conflict the reliability-layer brief didn't settle, and this file's resolution
The reliability change deleted standalone done and collapsed it into a
piggybacked number (lastDone). But motion-api §5.3
defines FOUR completion reasons: stop, timeout, estop,
aborted. A bare number loses the reason, and estop vs stop is a
distinction that matters (a program that ended normally vs. one that got
panic-stopped mid-leg are very different facts for a host to learn).
Resolved by piggybacking the reason too: ack <n> <lastDone> <reason>
and nack <n> <lastDone> <reason>, where reason describes lastDone.
This keeps the loss-tolerance property (a later ack re-carries it) and
the reason vocabulary, at one extra token per ack/nack. none is the
reason when lastDone is 0 (nothing has completed yet).
enum class DoneReason : uint8_t {
kNone, kStop, kTimeout, kEstop, kAborted,
};This is this library's own resolution of a conflict the reliability- layer discussion did not settle — flagged prominently so the stakeholder can find and overrule it. The alternative shapes considered and rejected:
-
A separate reason-only reply, not piggybacked (e.g. a standalone
done #<id> <reason>line reintroduced) — rejected because it brings back exactly the "the handler must remember outstanding ids and emit a reply later, on its own" problem §8.0/§9.9 already killed the olddonedesign over: it would need its own delivery guarantee, which is the whole thing this scheme already provides forlastDoneitself. Piggybacking the reason onto the SAME line that already survives loss costs one token and buys nothing to re-engineer. -
Fold the reason into a wider
lastDoneencoding (e.g. high bits of a 32-bit value) — rejected as needlessly clever for a text protocol whose whole design philosophy (§2) is "every wire value is a base-10 ASCII integer" with no bit-packing anywhere else in the grammar.
A dependent design point, also resolved here: does WHEELS_X/
MOVE_X/MOVE_V/GO_TO_R/GO_TO_W on DiffDriveAdapter answering
kUnknown (§5) rather than kUnimplemented cost anything for
lastDone/DoneReason? No — a merits rejection never touches
lastDone/lastDoneReason at all; those two fields only move when an
Adapter's own motion genuinely COMPLETES, which never happens for a verb
that was refused outright.
Ids run 1 .. 999999 by host-side convention. Modular wraparound is
explicitly out of scope and not implemented. A session must reconnect
(HELLO) before exhausting the id space. Comparing </> on a wrapping
counter is a classic bug source, and the space (999999 sequential
commands before a reconnect) is large enough in practice that wraparound
handling buys nothing but risk. The handler does not enforce the
999999 ceiling itself — ids are ordinary uint32_ts compared with plain
integer </==/>, so nothing stops a host from counting past it, but
nothing in the design analysis above holds once it does; this is a
host-side discipline, not a wire-enforced limit.
Stakeholder, verbatim: "I think a decode failure is a NAK. The goal here is that the movements don't work if you put them out of order... If you're driving a square and you've got eight movements you send, and you lose a turn, the whole square is wrong. The best thing to do there is to NAK and resend from that point on, so we need to make decode failures be NAK and err."
This reverses §8.2/§8.4's own pre-2026-08-22 behavior, where a decode failure acked and advanced the sequence exactly like a merits rejection. Two distinct cases, kept distinct on the wire:
| case | meaning | sequence | reply |
|---|---|---|---|
| decode failure — unknown verb, bad arity, unparseable field | the line did not arrive intact | does NOT advance |
nack <expectedNext_> <lastDone> <reason> (naming the SAME id, unchanged) and err <code> #<id>
|
| adapter rejection — decoded fine, refused on merit (e.g. out-of-range speed) | arrived intact, refused | advances |
ack <id> <lastDone> <reason> and err <code> #<id>
|
The distinction is the whole point: a corrupted line should be resent (resending the merits-rejected line would just be refused again, identically, so THAT case still advances and moves on).
Where the decode now happens. dispatch() still resolves the
mandatory id first (§8.1's three-way compare, unaffected) — but for the
id == expectedNext_ case, it now looks up the verb AND decodes its own
fields (arity + per-field parseability) before sending any reply at
all. Only once decoding succeeds does the sequence advance and the
ack go out; a decode failure at this point (verb not found, wrong
field count, an unparseable numeric field, an unrecognized TLM mode,
STOP's trailing token being anything other than the literal now, a
bare RUN with no function name, or more raw RUN tokens than the
handler's fixed storage can hold) instead calls a dedicated path that
NACKs expectedNext_ (still equal to the failed id, since it was never
accepted); a stalled stream keeps re-nacking because every subsequent
inbound line re-triggers nack <expectedNext_> (§8.1) exactly like a
numeric gap would, until a well-formed line finally supplies the same id. malformedCount() still
increments for a decode failure, exactly as it did before this change.
What does NOT change: a numeric gap (id > expectedNext_) is still
classified and replied to WITHOUT ever looking up the verb at all — that
row of §8.1's table is untouched. A stale retransmit (id < expectedNext_)
is also untouched. Only the id == expectedNext_ row's own internal
order changed (decode, then reply — not reply, then decode).
The hazard this creates, stated plainly: because a decode failure
never advances the sequence, a host that genuinely CONSTRUCTS a
malformed line (a real bug on the host side, not packet corruption) will
be NACKed forever on that same id, and the stream wedges. This is a
deliberate tradeoff — fail loud and stall rather than silently continue a
broken sequence — but it means the host needs its own give-up path:
a resend limit, a timeout, or an operator-visible stall detector that
eventually reconnects (HELLO) or aborts the sequence rather than
retrying the same bad line forever. This library does not, and structurally
cannot, supply that give-up path itself — it has no clock (§8.1) and no
notion of "how many times has this been retried," and inventing one would
reintroduce exactly the timer/pending-state machinery §8.0/§8.5 keep out
on purpose. Say this to every porter and every host implementation: a
decode failure is not a transient condition the wire protocol resolves on
its own; it needs an application-level backstop.
Everything above is design. This section is what the implementation actually found, and it is the part to read before extending any of it.
The first bullet below (the malformed-line #id recovery rule) is
historical, fully superseded by §8.4 (2026-08-21) — preserved for the
record, not current behavior. The second and third bullets (the
5000 ms ceiling's ownership, and the id's stricter numeric grammar) are
still current and unaffected by the reliability layer.
The malformed-line #id recovery rule is verb-agnostic, with one
deliberate exception. §2.3 (historical)'s own words — "if the line's last token is a
well-formed nonzero #id, reply err #<id> <code>" — carry no carve-out for
a verb whose own grammar has no id concept at all (HELLO/PING/ID/VER/
STATUS/HELP/GET/TLM in this library's scope), and "including unknown
verbs" confirms it fires even before a verb is identified. This handler
implements it that way, with exactly one exception: ESTOP, whose own rule
("never carries an id and is never acked … must not queue behind anything,
including an ack") is treated as the more specific rule winning over the
general one. This is a resolution the wire grammar does not spell out in one
place — it is this file's own call, recorded here and in
protocol_handler.h's own file-header ambiguity note.
WHEELS_V's 5000 ms ceiling (unchanged by the 2026-08-22 rename) is
prose at the verb-definition level with no stated owner in the grammar.
The handler holds no bounds table, so the adapter enforces it and
returns kRange above it. The same "handler holds no bounds table"
posture applies to every OTHER motion verb's own documented bound
(WHEELS_X/MOVE_X's timeout, etc.) — none of them are enforced by
ProtocolHandler either.
The id's own numeric grammar ('#' [0-9]+) is stricter than an ordinary
wire integer field. §2's general "every wire value is … optionally signed"
does not apply to the id itself — #+5 is not a well-formed id, even though a
+-prefixed ordinary field elsewhere might parse. Implemented with a
dedicated digit-only pre-scan (parseIdDigits()) rather than reusing the
general unsigned-field parser.
This whole subsection describes a decision that no longer stands. It is kept, unedited below, as the changelog record of what this library used to do and why §8 replaced it outright, rather than being deleted and leaving no trace of the reasoning that came before.
The 3× reply repeat. The wire's own design has an id-carrying
ok/err/donesent three times on consecutive cycles, so an outcome survives radio frame loss without a ring or an eviction policy. This handler does not do that, because "on consecutive cycles" needs a periodic entry point and pending state — exactly what the no-done-for-WHEELSdecision (§8) keeps out. The repeat is emission policy, owned by whatever drives a real per-cycle output loop — NOT a property of the line codec. This handler stays a pure function of its input bytes: it emits each id-carrying reply exactly once, and the repeat, if and when it is wanted, belongs to whatever drives a real cycle loop, not to this class.That is a real gap, not a rounding error: loss tolerance is currently unimplemented in this library. It belongs at the app or transport layer that owns a loop, or it comes back into the handler when
MOVEanddonedo. Worth deciding deliberately rather than discovering on a lossy link.
What actually happened: the gap this subsection flagged was real, and
the stakeholder closed it 2026-08-21 — not by implementing the 3×-repeat
idea this subsection was written against, but by replacing the entire
delivery model with the cumulative sequence-id ack/nack scheme in §8. The
3× repeat is deleted, not implemented late: it would have meant sending
the SAME ok/err three times per id, which cannot coexist with a scheme
where the id itself is the sequence number and re-sending an old ack for a
stale id must NOT look like accepting a new command (§8.1's "do not
re-execute" case exists precisely to keep those two ideas from colliding).
Loss tolerance is no longer a gap: §8 IS this library's answer to it.
The TLM projection is reduced, not a literal world-frame pose — see §5.2.
It deliberately does not reuse column names for different data than what
DiffDrive actually publishes.
The flags word uses a local bit layout, for the same reason — see §5.2.
Stakeholder direction: src/protocol/ is going to be an archetype,
ported to MicroPython and JavaScript by reading it and running its fixture,
so this pass focused entirely on the handler's own robustness, not new
features. Full detail lives in the fix-site comments in
protocol_handler.cpp and in tests/protocol/test_protocol_adversarial.py's
module docstring; this is the summary a future porter should read first.
Three real bugs, all fixed, none a wire-format change:
-
formatConfigValue()cast a NaN straight touint32_t— undefined behavior, confirmed live by UBSan. A NaN can never arrive over the wire (parseFloatFieldalready rejects it on input), so this was only reachable through theAdapterseam — an adapter's own stored config value being NaN, read back byGET. Fixed by clamping NaN to 0.0 before the cast;+Inf/-Infwere already handled correctly by the existing overflow clamp. -
Hex-float syntax (
SET name 0x1.8p3) bypassed "no exponents" entirely. The exponent guard only checked for'e'/'E'; a hex float's exponent marker is'p', gated behind a'0x'prefix the guard never looked for, sostrtofsilently accepted it (0x1.8p3→ 12.0). Archetype-relevant on its own: neither Python'sfloat()nor JavaScript'sNumber()/parseFloat()accepts hex-float syntax, so this was a C++-only divergence — a straight port would not have this bug at all, and would need to actively decide whether to add hex-float rejection or simply rely on its own numeric parser already refusing it. -
A leading-whitespace numeric field was silently accepted because
strtol/strtoul/strtofall skip leading whitespace per the C standard. Under the space grammar a literal LEADING SPACE can no longer reach a field decoder at all — the tokenizer collapses every run of' 'into one separator before a field pointer is ever produced. The guard is NOT dead code, though: the field grammar (field ::= any bytes except ' ' and '\n') still admits'\t','\v','\f', and'\r'as ordinary, legal field bytes, andstrtol/strtoul/strtofwould silently skip any of those too — so the guard survives, now targeted at a narrower set of bytes. Every language's numeric parser has its own leniency here regardless (Python'sint()/float()also strip whitespace and accept_digit separators; JavaScript'sNumber(" ")is0) — a port author should decide this deliberately per language, not inherit whichever behavior their host language's built-in parser happens to have.
One characterization finding, not fixed — read this before porting
dispatch(): every wire-touching comparison in this handler (verb
lookup, tokenizing) operates on NUL-terminated C strings, per the
no-allocation, no-std::string constraint. strcmp()/the tokenizer's own
forward scan both stop at the first NUL in a string, so PING extra compares
equal to "PING" and dispatches exactly like a bare PING — silently
discarding extra with no malformed-count increment. The grammar's verb
rule (verb ::= [A-Za-z][A-Za-z0-9_]*) does not admit NUL in a verb at all,
so the grammar-correct behavior would be rejection, not silent acceptance of
the truncated prefix. This is NOT reproduced by a length-aware host
language: Python bytes/JavaScript strings compare full length, embedded
NUL included, so b"PING extra" == b"PING" is False in Python. A faithful
line-by-line port of this C++ handler's logic would therefore behave
differently from this reference implementation on this one input class —
pinned as a characterization test
(test_embedded_nul_immediately_after_verb_matches_bare_verb) so it cannot
drift silently, not fixed, because a real fix means abandoning C-string
comparisons throughout the parser — a far larger, riskier change than this
pass's scope, and in tension with the explicit no-std::string firmware
constraint.
src/adapter/ — its own package, not inside src/protocol/ or
src/diffdrive/. It is the one component required to depend on both, and each
of those two has a standalone-build gate ("compiles with an include path of
exactly its own directory") that a cross-dependency would break.
Historical record, partially superseded by §8 (2026-08-21). This
section is preserved as it was written, describing the separator/id-
spelling migration as it stood the day before the reliability layer
existed. Where it says an id is "optional" or describes ok/bare-err/
#0-suppression as current behavior, read those as as of 2026-08-20
only — §8/§8.6 replace all of it: every sequenced verb's id is now
mandatory, ok is deleted, and #0 no longer suppresses anything. What
this section gets right and still holds: the SEPARATOR is still spaces,
the id is still spelled #<n> and still trails the line, and the
underlying object model (Adapter, Result, Sink, Snapshot/Column)
still did not change for THIS migration (§8 changed Result's wire
mapping, not its own shape — see §9.8).
Stakeholder decision, 2026-08-20: fields are separated by spaces, not
colons, and the correlation id returns to its historical #-prefix
spelling as a trailing, self-marking field. This section records what
changed in src/protocol/ and tests/protocol/, for the same reason §9.4
records the hardening sweep — a future porter reading this file needs the
"why", not just a diff.
What is a pure separator swap, and what is not. Every wire example in
this document uses the space grammar; the underlying OBJECT MODEL —
Adapter, Result, Sink, Snapshot/Column, the handler/adapter split
itself — did not change at all. This was a rewrite of
protocol_handler.{h,cpp}'s parsing and formatting internals, not a
redesign.
New mechanics this migration introduced:
-
Tokenizing, not colon-splitting.
tokenizeLine()collapses runs of' 'into one separator and trims leading/trailing line whitespace. A blank or all-whitespace line is now ignored SILENTLY (previously, under the colon grammar, an empty line dispatched as an unknown zero-length verb and counted malformed). -
The id is self-marking and line-trailing, not positional. Because it
announces itself with
#, an omitted optional field never shifts it into a data position — the reasonSET name valueandSET name value #9are both exactly two or three tokens, with no placeholder needed for the missing middle slot the old grammar would have required. -
Bare vs id-carrying replies are now genuinely different wire shapes.
An omitted id →
ok/err <code>(no#idtoken at all); an explicit nonzero id →ok #<id>/err #<id> <code>; an explicit#0(legal only where the id is optional) → no reply at all. This retired an old ambiguity where an omitted id and an explicit0looked equivalent but behaved differently with no sentence saying so. -
The malformed-line
#idrecovery rule is new capability, not a reformatting of old behavior. Under the old colon grammar an unknown verb's own arity was unknowable, so no field of its line could ever be trusted as an id; the new grammar's self-marking id makes it trustworthy regardless of whether the verb itself is known, or even well-formed.ESTOPis the one deliberate exception (§9.1). This inverted part of the oldtest_unknown_verb_no_replytest, which only covered the id-less case — now split intotest_unknown_verb_no_reply_when_no_recoverable_idandtest_unknown_verb_with_recoverable_id_gets_err_unknown. -
The id's own numeric grammar is stricter than an ordinary integer
field —
'#' [0-9]+, no sign at all, parsed with a dedicated digit-only pre-scan (parseIdDigits()) rather than reusing the general unsigned-field parser, so#+5is correctly NOT id 5.
What did NOT change: the Adapter interface (adapter.h) — every
method signature, Result/TlmMode/Column/Snapshot shape is untouched,
because none of them ever encoded a wire delimiter. mock_adapter.h and
protocol_shim.cpp needed zero changes for the same reason.
tests/adapter/test_diffdrive_adapter.py, which drives the real handler end
to end (not a mock), needed its wire literals updated for the same
mechanical reason golden_vectors.txt did.
Golden-vector fixture: every vector in golden_vectors.txt changed
SHAPE, not just separator — the old ok:0 id-less arm is gone; there is no
new equivalent single spelling, because "id-less" now means literally "no
#id token in the reply", i.e. a bare ok. The fixture also grew new
vectors for rules the colon grammar never had: space-run collapsing,
STOP #0 being malformed (required-id verb), and an unknown verb's
trailing #id recovering an err reply.
Two verbs added to the library: debug (robot-to-host only, §6.2) and
RUN (host-to-robot invocation by name, §6.3). Adapter gained one new
pure-virtual method, onRun() — both concrete implementations in this
repository (MockAdapter, DiffDriveAdapter) were updated; DiffDriveAdapter
registers nothing and answers every RUN with ERR_UNKNOWN (§6.3).
kCommandTable grew from 12 to 13 entries (RUN appended at the end), so
HELP's generated reply grew by five bytes ( RUN) — still comfortably
inside both its own local 160-byte formatting buffer and the wire's 240-byte
line cap; a porter should not assume this margin is infinite, only that it
held here.
A real bug found and fixed while implementing this: the first working
draft of sendDebug()/handleRun()'s final line-formatting buffer was sized
char buf[kMaxLineBytes] (240 bytes). kMaxLineBytes already counts the
wire content including the terminating '\n', so a line that legitimately
reaches the full 240 bytes needs a 241-byte buffer — snprintf() also
needs room for its own NUL terminator, and with only 240 bytes available it
silently truncated the last byte of the formatted string (the trailing
'\n' itself) to make room for the NUL it always writes. Caught by
test_send_debug_exactly_240_bytes_is_not_truncated, a boundary test
written specifically because the earlier hardening sweep (§9.4) had already
established that boundary-byte-count reasoning in this file is exactly where
bugs hide. Fixed by sizing the buffer kMaxLineBytes + 1. This is a
C++-only hazard in the same spirit as §9.4's hex-float finding: Python's
f-strings and JavaScript template literals have no equivalent "off by one
for a NUL the language forces you to reserve room for," so a straight port
would not reproduce this bug — but it WOULD need to get its own
line-length-cap arithmetic right by some other means, and this is exactly
the kind of boundary a porter should write a test for rather than reason
about by inspection.
Design decisions made here, recorded for a future reader instead of only
living in code comments (the middle two bullets below are, like §9.6,
partially superseded by §8 the next day — #0-suppression is gone and
RUN's id is no longer detected by content inspection — annotated inline
rather than rewritten, since the REASONING each bullet records is still
sound even though the mechanism it describes changed):
-
sendDebug("")andsendDebug(nullptr)are the same case (§6.2) — both emit the baredebug\nline. The alternative (making null a no-op that emits nothing at all) was rejected:sendBanner()/sendReady()never take a "should I even emit" argument, and givingsendDebug()a hidden suppression channel through its argument's nullness, distinct from the wire's own explicit#0suppression spelling used elsewhere in this file, would be a second, undocumented way to say "don't send this" with no wire vocabulary to describe it. (#0no longer suppresses anything as of §8/§2.2 — the point about not inventing a SECOND suppression channel still stands, it just has one fewer sibling to be consistent with now.) -
Sanitize, don't reject, for both
debug's text andRUN's returned value (§6.2/§6.3).sendDebug()isvoidwith no return channel at all;RUN's outcome channel (Result) is owned by the adapter's own resolution/conversion/invocation logic, not by whether its return value happens to contain a newline, so reusing that channel to signal "your return value had a bad byte in it" would conflate two unrelated failure modes. Stripping degrades gracefully; rejecting outright would silently drop legitimate content over one bad byte with no way for either caller to learn that happened. (Unaffected by §8 -- still current.) -
A last field beginning with
'#'is always the id slot, even againstRUN's own open arity (§6.3, as of 2026-08-20) — resolved by content inspection rather than by field count, becauseRUNis the one verb in this library whose arity the handler cannot know in advance. The consequence — a function's own final argument can never itself begin with'#'— is a genuine expressiveness limit, not an oversight, and is stated as such rather than left for a porter to discover by testing. As of §8 (2026-08-21), content inspection is gone too: the id is UNCONDITIONALLY the last token for every sequenced verb, RUN included, so this is no longer RUN-specific machinery — but the consequence for a function's own final argument is unchanged, and if anything more absolute now (see §6.3's own updated text). -
kMaxFieldTokensraised from 8 to 20, and a newkMaxRunArgs(16) added, both firmware resource limits with no wire meaning of their own.RUN's open arity meant, for the first time in this file, that a verb's own field count could legitimately exceed what the fixed-size token array was sized for — every other verb's fixed arity had always been comfortably inside the old cap, so this never mattered before.handleRun()checksfieldCountagainstkMaxFieldTokensbefore indexing the field array at all, which no other handler in this file needs to do (protocol_handler.h's own file-header ambiguity note #4 has the full reasoning). A line with more real arguments thankMaxRunArgsis rejected asERR_BADARGbefore the adapter is ever called — a resource ceiling, not a claim about any real function's arity.
What a MicroPython/JavaScript porter would get wrong, bluntly:
-
Under-building
onRun(), not over-building it. The natural instinct in a dynamic language isgetattr(module, name)/globals()[name]with no registration table at all — and that is correct for those languages, but it means "everything importable is remotely callable" unless the porter deliberately restricts it. §6.3's security framing ("the registration table IS the security boundary") is written for the C++ archetype, where building a table is unavoidable and therefore an obvious place to enforce an allowlist; a dynamic-language port has to choose to build that same restriction on purpose, because its own language's ergonomics actively work against it. - Assuming the id and the last argument can't collide. A JavaScript or Python port's own function-calling convention has no equivalent to "the last field might secretly be the correlation id" — a porter translating this handler's logic naively (e.g. "split on spaces, last token after the name list is the id if the caller says there's one") will get this wrong for a variadic function unless they re-derive the content-inspection rule from this document rather than from the C++ source's control flow alone.
- Reproducing the buffer-sizing bug in spirit, if not in fact. No dynamic language will hit an actual NUL-terminator off-by-one, but a port that computes "does this fit the 240-byte cap" by string concatenation length alone, without a boundary test at exactly 240 bytes, can still ship a fencepost error the same class of mistake produces — §9.4 already made this point about hex-floats and leading whitespace; this section's own finding is one more data point for the same lesson.
-
Forgetting
onRun()must return promptly. It is called synchronously fromfeed()in every implementation, dynamic or not; a JavaScript port built on an event loop is especially easy to get this wrong in, by registering anasyncfunction and awaiting it inline instead of deferring the actual work and returning immediately.
§8 is a stakeholder design brought to this library fresh; the brief settled the shape but not every corner. Each item below is a real fork this implementation hit, the choice made, and why — flagged explicitly rather than picked silently, per this project's own stated practice.
-
Does a handler-level decode failure (unknown verb, wrong arity, unparseable field) advance the sequence and get an
ack, the same as an adapter-level rejection? The brief's own worked example (WHEELS 99999 0 100rejected on range) only covers the adapter-Result case explicitly. Resolved: yes, uniformly. §8.2/§8.4 treat "the id was in order" as the ONLY gate for whether the sequence advances and anackis sent; what happens after that (recognized verb vs. not, right arity vs. not, adapter accepts vs. rejects) all funnels into the same "maybe also emiterr" step. The alternative — treating a handler-level decode failure as NOT consuming a sequence slot at all — would mean the host cannot tell "my malformed line was silently discarded, unsequenced" apart from "the wire ate it," which defeats the whole point of a delivery guarantee. Chosen for internal consistency: one rule, not two nearly-identical ones with a subtle carve-out. -
Does an out-of-order (
<or>expectedNext_) line incrementmalformedCount()? Resolved: no.malformedCount()tracks content/decode failures specifically (unchanged meaning from before 2026-08-21); an out-of-order line's content is never even inspected (§8.4), so there is nothing to call malformed — it is a normal, expected occurrence on a lossy or reordering transport, not a protocol violation. Counting it would makemalformedCount()climb on ordinary packet loss, which is precisely the noise this scheme exists to distinguish from real protocol violations. -
What does
lastDone_actually track in a library with no queue and no completion event? Addressed at length in §8.5.1: it is plumbed (correctly initialized, reset, and echoed) but never written past its initial 0, because nothing in this library's own verb set produces the asynchronous completion event the field exists to carry. This is flagged, not silently assumed, because it would be easy for a future reader to assumelastDone_tracks something (e.g. "the last acceptedWHEELSid") when it deliberately does not — accepting aWHEELSis not the same event as it completing, and this library has no way to observe the latter without the timer §8.1 rules out. -
Reply ordering: does
ackcome before or after a verb's own informational/error reply? Not stated in the brief. Resolved:ackfirst, always, emitted the instant an id is accepted as in order — before the verb is even looked up in the command table. Every other line the same command produces (get/pong/id/ver/status/help/ret/err) follows it. Chosen because it lets the accept decision and its wire evidence be emitted from one place (dispatch()), before delegating to per-verb logic that knows nothing about sequencing at all — simpler to implement and to reason about than threading "did I already ack this" state through every handler. -
Does
emitTelemetry()'s piggybacked reliability line come before or after thethdr/tframe it accompanies? Moot as of 2026-08-26 (§8.5):emitTelemetry()no longer emits any reliability line at all, so there is no ordering left to resolve. (Historical answer, kept for the record: after —thdrif due, thent, then theack/nackline.) -
Is
STATUS's wrong-arity case still recoverable against anerrthe way it implicitly was before? (Answer below is historical — superseded 2026-08-27.)STATUSis sequenced now, so a malformedSTATUS(extra fields) still getsack+err 2 #<id>as long as its id is in order (§8.4's item 3) — no special case needed, unlikeHELLO's.Re-resolved 2026-08-27:
STATUSis unsequenced, so the question dissolves rather than changing answer. It takesPING's maximally forgiving posture (§8.3), which means there IS no wrong-arity case for it any more —STATUS,STATUS 1 2 3andSTATUS #9are byte-identical. Nothing needs to be recoverable against anerrbecause nothing is ever refused. The same applies toHELP,IDandVER. -
Does a malformed
HELLO(wrong arity) get any reply at all? Resolved: no, same as before this change —HELLOis outside the sequence entirely (§8.3), so there is noackto anchor anerragainst, and inventing a bareerrfor an unsequenced verb would be a new, one-off wire shape with no counterpart anywhere else in this grammar. A malformedHELLOincrementsmalformedCount()and produces no reply, exactly like a sequenced verb whose id cannot be determined at all (§8.4 item 1/2).Scope widened 2026-08-27. This item now governs four more verbs (
HELP/ID/VER/STATUS), and it is precisely why they takePING's forgiving posture rather thanHELLO's strict one: under the strict posture, "no reply at all" would have swallowedID #1— the single most common thing a host or a human actually types — and traded one silent drop for another. See §8.3's arity-posture rule. -
Is
Result::kDuplicateId(ERR_DUPLICATE_ID, code 11) still reachable? No — flagged, not removed. §2.2 and §6.1 both call this out: the handler's own sequencing now guarantees the adapter is never handed a repeated id, so no code path in this library can produce code 11 any more.Result::kDuplicateIdstays declared inadapter.h(removing an enumerator from a stakeholder-owned wire-outcome type is out of this change's scope), but it is dead as of this commit. A future caretaker should not spend time trying to find a test for it — there isn't one, and there cannot be one against this handler as designed.
Preserved verbatim for the record — this is what §8 said before the 2026-08-21 reliability-layer design replaced it outright (§8.0):
The wire's outcome model has room for three reply verbs:
ok [#id]means accepted,err [#id] <code>means rejected, anddone #<id> <reason>would mean the thing you enqueued has now finished, wherereasonisstop(its stop condition was met) ortimeout(its backstop fired).
donewas designed forMOVE, which has a real stop condition — "drive until 400 mm of travel" genuinely finishes.WHEELShas no stop condition. It holds a wheel speed fordurationms and then the lease expires. So "finished" means only "the lease ran out", and the question is whether that is worth a reply at all:
- Emit
done #<id> timeout— the host can await completion instead of timing the wheels itself, andWHEELSbehaves like every other bounded command.- Emit nothing;
okis the whole story — the host already knows the duration it asked for, so the reply carries no information it lacks.The reason this is a design question and not a preference: emitting
donemeans the handler must remember outstanding ids and emit a reply later, on its own, which means it needs a periodic entry point and a notion of time. Withoutdone, the handler is a pure function of the bytes fed to it —feed()in, replies out, no state between calls beyond a partial line, nothing to tick.That is a real difference in what the class is, and it is much cheaper to decide now than to retrofit. Settled, 2026-08-20: no
doneforWHEELS— the handler stays stateless and pure for this library;donearrives withMOVE, which is the verb that actually needs it, whenever that is built.
The "handler stays a pure function of the bytes fed to it, nothing to
tick" property this text worried about protecting is, notably, the SAME
property §8.1 protects in the new design (expectedNext_ — and, before
their own later removals, lastDone_ and gapOutstanding_ — is
ordinary state, not a clock) — the reliability layer
answers a different question (did the bytes arrive?) than done was going
to (did the motion finish?), and does so without reintroducing the timer
this section was written to keep out.
Six changes in one pass: decode failure is a NAK (§8.9), ERR_DUPLICATE_ID
deleted (§2.2/§6.1), PING unsequenced (§8.3), lastDone/its reason move
to the Adapter (§8.8), the ack/nack reason token (§8.8.1), and the six
motion-api.md §9.1 verbs implemented at the wire/handler layer (§6,
WHEELS renamed WHEELS_V). Ambiguities this pass resolved on its own,
in the same spirit as §9.8's own list for the pass before it:
-
DiffDriveAdapter's five unimplemented motion verbs answerkUnknown, notkUnimplemented. The ticket driving this pass said so explicitly, and it matches an existing precedent this file already set:RUNon an adapter with an empty registration table answerskUnknownfor every name, "the same wire outcome any name a real registration table would not recognize" (§6.3) — notkUnimplemented, even though a real registration table conceivably COULD recognize the name someday. The same reasoning applies here:DiffDriveAdapterhas no planner at all, soWHEELS_X/MOVE_X/MOVE_V/GO_TO_R/GO_TO_Ware, from its own point of view, simply not verbs it knows anything about —kUnimplementedwould suggest a build flag or a half-wired feature, when the honest description is "this adapter has no planner, period." A future adapter that DOES have a planner but deliberately ships with one of the six verbs turned off is wherekUnimplementedwould be the correct choice instead. -
PING's own exemption is maximally forgiving, not strict zero-arity. The stakeholder's direction ("ESTOP, ping, and HELLO shouldn't require IDs") settled that PING is unsequenced but not whether it should tolerate trailing content the wayESTOPdoes or reject it the wayHELLOdoes. Resolved forgiving, matchingESTOP— see §8.3's own bullet for the reasoning (an old-style host still appending#<id>toPINGout of habit keeps working; liveness must not itself be refusable over a syntax nit, echoing exactly whyESTOPis forgiving). -
HELLO's reset no longer touches the Adapter's own completion state. The pre-2026-08-22 text hadHELLOresetlastDone_ = 0as part of the same call, back when it was handler state (§8.3's old text). Now that it lives on the Adapter (§8.8),HELLO's reset only touchesexpectedNext_(andgapOutstanding_, until its 2026-08-26 removal, §8.5) — see §8.8's own reasoning for why reaching across the seam would be the wrong call. -
ack/nackis sent BEFORE the verb's own execute step runs, even for a verb (likeSTOP) whose OWN execution can synchronously complete a motion. This meansSTOP's own ack reflectslastDone/lastDoneReasonas they stood immediately BEFORE that STOP executed — the completion IT just caused becomes visible starting with the NEXT reply (a later command's ack), not on its own ack. This preserves dispatch()'s uniform "ack always precedes verb execution, for every verb" structure (§8.2/§9.8 item 4 of the pass before this one) rather than special-casing STOP to execute-then-ack. Nothing is ever LOST — the value is not lost, only delayed by one reply — but a host readingSTOP's own ack literally should not expect to see the completion it just caused reflected there yet. -
A decode failure's
nackanderrshare the SAME id, by construction, not by a separate check. Because a decode failure is only reachable via theid == expectedNext_branch (§8.9), the id that failed to decode ISexpectedNext_at the moment the nack is formatted — there is no separate bookkeeping needed to keep the two numbers in sync, and no way for them to drift apart. -
The decode/execute split is per-verb, not global. Rather than a
single generic "parse this line, then maybe execute" pass,
ProtocolHandlerpairs each verb with its owndecodefunction (pure arity/field-parseability check, no adapter call, no sink write) andexecutefunction (runs only once decoding has already succeeded, and only afterdispatch()has already sent theack). The two re-derive the same fields independently rather than threading decoded values between them — a small, deliberate duplication (this is not a hot path) that keeps "what counts as decodable" defined in exactly one place per verb without a shared decoded-argument struct per verb.
Settled with the firmware implementation across a working session; that side is landed and hardware-verified, and the wire behaviour below was captured over a wired link (chosen deliberately — at the radio's measured rates "the reminder was absent" and "the reminder was dropped" are the same observation).
A stakeholder session on a raw relay link:
HELP -> (nothing)
ID -> (nothing)
ID #1 -> ack 1 0 none / id diffdrive vevov 1.0.10 vevov
ID #1 (resent) -> ack 1 0 none <- ack, NO id line
ID #2 -> ack 2 0 none / id ...
Read as one bug, this was five, and separating them is most of the value of this entry:
- A bare
HELPparsed as#0, fell belowexpectedNext_, and was silently dropped (§2.2) — the one verb whose job is orienting someone who does not know the grammar. - Query verbs were sequenced at all, so they were gated behind a counter a human at a keyboard has no way to track.
- The stale-retransmit re-ack (§8.1's middle row) answers a query with
a bare
ackand no payload. Correct for the machine case it was designed for; it reads as "accepted, then answered nothing." -
GETwith an unknown name did the same thing for an entirely different reason (below). - Most of the raw disappearance was the RF link, not the protocol at all (§8.0).
-
The stale re-ack stays (§8.1's middle row). A resent
WHEELS_Vwhose ack was lost needs to hear "I have everything through here, stop resending"; a nack would say the opposite, telling the host to resend what it just sent — a resend loop on a lossy link, which is the exact failure §8.1 exists to prevent. It would also makenack <n>ambiguous between "I need n" and "I already have n" — two states demanding opposite actions — since §8.9 already gavenacka second meaning. Unsequencing the query verbs dissolves the human-facing half of the complaint without touching the machine-facing guarantee. - No reply cache. Re-sending a cached payload on retransmit would satisfy "I asked, I got an answer," but §8.1's whole claim is that the receiver-side state is a single number — no ring, no per-id storage, no eviction policy. Not worth trading away.
-
Held in reserve, pre-agreed: if
GET's stale case proves annoying in practice, the fix is to re-execute a stale retransmit of a read-only sequenced verb (harmless by definition, and it needs no cache — you answer again rather than remembering the answer) rather than to add one. §8.3's rule already supplies the per-verb read-only property this would key off.
execGet() set errCode = 0 deliberately, citing §7 and §8.2, and
discarded the false that Adapter::onGet already returns. So the
handler detected the unknown name and declined to report it — while
SET, on the same config plane, for the same class of typo, returned
err 1. The asymmetry was documented in §6's table but never argued for.
On a lossy link the silence is also ambiguous in the worst way: "ack, no
get line" has two live explanations — wrong name, or the get line was
eaten in flight — and the operator cannot tell them apart. The fix costs
one short line, which is also the line most likely to survive.
Worth recording how it was found: it appeared as line 8 of a conformance capture taken for an unrelated purpose, was read as normal, and was scrolled past. The failure mode is that it does not look like a failure.
The first version of this change removed the reliability line from the unsequenced verbs entirely, on the reading that §8.5's "reply-only" meant "sequenced-only." It never said that. The stakeholder's objection there was to periodicity — a beacon several times a second on an idle link — and a line replying to an inbound unsequenced verb is still a reply.
The correction: sequence gating and reply emission are separable, and
only the first was ever objected to. Verbatim: "I don't actually mind
if ID, VER, and help also return an ACK/NAK. What I mind is that they
require an ACK/NAK." Hence §8.3's conditional reminder — a HELP you can
type any time, which still tells you your last command didn't land.
Conditional rather than unconditional was chosen over always-appending an
ack: the stakeholder's framing is a reminder, not a receipt, and on a
link dropping a third of its lines a second line per query is a real cost.
That choice is what requires gapOutstanding_ back (§8.1) — a partial
reversal of a deletion the stakeholder himself directed the day before,
put to him explicitly in those terms and approved. Its scope is written
into §8.1 and §8.5 precisely so that "we re-added gapOutstanding_" is
not read later as licence to restore the barrage.
The measurement work that ran alongside this is recorded because the reasoning is the reusable part.
- Draft 1: "ch4 degraded since this morning."
- Draft 2 retracted the size. The morning figure had been characterised partly by counting beacon keepalives — an instrument deleted along with the beacon. A link measured by beacon count is not commensurable with one measured by reply rate.
-
Draft 3 retracted the event. The morning data never showed a
regression: 8/8 successes has an exact two-sided 95% interval of
[63.1%, 100%], and 6/6 of [54.1%, 100%] — eight in a row cannot
establish a rate above about 63%. The morning's own
STATUSfigure (5/6 = 83.3%) is numerically identical to the best measurement taken that evening (50/60 = 83.3%). Nothing had changed; small samples had made a mediocre link look fine.
Two mechanical candidates were eliminated: relay-unit identity (74.2% vs
75.8%, n=120 each) and the beacon deletion — the latter tested without
a firmware flash, by using TLM POSE to produce continuous robot
transmission from the same radio (83.3% quiet vs 76.7% busy, p=0.361,
point estimate against the hypothesis). Substituting a configuration
change for a firmware change avoided leaving a beaconing build on a robot
whose stakeholder had deliberately removed the beacon.
Three methodological rules earned here, worth keeping:
- Interleave, always. Delivery drifted 17 points inside two hours, so sequential A/B on this link is worthless. The length experiment's result survived only because its arms were round-robin interleaved against a common drifting channel — that design choice did more work than its p-value did.
- A static cause cannot produce a time-varying failure. That, not the 1.6-point agreement between two relay units, is what rules out relay identity, antenna, and placement in one stroke. The clustering argument would have sent the next reader off to test the two units that never came up.
- Do not replace one unsupported number with another. §8.0's "~5%" is withdrawn rather than restated as "~75%": three runs spanning 17 points do not establish a rate either.
The open question is no longer "what broke ch4" but "why has ch4 apparently always been around 75%, and why did we believe otherwise?" — with the drift itself as the phenomenon to characterise, since a period would be the strongest available clue to what is duty-cycling.
Stakeholder-directed, out of process. §6.3 shipped RUN with the
registration table framed as "the security boundary… an explicit
allowlist," and then gave the wire no way to read that allowlist. A host
could only discover a callable name by guessing it and reading the
ERR_UNKNOWN. FUNCS (§6.5) closes the gap, together with three new
Adapter methods (runCount/runName/runSignature) that let a
concrete adapter declare its registry the same way fieldCount/
fieldName already let it declare its config surface.
Four decisions, all settled before implementation:
| question | decision | why |
|---|---|---|
| verb name |
FUNCS / reply funcs
|
LIST was the stakeholder's own word but reads as "list what?" beside HELP, which also lists things. FUNCS names its subject. |
| reply shape | one line per registered function | a rest-of-line help-style reply silently truncates past the 240-byte cap, which for an allowlist means a host reading a short list as the whole list; and it has nowhere to put a signature (§6.5) |
| sequenced? | yes | a variable number of reply lines needs a terminator, and the ack is the one this grammar already uses for exactly that, in bare GET. This is why FUNCS does not join the four query verbs §9.11 moved off the sequence. |
| scope | wire + handler + adapter seam | the registry itself stays each concrete adapter's business, unchanged from §6.3's division of responsibility — the handler gained a way to walk a table, not a table |
The signature field is deliberately unspecified beyond "one token."
Making it a parsed type language was considered and rejected: nothing in
this protocol would consume it, every consumer of it is a human or a
tool outside this contract, and a format the wire promises but never
validates is a guarantee waiting to be wrong — the same reasoning §6.4
used to make ID's name mandatory rather than optional.
Restored 2026-08-23. Commit 34d12c2 folded docs/protocol-v6-spec.md
into this file but dropped its entire §6 Telemetry chapter, even though
thdr/t frames, six TLM modes, and "the telemetry cadence" are
referenced throughout this document — line 33, §5.2, the §6 verb table's
TLM row, §8.5 — with no section anywhere defining the frame grammar, mode
semantics, or column layout. This chapter restores that definition,
appended as a new top-level chapter rather than reclaiming the old spec's
"§6" numbering, which would mean renumbering every section from current §6
onward — and every cross-reference to them — in an already heavily
cross-referenced document.
Not a verbatim restoration. The recovered source (git show 34d12c2^:docs/protocol-v6-spec.md, its own §6) describes a full robot:
world-frame pose fused from OTOS and encoder odometry, line sensors,
colour. None of that lives in this library. What follows is written
against DiffDriveAdapter's actual projection (§5.2) and
ProtocolHandler's actual emission code, not a transcription of the old
tables — §10.4 points at the archived full-robot tables for a port that
grows beyond DiffDrive.
| mode | wire token | effect |
|---|---|---|
| off | TLM OFF |
DiffDriveAdapter::telemetryEnabled() returns false; the calling app is expected to stop calling buildSnapshot()/emitTelemetry() for this connection — neither ProtocolHandler nor the adapter suppresses frames on its own (§10.2) |
| pose | TLM POSE |
7-column projection, §10.3 |
| full | TLM FULL |
11-column projection (POSE's 7 plus 4 more), §10.3 |
| now | TLM NOW |
decoded and acked like any other TLM command, but deliberately never stored into mode_ (onTlm(), diffdrive_adapter.cpp:287-295) — the current mode is left exactly as it was |
| auto | TLM AUTO |
stored into mode_, reported by STATUS's tlm= field; DiffDriveAdapter has no scheduler of its own, so the "silent while parked" cadence this token implies is not implemented here — see below |
| buffer | TLM BUFFER |
stored into mode_; same story — the REPL-side accumulate-and-drain behavior this token implies is an application concern, not implemented in ProtocolHandler or DiffDriveAdapter
|
Wire tokens are OFF/POSE/FULL/NOW/AUTO/BUFFER, decoded
case-sensitively by parseTlmMode(); an unrecognized token is a decode
failure (§8.9), not an application-level rejection.
Mode is state on the Adapter, not the handler. DiffDriveAdapter::mode_
defaults to OFF at construction — this library has no per-port default the
way the old spec's full robot did (POSE on radio/UDP, BUFFER on the
REPL); every connection starts silent until a host explicitly subscribes.
HELLO's reset (§8.3) does not touch it: handleHello() only resets the
reliability layer's own sequencing state (expectedNext_), so a
reconnecting host that does not resend TLM keeps whatever mode was last
set.
NOW, AUTO, and BUFFER do not carry their nominal behavior on this
adapter. All three are accepted, and (except NOW) persisted as the
current mode — but buildSnapshot() branches on exactly one condition,
mode_ == TlmMode::kFull; every other mode, including AUTO and BUFFER,
produces the identical 7-column POSE shape. NOW's "emit one frame
immediately" is not implemented at all: nothing observes a TLM NOW
command past onTlm() returning kOk, so it triggers no extra push — the
app-driven emitTelemetry() cadence (below) is the only thing that ever
emits a frame. A future adapter that DOES have its own scheduler
(parked-detection for AUTO, an accumulate/drain queue for BUFFER) is
where those tokens would grow real behavior; on DiffDriveAdapter they are
accepted, remembered, and otherwise inert.
No re-emit is triggered by the mode token itself. §10.2 states the
actual re-emit rule — a column-set change, not a mode change — and
because DiffDriveAdapter's own column set depends only on the FULL
boundary, switching among OFF/POSE/NOW/AUTO/BUFFER never changes
the header at all (they all share POSE's 7 columns); only a transition
across POSE↔FULL does.
No built-in rate floor. ProtocolHandler has no timer or clock of any
kind (§8.0/§8.5) — emitTelemetry() is an "unsolicited emission the app
drives, not the wire" (§3), so it emits exactly once per call, whenever the
caller calls it. DiffDriveAdapter::kCyclePeriod (24 ms, the kernel
fiber's own cadence) is the reference cycle time if an app calls
buildSnapshot()/emitTelemetry() once per kernel cycle, but that is a
kernel constant, not a floor this library enforces on the wire.
thdr seq now flags posl posr vell velr
t 5 1080 3 120 118 250 248
t 6 1104 3 122 120 250 248
emitTelemetry() (protocol_handler.cpp:1071-1093) calls emitHeader()
before emitFrame() whenever headerChanged() (protocol_handler.cpp: 1003-1015) says the remembered header is stale, then always calls
emitFrame(). headerChanged() is true when:
- this is the first frame this handler has ever emitted
(
everEmittedHeader_still false), OR - the column count differs from the remembered header, OR
- any column's name or hex-ness differs from the remembered header, at
any position up to
kMaxHeaderColumns(40,protocol_handler.h:301).
Nothing about the mode TOKEN is consulted directly — only the shape of the
Snapshot the adapter hands in (§10.1's last point is this rule applied to
DiffDriveAdapter specifically).
Value encoding (emitFrame(), protocol_handler.cpp:1045-1069): every
column prints in header order, one value per t line. An ordinary column
prints as a signed base-10 integer (%ld); flags — the one column with
Column::hex == true — prints lowercase hex with no 0x prefix (%x). A
reader zips thdr against each t positionally; nothing needs a schema, a
field table, or version negotiation.
DiffDriveAdapter::buildSnapshot() (diffdrive_adapter.cpp:297-335) is the
entire projection — §5.2 already states the principle ("a projection, not a
computation"); these are the actual columns it emits.
POSE — 7 columns, the projection for any subscribed mode except
FULL:
| col | unit | meaning |
|---|---|---|
seq |
— | increments per emitted frame, wraps at 128 (7-bit) |
now |
[ms] |
robot clock at frame assembly (DifferentialDrive::Output::now) |
flags |
hex | local bit layout — see below |
posl |
[mm] |
left wheel position, counts → mm through countsPerLength
|
posr |
[mm] |
right wheel position |
vell |
[mm/s ×10] |
left wheel velocity |
velr |
[mm/s ×10] |
right wheel velocity |
FULL adds 4 more (11 total):
| col | unit | meaning |
|---|---|---|
lambda |
[×1000] |
authority scale currently applied (Output::lambda, a [1] dimensionless 0–1 value, ×1000 wire quantum — Stage C's own bench-visible learned parameter) |
biasl |
[counts/s] |
Stage C's adapted left-wheel trim (Output::biasLeft) |
biasr |
[counts/s] |
Stage C's adapted right-wheel trim (Output::biasRight) |
cyc |
— | kernel heartbeat cycle count (Output::cycleCount) — the same sentinel RobotLoop watches for advance |
×10/×1000 mean the wire integer is the value times that factor, the
same convention the archived full-robot table uses for elv/erv (§10.4)
— no precision is lost relative to the float source.
flags — local bit layout (computeFlags(), diffdrive_adapter.cpp: 111-122):
| bit | meaning | bit | meaning |
|---|---|---|---|
| 0 | ready | 4 | left wheel connected |
| 1 | estopped | 5 | right wheel connected |
| 2 | lease expired | 6 | left wedge |
| 3 | stall halted | 7 | right wedge |
This is not the archived spec's §6.5 bit numbering, deliberately —
diffdrive_adapter.h's own header comment states why: this library has no
OTOS, line sensor, colour sensor, or planner, so reusing bit numbers that
meant those things elsewhere would misrepresent what they mean here to a
reader with the old spec open. §10.4 points at where that numbering still
applies.
The archived spec's own column tables survive in code form at
src/archive/protocol-v6/wire_v6_telemetry.h — auto-generated
(scripts/wire_v6_tables.py) kPoseColumns/kFullColumns name arrays: 9
columns for POSE (seq now flags x y h ox oy oh — world-frame pose fused
from encoder and OTOS odometry) and 35 for FULL (adds per-wheel encoder
detail, OTOS velocity, body twist, line/colour sensor channels, cycle
timing). These are the reference layout — and, for flags, the reference
bit numbering — for a port that grows beyond DiffDriveAdapter: a robot
that DOES carry OTOS, line sensing, or a real motion planner should extend
toward that table's column names and bit assignments rather than inventing
a third scheme, so an existing host-side decoder written against the full
spec keeps working unmodified. This library's own reduced projection
(§10.3) is not a subset of that table by column position — posl/posr/
vell/velr have no full-robot equivalent name, because DiffDrive
publishes wheel counts, not a fused pose.
Forward-specified here; implemented by ticket 003, not yet in code as
this chapter lands. The frame is self-describing only if the host
actually holds the header (§10.2) — if a thdr line is dropped on the
radio, or a host reconnects mid-stream, every t frame after it is
unparseable: values with no column names. The host knows this (no
remembered header, or a field count mismatch against the header it has)
but has had no way to ask for a fresh one.
TLM HDR #<id> — reuses the existing TLM <mode> #id slot exactly
like every other mode token, so it is sequenced like any other TLM form
with no new grammar rule. It forces the next emitTelemetry() call to
re-emit thdr before its next t frame, and does not change the
current subscription mode — STATUS's tlm= field reads exactly as it
did before the request. Mechanically (per ticket 003's plan): execTlm()
special-cases this token by clearing the handler's own remembered-header
state directly (everEmittedHeader_ = false, the same field
headerChanged() already checks) rather than forwarding it to
Adapter::onTlm() — a header-recovery request is not a subscription
change, and the handler already owns the state that needs clearing.
This is the correct recovery path — not TLM NOW. The old spec
claimed a host that missed the header "sends TLM NOW" to recover it;
that was never true of this implementation (§10.1: NOW triggers no extra
emission of any kind, header or otherwise) and is not restored here.
TLM HDR is the mechanism this chapter specifies instead.