-
Notifications
You must be signed in to change notification settings - Fork 0
Reliability FEC and NACK
The voice protocol uses a two-stage reliability scheme:
-
Reed-Solomon FEC absorbs sub-
parity_countlosses with no round-trip. - Selective NACK with bitmap recovers larger losses in a single round-trip.
The combination keeps best-case latency low while bounding worst-case airtime.
Implementation: reed-solomon-erasure
crate over GF(2⁸), with (total_data, parity_count) shards.
- Split audio into
total_datachunks ofchunk_sizebytes (zero-pad the last chunk). - RS-encode
parity_countparity shards. - Send all
total_data + parity_countshards (padding stripped on the final DATA frame; receivers re-pad for FEC math).
A receiver MUST be able to reconstruct the message if it has any
total_data shards out of the total_data + parity_count total —
any combination of DATA and PARITY shards counts toward the threshold.
parity_count is sender policy, expressed as a percentage of total_data:
| Mesh profile | parity_count |
|---|---|
| Short / quiet | 10 % |
| Medium / mixed | 20 % |
| Long / lossy | 33 % |
| Broadcast (no NACK feedback channel) | 50 % |
parity_count = 0 is allowed — it disables FEC entirely. NACK still
works.
Byte-aligned, 256-shard ceiling fits perfectly in u8 indices. No
bit-shuffle pre/post processing. Throughput on commodity hardware is well
above what LoRa airtime can deliver.
When loss exceeds parity_count, the receiver issues a NACK after a
quiet period of NACK_WINDOW_MS (default 3000 ms) since the last
chunk arrived for that message.
A bitmap of length ⌈total_data / 8⌉ lists missing DATA chunks (one bit
per data shard). The sender retransmits only those chunks. On a single
NACK round, all missing chunks are listed in one bitmap — there's no
chunk-by-chunk retry.
NACK_MAX_ROUNDS = 400 per message of consecutive rounds without
progress - the counter is reset every time a new shard lands, so a
sender that's still actively servicing every NACK round keeps the
assembly slot alive indefinitely (capped only by message_timeout,
default 1200 s). After 400 NACK rounds in a row with zero new chunks the
receiver gives up and either emits a partial message
(partial_play_on_timeout = true, the default) or discards the work.
At a nack_window of 3000 ms that's a ~1200 s ceiling on a truly silent
sender.
The 400 is only the seed default. Hosts that change message_timeout
or nack_window (a GUI slider, a CLI flag, a LoRa preset change) call
AssemblerConfig::sync_nack_cap_to_timeout,
which re-derives max_nack_rounds = ⌈message_timeout / nack_window⌉ so
the user-configured per-message timeout - not the round cap - is always
the practical ceiling. 1200 s / 3 s = 400.
Earlier revisions used a cumulative counter that never reset. That turned out to be indistinguishable from genuine silence on healthy slow-trickle messages — a 67-chunk Long-Slow broadcast could rack up 32 productive rounds long before delivery and surface a phantom
partial: 47/51 chunksline. The consecutive semantic preserves the protection against a chatty bad sender (message_timeoutis still the absolute upper bound) without the false positive.
NACK rounds do not continue all the way to message_timeout if the
sender goes quiet. If no real data (a DATA or PARITY frame) arrives for
dead_sender_timeout (default 120 s), the receiver presumes the
sender has dropped off the mesh and stops emitting NACKs for that
message, letting it sit until the hard message_timeout fires (and then
finalize partial or discard). This stops a long tail of cap-multiplied
NACKs being fired at a peer that will never answer. The timeout is
validated to be < message_timeout. This is an experimental
flood-control limit (see Constants and Limits).
A NACK whose bitmap is all zeros means "all chunks received, stop sending parity". Parsers MUST accept this; the reference implementation doesn't currently emit it (natural completion + the recently-finalized blacklist already handle late parity frames).
flags & 0x01 = the receiver has timed out. Senders SHOULD discard any
remaining queued chunks for this message — keep transmitting and you
just waste airtime.
Sender Receiver
────── ────────
build_message(audio) │
├─ split into N data chunks │
└─ RS-encode P parity chunks │
│ │
├─ DATA[0] ─────────────────────────────────► │
├─ DATA[1] ─────X (lost) │
├─ DATA[2] ─────────────────────────────────► │ pending
├─ PARITY[0] ─────────────────────────────────► │ FEC reconstructs DATA[1]
│ │ ✓ Complete
│ │
│ --- or, on heavier loss --- │
│ │
├─ DATA[0] ─────X │
├─ DATA[1] ─────X │
├─ DATA[2] ────────────────────────────────► │
├─ PARITY[0] ─────X │
│ │ quiet 3000 ms
│ ◄────────────────────────── NACK [bitmap=0xC0] │ (chunks 0 & 1 missing)
├─ DATA[0] ────────────────────────────────► │
├─ DATA[1] ────────────────────────────────► │ ✓ Complete
The sender side of NACK handling lives in
OutgoingVoiceRegistry
and is driven by VoiceSender.
[VoiceSender::send] registers an entry at burst start and the
background NACK-listener task feeds inbound NACKs into
take_retransmit, which enforces three caps:
-
Pending-chunk dedup. Every DATA index is marked pending at
register()time. The burst loop callsmark_chunk_sent(i)per frame as it leaves the worker. A NACK arriving while the initial burst is still draining is filtered againstpending_chunks, so chunks already queued up are not re-enqueued. -
Per-message cooldown. After each retransmit batch, the entry is
parked for
pacing × frames.len(), clamped to[1 s, 30 s]. The 30 s ceiling sits comfortably below the receiver's ~1200 s NACK budget so the sender always responds before the receiver gives up. NACK rounds during cooldown are dropped; the receiver re-NACKs after the next quiet window. -
Per-message retransmit budget.
MAX_RETRANSMITS_PER_MESSAGE = 2_400comfortably exceeds the widened receiver round cap (400 × max 30 s cooldown). Beyond that, the sender drops further NACKs.
After the initial burst, the sender lingers for SendRequest::linger
(default 600 s, see DEFAULT_LINGER)
before emitting SendStatus::Complete and releasing registry state.
A stale NACK arriving after that point finds nothing to retransmit
against. The previous value of 60 s was too short for slow modem presets
(e.g. LongFast at 900 ms pacing: a 155-frame burst alone takes ~140 s,
leaving insufficient linger time for NACK-driven retransmit rounds).
The absolute outer envelope is OutgoingVoiceRegistry::set_retain_ttl
(default 1200 s, covering max_burst_duration + linger on all
modem presets the receiver can still hear, so a NACK never finds its
registry entry expired while the sender is still alive).
Receivers MUST NOT emit NACKs for broadcast messages. With multiple
listeners on the same channel, every receiver NACK'ing the same missing
chunks would saturate the sender's radio, and the sender has no way to
pick a single retransmit target — broadcast is point-to-multipoint by
definition. The reference receiver short-circuits the NACK emission
branch in tick() when the in-progress entry's to is the broadcast
address; the state machine still drives timeouts and partial finalize.
Broadcast voice relies on FEC + partial-on-timeout: pre-pay the
airtime for extra parity shards (the default voice.fec_mode = auto
picks 50 % for broadcast), and accept that the message either completes
via Reed-Solomon recovery or finalises as partial.
The reference receiver schedules NACK rounds at
nack_window × backoff_base^min(round, 4). Both fields are
host-configurable on AssemblerConfig and are typically set via the
voice.nack_mode setting:
| Mode |
nack_window base |
backoff_base |
max_rounds |
|---|---|---|---|
Off |
(NACK disabled) | 0 |
0 |
Auto short/med fast |
1.5 s | 2 |
800 |
Auto med slow / LongFast |
3 s | 2 |
400 |
Auto long-range / unknown |
pacing × 4 min 4 s |
3 |
200 |
Conservative |
pacing × 4 min 4 s |
3 |
200 |
Aggressive |
1.5 s | 2 |
800 |
backoff_base = 0 is the disabled signal; the assembler skips both the
quiet-window and give-up NACK branches when it sees a zero base. The
broadcast short-circuit above is independent — broadcasts get no NACK
regardless of the configured backoff_base.
This protocol does not authenticate NACK frames; like DATA / PARITY, any
peer with the channel PSK (i.e. anyone Meshtastic lets join the channel)
can fabricate a give_up NACK and abort an in-flight transmission. This
matches Meshtastic's threat model for text traffic and is documented as
a non-goal in the spec. Mitigations available to senders:
- Treat
give_upas advisory; if airtime budget allows, retry under a freshmessage_idafter a backoff. - A future revision MAY reintroduce a keyed MAC field if a concrete threat model warrants it. See git history at the v2 line for the previous design.
| Setting | Default | Effect of increasing |
|---|---|---|
parity_count (sender) |
10-50 % | Better loss tolerance, more airtime |
NACK_WINDOW_MS |
3000 | Fewer spurious NACKs on jittery links |
NACK_MAX_ROUNDS (consecutive) |
400 † | Higher completion rate, longer worst-case (re-derived from message_timeout / nack_window) |
MAX_RETRANSMITS_PER_MESSAGE |
2_400 ‡ | Sender-side counterpart to NACK budget |
AssemblerConfig::message_timeout |
1200 s | Larger messages allowed, more state held |
AssemblerConfig::completion_memory |
= message_timeout
|
How long a finalized (from, message_id) is remembered so late chunks can't resurrect a phantom partial. Must be >= message_timeout. |
AssemblerConfig::dead_sender_timeout |
120 s ‡ | Silence after which the sender is presumed dead and NACKs are suppressed until message_timeout. Must be < message_timeout. |
OutgoingVoiceRegistry::retain_ttl |
1200 s ‡ | Sender remembers frames longer |
SendRequest::linger |
600 s | Sender stays subscribed to NACKs longer |
partial_play_on_timeout |
true |
Always emits something on timeout |
† Derived, not fixed.
NACK_MAX_ROUNDSis re-derived frommessage_timeout / nack_windowwhenever a host changes either;400is just the default-config seed.‡ Experimental flood-control / resource-bounding limits. These are heuristic safety valves, not wire-format values: they cap NACK and retransmit storms and bound reassembler/registry state so a dead, slow, or misbehaving peer can't flood the channel or exhaust local memory. Tuned empirically and may change between releases without a
PROTOCOL_VERSIONbump. See Constants and Limits.
→ Continue to Encryption.