Skip to content

smokeng v0.18.0

Choose a tag to compare

@Timdebruijn Timdebruijn released this 02 Sep 07:22
· 16 commits to main since this release

v0.18.0 — probes that were never sent are no longer loss

An irtt target reported 19 of 20 about a quarter of the time, with a
send-failure flag and five percent loss. Nothing was lost and nothing failed:
the twentieth packet was never scheduled.

irtt's sender does not count packets. It runs until a duration elapses and
keeps to an interval grid, so a packet sent past the halfway point of its slot
makes the client aim at the slot after next — costing a whole interval. The
margin was half a step, so one late send anywhere in the train ended the
session a packet short.

The window is now a full interval past the last slot, clamped to what the
bucket has left. Both halves matter: unclamped, spread mode asks for 315
seconds of a 300-second interval and every session is cancelled, which reads as
total loss for ever.

But the window cannot be the fix, because no duration yields exactly the
requested count both with and without a skipped slot. So a session that ends
early is no longer treated as a session that broke. The tail of a broken
session was owed and never went out: it counts, and says why — for a receive
failure as well as a send failure, since either stops the client. The tail of a
short session was never owed, and is left uncounted, so the interval reads 19
of 19: a narrower distribution than was configured, which the sent count
records and the graph shows, and no loss, because none occurred.

Reporting loss that did not happen is the same defect as reporting a value that
was not measured, on the one probe type where loss is the whole point.

SmokePing does not solve this. Its IRTT probe uses pings*interval, which is the
same arithmetic and fails the same way — and its graphs for the same target
show sixteen-of-twenty losses at moments this prober measures twenty of twenty
on the same host, port and second. Its three IRTT targets fail identically to
the packet, which is its own contention rather than the path, and it pads the
jitter metrics with an invented median so the loss count comes out right.

Three tests on this release passed with the code they guarded removed. Each was
found by mutating the code rather than by reading it, which is the only reason
to run a test suite.