# The #1 measurement for a Leios testnet campaign ## Summary If a testnet-instrumentation campaign could deploy exactly one measurement today, it should be **cross-node per-EB arrival correlation** — a per-EB, per-node arrival-time log joined with certification outcomes across all measurement nodes in the mesh. Companion documents for the surrounding analysis: [`leios-overlap.md`](leios-overlap.md), [`leios-testnet.md`](leios-testnet.md), [`leios-report-gaps.md`](leios-report-gaps.md), [`leios-conditional-diffusion.md`](leios-conditional-diffusion.md), [`leios-fetch-scheme.md`](leios-fetch-scheme.md). ## Why it's #1 Every other candidate — TCPInfo scraping, `estBurstS` per-hop measurement, topology snapshot, per-mini-protocol tx-submission DeltaQ — produces an *input to a model*. This one produces the **output of the model directly measurable**. The single number it yields: $$\hat{F}_{\text{full}\mid C}(14) = \frac{\left|\{j : C_j \text{ AND all nodes received EB body by 14 s}\}\right|}{\left|\{j : C_j\}\right|}$$ That's Yves Hauser's §5.6 conditional probability, but *empirically* — no assumptions about loss rate, no assumptions about which cwnd law is running, no assumptions about path length, no i.i.d. assumption. Just count. Two possible outcomes: - **$\hat{F}_{\text{full}\mid C}(14) = 1.000$ across the campaign.** Yves's *body-only, honest-case* headline is validated on real conditions — necessary, but not the whole safety story. Closure coverage and adversarial diffusion ($G_{\text{adv}}$, the T22 targeted-release case) are separate questions this body-arrival count does not answer (see [`leios-conditional-diffusion.md`](leios-conditional-diffusion.md)). For what it covers, confidence in the CIP-0164 safety property increases. - **$\hat{F}_{\text{full}\mid C}(14) < 1$.** A direct observation of the Coretti-et-al.-Lemma-7-analog safety property failing on the mesh. Certification succeeded, but some honest node didn't have the EB body by the round deadline. That's not a modelling concern — it's an empirical safety-property violation. CIP-0164 has to decide whether to revise parameters, revise topology assumptions, or explicitly accept the observed failure rate. Neither outcome is diagnosable without this measurement. ## What comes free from the same instrumentation The per-EB arrival log is a small piece of instrumentation. Once you have it, eleven analyses fall out with no additional collection — $F_{\text{full}\mid C}(14)$ (the primary safety-property answer) plus ten aggregate by-products: | Analysis | What it answers | Gap / caveat closed | |---|---|---| | Empirical $F_{\text{full}\mid C}(14)$ | Does Yves's safety property hold on the real mesh? | Yves's caveat 6 (no MC validation) + §5.6 headline | | Slowest-node identity per EB | Does $\arg\max_i T_j^{(i)}$ follow the persistent, rotating, or bimodal (churn-induced) regime across rounds? Determines which intervention shape applies — per-node targeting only pays off under the persistent regime. | Adaptive lever #1 in `leios-conditional-diffusion.md` (three-regime framing under "Slowest-node identity and its temporal structure") | | Cross-node arrival correlation matrix | Is the i.i.d. assumption pessimistic or realistic? | Yves's caveat 1 + gap 2 (cross-node correlation) | | Empirical vs modelled $G(t)$ | Where does the report's CDF differ from reality? | gap 3 (path-length distribution) sanity check | | Reconstructed path-length distribution | Do EBs actually take shortest-graph-paths? | gap 3 reconstruction (complements pseudo-mainnet topology survey) | | Empirical $\alpha$ per node | How often does pre-diffusion actually work? | gap 2 (α as time-varying observable, not scalar) | | Seed-hop delivery rate | Was the winning path a big-ledger seed on each certified round? | Memo's 76% one-hop coverage claim, empirically validated per [`leios-fetch-scheme.md`](leios-fetch-scheme.md) | | Orphan-set regime (persistent / rotating / bimodal) | Is the 24% orphan set the same nodes every round (persistent), a different rotating subset (rotating), or a per-round-type pattern where sparse-origin rounds systematically orphan more of the mesh (bimodal)? | Determines whether "slowest-node targeting", Improvement #7 (historical seed-node preference), or Improvement #8 (diversity-aware churn) is the right intervention; see the "seed-node coverage is 76% marginal" note in `leios-fetch-scheme.md` | | Which-channel-wins per EB (big-ledger vs NBL) | If the LeiosFetch asymmetric scheme is deployed, which channel actually delivered the winning arrival? | Validates the memo's two-channel design; identifies where the seed-hop coverage lives | | Seed-node identity stability | The persistently-fastest-arriving node — a proxy for genuine seed peers vs lottery winners | Feeds Improvement #7 (historical seed-node preference) and the three-regime (persistent / rotating / bimodal) orphan-set analysis | | Per-block-origin coverage rate | Does mesh-wide seed-hop coverage depend on the block-producing region? What's the bimodal-round-coverage pattern ($c_D$ vs $c_S$)? | Quantifies the churn tendency to concentrate big-ledger peers around block-production regions: sparse-origin blocks see low mesh-wide coverage while dense-origin blocks see high coverage. See [`leios-fetch-scheme.md`'s Improvement #8 and appendix](leios-fetch-scheme.md#improvement-8-diversity-aware-big-ledger-churn) — the tendency's actual strength on real networks is what this measurement pins down | One instrumentation → eleven analyses. No other single measurement has this leverage. ## Minimum viable setup - **Per-EB arrival trace** at every measurement node: `(EB_id, arrival_ts)`. `cardano-tracer` already has block-arrival events; the analog for EBs is a small extension to the tracing plumbing, not new mux machinery. - **Certification outcome per EB** from `cardano-tracer`'s existing EB-certification traces. - **Cross-node join**: 100 measurement nodes × ~30 k EBs/week × ~100 bytes per event ≈ 300 MB/week of raw log. Trivial to ship, trivial to store. Doesn't require the branch's DeltaQ/RTT machinery on production nodes. Doesn't require the `tracetcpinfo` cabal flag. Works on `main` today with a small trace extension. ## Do we need edge-level tracking too? Arrival times alone are sufficient for the primary goal and for the majority of the by-product analyses. Edge-level tracking (from-whom-to-whom) buys refinement, not the headline number. Most of the analyses need nothing more than `(EB_id, node, arrival_ts)`, optionally joined with per-round metadata (block-producer identity, per-connection big-ledger flag): - Empirical $F_{\text{full}\mid C}(14)$ — pure count over EB population. - Slowest-node identity per EB — $\arg\max_i T_j^{(i)}$. - Cross-node arrival correlation matrix — $\mathrm{Cov}(T_i, T_j)$ over EB population. - Empirical vs modelled $G(t)$ — CDF of arrival times. - Empirical $\alpha$ per node — join with mempool state, not path. Only path-length reconstruction (gap 3) benefits from explicit edge tracking. Even there, arrival-time-only inference gives a defensible answer. ### Statistical inference from arrival times + peer graph With the peer graph in hand (from topology survey, e.g. pseudo-mainnet or a live-testnet snapshot), we can infer the most likely propagation tree per EB by matching arrival-time ordering against peer-graph adjacency: each node's most-likely forwarder is the earliest-arriving peer that arrived before it. Aggregated over the EB population, this yields the *marginal* distribution of path lengths — the aggregate shape, ignoring per-EB detail — which is what gap 3 (`leios-report-gaps.md`) asks about. *What "marginal convergence" means.* The marginal distribution is the aggregate over the whole EB population, ignoring per-EB conditioning. Convergence means the empirical estimator approaches its population counterpart as sample size grows. So even if we mis-infer the actual path for any single EB, the empirical frequency of 3-hop, 4-hop, 5-hop paths across many EBs approaches the true frequencies - *provided the inference errors are unbiased*. Arrival-time-based path inference tends to construct spanning trees along fastest routes. If real EB propagation actually followed a slower route (e.g., the fastest-graph route was congested and delivery came via a longer path), inference *underestimates* the true path length. Direction: inferred path length ≤ true path length. Gap 3 is directionally useful even without correcting the bias: - If inferred paths are longer than 3.6 hops in aggregate, the true paths are at least that long → the report is optimistic. Clear finding. - If inferred paths are close to 3.6 hops, we've established a lower bound but haven't proven the report right — true paths could be longer. - If inferred paths come out *shorter* than 3.6 hops, this is the confounded direction: the one-sided downward bias means a short reading cannot distinguish "genuinely short paths" from "longer paths the fastest-route inference shortened". *Not* a clear finding on its own — resolve it by recording `arrived_from_peer` (below), which removes the inference entirely. The bias direction favours detecting *worse-than-report* outcomes — convenient for a validation frame focused on catching unpleasant tail surprises. ### What explicit edge tracking adds Records for every EB, at every node, `(EB_id, arrived_from_peer, arrival_ts)`. Gives per-EB exact paths with no inference, detects forwarding pathologies (loops, redundant retransmissions), and enables per-EB diagnosis when $F_{\text{full}\mid C} < 1$ (not just "node X is generally slow" but "node X → node Y transit was slow on this specific EB"). *Cost comparison.* Arrival-time-only: ~300 MB/week (as above). Edge-level: 5–10× that, with cross-node join complexity (matching sender-receiver-EB triples across nodes). ### Phased recommendation - **Phase 1**: arrival-time-only logging (with block-producer identity and per-connection big-ledger flag joined per round). Delivers $F_{\text{full}\mid C}$ and the aggregate by-product analyses. Path-length inference via arrival-time ordering + peer graph — enough to answer gap 3's headline directionally. - **Phase 2**: add `arrived_from_peer` to each event if Phase 1 shows $F_{\text{full}\mid C} < 1$, or if per-EB path fidelity becomes important. Turns statistical inference into direct observation. - **Free enrichment**: if `arrived_from_peer` is already in the trace stream at the mux layer (it usually is), persist it in Phase 1 — no new instrumentation needed, just enable the extra field in the log schema. Turns Phase 2 into "no code changes required, enable additional log field". Prefer this whenever the field is available: recording the forwarder observes path length directly and so **removes the fastest-route inference bias** entirely, rather than leaving it to be reasoned around. *Why the block-producer-identity join.* The producer is each EB's diffusion root (hop 0): path length and per-origin $G(t)$ are measured *from* it, so gap-3 reconstruction needs it; its region classifies the round for the dense-vs-sparse bimodality ($c_D$ / $c_S$, Improvement #8); and it lets per-producer anomalies surface (a persistently-slow producer, or the closure-poisoning case where mesh-wide cache-misses concentrate on one producer — threat Proposal G). ## What it doesn't give (for balance) - **Per-hop CDF.** For that you still need `estBurstS`, ideally on the branch's `MiniProtocolNum`-enriched build. The per-EB arrival log gives end-to-end, not per-hop. - **Loss rate `p` distribution.** For that, TCPInfo (Linux-only, testnet-only, `tracetcpinfo` compile flag) is still the right tool. - **Vote diffusion.** The log captures EB body arrivals and certification outcomes, not vote arrivals. The certification-timing budget's second leg — 450 votes reaching the certifying producer within the 7 s deadline — needs separate per-vote instrumentation; without it a failed round can't be attributed (body-late vs votes-late vs vote head-of-line-blocked behind bulk). See [`leios-testnet.md`'s vote-diffusion section](leios-testnet.md#vote-diffusion-head-of-line-blocking-and-praos-priority). - **Closure coverage.** The log as specified records EB *body* arrival; the safety-binding question is *closure* completion (the up-to-12 MB tx set). Add a closure-complete timestamp per node per EB — see [`leios-conditional-diffusion.md`'s closure-coverage section](leios-conditional-diffusion.md#closure-coverage). - **Multi-round feedback dynamics.** Yves's formula and this measurement are single-round. The whole of round-to-round coupling (gap 4) is out of scope: positive-feedback autocorrelation across consecutive rounds — whose runaway extreme is the [cascade to livelock](leios-testnet.md#cascade-to-livelock-bounding-sustained-overload-failure) — plus recovery signatures (monotone vs oscillatory) and round concurrency (gap 18). Each needs a distinct multi-round analysis; the livelock is one worst case within it, not the whole of it. - **Runtime signal.** This is a validation measurement, not a production runtime signal. Adaptive levers derived from it (slowest-node targeting, adaptive body sizing) require the runtime `PeerRTT` reader-view machinery. But those are refinements. The primary safety property is measurable *first*, and if it fails, that alone dictates the next steps. ## Runner-ups, briefly - **Per-mini-protocol tx-submission DeltaQ.** Second place. Answers whether the 1-hop pre-diffusion approximation actually holds. High payoff, but focused on one aspect of the pipeline (tx-gossip health), and only partially addresses the network-safety question. - **TCPInfo scraping** on Linux measurement nodes. Third place. Answers Recommendation 5 of the report (loss rate, rwnd state) with high confidence. But testnet-only, and doesn't give any runtime lever. - **Topology snapshot** via pseudo-mainnet extension. Cheap and quickly available (a day's work). Answers gap 3's structural component. But subsumed by the per-EB arrival log's path-length reconstruction, which measures actual diffusion paths rather than shortest-graph-paths. - **`estBurstS` cross-region sweep.** Fourth place. Falsifies the Mathis-vs-CUBIC choice with high precision but doesn't answer the safety-property question directly. ## The framing For protocol **robustness**: cross-node per-EB arrival correlation is the single measurement that would empirically confirm or falsify the network-safety property Linear Leios inherits from Full Leios. Every other empirical concern is downstream of "does this property hold?" For protocol **performance**: the same log identifies the slowest nodes per round and classifies the temporal regime (persistent / rotating / bimodal — see `leios-conditional-diffusion.md`). Per-node operational intervention is high-leverage only under the persistent regime; under rotating it doesn't apply, and under bimodal (churn-induced) a per-round-type intervention like Improvement #8's diversity-aware churn is the right response. The log quantifies how much any of those interventions moves the aggregate. For **protocol-design validation** (specifically, evaluating candidate LeiosFetch schemes): the same log answers the four-configuration experiment in [`leios-fetch-scheme.md`](leios-fetch-scheme.md#suggested-experimental-design) — empirical $F_{\text{full}\mid C}$ under each configuration, which channel wins per EB (big-ledger vs NBL), redundant egress per round, and orphan-set behaviour. The primary measurement is thus the shared substrate for both validating the report's model *and* evaluating concrete protocol-design proposals against it. One log; eleven analyses; three distinct use cases (safety-property validation, performance operational lever, LeiosFetch-scheme comparison). That's #1.