Repository navigation
Analytics
Analytics measure a trained Profile rather than just trust it. They live in
src/analytics.rs; the sampled (Monte-Carlo) variants live in src/sampled.rs.
The central rule: for 2-player zero-sum games exploitability is a hard Nash proof; for N-player or general-sum games it is only a quality signal.
-
best_response(game, profile, player): value toplayerwhen it plays a best response to everyone else's fixed strategies, respecting information sets (a single action per information set, maximizing the reach-weighted counterfactual value). -
best_responses(game, profile): per-player vector; the value each player secures by best-responding. -
exploitability_2p_zerosum(game, profile, game_value): the Nash-proof metric:(BR₀ − gv₀) + (BR₁ − gv₁), guaranteed non-negative, tends to 0 at Nash. Pass the known analytic value (Kuhn[-1/18, 1/18], RPS[0, 0]). -
exploitability_symmetric(game, profile): assumes a symmetric zero-sum game with value[0, 0]. -
exploitability_zerosum(game, profile): N-player zero-sum sum of best responses. Not a Nash proof: the sum can be negative, and even 0 is only necessary, not sufficient, for Nash. -
expected_payoffs(game, profile): exact full-tree expected payoff to each player when everyone plays the profile. Panics onsample_chancefailure; useexpected_payoffs_sampledwhen you need error handling. -
nash_conv(game, profile):Σᵢ (BRᵢ − uᵢ). Always non-negative.NashConv = 0is necessary but not sufficient for Nash; it characterizes the weaker coarse-correlated equilibrium (NashConv = 0 ⇔ CCE). Use it as the convergence signal for N-player / general-sum games. -
cfr_br(...): Counterfactual Best Response (Johanson et al. 2012): stochastic response sampled from a Boltzmann distribution over Q-values, and the estimate averages over draws. Lower variance than the deterministic argmax on sampled games. -
exploitability_descent(...): the gradient directiong[I][a] = π⁻ⁱ(I)· (br_policy[I][a] − σ[I][a])for smooth-best-response / exploitability-descent updates.
For games with enumerable chance (chance_outcomes returns Some
everywhere), every walk is an exact full-tree traversal. For games with
non-enumerable chance, the exact walk would not terminate, so a Monte-Carlo
estimate of the chance reach is used: a stochastic estimate. Because
decide_infoset argmaxes over noisy estimates, the returned best-response value
is biased high (maximization bias). Use best_response_sampled to control
the sample budget; larger budgets shrink the variance and the bias.
The one entry point when you do not know the game shape:
- 2-player zero-sum: if
nash_value()returnsSome(v), callexploitability_2p_zerosum(v); elseexploitability_symmetric. - N > 2 / general-sum:
nash_conv.
For a 2-player general-sum game where nash_value() is None, the dispatcher
silently assumes value [0, 0] (exploitability_symmetric). That is wrong for a
genuine general-sum game. Call nash_conv directly instead.
These functions return Result (not panics) on a malformed game
(sample_chance failure). expected_payoffs panics on failure because the rest
of the crate relies on its bare Vec<f64> return; use expected_payoffs_sampled
when you need error handling.
For games too large for the exact full-tree analytics: the same games where MCCFR shines.
-
SampleConfig { playouts, posterior_samples, seed }: defaultplayouts = 2000, posterior_samples = 64. -
sampled_best_response/sampled_best_responses: consistent, sampled estimator with a confidence interval. -
sampled_exploitability_2p_zerosum(game, profile, game_value, cfg): sampled proxy for the exact 2p zero-sum metric.
SampledEstimate { mean, std_error, samples } reports ci95() = 1.96·std_error,
lower(), upper(). Treat mean ± ci95() as an approximate 95% CI.
Even for 2-player zero-sum games, the sampled estimator is not an exact bound
and not a Nash proof. It is a quality signal with an interval. The argmax
over noisy action-value estimates slightly overestimates the true best response
(the usual max-over-noise effect); bias shrinks as posterior_samples grows.
Calibrate on a small proxy game (confirm the mean matches
exploitability_2p_zerosum and the interval contains it) before trusting it on
the real large game.
The exact cfr_br and the gradient from exploitability_descent are the building
blocks for exploitability-descent / smooth-best-response training loops, where you
step σ ← σ + lr·g and renormalize per information set.