Skip to content

Analytics

Niclas edited this page Jul 18, 2026 · 3 revisions

Analytics

Analytics measure a trained Profile rather than just trust it. They live in src/analytics.rs; the sampled (Monte-Carlo) variants live in src/sampled.rs.

The central rule: for 2-player zero-sum games exploitability is a hard Nash proof; for N-player or general-sum games it is only a quality signal.

Exact analytics (analytics)

  • best_response(game, profile, player): value to player when it plays a best response to everyone else's fixed strategies, respecting information sets (a single action per information set, maximizing the reach-weighted counterfactual value).
  • best_responses(game, profile): per-player vector; the value each player secures by best-responding.
  • exploitability_2p_zerosum(game, profile, game_value): the Nash-proof metric: (BR₀ − gv₀) + (BR₁ − gv₁), guaranteed non-negative, tends to 0 at Nash. Pass the known analytic value (Kuhn [-1/18, 1/18], RPS [0, 0]).
  • exploitability_symmetric(game, profile): assumes a symmetric zero-sum game with value [0, 0].
  • exploitability_zerosum(game, profile): N-player zero-sum sum of best responses. Not a Nash proof: the sum can be negative, and even 0 is only necessary, not sufficient, for Nash.
  • expected_payoffs(game, profile): exact full-tree expected payoff to each player when everyone plays the profile. Panics on sample_chance failure; use expected_payoffs_sampled when you need error handling.
  • nash_conv(game, profile): Σᵢ (BRᵢ − uᵢ). Always non-negative. NashConv = 0 is necessary but not sufficient for Nash; it characterizes the weaker coarse-correlated equilibrium (NashConv = 0 ⇔ CCE). Use it as the convergence signal for N-player / general-sum games.
  • cfr_br(...): Counterfactual Best Response (Johanson et al. 2012): stochastic response sampled from a Boltzmann distribution over Q-values, and the estimate averages over draws. Lower variance than the deterministic argmax on sampled games.
  • exploitability_descent(...): the gradient direction g[I][a] = π⁻ⁱ(I)· (br_policy[I][a] − σ[I][a]) for smooth-best-response / exploitability-descent updates.

Enumerable vs. non-enumerable chance

For games with enumerable chance (chance_outcomes returns Some everywhere), every walk is an exact full-tree traversal. For games with non-enumerable chance, the exact walk would not terminate, so a Monte-Carlo estimate of the chance reach is used: a stochastic estimate. Because decide_infoset argmaxes over noisy estimates, the returned best-response value is biased high (maximization bias). Use best_response_sampled to control the sample budget; larger budgets shrink the variance and the bias.

exploitability dispatcher

The one entry point when you do not know the game shape:

  • 2-player zero-sum: if nash_value() returns Some(v), call exploitability_2p_zerosum(v); else exploitability_symmetric.
  • N > 2 / general-sum: nash_conv.

For a 2-player general-sum game where nash_value() is None, the dispatcher silently assumes value [0, 0] (exploitability_symmetric). That is wrong for a genuine general-sum game. Call nash_conv directly instead.

These functions return Result (not panics) on a malformed game (sample_chance failure). expected_payoffs panics on failure because the rest of the crate relies on its bare Vec<f64> return; use expected_payoffs_sampled when you need error handling.

Sampled analytics (sampled)

For games too large for the exact full-tree analytics: the same games where MCCFR shines.

  • SampleConfig { playouts, posterior_samples, seed }: default playouts = 2000, posterior_samples = 64.
  • sampled_best_response / sampled_best_responses: consistent, sampled estimator with a confidence interval.
  • sampled_exploitability_2p_zerosum(game, profile, game_value, cfg): sampled proxy for the exact 2p zero-sum metric.

SampledEstimate { mean, std_error, samples } reports ci95() = 1.96·std_error, lower(), upper(). Treat mean ± ci95() as an approximate 95% CI.

This is not a Nash proof

Even for 2-player zero-sum games, the sampled estimator is not an exact bound and not a Nash proof. It is a quality signal with an interval. The argmax over noisy action-value estimates slightly overestimates the true best response (the usual max-over-noise effect); bias shrinks as posterior_samples grows. Calibrate on a small proxy game (confirm the mean matches exploitability_2p_zerosum and the interval contains it) before trusting it on the real large game.

CFR-BR and exploitability descent

The exact cfr_br and the gradient from exploitability_descent are the building blocks for exploitability-descent / smooth-best-response training loops, where you step σ ← σ + lr·g and renormalize per information set.

Clone this wiki locally