Skip to content

Evaluating agents

Niclas edited this page Jul 18, 2026 · 2 revisions

Evaluating agents

src/eval.rs answers the practical question the analytics tools don't: if two agents actually played, who would win, and by how much? All three functions panic if profiles.len() != num_players() (and play_match also requires game.num_players() == start.num_players()), if the game is malformed (sample_chance / step fails), or if a subgame fails to terminate within DEFAULT_MAX_STEPS (10,000,000). The try_* variants return Result instead of panicking.

  • play_out(game, profiles, rng): sample one trajectory of a single subgame under the per-player policies; returns the terminal payoffs. Stochastic. Repeat and average for a win rate. Panics on a player-count mismatch.
  • play_match(game, profiles, start, transition, rng): chain subgames via a caller-supplied transition and play the full match (many rounds), accumulating terminal returns. transition(&terminal, &actions) -> Option<G> encodes how one resolved subgame advances to the next (None ends the match). Panics on a player-count mismatch, and also requires game.num_players() == start.num_players().
  • head_to_head(a, b): the exact expected payoff of profile a (player 0) versus b (player 1) via full-tree expectation over the joint profile, with no sampling noise. For non-enumerable-chance games it falls back to a fixed-seed 64-sample Monte-Carlo walk. Returns (payoff_to_0, payoff_to_1). Panics on a player-count mismatch.

A single play_out is a strict subset. It is equivalent to play_match with transition returning None immediately.

Fallible variants

  • try_play_out / try_play_match / try_head_to_head return Result instead of panicking. Use these when you must not panic, including when the player count is not known to be correct ahead of time.

Clone this wiki locally