Repository navigation
Getting Started
regret trains a strategy for an imperfect-information game by running
Counterfactual Regret Minimization. You do not need to understand the algorithm
to get a usable strategy out: pick a game, train, read the result. This page uses
the bundled KuhnPoker so the example compiles against the current crate without
writing a game first.
regret is published on crates.io. regret-games (which holds KuhnPoker and
the other reference games) is not published, so pull it from the git source:
[dependencies]
regret = "2.2.1"
regret-games = { git = "https://github.com/sqlongithub/regret" }use regret::prelude::*; brings the types you need: Game, Trainer,
Profile, Algorithm, RegretVariant, and the analytics functions.
Profile::train(game, iters) runs the recommended default (external-sampling
MCCFR with DCFR+) and returns a Profile. This is the common case:
use regret::prelude::*;
let game = regret_games::KuhnPoker::new();
let profile = Profile::train(game.clone(), 20_000);KuhnPoker::new() starts at the deal (chance) node. Training is deterministic
for a given iteration count on a single worker; pass .seed(...) for a
reproducible run when you tune the Trainer directly.
When you want a specific algorithm, regret variant, seed, or pruning, use the
Trainer builder and finish with .train(iters).unwrap().profile():
use regret::prelude::*;
let game = regret_games::KuhnPoker::new();
let trainer = Trainer::new(game)
.algorithm(Algorithm::ExternalSampling)
.regret(RegretVariant::DCFRPlus { alpha: 1.5, gamma: 4.0 })
.seed(0xC0FFEE);
trainer.train(20_000).unwrap();
let profile = trainer.profile();.train returns a Result; .unwrap() is fine here because training only fails
on a malformed game or a worker pool error. Keep the game value alive if you
want to query the profile against it afterward (as on this page).
Profile maps each information set to a probability distribution over actions.
Read it with policy(game, player), which returns a Vec<f64> aligned to
game.legal_actions(). For an information set never visited during training,
policy returns the uniform distribution (a safe fallback, not an error). Use
policy_strict when a wrong move is worse than a loud failure.
act and act_stochastic panic unless game is a decision node for player
(i.e. game.player_to_act() is PlayerToAct::Player(player)). Call them only at
a node where player is to move; otherwise the library panics. They are for
playing the learned strategy, not for inspecting it.
To inspect a distribution, drive the game to a decision node and read policy:
use regret::prelude::*;
let game = regret_games::KuhnPoker::new();
let profile = Profile::train(game.clone(), 20_000);
// Player 0 holds the Jack (card 0) at the opening decision (no bet outstanding).
let mut g = game.clone();
g.step(®ret_games::KuhnAction::Deal(0, 1)).unwrap();
let dist = profile.policy(&g, 0);
println!("{:?}", dist); // one probability per legal action
println!("{}", profile.pretty(&g, 0)); // "C: 0.333, B: 0.667" styleThe indices of dist line up with g.legal_actions(), which here is
[Check, Bet]. pretty prints each action paired with its probability, so you
do not have to zip the vector against legal_actions() yourself.
For a 2-player zero-sum game, exploitability_2p_zerosum is a hard Nash proof:
it is non-negative and tends to 0 as the profile approaches equilibrium. Pass the
known analytic Nash value per player. For Kuhn poker that value is
[-1.0 / 18.0, 1.0 / 18.0] (player 0 is the disadvantaged first player):
use regret::prelude::*;
let game = regret_games::KuhnPoker::new();
let profile = Profile::train(game.clone(), 20_000);
let br = exploitability_2p_zerosum(&game, &profile, [-1.0 / 18.0, 1.0 / 18.0])
.unwrap();
assert!(br < 0.01, "strategy should be near-Nash, got {br}");For N > 2 or non-zero-sum games, exploitability is a quality signal only, not a
Nash proof. Use exploitability when you do not know the game shape in advance;
it dispatches to the tightest available metric. See Analytics for
the dispatcher logic and its caveats.
- Implementing your own game: Implementing a game
- Tuning the trainer (algorithms, regret variants, pruning, parallelism): Training
- Reading strategies, the checkpoint format, and provenance: Profiles
- Measuring quality and the exploitability limits: Analytics