Skip to content

Getting Started

Niclas edited this page Jul 18, 2026 · 4 revisions

Getting started

regret trains a strategy for an imperfect-information game by running Counterfactual Regret Minimization. You do not need to understand the algorithm to get a usable strategy out: pick a game, train, read the result. This page uses the bundled KuhnPoker so the example compiles against the current crate without writing a game first.

Add the dependencies

regret is published on crates.io. regret-games (which holds KuhnPoker and the other reference games) is not published, so pull it from the git source:

[dependencies]
regret = "2.2.1"
regret-games = { git = "https://github.com/sqlongithub/regret" }

use regret::prelude::*; brings the types you need: Game, Trainer, Profile, Algorithm, RegretVariant, and the analytics functions.

Train with one call

Profile::train(game, iters) runs the recommended default (external-sampling MCCFR with DCFR+) and returns a Profile. This is the common case:

use regret::prelude::*;

let game = regret_games::KuhnPoker::new();
let profile = Profile::train(game.clone(), 20_000);

KuhnPoker::new() starts at the deal (chance) node. Training is deterministic for a given iteration count on a single worker; pass .seed(...) for a reproducible run when you tune the Trainer directly.

Train through the builder

When you want a specific algorithm, regret variant, seed, or pruning, use the Trainer builder and finish with .train(iters).unwrap().profile():

use regret::prelude::*;

let game = regret_games::KuhnPoker::new();
let trainer = Trainer::new(game)
    .algorithm(Algorithm::ExternalSampling)
    .regret(RegretVariant::DCFRPlus { alpha: 1.5, gamma: 4.0 })
    .seed(0xC0FFEE);

trainer.train(20_000).unwrap();
let profile = trainer.profile();

.train returns a Result; .unwrap() is fine here because training only fails on a malformed game or a worker pool error. Keep the game value alive if you want to query the profile against it afterward (as on this page).

Read the profile

Profile maps each information set to a probability distribution over actions. Read it with policy(game, player), which returns a Vec<f64> aligned to game.legal_actions(). For an information set never visited during training, policy returns the uniform distribution (a safe fallback, not an error). Use policy_strict when a wrong move is worse than a loud failure.

act and act_stochastic panic unless game is a decision node for player (i.e. game.player_to_act() is PlayerToAct::Player(player)). Call them only at a node where player is to move; otherwise the library panics. They are for playing the learned strategy, not for inspecting it.

To inspect a distribution, drive the game to a decision node and read policy:

use regret::prelude::*;

let game = regret_games::KuhnPoker::new();
let profile = Profile::train(game.clone(), 20_000);

// Player 0 holds the Jack (card 0) at the opening decision (no bet outstanding).
let mut g = game.clone();
g.step(&regret_games::KuhnAction::Deal(0, 1)).unwrap();
let dist = profile.policy(&g, 0);
println!("{:?}", dist);                 // one probability per legal action
println!("{}", profile.pretty(&g, 0));  // "C: 0.333, B: 0.667" style

The indices of dist line up with g.legal_actions(), which here is [Check, Bet]. pretty prints each action paired with its probability, so you do not have to zip the vector against legal_actions() yourself.

Check the strategy quality

For a 2-player zero-sum game, exploitability_2p_zerosum is a hard Nash proof: it is non-negative and tends to 0 as the profile approaches equilibrium. Pass the known analytic Nash value per player. For Kuhn poker that value is [-1.0 / 18.0, 1.0 / 18.0] (player 0 is the disadvantaged first player):

use regret::prelude::*;

let game = regret_games::KuhnPoker::new();
let profile = Profile::train(game.clone(), 20_000);

let br = exploitability_2p_zerosum(&game, &profile, [-1.0 / 18.0, 1.0 / 18.0])
    .unwrap();
assert!(br < 0.01, "strategy should be near-Nash, got {br}");

For N > 2 or non-zero-sum games, exploitability is a quality signal only, not a Nash proof. Use exploitability when you do not know the game shape in advance; it dispatches to the tightest available metric. See Analytics for the dispatcher logic and its caveats.

Where to go next

  • Implementing your own game: Implementing a game
  • Tuning the trainer (algorithms, regret variants, pruning, parallelism): Training
  • Reading strategies, the checkpoint format, and provenance: Profiles
  • Measuring quality and the exploitability limits: Analytics

Clone this wiki locally