Skip to content

Deep CFR

Niclas edited this page Jul 18, 2026 · 3 revisions

Deep CFR (optional deep feature)

Experimental. The deep-CFR surface (DeepTrainer, EscherTrainer, DreamTrainer, PredictiveDeepTrainer, RobustDeepTrainer, Sample, ReservoirBuffer, TrainCadence, Approximator, and friends) is not covered by the crate's stability guarantees and may change in a minor release. Pin an exact version if you depend on it. This path is gated behind the deep feature and is single-threaded: it does not use rayon, so for parallelism run independent processes and ensemble the output profiles.

The deep-CFR path exists for games too large to tabulate exactly: instead of one regret entry per information set, it samples trajectories and trains a function approximator to predict each information set's regret (and average strategy) from its features. The library imposes no ML dependency. Supply an Approximator or use the provided dependency-free TabularApproximator (exact in the limit of sampling) or LinearApproximator.

Schemes

  • DeepTrainer: external-sampling Deep CFR (Brown et al. 2019). Traverse the tree, record counterfactual regret and average-strategy mass per information set, then fit the regret approximator R and the average-strategy approximator Sum to the sampled (I, r) / (I, σ) pairs. Return Sum(I).
  • EscherTrainer: ESCHER (McAleer et al. 2023). The updating player samples from a fixed policy; the counterfactual regret is read from a learned history-value function qᵢ, no importance-sampling ratio.
  • DreamTrainer: DREAM (Steinberger et al. 2020). Outcome-sampling deep MCCFR whose sampled counterfactual values are centered on a learned advantage (Q) baseline, with epsilon-exploration keeping importance weights finite.
  • PredictiveDeepTrainer (RD-002): DeepPDCFR+ (Xu et al. 2025). Mirrors Predictive Discounted CFR+ at the sampling level: the traversing player's strategy is the regret match of the one-step-ahead predicted regret, and the cumulative-advantage buffer bootstraps with the DCFR+ discount plus zero floor.
  • RobustDeepTrainer (RD-003): Robust Deep MCCFR (El Jaafari 2025). Outcome-sampling deep MCCFR made robust via a delayed target network, epsilon-exploration mixing, and trust-region regret / importance-weight clipping.

Approximators

Approximator is the plug-in trait: regret_strategy, avg_strategy, observe_regret, observe_avg. TabularApproximator is the crate-local faithful Deep CFR; LinearApproximator is a linear baseline. infer_max_actions(game, trajectories, seed) heuristically infers the max legal-action count by sampling trajectories (an upper bound, not a proof) so the caller need not hand-derive the max legal-action count.

Single-threaded

DeepTrainer and the deep re-solving path are single-threaded (no rayon). For parallelism, run independent processes and ensemble the output profiles.

Clone this wiki locally