Repository navigation
Deep CFR
Experimental. The deep-CFR surface (
DeepTrainer,EscherTrainer,DreamTrainer,PredictiveDeepTrainer,RobustDeepTrainer,Sample,ReservoirBuffer,TrainCadence,Approximator, and friends) is not covered by the crate's stability guarantees and may change in a minor release. Pin an exact version if you depend on it. This path is gated behind thedeepfeature and is single-threaded: it does not use rayon, so for parallelism run independent processes and ensemble the output profiles.
The deep-CFR path exists for games too large to tabulate exactly: instead of one
regret entry per information set, it samples trajectories and trains a function
approximator to predict each information set's regret (and average strategy) from
its features. The library imposes no ML dependency. Supply an Approximator or
use the provided dependency-free TabularApproximator (exact in the limit of
sampling) or LinearApproximator.
-
DeepTrainer: external-sampling Deep CFR (Brown et al. 2019). Traverse the tree, record counterfactual regret and average-strategy mass per information set, then fit the regret approximatorRand the average-strategy approximatorSumto the sampled(I, r)/(I, σ)pairs. ReturnSum(I). -
EscherTrainer: ESCHER (McAleer et al. 2023). The updating player samples from a fixed policy; the counterfactual regret is read from a learned history-value functionqᵢ, no importance-sampling ratio. -
DreamTrainer: DREAM (Steinberger et al. 2020). Outcome-sampling deep MCCFR whose sampled counterfactual values are centered on a learned advantage (Q) baseline, with epsilon-exploration keeping importance weights finite. -
PredictiveDeepTrainer(RD-002): DeepPDCFR+ (Xu et al. 2025). Mirrors Predictive Discounted CFR+ at the sampling level: the traversing player's strategy is the regret match of the one-step-ahead predicted regret, and the cumulative-advantage buffer bootstraps with the DCFR+ discount plus zero floor. -
RobustDeepTrainer(RD-003): Robust Deep MCCFR (El Jaafari 2025). Outcome-sampling deep MCCFR made robust via a delayed target network, epsilon-exploration mixing, and trust-region regret / importance-weight clipping.
Approximator is the plug-in trait: regret_strategy, avg_strategy,
observe_regret, observe_avg. TabularApproximator is the crate-local faithful
Deep CFR; LinearApproximator is a linear baseline. infer_max_actions(game, trajectories, seed) heuristically infers the max legal-action count by sampling
trajectories (an upper bound, not a proof) so the caller need not hand-derive the
max legal-action count.
DeepTrainer and the deep re-solving path are single-threaded (no rayon). For
parallelism, run independent processes and ensemble the output profiles.