Repository navigation
Interior perturbed strategies
The rtcfr module produces a near-Nash strategy in which every action keeps a
probability mass of at least γ. It trains a plain CFR+ solver (external-sampling
MCCFR with the CFR+ regret and averaging schedule), reads the CFR+ average, and
nudges every information set's distribution a small amount toward uniform so the
result lands strictly inside the γ-interior of the probability simplex. The
perturbation strength γ is annealed toward 0 over the training rounds, so the
returned profile stays close to the raw CFR+ average.
Despite the module name, this is not a last-iterate "Reward-Transform CFR+" algorithm. It exposes no last iterate, keeps no auxiliary regularizer or reference point, and applies no reward transform to the regret accumulation. It returns the CFR+ average (the near-Nash object CFR+ converges on for 2-player zero-sum games) with a controllable, annealed γ-interior perturbation applied to the final output vectors.
A raw CFR+ average can drive some actions to zero probability. That is correct for equilibrium play, but a downstream consumer sometimes needs every action to retain non-zero probability:
- A robust behavior policy that never fully commits, so it keeps responding when the opponent leaves the equilibrium path.
- A sampling policy with bounded importance weights, where a zero-probability action would make a weight blow up or undefined.
For plain equilibrium solving, use the Trainer (see Training) or
Profile::train. Reach for rtcfr only when the strictly-interior property is
the point.
rtcfr_solve(game, iters_per_round, rounds, seed) runs CFR+ for rounds rounds
of iters_per_round iterations each (the trainer accumulates across rounds) and
returns the perturbed average. Both iters_per_round and rounds must be
positive; passing zero panics. The result is near-Nash for 2-player zero-sum
games, with every action retaining strictly positive probability.
use regret::rtcfr::rtcfr_solve;
let profile = rtcfr_solve(game, 10_000, 20, 0xABCD);
let dist = profile.policy(&game, 0); // every entry is strictly positivertcfr_solve_with(game, iters_per_round, rounds, seed, gamma_init, anneal_every)
adds explicit control over the perturbation. gamma_init is the initial strength
γ₀; every anneal_every rounds γ is halved. A smaller gamma_init or a shorter
anneal period biases the profile closer to the raw CFR+ average; a larger value
keeps every action further inside the simplex interior. rtcfr_solve uses the
defaults rtcfr_default_gamma (0.03) and rtcfr_default_anneal_every (5).
perturb_uniform(profile, gamma) applies the γ-interior nudge on its own, to any
Profile. It is only valid for gamma < 1/|A(I)| at an information set with
|A(I)| actions; a larger γ would drive the scale 1 − γ·|A(I)| negative and
the result would no longer be a distribution. To stay total, the function clamps
γ per information set to just below 1/|A(I)|, so the output is always a valid
γ-interior point. Information sets never visited during training are absent from
the map and fall back to the uniform strategy from Profile::policy.
For an information set I, each action's perturbed probability is
σ̂[a] = (1 − γ·|A(I)|)·σ[a] + γ
The added γ terms sum to γ·|A(I)|, exactly cancelling the γ·|A(I)| subtracted
from the scaled σ, so the vector is already a distribution; it is renormalized
defensively against float round-off. Every entry stays at least γ, strictly
inside the interior.
The whole module is built on the public Trainer, Profile, and Game API and
touches no trainer internals, so it is a faithful example of composing the crate
rather than a special path.