Skip to content

Interior perturbed strategies

Niclas edited this page Jul 18, 2026 · 3 revisions

Interior-perturbed strategies (rtcfr)

The rtcfr module produces a near-Nash strategy in which every action keeps a probability mass of at least γ. It trains a plain CFR+ solver (external-sampling MCCFR with the CFR+ regret and averaging schedule), reads the CFR+ average, and nudges every information set's distribution a small amount toward uniform so the result lands strictly inside the γ-interior of the probability simplex. The perturbation strength γ is annealed toward 0 over the training rounds, so the returned profile stays close to the raw CFR+ average.

Despite the module name, this is not a last-iterate "Reward-Transform CFR+" algorithm. It exposes no last iterate, keeps no auxiliary regularizer or reference point, and applies no reward transform to the regret accumulation. It returns the CFR+ average (the near-Nash object CFR+ converges on for 2-player zero-sum games) with a controllable, annealed γ-interior perturbation applied to the final output vectors.

When to use it

A raw CFR+ average can drive some actions to zero probability. That is correct for equilibrium play, but a downstream consumer sometimes needs every action to retain non-zero probability:

  • A robust behavior policy that never fully commits, so it keeps responding when the opponent leaves the equilibrium path.
  • A sampling policy with bounded importance weights, where a zero-probability action would make a weight blow up or undefined.

For plain equilibrium solving, use the Trainer (see Training) or Profile::train. Reach for rtcfr only when the strictly-interior property is the point.

Solving

rtcfr_solve(game, iters_per_round, rounds, seed) runs CFR+ for rounds rounds of iters_per_round iterations each (the trainer accumulates across rounds) and returns the perturbed average. Both iters_per_round and rounds must be positive; passing zero panics. The result is near-Nash for 2-player zero-sum games, with every action retaining strictly positive probability.

use regret::rtcfr::rtcfr_solve;

let profile = rtcfr_solve(game, 10_000, 20, 0xABCD);
let dist = profile.policy(&game, 0); // every entry is strictly positive

rtcfr_solve_with(game, iters_per_round, rounds, seed, gamma_init, anneal_every) adds explicit control over the perturbation. gamma_init is the initial strength γ₀; every anneal_every rounds γ is halved. A smaller gamma_init or a shorter anneal period biases the profile closer to the raw CFR+ average; a larger value keeps every action further inside the simplex interior. rtcfr_solve uses the defaults rtcfr_default_gamma (0.03) and rtcfr_default_anneal_every (5).

The perturbation

perturb_uniform(profile, gamma) applies the γ-interior nudge on its own, to any Profile. It is only valid for gamma < 1/|A(I)| at an information set with |A(I)| actions; a larger γ would drive the scale 1 − γ·|A(I)| negative and the result would no longer be a distribution. To stay total, the function clamps γ per information set to just below 1/|A(I)|, so the output is always a valid γ-interior point. Information sets never visited during training are absent from the map and fall back to the uniform strategy from Profile::policy.

For an information set I, each action's perturbed probability is

σ̂[a] = (1 − γ·|A(I)|)·σ[a] + γ

The added γ terms sum to γ·|A(I)|, exactly cancelling the γ·|A(I)| subtracted from the scaled σ, so the vector is already a distribution; it is renormalized defensively against float round-off. Every entry stays at least γ, strictly inside the interior.

The whole module is built on the public Trainer, Profile, and Game API and touches no trainer internals, so it is a faithful example of composing the crate rather than a special path.

Clone this wiki locally