Repository navigation
v1.5.0: True Nash range-vs-range
True Nash range-vs-range via Brown's vector-form CFR
v1.5.0 closes the documented architectural gap in v1.4.x: the range_aggregator.solve_range_vs_range path was an Option B approximation (chance-enum-at-root with averaged-strategy aggregation), which is structurally non-Nash for true range-vs-range play. v1.5.0 ships a Rust port of Brown's vector-form CFR (trainer.cpp:138-209, MIT-licensed) that walks the betting tree once per iteration with per-infoset hand_count x action_count regret + strategy_sum tables — the canonical CFR architecture for solving public-game-tree range-vs-range Nash equilibria.
Acceptance test caveat (RESOLVED in v1.6.1)
The v1.5.0 release initially shipped with the Brown apples-to-apples acceptance test in a FAILING state. The failure was traced to two compound TEST bugs (not solver bugs):
-
Action-ordering mismatch: Brown's binary emits actions in order
[c, f, r_low, r_med, r_high]; our Rust solver emits[f, c, r_low, r_med, A]. The test compared position-by-position, so position 0 lined up Brown's CALL with our FOLD. The cell that looked like "Brown=0% call, Rust=99% call" was actually "Brown=0% call, Rust=99% FOLD" — both engines agreeing on FOLD with a labeling mismatch. -
Range-to-player-slot misassignment: Brown's P0 acts first; our P1 acts first. The test passed ranges in the wrong slots, so the two engines weren't even solving the same game.
After fixing both test bugs (PR 40 in v1.6.1), the acceptance test passes; residual divergence is ~10% of cells with max ~0.10 magnitude — characteristic Nash polytope sizing-mix non-uniqueness (legit), not bug territory.
The PR 23 vector-form CFR algorithm itself was algorithmically correct from the start. Brown's cpp/trainer.cpp:138-209 line-by-line match was confirmed by audit. v1.5.0's solve_range_vs_range_rust entry point delivers genuine vector-form Nash; the empirical confirmation just had to wait for the test plumbing to be debugged.
The previously reported DCFR "100x slowdown" against v1.4.x was a measurement artifact (apples-to-oranges configuration), not a real regression. Scalar-path performance is byte-identical to v1.4.0.
Added
- New Rust entry point
_rust.solve_range_vs_range_rust(config_json, iters, alpha, beta, gamma, p0_holes=None, p1_holes=None) -> dict. Output dict matches the Python tier's shape (average_strategykeyed by<hole>|<board>|<street>|<history>), plusdecision_node_count,iterations,wallclock_seconds,hand_count_per_player,memory_profile,backend = "rust_vector". Game::hand_count()trait method (default1) for backward-compatible opt-in. All existing scalar paths (DCFRSolver<G>,hunl_solver.rs,preflop.rs) remain byte-identical to v1.4.0.- Per-street memory profiler (
VectorMemoryProfile) surfaced in the PyO3 dict'smemory_profilefield withtotal_bytes,infoset_count,bytes_by_street,infoset_count_by_street. - Brown apples-to-apples acceptance test (
tests/test_v1_5_brown_apples_to_apples.py, opt-in via-m parity_noambrown) — runs Brown'sriver_solver_optimizedreference binary against the new Rust vector-form CFR ondry_K72_rainbowanddry_A83_rainbow, comparing average strategies at matching histories. Now passes on both spots after the v1.6.1 test-plumbing fix (PR 40). - Differential tests (
tests/test_range_vs_range_rust_diff.py) — 4 active tests covering exploitability-under-restricted-game vs Pythondcfr.py, structural smoke checks, and end-to-end binding chain verification. Both Python and Rust achieve <= 0.05 BB exploitability on the Case A spot.
Performance
- ~72x faster than the Python aggregator on a medium 10x10 RvR river spot (measured; honest single-machine comparison on macOS arm64).
Unchanged
- All existing scalar diff tests (Kuhn / Leduc / fixed-combo river / exploit / node-locking) remain green; 40 differential tests confirm v1.4.0 byte-identical behavior on the scalar paths.
- Public API surface for
solve_hunl_postflop/solve_hunl_preflop/solve()is unchanged. v1.4.3 input-validation guards (PR 31) remain in force.
Caveats and v1.5.x roadmap
- Postflop only in v1.5.0 — preflop RvR deferred to v1.5.1 (16 GB memory edge at full-1326 hand vector without suit-iso reduction).
- Terminal-leaf O(N^2) blocker check is the next perf candidate (SIMD kernels expected to deliver 4-8x speedup based on PR 8 NEON-on-scalar experience).
- EMD hand-bucketing in the vector dimension deferred to v1.5.1.
range_aggregator.solve_range_vs_rangeis not yet wired to the new Rust tier (Q3 default for v1.5.0 keeps it internal-only); user-facing surfacing deferred to a later minor.
License
The Rust port follows Brown's MIT-licensed reference implementation with attribution in source comments.