Skip to content

Connect-4 n-tuple agent trained by self-play (TD(λ), 25,000 steps)

Latest

Choose a tag to compare

@MarkusThill MarkusThill released this 26 Sep 12:46

Weights of an n-tuple network (200 tuples of eight cells, two tables per tuple) trained with
lab2/1_train_ntuple_net.ipynb (commit c18a0b5) on an NVIDIA A100: 50,000 boards in parallel, ε = 0.1,
truncated λ-return with λ = 0.7 and n = 5, target network with τ = 0.15, Adam with learning rate 3·10⁻⁴,
25,000 steps (8.6 minutes of training). It is the strongest of ten independent runs in a tournament of 200
games per condition against BitBully opponents of increasing strength.

On 200 fresh games per condition, moving first against BitBully at full strength (perfect play), it wins 182
and draws 18, without a loss.

from bitbully import Board
from techdays26.td_agent import TDConnect4AgentTorch

agent = TDConnect4AgentTorch(model_path="connect4-ntuple-agent.pt")  # runs on the CPU
print(agent.best_move(Board()))  # 3

SHA-256: b84b954820842b1f71983f91d1bac28cbe8fb652eeb4ea44afb1586eb6af8f1c

Details: https://markusthill.github.io/blog/2026/near-perfect-connect4-in-five-minutes/