Skip to content

Hard Level Bot AI Logic

m8rvin edited this page Apr 26, 2026 · 6 revisions

Hard Level Bot AI Logic

This page explains how the EvilBot makes decisions. Instead of Minimax, this bot uses a powerful algorithm called Monte Carlo Tree Search (MCTS) combined with a custom heuristic bias to navigate the Game of Y.


1. How the Bot Thinks (MCTS)

The bot doesn't evaluate every single possible future like Minimax. Instead, it plays thousands of rapid, simulated games in its head to figure out which moves lead to a win most often.

  • Time-Based Thinking: Instead of looking a fixed number of moves ahead, the bot is given exactly 4.8 seconds to think. It runs as many simulations as possible in that time limit.

  • The UCB Formula (Smart Exploration): When deciding which moves to simulate, the bot uses a mathematical formula (Upper Confidence Bound). This balances two things:

    • Exploitation: Playing moves that already have a high win rate.

    • Exploration: Trying out new moves that haven't been tested enough yet.

  • Playing out the Future (Simulation): Once it picks a path, the bot rapidly fills up the rest of the board to see who wins.

  • Learning (Backpropagation): After a simulated game ends, the bot traces its steps back and records whether that specific path resulted in a win or a loss, updating its statistics.


2. How the Bot Scores a Move

Unlike Minimax, MCTS doesn't need a complex scoring system to evaluate half-finished boards—it plays until the game is actually over. However, to make the random simulations smarter, EvilBot uses a Heuristic Bias to prioritize good cells while filling the board.

A. Touching the Sides

The bot knows that edges are valuable for winning.

  • If an empty cell touches Side A, B, or C, it gets +50 points for each side it touches.

B. Connectivity (Neighbors)

Pieces are stronger when they have many options to connect.

  • The bot counts how many valid neighboring cells an empty spot has.

  • It gets +5 points for every connected neighbor.

C. The "Biased" Playout

During the simulation phase, instead of placing pieces completely randomly, the bot sorts the remaining empty cells based on the heuristic scores above.

  • Result: The bot heavily prefers to simulate futures where players fight over the edges and highly connected central spots, rather than wasting time playing in useless corners.

  • At the very end of the simulation, it runs a fast pathfinding check (has_path) to see if the player successfully connected all three sides.


Technical Details

Component Logic
Search Limit Time-based: Exactly 4.8 seconds (4800 ms) per turn.
Memory / Tree Size The bot remembers up to 300,000 board states (nodes) in a single turn to optimize speed.
Win Condition Evaluates a true win state at the end of every simulation (connecting Side A, B, and C).
Selection Strategy Uses the UCB1 algorithm with an exploration constant of 1.414.
Final Move Choice Picks the move that was visited the most during the 4.8 seconds of simulation, as this represents the most reliable path to victory.

Clone this wiki locally