With most information hidden, the game Stratego had stumped AI—until now

A research team from Carnegie Mellon, MIT, NYU and Stanford built an AI called Ataraxos that defeats top Stratego players. Ataraxos beat Pim Niemeijer 15–1 with four draws and was trained on 16 GPUs for a cost of a few thousand dollars.

By AI Newsroom· Reviewed by Pranav, Founder & Editor-in-ChiefPublished 30 minutes agoUpdated 30 minutes ago0 views
With most information hidden, the game Stratego had stumped AI—until now

Why It Matters

Stratego resisted prior AI dominance because it combines extremely large hidden-state space with long game horizons and strategic bluffing. Cracking it demonstrates new progress in imperfect-information game solving techniques that could transfer to other domains with prolonged, high-uncertainty decision processes.

Key Facts

  • AI name: Ataraxos
  • Research institutions: Carnegie Mellon, MIT, New York University, Stanford
  • Match result vs Pim Niemeijer: 15 wins for Ataraxos, 1 loss, 4 draws
  • Compute used to train: 16 GPUs
  • Training cost: a few thousand dollars

Stratego presents a different challenge from earlier board- and card-game milestones in AI because most of the game state is hidden and unfolds slowly. Each player fields 40 pieces with ranks, bombs, and a flag; identities are only revealed when pieces clash, and a full board can represent more than a decillion possible piece arrangements. Games can also run far longer than typical chess matches, sometimes extending to thousands of moves, and human play involves deliberate bluffing to manipulate an opponent’s beliefs.

Those features helped keep Stratego out of reach for earlier systems, including DeepMind’s DeepNash, which the researchers say could not reliably beat the best human players. To address the problem, the Ataraxos team added a second neural network whose role is to infer the likely identities of an opponent’s hidden pieces during play. That inference network works alongside the policy network to guide decisions under uncertainty, enabling Ataraxos to represent and update beliefs about the unseen board as the game progresses.

Using this architecture, the researchers trained Ataraxos with relatively modest resources: 16 GPUs and a total training cost on the order of a few thousand dollars. In head-to-head play, the system decisively outperformed Pim Niemeijer, regarded as one of the strongest living Stratego players, winning 15 games while losing once and drawing four times.

The result highlights a technical advance in handling vast hidden-state spaces and long-horizon decision making in imperfect-information settings. By combining a belief-estimation component with conventional policy learning, the team showed an approach that scales to the particular difficulties of Stratego’s combination of secrecy, duration, and strategic deception.

Keep Reading