A team of researchers from Carnegie Mellon University, New York University, Stanford University, and MIT has developed an artificial intelligence called Ataraxos. According to a paper published in Nature, it is the first AI to reach superhuman performance in the board game Stratego. In an official 20-game series, Ataraxos defeated Dutch player Pim Niemeijer with 15 wins, one loss, and four draws. Niemeijer is widely considered the most decorated player in the history of the game. He has won four world championships, 15 Dutch national titles, and two online world championships, while spending more than 600 weeks at the top of the world rankings. George Franka, the only player who has competed in every world championship since 1997, calls Niemeijer the best Stratego player of all time. Stratego poses a difficult test for artificial intelligence because both sides set up their 40 pieces face down. The paper states there are more than 10 to the 33rd power possible setups. In games with hidden information, the value of a move depends on past and future decisions.
Academic system achieved low training costs
Methods derived from poker artificial intelligence struggled with this level of hidden information because computational effort increases with hidden data. DeepMind previously attempted to master Stratego with its DeepNash system. According to the paper, DeepNash trained on 1,024 TPU nodes for two to three months, which the authors estimate would cost between $3 million and $4.5 million at 2025 prices. By contrast, Ataraxos required one week on 16 Nvidia H100 GPUs, alongside four more days on four GPUs for its belief network. At 2025 prices, that training cost less than $8,000. The researchers attribute this to a custom GPU simulator and higher sample efficiency, using roughly one five-hundredth of the compute cost of prior methods. There was no direct head-to-head comparison because DeepMind stated its DeepNash code no longer works. At the 2023 world championship, DeepNash won 19 of 28 games but lost to top players including Niemeijer.
Regularization kept the learning process on track
Ataraxos learns without human data by playing against itself. The key mechanism is an extra condition during training that forces the system to vary its setups and moves instead of locking into a fixed strategy early. The researchers gradually loosen this pressure and adjust learning step sizes over time. This regularization prevents the learning process from going in circles or turning chaotic. Ataraxos also uses a belief network that predicts opponent hidden pieces to evaluate game states before making a move.
AI won despite facing a clear structural handicap
The series against Niemeijer spanned three weeks to prevent fatigue and provide preparation time. Niemeijer received $1,000 for participating, plus $100 per win and $50 per draw. The effective win rate reached 85 percent when counting draws as half a win. Three-time world champion Vincent de Boer noted that Ataraxos played at a structural handicap because Niemeijer could adapt to the AI over many games while the AI could not adapt to him. During an exhibition at the 2025 Stratego World Championship, Ataraxos won 38 of 40 games against tournament players.
The researchers applied the same methodology to other environments. In the Barrage Stratego variant, their AI won four 50-game series against three of the four top-ranked players. In the cooperative card game Hanabi, it set new records across every variant. In the Chinese card game Dou dizhu, it outperformed previous benchmark systems. The authors argue that large amounts of hidden information are no longer an obstacle for reinforcement learning and search. The team has made the code for Ataraxos publicly available.



