The cost gap
Stratego gives each player 40 pieces, placed face down before play. A paper in Nature puts the number of possible starting layouts above 10^33. That hidden information makes each move depend not just on the current board, but on what may have happened earlier.
Earlier methods borrowed from poker AI struggled as hidden information grew. DeepMind’s DeepNash system had also tried to beat top human players, training for two to three months on 1,024 TPU nodes. Ataraxos trained for a week on 16 Nvidia H100 GPUs, then spent four more days on four GPUs training its belief network.
At 2025 prices, the researchers estimate DeepNash’s training cost at $3–4.5 million and Ataraxos’s at less than $8,000. They put Ataraxos’s compute cost at about 1/500 of DeepNash’s, with 1/30 as many self-play games and 1/100 as many training examples.
The researchers attribute the savings to a GPU simulator they built and more efficient use of training examples. The systems could not be compared directly: the researchers proposed building the necessary infrastructure, but DeepMind said DeepNash’s code no longer worked. At the 2023 world championship, DeepNash won 19 of 28 games but lost to most leading players, including Neimeijer.
Learning not to be predictable
Ataraxos learned entirely through self-play, without human data. Its training includes regularization that pushes it to vary its setups and moves rather than settling too early on one strategy. That matters in Stratego, where randomness is part of strong play: an opponent who becomes predictable can be exploited.
The researchers gradually weaken the constraint and adjust the size of the system’s learning steps. Early on, it changes strategy often and takes larger steps; later, it makes smaller adjustments. Ataraxos also uses a belief network to predict the opponent’s hidden pieces. Before each move, it considers possible game states, tests candidate moves and performs an additional learning step for that decision.
Both players set up their pieces in secret. A piece's rank is revealed only when it runs into an opposing piece. Whoever captures the enemy flag wins. | Image: Sokota, S., Vinitsky, E., Hu, H. et al.
Source: the-decoder.com
A win with a built-in handicap
The series against Neimeijer ran for three weeks so he could rest and prepare. He knew Ataraxos would not adapt to his style. He received $1,000 for taking part, plus $100 for each win and $50 for each draw.
The authors calculate Ataraxos’s effective win rate at 85%, counting a draw as half a win. They call the result unprecedented at the top level. But Neimeijer could adjust to the system over the series, while Ataraxos could not adjust to him. Three-time world champion Vincent de Boer considers that a serious disadvantage for the AI.
At the 2025 Stratego world championship, Ataraxos also won 38 of 40 exhibition games against tournament participants. Players described it as hard to read: it bluffs where humans see excessive risk and, when behind, stubbornly prolongs games.
My read is that the match demonstrates both strength and a limit. Ataraxos won despite being unable to adapt, but one series cannot show how much that handicap mattered. I’d want to know how the result changes when both sides can adjust—or when the same system faces opponents who have studied its play.
Beyond Stratego
The researchers applied the method to other games. In Barrage, a Stratego variant, their AI won four series of 50 games against three of the four highest-rated players; all three had twice been Barrage world champions. In Hanabi, it set new records across all variants and used 100 times less compute than the previous best system in the two-player version. In Dou dizhu, it beat PerfectDou and DouZero, the previous main benchmarks.
The paper argues that large amounts of hidden information need no longer be an insurmountable barrier to reinforcement learning and search. It points to financial markets, military conflicts and negotiations as possible applications, provided fast, accurate simulators can be built.
There is a limit to the method: its search simulates only one learning step, so simply adding more compute will not keep improving results. The researchers say a more complex search method could change that. Ataraxos’s code is publicly available.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X