Ataraxos beats the best Stratego player 15-1, trained on 16 GPUs and a few thousand dollars
- Ataraxos, built by researchers at Carnegie Mellon, MIT, NYU and Stanford, beat Pim Niemeijer 15 games to one with four draws across 20 online games; Niemeijer has four world championships and more than 600 weeks as the top-ranked player, and earned $100 for each win.
- Training ran 163 million self-play games on 16 GPUs for a few thousand dollars. DeepMind's DeepNash, introduced in 2022, could not reliably beat top humans despite a much larger budget.
- The main new mechanism is a second neural network, a belief model that infers the opponent's hidden pieces from how they move, so Ataraxos samples plausible board setups and searches ahead instead of iterating over more than a decillion arrangements.
- Stratego's hidden information is what stalled earlier bots: 40 pieces can sit in any order, a game can run 2,000 moves against roughly 40 in chess, and bluffing forces a balance between threats that are credible and threats that are predictable.
- At the 2025 Stratego World Championship, Ataraxos won 38 of 40 games against attendees. The team says its calm playstyle lets it climb back from a 2 percent win probability and ignore secrets a human would leak by reacting to them.
Hacker News opinions
Wait, didn't DeepMind already have a strong Stratego bot back in 2022? I thought this one was solved.
The article covers that. DeepNash needed way more training and this new one plays better with something like two orders of magnitude less data.
That 2022 paper was literally called "Mastering the Game of Stratego" but apparently mastering wasn't quite there. Four years later this one actually beats humans.
This belief model plus search approach should also work for Hanabi, where you can't see your own cards but everyone else can. I bought the game after an RL podcast episode mentioned it.
I played this as a preschooler and it was mostly psychology and bluff.
Back then it was hard partly because you couldn't remember every opponent piece you'd seen. An AI never forgets.
Oh no, Stratego was on my list of games nobody had properly cracked, I was going to build the first bot. Now you just ask Claude about the Riemann hypothesis and it tells you there's a counterexample at 1 + pi*i nobody noticed.
We just got a StarCraft 1 RL bot and an Advance Wars bot at upper human level, so the bar keeps moving.
How does it do on the silent defense variant, where the defender doesn't reveal their piece in a fight? The hidden information is even bigger there.
A friend had Electronic Stratego as a kid, where neither side reveals strength during a fight. Much more interesting game. We banned the probing feature.
I think what makes these games beatable is that they're static. If you ship new cards weekly or biweekly, a single card can invalidate the whole training set and a human adapts faster than a compute-bound agent.
New cards don't invalidate training data like that. Scanning new cards for interactions with thousands of existing ones is far easier for an agent than for a human to remember them all.
The real reason there's no Stockfish for MTG is that implementing the rules is a pain and nobody cares enough. Nothing in principle separates it from poker or bridge, and we have superhuman engines for both.
Hearthstone is absolutely swarming with bots that beat humans regularly.
Just 16 GPUs and a few thousand dollars? Those were researchers from Carnegie Mellon, MIT, NYU and Stanford, this wasn't just anyone.
The contrast is against DeepMind's budget, not the talent. The preceding sentence is about DeepMind having an exceptional budget and still failing.
There's a game archive up at ataraxosai.github.io if you want to look at the matches.
I crushed everyone at Stratego as a kid, including my dad who was way better than me at chess, so I never guessed models would struggle with it.
We played live Stratego at camp with index cards in our socks. My strategy was probing attacks to find where the ranks were, then channels to move up a piece that outranked the target by one.