‹ BackHN Continuity

Thread

With most information hidden, the game Stratego had stumped AI until now

288 points · 149 comments · PaulHoule

  1. peter_d_sherman · · focus · HN ↗
    >"DeepNash was trained for two to three months on 1,024 of Google’s specialized chips, a run the Ataraxos team estimates would cost $3 million to $4.5 million at 2025 prices. Ataraxos, in contrast, needed 16 GPUs for a week, plus an additional four GPUs for four days to train the belief model. Farina and lead author Samuel Sokota achieved this efficiency by

    writing a simulator that runs millions of moves per second on graphics cards.

    “At the scale that we are in academia, we don’t really have access to an entire field of GPUs,” Farina said.

    The algorithm also learned far faster—it played about 34 times fewer games than DeepNash, and still ended up much stronger."

    Isn't that a case of David vs. Goliath!

    Also, I like the idea of the simulator/training software/aka "oracle" (source of truth for training data) running adjacently on the same graphics card(s) or AI accelerator(s) that the training is taking place on... That approach moves way more data way faster than say, having the training have to interact with a game running on a CPU, and having to screen scrape and interpret screen data (and wait!) with every new move...

    Anyway, great article!

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.