‹ BackHN Continuity

Thread

With most information hidden, the game Stratego had stumped AI until now

288 points · 149 comments · PaulHoule

  1. cubefox · · focus · HN ↗
    There is no information in this article on how modern LLMs like GPT-6 Astra or Claude Opus 5.5 perform at this game. I have a hard time believing they are bad at it.
    1. abstractcontrol · · focus · HN ↗
      The intelligence of LLMs is a mirage masked by their utility. It makes more sense to assume that ML models not fitted to a particular game would be bad at it than otherwise. If you're relating how good a human with inate capabilities of the modern LLMs would be, then yes, it might make sense to assume that they'd be good at anything, but that isn't the situation in reality.
      1. cubefox · · focus · HN ↗
        > It makes more sense to assume that ML models not fitted to a particular game would be bad at it than otherwise.

        No, it doesn't make sense. GPT-6 Astra solves ARC-AGI-3, which consists of a large number of small games which the LLM doesn't know and which it has to solve on the fly with a limited number of turns.

        1. abstractcontrol · · focus · HN ↗
          I wouldn't take those benchmarks too seriously; the companies are benchmaxing as there is a lot of money at stake.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.