‹ BackHN Continuity

Thread

With most information hidden, the game Stratego had stumped AI until now

288 points · 149 comments · PaulHoule

  1. hnedeotes · · focus · HN ↗
    I think that what makes these games beatable repeatedly is that they&#x27;re static. Not saying an algorithm properly trained won&#x27;t play better than the average player a game like MtG, or my own <a href="https:&#x2F;&#x2F;aethersummon.com" rel="nofollow">https:&#x2F;&#x2F;aethersummon.com (specially now while it has under 90 possible scrolls only) but if you have a regular release cadence (say weekly or bi-weekly) of relevant new &quot;cards&quot;, then I think the playing field is much more even for humans.

    Those new additions can invalidate the whole training data by a single new &quot;card&quot; that changes completely the dynamics and would be easy for a player to understand and incorporate but not for an algorithm (perhaps with enough compute to re-train it regularly it could) - that along with the decision trees being orders of magnitude deeper, wider and with more conditionalities than go, chess or stratego - even through the same turn with the same cards available and same table state - would probably pose much harder problems for a compute bound algo.

    1. Arainach · · focus · HN ↗
      &gt; Those new additions can invalidate the whole training data by a single new &quot;card&quot; that changes completely the dynamics

      This doesn&#x27;t follow. You&#x27;re basically proposing that new combo decks be added all the time, and it&#x27;s far simpler for an agent to scan the new cards for potential interactions with the thousands of other cards in circulation than for a human to remember all of them.

      Your analogy is akin to saying that all you have to do is keep landing new code all the time, and since the agents weren&#x27;t trained on the code they won&#x27;t be able to identify and respond to security vulnerabilities in it as fast as humans, which hasn&#x27;t turned out to be correct

      1. [deleted] · · focus · HN ↗

        [deleted]

      2. hnedeotes · · focus · HN ↗
        No, well, in MtG you could interpret it as meaning such but what I mean is that if in the training set sequence A-B-B-A when state is C-A-X-Y is the play 80% of the time, then you have a new card (that doesn&#x27;t need to be combo) that by sheer mechanics thwarts that then that strategy won&#x27;t stick by the addition of that single card to the opposing deck (that you can&#x27;t know if your opponent is playing or not) and having one or 2 or 3 or 10 different cards renders every calculation very problematic as a play can be the best or the worst depending on such simple things diluting further the best play as the pool grows. Then you need to take into account in MtG shuffling and drawing. I think it&#x27;s fair to say it&#x27;s much more difficult to model... And while an agent can learn new combos, you just need to read the card once, the agent needs to be retrained.
        1. ironSkillet · · focus · HN ↗
          Doesn&#x27;t this entirely depend on the latent embeddings of strategies and game space in the AI model, which may not be so concrete and explicit as you&#x27;ve described? That&#x27;s kind of the magic of LLMs with coding, they can generalize because the abstract patterns are encoded in latent space, not the specifics.
          1. hnedeotes · · focus · HN ↗
            I might be wrong but what I was thinking was that in chess (or even imperfect information games with a much smaller &quot;range&quot; such as Stratego,) a model can calculate all possibilities for all moves and following moves, by itself and opponent up to a depth that the human cannot. So it can see everything that can happen if it does move X-Y, then Y-Z, then A-C and figure out one that is unbeatable no matter what (or at worse leads to a draw).

            But on MtG in particular that never really applies in full due to drawing new cards. You can play perfectly and still lose due to sheer randomness of draws.

            The latent space I&#x27;m not sure how it translates to a game playing bot, but I would imagine that it would open it up to fail in the same ways a human fails.

            On the game I&#x27;m designing it could do that (calculate all possibilities up to X depth, for all possible scrolls and table states) but it would be extremely expensive to do so (not a very good argument if compute power keeps increasing), but more than that, in contrast to something like chess, there can be many more paths and decision points where a bad decision turns into a loss, so if it assumes that the best play is X at some point, a sequence that it discarded due to not being the most probable can exist and the bot can never be sure, so if it makes a decision that plays into a &quot;trap&quot; he can&#x27;t undo to a favourable position. While in Chess it&#x27;s much clearer what is possible from a given state, it&#x27;s unambiguous and the rules are fairly limited.

            In stratego you have a 10x10 board game, a very clear objective and at most 40 pieces (with repeated pieces and simple mechanics amongst them), while in MtG and similar games a single piece (card) can have probably hundreds of different interactions depending on everything else going (and everything else hidden), at many points of decision. In stratego it also seems that for humans at least, most moves are &quot;inconsequential&quot;, as it probably plays more at the psychological&#x2F;bluff level. Maybe a human player that was given the same budget for training could spend a month training against bots might fare better as the strategies might be then better understood (by the article it&#x27;s mentioned that the agent recovered from bad positions, so it seems that it was mostly human error, as the human was playing better up to that point).

            While on MtG or Asummon, although there can be inconsequential moves (they don&#x27;t matter given the context&#x2F;stage of the game), every move carries with it a possibility of being consequential in unpredictable ways. Anyway, there should be ways of training models with just a rule abiding client for these games, without codifying all rules, that they can just keep playing to figure out the interactions, so if that theory is true then it should be possible to create an unbeatable bot - I&#x27;m just not sure it is without infinite time&#x2F;compute and less so if the &quot;meta&quot; keeps changing rendering possible training inconsequential regularly.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.