With most information hidden, the game Stratego had stumped AI until now
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
With most information hidden, the game Stratego had stumped AI until now
Unofficial Hacker News client; not affiliated with Y Combinator.
hnedeotes · · focus · HN ↗
Those new additions can invalidate the whole training data by a single new "card" that changes completely the dynamics and would be easy for a player to understand and incorporate but not for an algorithm (perhaps with enough compute to re-train it regularly it could) - that along with the decision trees being orders of magnitude deeper, wider and with more conditionalities than go, chess or stratego - even through the same turn with the same cards available and same table state - would probably pose much harder problems for a compute bound algo.
Arainach · · focus · HN ↗
This doesn't follow. You're basically proposing that new combo decks be added all the time, and it's far simpler for an agent to scan the new cards for potential interactions with the thousands of other cards in circulation than for a human to remember all of them.
Your analogy is akin to saying that all you have to do is keep landing new code all the time, and since the agents weren't trained on the code they won't be able to identify and respond to security vulnerabilities in it as fast as humans, which hasn't turned out to be correct
hnedeotes · · focus · HN ↗
ironSkillet · · focus · HN ↗
hnedeotes · · focus · HN ↗
But on MtG in particular that never really applies in full due to drawing new cards. You can play perfectly and still lose due to sheer randomness of draws.
The latent space I'm not sure how it translates to a game playing bot, but I would imagine that it would open it up to fail in the same ways a human fails.
On the game I'm designing it could do that (calculate all possibilities up to X depth, for all possible scrolls and table states) but it would be extremely expensive to do so (not a very good argument if compute power keeps increasing), but more than that, in contrast to something like chess, there can be many more paths and decision points where a bad decision turns into a loss, so if it assumes that the best play is X at some point, a sequence that it discarded due to not being the most probable can exist and the bot can never be sure, so if it makes a decision that plays into a "trap" he can't undo to a favourable position. While in Chess it's much clearer what is possible from a given state, it's unambiguous and the rules are fairly limited.
In stratego you have a 10x10 board game, a very clear objective and at most 40 pieces (with repeated pieces and simple mechanics amongst them), while in MtG and similar games a single piece (card) can have probably hundreds of different interactions depending on everything else going (and everything else hidden), at many points of decision. In stratego it also seems that for humans at least, most moves are "inconsequential", as it probably plays more at the psychological/bluff level. Maybe a human player that was given the same budget for training could spend a month training against bots might fare better as the strategies might be then better understood (by the article it's mentioned that the agent recovered from bad positions, so it seems that it was mostly human error, as the human was playing better up to that point).
While on MtG or Asummon, although there can be inconsequential moves (they don't matter given the context/stage of the game), every move carries with it a possibility of being consequential in unpredictable ways. Anyway, there should be ways of training models with just a rule abiding client for these games, without codifying all rules, that they can just keep playing to figure out the interactions, so if that theory is true then it should be possible to create an unbeatable bot - I'm just not sure it is without infinite time/compute and less so if the "meta" keeps changing rendering possible training inconsequential regularly.