With most information hidden, the game Stratego had stumped AI until now
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
With most information hidden, the game Stratego had stumped AI until now
Unofficial Hacker News client; not affiliated with Y Combinator.
hnedeotes · · focus · HN ↗
Those new additions can invalidate the whole training data by a single new "card" that changes completely the dynamics and would be easy for a player to understand and incorporate but not for an algorithm (perhaps with enough compute to re-train it regularly it could) - that along with the decision trees being orders of magnitude deeper, wider and with more conditionalities than go, chess or stratego - even through the same turn with the same cards available and same table state - would probably pose much harder problems for a compute bound algo.
qsort · · focus · HN ↗
There is nothing that, in principle, makes MTG different from poker or bridge, and we have superhuman engines for both.
askjdfksdbfhk · · focus · HN ↗
We don't have superhuman play for bridge.
Poker and bridge are quite different from each other in terms of solving them. Among other things, the hidden information space in poker (at least, in hold'em) is far smaller than in bridge (or Stratego, for that matter, as discussed in the linked paper). This makes hold'em solvable using CFR, an algorithm which essentially optimizes play by considering all the possible holdings than the opponent might have and their best strategy with each one. Even going from two to four hidden cards per player (Omaha) requires a slightly different approach although you can still use CFR as the basis for the search algorithm.
Bridge has 13 hidden cards per player which makes CFR basically impossible to apply, at least in any obvious way--just way too many states. Similarly you see it's not used at all in this Stratego paper.
qsort · · focus · HN ↗
The point is that, especially for games perceived as being lower-status like MTG and other board games, I'm more inclined to believe the answer is closer to "nobody is willing to pour in the resources to seriously try" as opposed to "we definitively cannot with current science and technology."
vmilner · · focus · HN ↗
askjdfksdbfhk · · focus · HN ↗
I don't know what this is referring to. The ubiquitous dds library, which is AFAIK basically the only double dummy solver, has not seen any real improvements. Neither of those numbers look right to me, I think it's more in the range of ~10ms.
There's a crank who was claiming some magical improvements a little while ago that really just boiled down to AI psychosis (and having absolutely no understanding of what he was claiming). I hope that's not what you're referring to.
vmilner · · focus · HN ↗
I believe this was used in his 'ben' bridge engine, <a href="https://github.com/lorserker/ben/" rel="nofollow">https://github.com/lorserker/ben/ now maintained by ThorvaldAagaard, though I have to admit it now seems to be heavily dds focussed, so there may have been a rollback along the lines you outlined.
I'm attempting to recreate the concept myself, so should soon have an idea whether its moonshine or not,
askjdfksdbfhk · · focus · HN ↗
I wasn't aware of this although this project had a substantial error rate. Flipping through the YouTube video it only predicted the correct number of tricks 70% of the time--so I'm not sure if your 99.9% is referring to a different project that I'm unable to find, or if you misremembered. I don't think anything like this was ever used in Ben; looking at the commit history, I think Ben has always used the standard dds library.
FWIW I'm pretty confident that you could beat the performance of the project I linked with a fairly straightforward transformer architecture.
vmilner · · focus · HN ↗
(There was some Ben code to use a neural net double dummy evaluator, because Ive used it, but it was removed. Ill try and find the git point it was removed.)
vmilner · · focus · HN ↗
also the keras model ben/models/TF2models/RPDD_2024-07-08-E02.keras is still present in the current repo but not apparently used. It was trained on ten million deals where all double dummy info is known.
askjdfksdbfhk · · focus · HN ↗
I agree that many games could make progress if people were actually inclined to try.
I think it will get a bit better in the coming decade thanks to continued hardware improvements & powerful LLM coding agents making it more feasible for amateurs to tackle these things at home. Personally I've been working on a game AI project for the last month at home based around published techniques for a similar game, using my 5090 for training and Opus for implementation and orchestrating tasks and so on. It's going quite well and it looks like I'm on track for a SOTA, superhuman AI at the end. Doing this ten years ago would have been incomparably harder.