I want to see AI beat a 4x strategy video game or a roguelike. Let me see a 20+ win streak on the hardest difficulty in Balatro or Slay the Spire. Let me see AI beat Civilization against the best human players.
Every time I mention this someone assures me that it's possible, and they point to simple board games that computers excel at, or real time strategy games where proper use of APM and clicking accurately go a long way. But I haven't seen AI succeed at any decision-focused game where describing the rules requires more than one minute.
I want to see what AI can do in, not simple, and not complex, but complicated toy environments, where decisions are all that matter.
I'm a bit skeptical of their "2 seeds in a row!" boast. Last time I investigated a claim like that I found the seeds were cherry picked. This was way back in the OpenAI Gym days though (remember when OpenAI was open and just doing goofy research like OpenAI Gym?), their leader boards had some amazing claims about certain RL solutions, but when I ran them myself on new seeds they were far worse than claimed.
Buttons840 · · focus · HN ↗
Every time I mention this someone assures me that it's possible, and they point to simple board games that computers excel at, or real time strategy games where proper use of APM and clicking accurately go a long way. But I haven't seen AI succeed at any decision-focused game where describing the rules requires more than one minute.
I want to see what AI can do in, not simple, and not complex, but complicated toy environments, where decisions are all that matter.
IanCal · · focus · HN ↗
<a href="https://github.com/Attol8/balatro-ai" rel="nofollow">https://github.com/Attol8/balatro-ai
Buttons840 · · focus · HN ↗
I'm a bit skeptical of their "2 seeds in a row!" boast. Last time I investigated a claim like that I found the seeds were cherry picked. This was way back in the OpenAI Gym days though (remember when OpenAI was open and just doing goofy research like OpenAI Gym?), their leader boards had some amazing claims about certain RL solutions, but when I ran them myself on new seeds they were far worse than claimed.