‹ BackHN Continuity

Thread

Yes, Claude can do nine loops

103 points · 62 comments · tzury

  1. Buttons840 · · focus · HN ↗
    I want to see AI beat a 4x strategy video game or a roguelike. Let me see a 20+ win streak on the hardest difficulty in Balatro or Slay the Spire. Let me see AI beat Civilization against the best human players.

    Every time I mention this someone assures me that it's possible, and they point to simple board games that computers excel at, or real time strategy games where proper use of APM and clicking accurately go a long way. But I haven't seen AI succeed at any decision-focused game where describing the rules requires more than one minute.

    I want to see what AI can do in, not simple, and not complex, but complicated toy environments, where decisions are all that matter.

    1. IanCal · · focus · HN ↗
      I don’t know how much you find the use of numerical things or a tailored system an issue vs “here’s balatro let’s go” but this might be of interest

      <a href="https:&#x2F;&#x2F;github.com&#x2F;Attol8&#x2F;balatro-ai" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;Attol8&#x2F;balatro-ai

      1. Buttons840 · · focus · HN ↗
        This counts. Thanks!

        I&#x27;m a bit skeptical of their &quot;2 seeds in a row!&quot; boast. Last time I investigated a claim like that I found the seeds were cherry picked. This was way back in the OpenAI Gym days though (remember when OpenAI was open and just doing goofy research like OpenAI Gym?), their leader boards had some amazing claims about certain RL solutions, but when I ran them myself on new seeds they were far worse than claimed.

    2. Aurornis · · focus · HN ↗
      You can get an OpenAI subscription and have it play games right now, today.

      A lot of influencers are trying it. You can set it loose on games like Slay the Spire without any training and it can win: <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=9bDG0uuHM2w" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=9bDG0uuHM2w

      That&#x27;s a general purpose LLM. Actually training a model for a game like that would be old news.

      I do have to laugh at how the goalposts keep moving to higher and higher levels like &quot;I need to see it win 20X in a row at the highest difficulty! Why has nobody shown this?&quot;

      1. Buttons840 · · focus · HN ↗
        Humans can do that, so that was my goal post. I don&#x27;t think I&#x27;ve changed it since this is the first time I&#x27;ve proposed a goal post.

        Thanks for the video link. I tried to find one like this recently, but the video you linked wasn&#x27;t in the results I looked at.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.