‹ BackHN Continuity

Thread

Brood War Bench

346 points · 156 comments · benswerd

  1. dschuessler · · focus · HN ↗
    Somewhat related: In 2018, Google DeepMind had already created AIs that were capable of beating professional gamers in StarCraft 2 (the sequel to Brood War): <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=cUTMhmVh1qs" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=cUTMhmVh1qs
    1. benswerd · · focus · HN ↗
      I predict LLMs will reach superhuman level and beat even that model in the next 12 months
      1. orbital-decay · · focus · HN ↗
        Starcraft is APM-dependent. Unless the latency will improve greatly in frontier reasoning LLMs (which is unlikely), it will remain a bit like knitting with an excavator.
        1. benswerd · · focus · HN ↗
          I predict latency will improve greatly in the next 12 months to more than 4x speed on current frontier tasks
          1. loeg · · focus · HN ↗
            Yeah but Starcraft needs, like, 10-20x the APM these agents are doing.
            1. benswerd · · focus · HN ↗
              I’m not convinced a lot of it can’t be solved with code mode.

              Marine staggering for example seems like an ideal code mode task.

              1. loeg · · focus · HN ↗
                Yeah. Some of it may just be &quot;thinking&quot; less rather than faster token generation.
              2. jayd16 · · focus · HN ↗
                Do the humans get to use this auto-stagger too?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.