‹ BackHN Continuity

Thread

Brood War Bench

346 points · 156 comments · benswerd

  1. GodelNumbering · · focus · HN ↗
    A friend of mine created GoBench[1][2] that evaluates LLMs on 9×9 Go using KataGo opponents as Elo anchors, you see real capability differences there, like Astra Max substantially leading all other models. I think strategy is a generally interesting area to evaluate LLMs on

    [1] <a href="https:&#x2F;&#x2F;rolandgao.com&#x2F;blog&#x2F;gobench&#x2F;" rel="nofollow">https:&#x2F;&#x2F;rolandgao.com&#x2F;blog&#x2F;gobench&#x2F;

    [2] <a href="https:&#x2F;&#x2F;rolandgao.com&#x2F;gobench.pdf" rel="nofollow">https:&#x2F;&#x2F;rolandgao.com&#x2F;gobench.pdf

    1. nullc · · focus · HN ↗
      A better benchmark might be asking the LLMs to write GO ai and then comparing that-- the issue is that there will be a HUGE difference in performance that depends purely on this game being in the LLM&#x27;s training... but training a general LLM to directly play these games would be a waste of capacity and shouldn&#x27;t be encouraged for benchmaxxing sake.

      Programming an engine OTOH is a skill that is more general and they should all have.

      Might be useful to have the target of the engine be some specific virtual machine that gets a strict cycle budget-- e.g. execution runs so many cycles, and result is read out of a specific memory address at the end (or when it terminates early).

      1. theendisney · · focus · HN ↗
        This is a great idea. It could learn from its mistakes, repeat things that work, abstract complex situations. I would even want humans monitoring the project.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.