‹ BackHN Continuity

Thread

Vote on which of Hacker News' challenges for AI have been met

202 points · 271 comments · stabbles

  1. bmenrigh · · focus · HN ↗
    At least 1/3rd of these predictions aren't clear enough to determine exactly what is being claimed/predicted. Even after reading the full comment multiple times, on a lot of them I couldn't tell where the author had set the goalposts well enough to say whether we've crossed it or not.
    1. jerf · · focus · HN ↗
      Well, I can answer this one: <a href="https:&#x2F;&#x2F;stoppels.ch&#x2F;goalposts&#x2F;?c=40662140" rel="nofollow">https:&#x2F;&#x2F;stoppels.ch&#x2F;goalposts&#x2F;?c=40662140

      jerf, 2024: &quot;If it could be solved with a Math Overflow-post level of effort, even from Terence Tao, it isn&#x27;t what I was talking about as &quot;high level math&quot;.

      &quot;I also am not surprised by &quot;Consider a generation function&quot; coming out of an LLM. I am talking about a system that could solve that problem, entirely, as doing high level math. A system that can emit &quot;have you considered using wood?&quot; is not a system that can build a house autonomously.

      &quot;It especially won&#x27;t seem all that useful next to the generation of AIs I anticipate to be coming which use LLMs as a component to understand the world but are not just big LLMs.&quot;

      The voting gloss: &quot;An AI fully solves a research-level math problem on its own, not just suggesting an approach.&quot;

      Yes, I&#x27;m satisfied. I don&#x27;t even feel bad in hindsight. Coding assistants had a nice, gradual rise up the utility curve. Math went from &quot;lol, can&#x27;t add two six-digit numbers&quot; to research-math level almost overnight in comparison.

      1. 48844858 · · focus · HN ↗
        But it still makes mistakes when adding numbers
        1. CamperBob2 · · focus · HN ↗
          No, not really. Not unless you go out of your way to use an obsolete or extremely low-end model.
          1. yorwba · · focus · HN ↗
            If there is a model that never makes mistakes on simple arithmetic, the developers should really claim their 1.0000 crown on the GSM8k benchmark <a href="https:&#x2F;&#x2F;llm-stats.com&#x2F;benchmarks&#x2F;gsm8k">https:&#x2F;&#x2F;llm-stats.com&#x2F;benchmarks&#x2F;gsm8k (GSM is Grade School Math).
            1. CamperBob2 · · focus · HN ↗
              The GSM8k problems are not &quot;adding numbers.&quot; They are word problems, e.g. &quot;Katy makes coffee using teaspoons of sugar and cups of water in the ratio of 7:13. If she used a total of 120 teaspoons of sugar and cups of water, calculate the number of teaspoonfuls of sugar she used.&quot;

              They are the kind of problems that, if your teacher was anything like mine, were usually skipped in order to keep the slower students from bogging down the class as a whole. 0.996 (MiMo-V2.5) is substantially better than what the vast majority of humans would do.

              If you limited the question to adding arbitrary pairs of numbers of reasonable size, I imagine quite a few models could get to 1.000.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.