‹ BackHN Continuity

Thread

Vote on which of Hacker News' challenges for AI have been met

202 points · 271 comments · stabbles

  1. bmenrigh · · focus · HN ↗
    At least 1/3rd of these predictions aren't clear enough to determine exactly what is being claimed/predicted. Even after reading the full comment multiple times, on a lot of them I couldn't tell where the author had set the goalposts well enough to say whether we've crossed it or not.
    1. jerf · · focus · HN ↗
      Well, I can answer this one: <a href="https:&#x2F;&#x2F;stoppels.ch&#x2F;goalposts&#x2F;?c=40662140" rel="nofollow">https:&#x2F;&#x2F;stoppels.ch&#x2F;goalposts&#x2F;?c=40662140

      jerf, 2024: &quot;If it could be solved with a Math Overflow-post level of effort, even from Terence Tao, it isn&#x27;t what I was talking about as &quot;high level math&quot;.

      &quot;I also am not surprised by &quot;Consider a generation function&quot; coming out of an LLM. I am talking about a system that could solve that problem, entirely, as doing high level math. A system that can emit &quot;have you considered using wood?&quot; is not a system that can build a house autonomously.

      &quot;It especially won&#x27;t seem all that useful next to the generation of AIs I anticipate to be coming which use LLMs as a component to understand the world but are not just big LLMs.&quot;

      The voting gloss: &quot;An AI fully solves a research-level math problem on its own, not just suggesting an approach.&quot;

      Yes, I&#x27;m satisfied. I don&#x27;t even feel bad in hindsight. Coding assistants had a nice, gradual rise up the utility curve. Math went from &quot;lol, can&#x27;t add two six-digit numbers&quot; to research-math level almost overnight in comparison.

      1. 48844858 · · focus · HN ↗
        But it still makes mistakes when adding numbers
        1. shiandow · · focus · HN ↗
          Honestly that makes me more convinced it&#x27;s actually doing mathematics.
        2. CamperBob2 · · focus · HN ↗
          No, not really. Not unless you go out of your way to use an obsolete or extremely low-end model.
          1. yorwba · · focus · HN ↗
            If there is a model that never makes mistakes on simple arithmetic, the developers should really claim their 1.0000 crown on the GSM8k benchmark <a href="https:&#x2F;&#x2F;llm-stats.com&#x2F;benchmarks&#x2F;gsm8k">https:&#x2F;&#x2F;llm-stats.com&#x2F;benchmarks&#x2F;gsm8k (GSM is Grade School Math).
            1. CamperBob2 · · focus · HN ↗
              The GSM8k problems are not &quot;adding numbers.&quot; They are word problems, e.g. &quot;Katy makes coffee using teaspoons of sugar and cups of water in the ratio of 7:13. If she used a total of 120 teaspoons of sugar and cups of water, calculate the number of teaspoonfuls of sugar she used.&quot;

              They are the kind of problems that, if your teacher was anything like mine, were usually skipped in order to keep the slower students from bogging down the class as a whole. 0.996 (MiMo-V2.5) is substantially better than what the vast majority of humans would do.

              If you limited the question to adding arbitrary pairs of numbers of reasonable size, I imagine quite a few models could get to 1.000.

        3. mswphd · · focus · HN ↗
          see the Grothendiek prime

          <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;57_(number)" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;57_(number)

        4. sanex · · focus · HN ↗
          So do I, that it why both I and claude use calculators. :)
        5. rcxdude · · focus · HN ↗
          This correlates quite strongly with advanced mathematics capability in humans as well :P. In the list of people I would trust to add two two-digit numbers correctly, a maths PhD puts you in the bottom half of the list.
        6. gregsadetsky · · focus · HN ↗
          I was working with LLMs last year and asking them to do &quot;logical circuits in their heads&quot; ie

          &quot;there is a nand gate A connected to gate B through these wires, connected to another gate C, etc. what is the output value of gate C if i place a 1 at this gate&quot;

          Sort of like doing math &quot;in their heads&quot; (i.e. your &quot;when adding numbers&quot;), they would get it very right for simple&#x2F;small cases (although the answer could have been in their training), then some LLMs would get it for harder cases (which were clearly not in their training), and all would fail at some point. This was all without any tool calling.

          After a year of thinking about it, I made an eval [0] with ever-complexifying nand circuits - like, truly, bananas circuits [1] - and some models, do, in effect (through chain of thought? mostly?), get the right answer. ((what&#x27;s nice is that you can always make a circuit at the very edge of what all models can correctly solve))

          Tool-calling 100000% solves this problem for sure (evaluating a nand gate is trivial). But if you even prompt an llm to do math like a 5&#x2F;6th grader (i.e. do it digit by digit, carry the 1, etc.) - I am quite certain most llms can, in fact, add numbers.

          But yeah. These piles of weights are fascinating in how they seem flawed one day (&quot;how many r&#x27;s&quot;) and magical at once.

          [0] <a href="https:&#x2F;&#x2F;lockstep.greg.technology" rel="nofollow">https:&#x2F;&#x2F;lockstep.greg.technology

          [1] <a href="https:&#x2F;&#x2F;lockstep.greg.technology&#x2F;c&#x2F;?id=rand_s4161_g160_d8" rel="nofollow">https:&#x2F;&#x2F;lockstep.greg.technology&#x2F;c&#x2F;?id=rand_s4161_g160_d8

        7. jerf · · focus · HN ↗
          I&#x27;m not all that worried about it. If we humans try to add numbers the way we expect an LLM to do it, we&#x27;re generally terrible too. I don&#x27;t just mean the well-known propensity for mathematicians to actually be pretty bad at arithmetic, I mean, if you just recite two six-digit numbers to a human and ask them to add it together, on the spot, with no paper, no external tools, we&#x27;re pretty bad at it too. It can be done, certainly, but it takes deliberate practice and training. It&#x27;s not something we get for free just because we&#x27;re smart.

          Decades before the modern AI push I was marveling at the distinction between the sheer overwhelming computational power of the human brain, if measured from the perspective of how much math it is doing under the hood, and its utter ineptness at basic arithmetic compared to the tools we can build. The earliest, klunkiest, most garbage mechanical adding machines we ever produced, long before we improved them by literally over a dozen orders of magnitude, were still already way better than we are at basic arithmetic.

          There is something profound I still have not fully put my finger on in how basic arithmetic is so easy for a machine, yet the decisions we routinely make with our neural nets has been the laborious effort of decades with us still not arriving yet even with the trillions now poured into AI for machines. And vice versa. Even that practice I alluded to that allows you to train yourself to do this task on demand easily would incorporate mathematical advancements in representations that took our species thousands of years to come up with, rather than being something you get &quot;for free&quot; just for being smart. Our brains casually run an entire human body through an unbelievably complicated external universe, yet struggle with basic arithmetic.

          There&#x27;s so many places where we have one architecture that&#x27;s a bit better than another at one thing, and a bit worse at another, but in the end they can both do the job. Like, if we had to do all our programming in immutable languages and run all our imperative code through an O(n log n) worst-case penalty for immutable languages emulating imperative RAM, we&#x27;d survive just fine over all. But between neural architectures and conventional arithmetic on dedicated silicon is this dozen+ order of magnitude difference on tasks. It&#x27;s a pretty wild disparity. Some of the reason is somewhat obvious, I don&#x27;t want to make it sound like I&#x27;m completely mystified... I just think there&#x27;s probably, somewhere, an even more profound way to see it than the obvious differences that says something more powerful about the limits of computation than I&#x27;ve seen anyone say. Which is not to say somewhere out there someone already has had the idea I&#x27;m grasping for and written it in some brilliant paper or something. I&#x27;m just saying I haven&#x27;t seen it.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.