‹ BackHN Continuity

Thread

Vote on which of Hacker News' challenges for AI have been met

202 points · 271 comments · stabbles

  1. bmenrigh · · focus · HN ↗
    At least 1/3rd of these predictions aren't clear enough to determine exactly what is being claimed/predicted. Even after reading the full comment multiple times, on a lot of them I couldn't tell where the author had set the goalposts well enough to say whether we've crossed it or not.
    1. jerf · · focus · HN ↗
      Well, I can answer this one: <a href="https:&#x2F;&#x2F;stoppels.ch&#x2F;goalposts&#x2F;?c=40662140" rel="nofollow">https:&#x2F;&#x2F;stoppels.ch&#x2F;goalposts&#x2F;?c=40662140

      jerf, 2024: &quot;If it could be solved with a Math Overflow-post level of effort, even from Terence Tao, it isn&#x27;t what I was talking about as &quot;high level math&quot;.

      &quot;I also am not surprised by &quot;Consider a generation function&quot; coming out of an LLM. I am talking about a system that could solve that problem, entirely, as doing high level math. A system that can emit &quot;have you considered using wood?&quot; is not a system that can build a house autonomously.

      &quot;It especially won&#x27;t seem all that useful next to the generation of AIs I anticipate to be coming which use LLMs as a component to understand the world but are not just big LLMs.&quot;

      The voting gloss: &quot;An AI fully solves a research-level math problem on its own, not just suggesting an approach.&quot;

      Yes, I&#x27;m satisfied. I don&#x27;t even feel bad in hindsight. Coding assistants had a nice, gradual rise up the utility curve. Math went from &quot;lol, can&#x27;t add two six-digit numbers&quot; to research-math level almost overnight in comparison.

      1. 48844858 · · focus · HN ↗
        But it still makes mistakes when adding numbers
        1. gregsadetsky · · focus · HN ↗
          I was working with LLMs last year and asking them to do &quot;logical circuits in their heads&quot; ie

          &quot;there is a nand gate A connected to gate B through these wires, connected to another gate C, etc. what is the output value of gate C if i place a 1 at this gate&quot;

          Sort of like doing math &quot;in their heads&quot; (i.e. your &quot;when adding numbers&quot;), they would get it very right for simple&#x2F;small cases (although the answer could have been in their training), then some LLMs would get it for harder cases (which were clearly not in their training), and all would fail at some point. This was all without any tool calling.

          After a year of thinking about it, I made an eval [0] with ever-complexifying nand circuits - like, truly, bananas circuits [1] - and some models, do, in effect (through chain of thought? mostly?), get the right answer. ((what&#x27;s nice is that you can always make a circuit at the very edge of what all models can correctly solve))

          Tool-calling 100000% solves this problem for sure (evaluating a nand gate is trivial). But if you even prompt an llm to do math like a 5&#x2F;6th grader (i.e. do it digit by digit, carry the 1, etc.) - I am quite certain most llms can, in fact, add numbers.

          But yeah. These piles of weights are fascinating in how they seem flawed one day (&quot;how many r&#x27;s&quot;) and magical at once.

          [0] <a href="https:&#x2F;&#x2F;lockstep.greg.technology" rel="nofollow">https:&#x2F;&#x2F;lockstep.greg.technology

          [1] <a href="https:&#x2F;&#x2F;lockstep.greg.technology&#x2F;c&#x2F;?id=rand_s4161_g160_d8" rel="nofollow">https:&#x2F;&#x2F;lockstep.greg.technology&#x2F;c&#x2F;?id=rand_s4161_g160_d8

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.