‹ BackHN Continuity

Thread

Vote on which of Hacker News' challenges for AI have been met

202 points · 271 comments · stabbles

  1. hatthew · · focus · HN ↗
    My comment about humanity's last exam being a misnomer is included, and I proposed better ideas about what a last exam could look like. One of the things I said was "solve an open math problem" which has conclusively been done with Navier-Stokes (regardless of the controversy surrounding that). However, in the spirit of clarifying the goalposts, AI has only passed 1/6 of the tests I proposed. 17% is not a passing grade, so I'd say no, my challenge has not been met.

    Another thing to note is that the (presumably AI-generated) summary of my challenge does not accurately represent what I wrote, listing only half the things I said and saying "or" rather than "and".

    1. an0malous · · focus · HN ↗
      How can you say it’s conclusively been done if it might have been stolen from a math researcher and was aided in unknown ways by a whole team of math researchers? I find it mind boggling that HN just accepts these shenanigans with no transparency. At the very least, they could share the conversation / thinking trace easily and if their claims are true there shouldn’t be anything controversial or negative for their company in the trace.
      1. senordevnyc · · focus · HN ↗
        Yeah, they could easily just share the untold trillions of tokens the 10k agents generated over 88 hours, which would also be a goldmine for their competitors, no big deal.

        I find it mind boggling that anyone thinks these agents only solved this problem because they maybe could possibly have seen the unfinished work of researchers who were working on a simpler version of the problem, (also with AI).

        Looking forward to the cope when the next big problem falls.

        1. dormento · · focus · HN ↗
          > Yeah, they could easily just share the untold trillions of tokens the 10k agents generated over 88 hours, which would also be a goldmine for their competitors, no big deal.

          I hope I'm not being too blunt, but the other alternative is to "just trust me bro" the hyperscalers, who are pretty much locked into a battle for profitability and have all the incentives to make up things to prop up their stock, no? I don't think this is the way.

          1. pessimizer · · focus · HN ↗
            AI can't be a religion if you don't attack with fire the people for whom "just trust me bro" isn't good enough. "Just Trust Them, Bros!" is the central tenet.
          2. senordevnyc · · focus · HN ↗
            Then don't trust them?
            1. someonebaggy · · focus · HN ↗
              That's what they're asking for - by getting the tokens
        2. eviks · · focus · HN ↗
          Yeah, they could indeed easily just share the tokens. And you can ask your favorite AI to explain that it doesn't have to be a public release to all the competitors if you can't come up with a better alternative yourself. There is a big range between no-one and every-one
        3. _superposition_ · · focus · HN ↗
          Yeah they don't want to share the tokens because I guarantee it's a bunch of nonsense trial and error leaning on lean for verification. Those tokens will show the lack of intelligence, not the presence of it.
          1. someonebaggy · · focus · HN ↗
            A machine that solves open problems by brute force trial and error is still pretty cool.
      2. off_with_their_ · · focus · HN ↗

        [dead]

      3. olmo23 · · focus · HN ↗
        navier stokes is not the only example of an open problem in maths that was solved by AI (eg the counterexample to the jacobian)
        1. tsunamifury · · focus · HN ↗
          In the other hand why does anyone find it surprising at all that a computer solved a math problem.
          1. stronglikedan · · focus · HN ↗
            Because up until then, computers had not solved math problems. Humans solved them, often using computers as a tool.
            1. tsunamifury · · focus · HN ↗
              I'm pretty sure computers have been solving math problems for a very very long time through various techniques.
          2. ben_w · · focus · HN ↗
            To ask that question suggests you are unfamiliar with the difference between arithmetic (which computers are good at) and pure mathematics (which is like metaprogramming combined with formal methods, and Gödel's incompleteness theorem is worse than the Halting problem)?

            Simply put: For the same reason computers have not already solved all problems in mathematics.

            More concretely:

            Consider the Collatz conjecture. It's a very simple rule to write down:

              Take some positive integer: If the number is even, divide it by two; otherwise triple it and add one. With enough repetition, do all positive integers converge to 1?
            
            Trivial to write a program to test numbers starting at 1 and going up. We know it holds up to at least 2.36e21 (according to Wikipedia), but to prove it is true with such a program requires testing all of the infinite set of positive integers.

            But you may notice some things about the rules, that they suggest a subset of numbers will trivially always converge to 1, so that you don't need to even test them: any integer 2^n where n is also a positive integer.

            You may find other easy wins, or ways to simplify the test, e.g. once you know the numbers up to m will converge to 1, you can terminate your loop early if you test m+1 and it ever has an intermediate value less than or equal to m. You can combine that with applying one of the rules in reverse, and know that all even numbers between m and 2m will on their first move be halved, making them smaller than m, which means you know they'll eventually converge.

            But actually proving this is fully general? Nobody knows. You can't just throw arithmetic at the problem directly, you have to figure out patterns that would let you prove that it always holds, no matter what.

            Or, you may find many such patterns and directly calculate some number not in any of them, to find one which doesn't converge to 1.

      4. Jtariiiii · · focus · HN ↗
        >How can you say it’s conclusively been done if it might have been stolen from a math researcher and was aided in unknown ways by a whole team of math researchers?

        The solution to NS was categorically not stolen, nobody is alleging that OpenAI stole a complete solution to NS. The alleged theft was about a different set of related equations.

        1. someonebaggy · · focus · HN ↗
          If I want to solve an open problem, stealing a solution to half of it surely makes it much easier. Can I claim i did it myself? Well... Maybe.
      5. 3uruiueijjj · · focus · HN ↗
        Man I can't believe mathematicians were just about to solve dozens of significant problems all in the same year, and that's exactly the year LLMs come around to steal their solutions.

        Talk about bad luck!

        1. irthomasthomas · · focus · HN ↗
          No one ever spent tens of millions of dollars trying before. To evaluate the achievement we must see all the logs, know how much they spent, and how much help they had from humans.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.