‹ BackHN Continuity

Thread

Vote on which of Hacker News' challenges for AI have been met

202 points · 271 comments · stabbles

  1. joegibbs · · focus · HN ↗
    There's one of mine in there where I predicted in 2023 that it would be 20 years until AI would be reliably able to entirely build and deploy arbitrary applications from a prompt. I was off by about 18 years on that one!
    1. aleph_minus_one · · focus · HN ↗
      > There's one of mine in there where I predicted in 2023 that it would be 20 years until AI would be reliably able to entirely build and deploy arbitrary applications from a prompt. I was off by about 18 years on that one!

      If you take "arbitrary" seriously, we are still very far away from it.

      1. eleventen · · focus · HN ↗
        What is the strict definition of arbitrary you feel we haven’t reached? Right now the limit is scale and complexity, not domain or “type” of application.
        1. dwattttt · · focus · HN ↗
          > reliably able to entirely build and deploy arbitrary applications from a prompt

          You're maybe thinking that we can build and deploy arbitrary "kinds" of application. Being able to "build and deploy arbitrary applications" would mean I could ask for any scale or complexity in my application.

          1. dmd · · focus · HN ↗
            Find me even one human in the world who meets your criteria, then.

            edit: I guess I misread this. I thought you were saying "well, AI isn't smart until it can solve any arbitrary problem in the whole world"

            1. dwattttt · · focus · HN ↗
              ... I don't think there is? I don't see what bearing that has on the original question either.

              EDIT: to forestall further back and forth, I don't think there'd be any controversy if the problem statement said "common applications"

              1. aleph_minus_one · · focus · HN ↗
                > I don't think there'd be any controversy if the problem statement said "common applications"

                I would claim "common applications" is also controversial because what is a "common application" depends insanely on the area in which you work. Even if you exclude some highly advanced scientific applications (because you don't consider these to be common), in many industrial sectors there exist applications that have grown over multiple decades, and which encode an insane amount of knowledge about the respective sector and its workflows; this is a central reason why these applications are so hard to replace.

            2. zahlman · · focus · HN ↗
              Now that is moving the goalposts. The entire subthread has nothing to do with whether AI can match humans in any particular endeavour. It's simply about one user's past prediction about an AI capability.
        2. lukeschlather · · focus · HN ↗
          Well, it depends on how you define "complexity." LLMs are totally rewriting the notions of what's hard for a computer but easy for a human. LLMs have really bad spatial reasoning. I would have to play around a bit, but I'm very certain you could come up with a prompt where an AI is incapable of properly generating a "simple" app that has some important layout constraints.

          And just generally, anything that requires the LLM to understand something LLMs don't understand, it's going to fail. I'd hesitate to give an exact example without trying Astra/Fable but I'm sure they exist.

        3. pessimizer · · focus · HN ↗
          One-shot a minesweeper clone that isn't screwed up in an insane way. Bonus points if you choose a language/platform that LLMs aren't likely to have already seen a minesweeper clone done in/for already.

          Right now, they can't build anything without help that isn't buggy in unintelligible ways. If you push the thing feature by feature, have a lot of tests and a lot of instrumenting, and you check that it isn't cheating or lying after every step, you can get a lot of work done.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.