‹ BackHN Continuity

Thread

Vote on which of Hacker News' challenges for AI have been met

202 points · 271 comments · stabbles

  1. ben_w · · focus · HN ↗
    Very pleased one of my predictions was totally wrong: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=23252711">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=23252711

    Sure, sure, what LLMs make still isn&#x27;t &quot;efficient bug-free code&quot;: my prediction is falsified because while LLMs can write and train new models with machine learning, ML is fundamentally not advanced enough to throw arbitraty new tasks at like this.

    1. FabCH · · focus · HN ↗
      Somewhat appropriate the site the OP links to is called „goalposts“ because as far as I can see, people keep shifting theirs.

      In your case, the comment you link to says „business tasks“ and you expanded it now to „arbitrary new tasks“. Those are not the same. An LLM today sure can do many many many business-speak conversion tasks.

      1. tripleee · · focus · HN ↗
        &gt; An LLM today sure can do many many many business-speak conversion tasks

        Not reliably, and not without supervision. That&#x27;s the main point. I&#x27;m trying really hard to figure out a workflow that doesn&#x27;t require me to review the code and I just don&#x27;t see how it&#x27;s possible (yet)

        You either need a comprehensive test suite (which requires understanding the code in order to create) or you need to review the actual implementation code to make sure it does the right thing

        1. FabCH · · focus · HN ↗
          Code is a tiny part of &quot;business&quot;.

          Most business is correspondence with people who want money from you and people you want money from.

          1. rstuart4133 · · focus · HN ↗
            The issue is that correspondence is legally enforceable [0]. LLM&#x27;s are good and getting better, but if LLMs are giving enforceable undertakings, you want to very sure they are not going to promise something that will send the company broke.

            I&#x27;m not sure what risk a businessman is willing to accept, but I&#x27;d be asking for probabilities under once in a millennium. The latest round has improved considerably in their ability to follow instructions (thank $DEITY), but they aren&#x27;t anywhere near that yet.

            [0] <a href="https:&#x2F;&#x2F;www.bbc.com&#x2F;travel&#x2F;article&#x2F;20240222-air-canada-chatbot-misinformation-what-travellers-should-know" rel="nofollow">https:&#x2F;&#x2F;www.bbc.com&#x2F;travel&#x2F;article&#x2F;20240222-air-canada-chatb...

            1. sokoloff · · focus · HN ↗
              People who demand risks to be lowered to once-a-millennium are not the type to go start or even run businesses.

              There’s nothing wrong with that, but starting a business means fading several once-a-year risks of failure and running even an established one means facing several once-a-century risks every year.

        2. user43928 · · focus · HN ↗
          No, you don&#x27;t.

          I&#x27;ve stopped reviewing the code in my mobile app project months ago. I now only look at files changed and lines count in MRs. Functionality is best verified via manual QA testing.

          I know that people here are going to doubt the quality of my project and say that it is impossible, but they are clueless and have evidently not build a project in this way. Experiences from eg. corporate backend work are hardly relevant.

          It is clear to me that the fewer consequential mistakes people find during code review, their attention to code reviews is going to go down, to the point of also skipping them.

          I expect that for most development, not reviewing the code will be the standard by March of next year. Only critical code like authentication will be reviewed.

          1. suddenlybananas · · focus · HN ↗
            What&#x27;s the app?
            1. user43928 · · focus · HN ↗
              Not public.

              It&#x27;s a paid app, and after 6 months of work I expect to publish it this month.

              In other words, you will have to take my word for its quality.

          2. tripleee · · focus · HN ↗
            what do your prompts &#x2F; your workflow look like?
            1. user43928 · · focus · HN ↗
              Nothing special at all, it just works.

              One thing that I do is have each change reviewed by another model. If I implement with Codex, I would have Claude do the review, or the other way around.

              I did not benchmark this against a review from a subagent with the same model, so I don&#x27;t know if that in particular helps.

        3. sokoloff · · focus · HN ↗
          I have a task that I do once per year for a robotics team that I mentor: Roughly,

          Take this calendar of events and rank your preferences for the event lottery. Events are spread across 5 weeks, some are 20 minutes away, some are 4.5 hours away, some are Friday&#x2F;Saturday, some are Saturday&#x2F;Sunday, some are historically extremely competitive, some fill up in round 1, others don’t even fill after round 2. For the last two years, I’d written some scripts to scrape the event sites, find the addresses, ask Google Maps to give me driving distances and times, scrape prior year registration information to find which teams went and the strength of those teams, etc. It was several hours of effort.

          This year, ChatGPT was capable of doing almost all of that basic research and data conversion, filling out our internal spreadsheet. It probably still took 4 hours on the wall clock, but at 2% attention (5 minutes of human toil).

          I doubt I go a single workday without some kind of “I have an idea and I know there are disparate data sources out there; go find those and cross-correlate or extract the relevant data points.” question that is now 10-50x more efficient than 2 years ago.

      2. Dylan16807 · · focus · HN ↗
        You can&#x27;t ignore the rest of the sentence. &quot;every other task their business does&quot; &quot;everyone will be out of a job&quot;

        This means it has to handle basically all business tasks, so &quot;arbitrary&quot;. I&#x27;m not sure what percent you have in mind by &quot;many many many&quot; but I would say it can&#x27;t code half the things you need in an efficient and minimally buggy way.

        1. FabCH · · focus · HN ↗
          What code does a village vet clinic need? In all seriousness.

          Even IF they need code, they need at best a CRUD app to track patients, that&#x27;s it. There is no way Fable or Opus 5.5 can&#x27;t one-shot a village vet clinic app in 30 minutes, and only with &quot;I need a village vet clinic app&quot; as a prompt, and whatever questions it decides to ask along the way with it&#x27;s &quot;ask user&quot; tool.

          Or a florist, to use the example from a sibling comment.

          Code is tiny part of &quot;business&quot;.

          1. ben_w · · focus · HN ↗
            &gt; What code does a village vet clinic need? In all seriousness.

            Automated diagnostics, pharmacist, surgical robot, something to express anal glands without harming the patient.

            Dog-English machine translation.

            1. pixl97 · · focus · HN ↗
              And they pay a lot for a CRM that keeps track of pets, vaccinations, appointments, x-ray images, tests and charts, and pet deaths and sending information out to text or mail.

              I did support for around 15 independent vet clinics in the past.

          2. Dylan16807 · · focus · HN ↗
            Anything you can&#x27;t solve with code just means the AI is doing worse on the benchmark isn&#x27;t it? That&#x27;s why I didn&#x27;t go into detail on that aspect.

            And that one shot app is not going to be bug free.

      3. ben_w · · focus · HN ↗
        I&#x27;m not always precise with my language, but business tasks can be pretty broad, I think &quot;arbitrary new tasks&quot; is not an unreasonable rephrasing on my part?

        Consider I was replying to this:

        &gt; So are we all going to be out of a job?

        While your boss now has the capacity to ask Claude to train a new AI model to auto-balance a tower defence game&#x27;s mob, cost, and tower parameters (I know because I&#x27;ve done it), this only matters if you and your boss are working in a video games company.

        If you and your boss are actually florists, you care if your boss can get Claude to automate a rose pruning, dead-heading, and fertilising robot.

        People are trying, but I don&#x27;t think they&#x27;d be happy with 91.5% success rate: <a href="https:&#x2F;&#x2F;www.emerald.com&#x2F;ir&#x2F;article-abstract&#x2F;doi&#x2F;10.1108&#x2F;IR-04-2026-0198&#x2F;1398287&#x2F;Design-and-experimental-evaluation-of-an?redirectedFrom=fulltext" rel="nofollow">https:&#x2F;&#x2F;www.emerald.com&#x2F;ir&#x2F;article-abstract&#x2F;doi&#x2F;10.1108&#x2F;IR-0...

        1. FabCH · · focus · HN ↗
          Don&#x27;t get me wrong, we are all guilty of this.

          It&#x27;s just amazing how quickly we accept that models are good at something.

          My florist boss can&#x27;t get Claude to automate rose pruning. But she sure as hell doesn&#x27;t need to wait until Jacques is back in the shop to respond to that French supplier anymore. There is a lot of &quot;business tasks&quot; that are just paper being shuffled around no matter if you are a florist, baker, workshop owner, custom CNC shop, student offering lessons in extra time or whatever. And LLMs are already scary good at those.

          1. ben_w · · focus · HN ↗
            &gt; There is a lot of &quot;business tasks&quot; that are just paper being shuffled around no matter if you are a florist, baker, workshop owner, custom CNC shop, student offering lessons in extra time or whatever. And LLMs are already scary good at those.

            Yes indeed, but I was responding to &quot;So are we all going to be out of a job?&quot;, not &quot;Will AI radically change the jobs market?&quot;

            We got the thing I thought would make everyone unemployed (AI which can make AI), but it turned out the AI good enough to make AI, happened before we figured out the general problem of few-shot learning that would mean the AI made by AI puts us all out of jobs.

        2. [deleted] · · focus · HN ↗

          [deleted]

    2. vlyan · · focus · HN ↗
      so the conditions for your prediction simply haven&#x27;t been met yet.

      if&#x2F;when you can tell a model to do a thing and be confident that it did the thing, it&#x27;s joever for 90% of knowledge workers.

      1. ben_w · · focus · HN ↗
        The relevant condition was met; my misjudgement was that meeting it would require ML to be advanced enough to be able to train on arbitraty tasks from realistic (ie small) numbers of examples.
    3. tripleee · · focus · HN ↗
      &gt; reliably convert business-speak into efficient bug-free code

      I actually think this would take AGI to solve, which makes me optimistic about the future of software development.

      All the benchmarks are currently testing against automated tests the AI can use as an oracle

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.