‹ BackHN Continuity

Thread

OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005

738 points · 448 comments · sohkamyung

  1. mgaldys4 · · focus · HN ↗
    Even the Millennium Prize Problems have, in a way, become benchmarks for model companies to prove themselves. The smartest individuals among humans are becoming replaceable. Intelligence has become a product you can quantify and buy with electricity. That feels awful.
    1. applfanboysbgon · · focus · HN ↗
      Solving obscure puzzle samples that approximately ~0 humans on Earth ever attempted to solve, mostly by pattern matching known solutions to similar puzzles, is not intelligence. DeepBlue has been outperforming the best humans at a specific puzzle-like task since the last century.

      Do any of the people proclaiming this shit actually use these models? No matter how many headlines are coming out, every day I deal with reams of the most horrific code I've ever seen technically compile, with routine mistakes that any human would get fired for if they made.

      1. fidotron · · focus · HN ↗
        But humans have been confusing pattern matching against known solutions for intelligence for a hundred years!

        Seriously though, it ends up looking like that. To take a stupid example a couple of weeks ago I asked an agent to look at porting my hand written WebGL renderer (+ shaders etc) to WebGPU. It estimated a human would take 6-10 weeks, and I would agree. (Which is why I hadn't done it). 24 hours later it was deployed and live. This is classic tedious, difficult, low level if quasi mechanical work (rather like cracking an enigma message), and LLMs absolutely fly through it.

        1. trixn86 · · focus · HN ↗
          Time estimations by LLMs are hilariously incorrect all the time. It estimates very simple things that a human could do in an hour to take days or weeks and other things that are genuinely tedious and time-consuming it estimates taking a few hours. LLMs have no understanding of time and no world model that even allows them to make correct time estimates. They will always fail to provide decent time estimates unless the task is well-known, in the training data and they can extrapolate that with a simple math script.
          1. shmeeed · · focus · HN ↗
            My favourite story is of an elderly coworker with no AI experience who innocuously asked GPT-5-mini (at the time his framework's default model) to translate a 50 page engineering spec, and it kept him at arm's length for about a week about how that task would take just another 24-36 hours more. He kept asking like, you finished yet? and it just made up excuses, "so sorry, I got distracted, give me one more day", and he went "please finish, I need this", and 5-mini came back "I totally understand -- let me get to work immediately, I'll report back ASAP", end of conversation. He got annoyed but never even suspected anything wrong, because this is the kind of conversation he's used to. I wonder how long this would have continued, if I hadn't intervened by chance.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.