‹ BackHN Continuity

Thread

GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence

80 points · 102 comments · theanonymousone

  1. atkrista · · focus · HN ↗
    I would just LOVE to see all the behind-the-scenes shithousery both companies are employing to one-up the other in this, largely, 2-horse AGI race. Someone should make a mockumentary when all is said and done!
    1. TeMPOraL · · focus · HN ↗

      [dead]

      1. cindyllm · · focus · HN ↗

        [dead]

      2. f6v · · focus · HN ↗
        The sooner the better, brother.
      3. aeve890 · · focus · HN ↗
        "They trusted me. Dumb fucks" -evil ASI probably
        1. cindyllm · · focus · HN ↗

          [dead]

    2. petesergeant · · focus · HN ↗
      > largely 2-horse

      The absolute frontier is largely 2-horse, but the rest of the pack is very close behind, which I'm grateful for. Grok, Facebook, and the Chinese vendors are producing excellent models.

      1. Bluestein · · focus · HN ↗
        ... and, must be said a plethora of largely unsung, small, unknown "labs", outfits, "researchers" and the like. There is a long tail of smart people having at this. I guess, sheer compute aside, I think much progress - or, at least, important pieces thereof, will come from there.-
        1. bayindirh · · focus · HN ↗
          There are some niche research areas where bog standard machine learning algorithms make miracles. LLM is just the poster child. AI/ML is a much larger and wider research area.
          1. Bluestein · · focus · HN ↗
            Your "bog" if unrelated, put me in mind of a "Cambrian explosion" (of intelligence) in this case, ongoing.-

            In a way: We have had "recursive self improvement" (RSI) - that's what genetics is, giving rise to ... us.-

            This time around, I think the truly worrying thing is that, having transcended the biological substrate, the pace is unlike anything previously seen.-

            Just a thought.-

      2. bayindirh · · focus · HN ↗
        Gemini is also pretty nice for researching things. It turns out that having the whole internet indexed and having unlimited access to YouTube is a force multiplier of some kind.

        Since Google has their own TPUs, TPS is also pretty high w.r.t. Claude, for example.

        1. thunfischtoast · · focus · HN ↗
          I think Gemini could be a great product, if they didn't decide to stuff it down my throat at every possible occasion.

          Recent example: on my e-reader, tapping a word I don't now and clicking "Translate" pulls up the possible translations from a dictionary, a local file just a couple of megabytes big, near instantly, on this tiny processor.

          Doing the same on my Android phone starts a Gemini-chat with the prompt "Translate the word x into y". Takes forever, internet access needed, results vary, burns who knows who much energy.

          Why? Just why?

          1. selestify · · focus · HN ↗
            So that some team at Google can hit their OKRs for user adoption and get promoted.
          2. bayindirh · · focus · HN ↗
            That kind of shoving down is pretty bad, I agree.

            I neither use Android devices nor Google Search, so the only Gemini thing I see is the Gemini chat interface.

            I understand the pain, though.

        2. lxgr · · focus · HN ↗
          Gemini is unbelievably bad for research in my experience. It hallucinates like it's 2023, doesn't use its own search, makes up fake rationalizations for why it didn't need to etc.

          It's baffling that OpenAI managed to get better at web search than the company literally synonymous with web search.

          1. bayindirh · · focus · HN ↗
            Pretty interesting. It one-shots correct information with references and a further reading list 99% of the time for me.

            I don't want it to fill in the gaps though, but make it reference anything and everything it brings, hence it doesn't hallucinate much.

            If something feels off, I ask it to back it with concrete data, and if it can't, I don't consider that information correct. That happened once, though, and web doesn't have any information on that thing either. So in that case, not only it had no information on the web, the training data had no information on that thing either short of feeding confidential design documents if they were ever present in the first place.

            The question was about an instrument preamp though, so nothing crucial.

            1. lxgr · · focus · HN ↗
              I guess there's one thing that could really bias my anecdata: I almost exclusively use Gemini in a professional/niche domain and ChatGPT in a personal/idle curiosity context, so there's a chance that ChatGPT is just as wrong but I am not knowledgeable in a given domain to tell.
        3. lp92 · · focus · HN ↗
          Since Gemini 3.8 Flash release, I've been largely using that for my coding ttasks. The architectureand design work I use Opus, but the rest of the work is done via Gemini 3.8. I don'thave to worry aboutweekly limits even on the $20 plan. I get down to about 10% weekly limit before it resets. Currently I'm building my own generative modeling + aero sim package largely using Gemini and it's been working great.
    3. nbardy · · focus · HN ↗
      I think it's weirdly just a choice of deciding to cut releases.

      We already know OpenAI has "bel" that is MUCH better than astra and is being used internally

      1. mFixman · · focus · HN ↗
        Any strong enough model with weak enough safeguards can cause an AI Chernobyl event that will make people and governments against AI development and deployment, just like Chernobyl did for nuclear energy.
        1. ChromeUltron · · focus · HN ↗
          tell me "I drank the kool aid" without telling me you drank the kool aid.
          1. mFixman · · focus · HN ↗
            The US government and most large companies drank the kool aid, and they will be the ones blaming Big AI if things go very wrong.
            1. esseph · · focus · HN ↗
              I think there's going to be a brutal backlash against the entire technology sector.

              Gov and Corp will throw their hands up and explain how it's not their fault.

      2. iLoveOncall · · focus · HN ↗
        > We already know OpenAI has "bel" that is MUCH better than astra and is being used internally

        You're just believing their own bullshit. There's no indication that this is true except from claims from people working at OpenAI.

        If they really had a much more powerful model, it would make absolutely no sense to sit on it.

        1. adamzenith · · focus · HN ↗
          You don't think having a more intelligent model they can use internally that others can't is an advantage?
          1. iLoveOncall · · focus · HN ↗
            No? The top of human engineers are much better than any model would be, so AI models really aren't a big advantage when you're trying to develop anything that is SOTA.
            1. ceejayoz · · focus · HN ↗
              Not every problem is best addressed by a top engineer.

              Plenty of the tasks that keep a company running can benefit from good-enough (and better than the competition).

              1. iLoveOncall · · focus · HN ↗
                Yes, and none of those tasks require even the current SOTA models.
            2. x187463 · · focus · HN ↗
              You can't fathom how 'top human engineers' could take advantage of exclusive access to a frontier model at ultrafast speeds and limitless token budgets to conduct research?
              1. iLoveOncall · · focus · HN ↗
                I have all that and LLMs invariably produce garbage, so, no.
          2. owebmaster · · focus · HN ↗
            If that's the case, do Anthropic have an even better one helping them? The Chinese labs too? Where this OpenAI "advantage" is taking them?
        2. howdareme9 · · focus · HN ↗
          its not done training, why would they release a model that hasn't finished training?

          besides, we know anthropic are sitting on models too

          1. iLoveOncall · · focus · HN ↗
            > its not done training, why would they release a model that hasn't finished training?

            Because clearly they have no problem with releasing newer versions of models even just a week apart.

            > besides, we know anthropic are sitting on models too

            It is from your crystal ball or from other bullshit you heard from Anthropic employees on Twitter?

            We all know Anthropic had Mythos and Fable, and they turned out to be completely normal models, entirely in line with the capability of their predecessors.

            All they do is lie, and you're believing their lies.

            1. meowface · · focus · HN ↗
              The poster is not claiming Bel is a secret AGI. Just that it exists and only exists internally at the moment.

              It's rumored to be over 10T parameters. When released it'll probably be very good at certain tasks, albeit slow and expensive and not necessarily "wiser". You don't have to make this a binary.

              Also, Mythos was in fact a significant step-up in several ways. It fits the trend line, but only because the trend line for LLMs is quite steep. Plus Astra is still in many ways less intelligent than Fable/Mythos despite being released much later.

            2. azan_ · · focus · HN ↗
              > its not done training, why would they release a model that hasn't finished training?

              > Because clearly they have no problem with releasing newer versions of models even just a week apart.

              You can see how it is pure non-sequitur, right?

              > We all know Anthropic had Mythos and Fable, and they turned out to be completely normal models, entirely in line with the capability of their predecessors.

              Fable was absolutely not in line with capabilities of other models when released. For cybersec work it was much, much better.

        3. meowface · · focus · HN ↗
          With all due respect, you do not have a single clue what you're talking about.
          1. owebmaster · · focus · HN ↗
            With the same respect, you don't either. Simping for openai don't make you part of their in group
            1. meowface · · focus · HN ↗
              I definitely do not know what I'm talking about, but that individual doesn't know what they're talking about even more than I don't know what I'm talking about.
        4. simonw · · focus · HN ↗
          It makes sense for them to sit on it until they've finished testing it. More powerful but also more likely to delete all your email by mistake = you shouldn't release it yet.
        5. 233mhz · · focus · HN ↗
          > If they really had a much more powerful model, it would make absolutely no sense to sit on it.

          Makes total sense if they don't have the compute and can't serve it in an economically viable way. Also it lets you build things no one else in the world can build as fast as you until it's released

          1. iLoveOncall · · focus · HN ↗
            > Also it lets you build things no one else in the world can build as fast as you until it's released

            "We can generate slop faster than anyone in the world" :evil_emoji:

          2. owebmaster · · focus · HN ↗
            This theory would be easy to prove IF openai was pushing high quality software. It's not.
        6. _davide_ · · focus · HN ↗
          > it would make absolutely no sense to sit on it.

          Yeah, it does, it might be misaligned, a snapshot of the going on training, bigger than they can serve publicly, not yet completed the full training pipeline.

          If you see the knowledge cutoff you can see that sol 5.6 finished the initial main (+stage edited) of the training pipeline on Feb 16, 2026 but it was publicly released on July 9.

          The opposite would be weird: if they do NOT have an unreleased in-house model that would be really odd.

          1. iLoveOncall · · focus · HN ↗
            > The opposite would be weird: if they do NOT have an unreleased in-house model that would be really odd.

            This isn't at all what I said. The original commenter mentioned a model MUCH stronger than Astra.

          2. nextaccountic · · focus · HN ↗
            > The opposite would be weird: if they do NOT have an unreleased in-house model that would be really odd.

            Indeed it would be really odd if OpenAI were actually open

      3. 233mhz · · focus · HN ↗
        > We already know OpenAI has "bel" that is MUCH better than astra and is being used internally

        I mean, isn't it almost a guarantee that what we get is a gimped version of what they use internally? They probably already serve themselves next gen level models at 1k+ tps from cerebras machines hosted on perm while we get quantized astra/opus at 50tps on a good day

        1. arcfour · · focus · HN ↗
          Was it not essentially confirmed that all of the frontier labs have much stronger models internally that don't make sense to serve at scale yet, which they use for development and to train models that are able to be served at scale?
          1. Alifatisk · · focus · HN ↗
            Where was this confirmed?
    4. michelb · · focus · HN ↗
      I REALLY need new seasons of 'Silicon Valley'
      1. tom1337 · · focus · HN ↗
        Still waiting for the moment where GPT either orders 4,000 pounds of raw beef or just deletes the whole OpenAI repo because it "thought the easiest way to remove all bugs is by deleting the whole repository"
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.