‹ BackHN Continuity

Thread

Grok 4.7

609 points · 541 comments · meetpateltech

  1. moojacob · · focus · HN ↗
    Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.

    Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.

    However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

    My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.

    1. smashers1114 · · focus · HN ↗
      FYI a quick fix for claudish is to ask for the response to be in ASD-STE100 (Simple Technical English). Then it is far more readable. But I would agree that this is an annoyance and shouldn't require user workaround to get something readable.
      1. a2dam · · focus · HN ↗
        I think this is more a meme than anything else, for a couple reasons:

        First, after a while it's just as grating as Claudeish. Second, my hunch is that it constricts the actual thinking of the LLM, like the same way that Newspeak does in 1984. It shrinks the range of thought that can be expressed if used as an input.

        I think the real way to do it is to have another Claude entirely deal with the user as a liaison, but to keep the thinking in whatever format it came in.

        Latent space reasoning, if you think about it, is exactly this to a crazy degree: why even formulate a thought as words if you can just keep it as matmuls until the user needs it? And then, if the user needs it, have it always specifically formulated for the user by another LLM rather than constrict its range of thought? Anyway, that's my take.

        1. TuxMark5 · · focus · HN ↗
          This is the same reason why I am a bit skeptical of LLM superintelligence. LLMs in the end have to operate in natural language concepts and the complexity of natural language is bounded by limits of human cognition. I'm sure super advanced AI could use concepts that humans not only have no words for, but might not be able to understand alltogether. As such if my thesis is correct, the only way forward for true superintelligence may be getting rid of natural language COTs.
          1. nomel · · focus · HN ↗
            > LLMs in the end have to operate in natural language concepts and the complexity of natural language is bounded by limits of human cognition.

            I don't think this is true.

            They have to express themselves as tokens. The meaning of those tokens doesn't have to be text. See any model that can handle images/video. Also, I don't think math, svg, etc, are "natural" language.

            And, only the final expression is tokens. The intermediate layers, with the encoded concepts, aren't "natural language".

            But, to address your concern (which nobody can disagree with, since even humans can&#x27;t fully express through text&#x2F;pictures), potentially: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49758615">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49758615

            1. _puk · · focus · HN ↗
              Yeah, as I understand it, natural language is tokenised and vectorised, and then maths takes ahold.

              The model isn&#x27;t limited to concepts that can be expressed in natural language.

              It&#x27;s only once the AI gets to the output layers that natural language comes back into play.

              After all, they&#x27;re all made out of weights[0].

              0: <a href="https:&#x2F;&#x2F;maxleiter.com&#x2F;blog&#x2F;weights" rel="nofollow">https:&#x2F;&#x2F;maxleiter.com&#x2F;blog&#x2F;weights

              1. lelanthran · · focus · HN ↗
                &gt; The model isn&#x27;t limited to concepts that can be expressed in natural language.

                How do we know for sure? We don&#x27;t even know how the emergent properties we see actually emerged?

                For humans we know for sure that people sometimes have concepts that they have no word for (the reason the phrase &quot;It&#x27;s on the tip of my tongue&quot; is a phrase, after all).

                We don&#x27;t know this for LLMs. When it makes new phrases, it&#x27;s always a mixup of two existing words hyphenated (aside, that also seems to be the limits of SOTA models creativity - join two unrelated words together with a hyphen).

                LLMs never respond with &quot;It&#x27;s on the tip of my tongue&quot; type responses, indicating it has a concept but cannot remember (or does not have) a word for that concept. Every human, pre-speech-age, has managed to express or convey concepts that they had no word for.

                So, no. I&#x27;d need a citation, preferably multiple, that did the trials and found that a model can generate concepts for which it does not have any words for.

            2. timacles · · focus · HN ↗
              Reality cannot be reduced to tokens
              1. Dylan16807 · · focus · HN ↗
                Thoughts are a poor reflection of reality to begin with.
              2. nomel · · focus · HN ↗
                Can it be reduced to ion concentrations? Because that&#x27;s how we perceive it. A useful perception is all that really matters.
            3. helloplanets · · focus · HN ↗
              And most LLMs have been multimodal for years at this point. It&#x27;s vectors all the way down.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.