‹ BackHN Continuity

Thread

Grok 4.7

609 points · 541 comments · meetpateltech

  1. moojacob · · focus · HN ↗
    Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.

    Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.

    However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

    My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.

    1. imron · · focus · HN ↗
      > My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish.

      Grok has its own feel too. It's not as bad as Claude, but one of the things that bugs me is that it is far too terse.

      It regularly seems to come up with terms and descriptions for things in its chain of reasoning and then uses these terms in its output assuming you understand what it's talking about.

      I find I often have to ask it to re-explain what it means.

      1. iamflimflam1 · · focus · HN ↗
        It’s frustrating that we can’t see the “thinking” - it’s like we only have access to half the conversation.
        1. bilbo-b-baggins · · focus · HN ↗
          Devin shows model thinking.

          I’m pretty sure the big bois don’t do it because it would undermine “confidence”.

          Seeing a model output “Oh I should just delete blah. Wait blah is a production service, I shouldn’t touch that. Maybe I can gain access to blah? Oh the aws cli isn’t signed in to blah. I see kubectl has access to blah though! Wait, I should ask user permission first.”

          Yeaaaaah. Thinking tokens are fuckin’ wild.

          1. tyre · · focus · HN ↗
            I doubt most users would look at them if they were available. More likely they don’t want to stream distillation material.
            1. glub · · focus · HN ↗
              I stay much more hands-on when I'm using models that display full reasoning traces. And I tend to get more things done as a result, because I know exactly when it thought of a good solution that it talked itself out of because of some invalid assumption.

              Could be something very stupid like - "I don't have ffmpeg available here. Should I install it? No, I can't. I'll proceed doing something that will take me 100x more tokens and wall clock just to avoid adding a dependency." I can then just stop and say - you've got nix flake there, just add it.

              That's impossible with western models. The only way is to ask why it did something stupid when it already spent your 50% of your weekly quota.

              1. tyre · · focus · HN ↗
                Oh yes, I’m 100% with you. I wish they’d keep it. It’s better for users but perhaps untenable for the business.
          2. geetee · · focus · HN ↗
            Running some models locally and seeing these thinking tokens was quite the experience. I never saw an LLM so "unsure" about virtually everything.
          3. demibabs · · focus · HN ↗
            Idk I feel like the more likely answer is to prevent distillation. Having the thinking is definitely better UX (oftentimes, I don’t know if Codex is just hanging, which it often does, or working in silence).
        2. greenavocado · · focus · HN ↗
          Ah, but you CAN see the thinking if you are willing to risk your account being banned. You just have to expose a "tool" with a specially crafted definition.
        3. imron · · focus · HN ↗
          You can double click on the 'thinking' text and it will expand and you can read it. The problem is that it will often have multiple thinking/tool call sections and it can be a needle/haystack problem to find the one with the thinking you are interested in.
          1. tempay · · focus · HN ↗
            We don’t have access to the real reasoning text for most closed models these days, mostly due to distillation threats
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.