‹ BackHN Continuity

Thread

Grok 4.7

609 points · 541 comments · meetpateltech

  1. moojacob · · focus · HN ↗
    Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.

    Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.

    However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

    My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.

    1. smashers1114 · · focus · HN ↗
      FYI a quick fix for claudish is to ask for the response to be in ASD-STE100 (Simple Technical English). Then it is far more readable. But I would agree that this is an annoyance and shouldn't require user workaround to get something readable.
      1. _boffin_ · · focus · HN ↗
        Does not work for Claude, at least for me and I put it as the system prompt
        1. LPisGood · · focus · HN ↗
          I don’t think system prompts are particularly reliable way to do much at all. It’s better to put it as a hook after each response, or a skill at least so you can trigger it at will if you don’t want it everytime.
          1. kekebo · · focus · HN ↗
            Do you think they're unreliable based on the position in the conversation or other factors?
            1. MisterMunchkin · · focus · HN ↗
              Anthropic has probably RL’d the system prompt into nothing because of their fear of the user being able to control the model. If it listened to you about the slop language, it might listen to you if you asked it to help you with no-no tasks.
        2. ffsm8 · · focus · HN ↗
          It does work, you however have to put it into every single prompt in which you didn't want a rubbish response

          Literally every one, even 1-2 prompts later it starts to go back

        3. bel8 · · focus · HN ↗
          For me it works at first but Claude models forgets it after some prompts, despite only using like 100k tokens.
        4. tempest_ · · focus · HN ↗
          Your best bet is to use hooks and inject it after every file edit / response by first running the content through haiku and asking if it is asd 100 ste.

          It burns more tokens but is the only way to get tolerable text.

          1. hungryhobbit · · focus · HN ↗
            Doesn't it just get attenuated and start ignoring those commands?
            1. tempest_ · · focus · HN ↗
              The hook sends the text to another agent/context with a request to validate and return a good or bad + reason response. Every request is a fresh context.

              <a href="https:&#x2F;&#x2F;code.claude.com&#x2F;docs&#x2F;en&#x2F;hooks-guide#agent-based-hooks" rel="nofollow">https:&#x2F;&#x2F;code.claude.com&#x2F;docs&#x2F;en&#x2F;hooks-guide#agent-based-hook...

              1. hungryhobbit · · focus · HN ↗
                Yes but Claude starts ignoring messages when it keeps getting told the same thing over and over.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.