‹ BackHN Continuity

Thread

Grok 4.7

609 points · 541 comments · meetpateltech

  1. simonw · · focus · HN ↗
    <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F58dc4e8ec482330856fce89dac670727" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht... - default reasoning level.

    Here&#x27;s reasoning level high: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F8ab126bda2b384264b3ad931e3ebb8b4" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    For some reason reasoning effort low and medium used similar numbers of tokens, and xhigh used less than high. I think I need to try without OpenRouter in the middle.

    UPDATE: I tried again with the xAI API directly: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F5a1819a2bd24bb642f38c4bc6733090f" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht... - not a great deal of difference between reasoning levels, and this time xhigh and low used the same number of reasoning tokens for some reason.

    For comparison here&#x27;s a fresh run against Grok 4.6: <a href="https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ffedc404b9aa8e6fca31d59c898dabba0" rel="nofollow">https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht...

    1. MattDamonSpace · · focus · HN ↗
      Are there good tools for doing context audits? I feel I have no good way to visualize what a new session is getting by default in a given repo without crawling through every potentially included markdown file
    2. datsci_est_2015 · · focus · HN ↗
      Poor fella doesn’t have a seat. Intriguing design where both pedals are on the same side of the frame. Balancing must be a challenge.
    3. kiliancs · · focus · HN ↗
      What is the default reasoning level?
    4. forgot-my-pw · · focus · HN ↗
      I tried in Cursor and see a lot of improvements over Grok 4.6 svgs. The AA numbers indicate it&#x27;s not very token efficient though: <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;agents&#x2F;coding-agents?agents=codex-deepseek-v4-pro-0813-max%2Ccodex-gpt-6-astra-max-reasoning-effort-max%2Cclaude-code-opus-5-max%2Ccodex-gpt-5-6-sol-max-reasoning-effort-max%2Cdevin-fusion-cli-claude-fable-5-1-xhigh-swe-2-medium%2Cclaude-code-qwen3-8-max%2Cdevin-fusion-cli-gpt-6-astra-xhigh-swe-2-medium%2Cantigravity-sdk-gemini-3-8-flash-high%2Cmuse-code-muse-spark-1-3-max%2Ckimi-code-cli-kimi-k3%2Cclaude-code-fable-5-1-max-with-fallback%2Copencode-glm-5-3-reasoning-effort-max%2Cgrok-build-grok-4-7-xhigh%2Cgrok-build-grok-4-6-xhigh#coding-agents-token-usage-chart-tabs" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;agents&#x2F;coding-agents?agents=co...
    5. TomGarden · · focus · HN ↗
      I think these are the worst I&#x27;ve seen, at least in some time. It&#x27;s a silly benchmark though, not sure what to make of it
      1. Mashimo · · focus · HN ↗
        If you think this is bad, look up mistral.
      2. athrowaway3z · · focus · HN ↗
        I think the result is fine. The benchmark is silly to the point of being useless nowadays.

        It used to be a mess in various interesting ways. Now, almost every big release can draw something perfectly functional.

        So the question - without a correct answer - given the prompt &quot;Generate an SVG of a pelican riding a bicycle&quot;:

        Does the user want the least lines of code to make it functional, or the best looking version?

        1. wolttam · · focus · HN ↗
          The user at the very least expects the bike to have bike geometry; Grok seems to struggle with that
        2. peder · · focus · HN ↗
          exactly, the benchmark just needs to be downvoted into oblivion each times it&#x27;s posted. The outcome is not deterministic and the model needs to determine what level of detail is appropriate for an svg. There is no wrong answer to this unless it&#x27;s obviously un-Pelican-like.
      3. paimapi · · focus · HN ↗
        it&#x27;s not truly tested until it plays a match or ten in Brood War imo
    6. daveguy · · focus · HN ↗
      Hahaha. I remember when musk and his merry band of sycophants were bragging about grok producing the only physically accurate bicycle. What happened?
      1. simonw · · focus · HN ↗
        Yeah, he had opinions: <a href="https:&#x2F;&#x2F;twitter.com&#x2F;elonmusk&#x2F;status&#x2F;2023833496804839808" rel="nofollow">https:&#x2F;&#x2F;twitter.com&#x2F;elonmusk&#x2F;status&#x2F;2023833496804839808
    7. nicolamanzini · · focus · HN ↗

      [dead]

      1. danappelxx · · focus · HN ↗
        Wow, thanks for sharing, fun benchmark!
      2. iankp · · focus · HN ↗
        Would be amazing to see this for frontend.
    8. mrtesthah · · focus · HN ↗

      [dead]

    9. jfoster · · focus · HN ↗
      The default reasoning level seems better than the high reasoning level:

      Has a shadow

      Better shaped beak

      Leg position more realistic for bicycle riding

      Better feathers

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.