‹ BackHN Continuity

Thread

Grok 4.7

609 points · 541 comments · meetpateltech

  1. MuffinFlavored · · focus · HN ↗
    If the CursorBench 4.0 score diagram is the headline, I read it as "Grok 4.7 xHigh is almost the same as Fable5.1 on low".

    Is there a metric for like... time taken when comparing these two? I see score and cost.

    If Fable5.1 can knock it out more quickly on low but Grok4.7 might take twice as long to stumble through a problem (and leave behind a bunch of yucky comments or un-needed extra unit tests), are they really comparable?

    Or like... the "quality" of the solution? "It works" versus "it's unmaintainable/very messy/hacky".

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.