‹ BackHN Continuity

Thread

Grok 4.7

609 points · 541 comments · meetpateltech

  1. Tsarp · · focus · HN ↗
    Waiting on simonw "Generate an SVG of a pelican riding a bicycle " benchmark to judge this model
    1. rvz · · focus · HN ↗

      [dead]

      1. user43928 · · focus · HN ↗
        You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for?
        1. TylerE · · focus · HN ↗
          Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes.
          1. lumirth · · focus · HN ↗
            Have you considered that the single most impressive breakthrough of LLMs as a technology is their ability to generalize beyond what they were explicitly trained on? Great analogy, pal, but LLMs aren't cars.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.