‹ BackHN Continuity

Thread

Introducing System One Models and Jev

1989 points · 520 comments · albelfio

  1. pennomi · · focus · HN ↗
    > Extraordinary claims require extraordinary evidence so see below for the receipts.

    Yes, that’s the kind of attitude I want to see in these model releases

    1. ramon156 · · focus · HN ↗
      But the evidence is not there...
      1. pennomi · · focus · HN ↗
        Indeed, they talk as skeptics but don’t offer a ton of evidence, other than a couple videos of demos. A live demo would be far more convincing.
        1. simianwords · · focus · HN ↗
          They gesture at not using benchmarks for some reason...
          1. meric_ · · focus · HN ↗
            <a href="https:&#x2F;&#x2F;typesafe.ai&#x2F;blog&#x2F;antibenchmaxxing" rel="nofollow">https:&#x2F;&#x2F;typesafe.ai&#x2F;blog&#x2F;antibenchmaxxing

            But also effectively this is a classification model. It excels at specific certain types of workloads, and obviously will fail at others. Not really sure how one benchmarks this tbf. I can see their argument on why this requires a novel specific eval for whatever your usecase is. A consistent &quot;global&quot; benchmark might be hard to do

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.