‹ BackHN Continuity

Thread

Introducing System One Models and Jev

1989 points · 520 comments · albelfio

  1. zmmmmm · · focus · HN ↗
    The eval is baffling me

    > we assume there is a correct compute graph (a “workflow” represented in code) and use the predictions of the largest, smartest, and most expensive external models as reference probabilities. ... Rephrased: every model gets the same workflow. We test how they compare to the average of the smartest models (in this case, Astra and Fable).

    They assume there is a correct graph, but they don't compare to that, they compare to the average of the smarts models? So the smartest models are getting it wrong but you compare that anyway as a benchmark? So the outcome is "how much of a Fable am I getting" etc. Why not compare the actually correct thing?

    But then even on this hand constructed eval, the first plot is showing Jev at less than Sonnet 5 accuracy. It is barely better than Luna. There are two Opus 5's and two Sonnet 5's without explanation. What is the plot showing?

    I gave up.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.