‹ BackHN Continuity

Thread

Ollaya – Ollama for open-source, Jev-style decision models

618 points · 145 comments · Ardakilic

  1. george_max · · focus · HN ↗
    Has anyone actually seen better or the same results with Laya compared to Jev? From my experience, Laya performs significantly worse. It's less confident and often makes wrong decisions with more complex queries.
    1. jonmagic · · focus · HN ↗
      I've been following jevbench twice a day for the past week and that's been a lot of fun. Latest update:

      Rank System Score Public / sealed accuracy Evidence

      1 decider-4b v2 64.13 83.5% / 34.7% Evaluator-run, offline

      2 Jev 1.13 63.29 86.6% / 36.7% Evaluator-run API

      3 JevK5 v0.2 62.04 85.3% / 33.1% Evaluator-run

      4 Cygnet 12B 61.76 87.9% / 33.8% Evaluator-run, offline

      5 Hopper 59.43 82.3% / 34.1% Evaluator-run

      28 Kev 4B 36.14 66.2% / 22.4% Evaluator-run

      41 Laya 421M 30.25 58.4% / 30.8% Evaluator-run

      <a href="https:&#x2F;&#x2F;benchmarkheaven.com&#x2F;jev-models" rel="nofollow">https:&#x2F;&#x2F;benchmarkheaven.com&#x2F;jev-models

      1. philipodonnell · · focus · HN ↗
        What the best way to see how a homegrown version compares?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.