‹ BackHN Continuity

Thread

Ollaya – Ollama for open-source, Jev-style decision models

618 points · 145 comments · Ardakilic

  1. george_max · · focus · HN ↗
    Has anyone actually seen better or the same results with Laya compared to Jev? From my experience, Laya performs significantly worse. It's less confident and often makes wrong decisions with more complex queries.
    1. cobanov · · focus · HN ↗
      Developer here. You're right, Laya is a lot weaker than Jev, especially on harder queries. It's a small model, so it's fast, but that's the trade-off. The open models that get close to Jev are much bigger, and running those is what I'm working on next.
      1. mikodin · · focus · HN ↗
        What are the models? I am super curious in these as well
        1. simcop2387 · · focus · HN ↗
          Probably Kev and/or the decider models. Kev is trained on one of the 4B qwen models, similar for decider but it ranges from 0.8B through to the 35B-A3B model so far I believe.
        2. cobanov · · focus · HN ↗

          [dead]

      2. adinb · · focus · HN ↗
        It doesn’t to be a ton bigger, 16k and reliable 8k would be a godsend. (I run at 2k)
    2. scronkfinkle · · focus · HN ↗
      Yes. JEV generalizes better because they probably have an enormous corpus and trained on it for a long time. Laya's out of the box model is much weaker. However, in the age of LLM's it's incredibly easy and cheap to generate large datasets to fine tune laya for your task, and the training loop is pretty quick and cheap too.

      It's so easy that I question why I would ever pay for JEV when eventually I'll have done enough random things that I will also have a large corpus and likely a general model as well.

      1. mtkd · · focus · HN ↗
        Isn't the point of Jev that it generalises better?

        It's a fast classifier you can use out-the-box, ~1.5bn tokens is about $40 (I've been hammering it)

        It just works ... a whole bunch of low-level/low-importance workflow stuff that was getting farmed out to small/fast LLM models now has a competitive alternative ... and bits that hadn't even been considered to go into some external descision/classifier service can be tested/deployed at ~$0.00003/req

        I don't get this wall of negativity on it, it's genuinely innovative/useful tech ... would expect HN to be more positive, regardless of whether it's the absolute best execution

        1. shepardrtc · · focus · HN ↗
          It really does just work. And it works so well I already integrated it into my product. Saves me about 75% of costs for the section its working in, which isn't a small amount. I see a lot of negativity and I don't really get it either. Its so cheap and so fast, why not give it a try?
          1. DenisM · · focus · HN ↗
            I think it’s the infamous Dropbox reaction - anyone can wrap an FTP server, where the innovation?

            Starting from a business POV one should inflate terminology, hack together an MVP, and see if the market demands it before doing hardcore R&D.

            But starting from technical/craftsman POV all you see is a hack and a lot of big words, so it’s easy to become jaded.

        2. not_a_bot_4sho · · focus · HN ↗
          I didn't see any negativity in the post you replied to.

          I think the point being made is that Jev is great but it has no competitive moat, and open source versions will very soon catch up if their secret sauce is just synthetic data.

          (Whether or not that is true, I don't know.)

          1. killingtime74 · · focus · HN ↗
            It's probably true because even for Frontier LLM models there are many competitors now.
        3. digitaltrees · · focus · HN ↗
          I think your point is valid but many are annoyed that it is presented as groundbreaking, revolutionary, novel frontier tech when it is a known classification system. It’s the hype that feels undeserved. Honestly it was one of the best marketing campaigns I’ve seen.
        4. ichorio · · focus · HN ↗
          If you don't mind me asking, what are you using it for?

          I've been unable to find a good use case for now.

          1. taylorfinley · · focus · HN ↗
            You should try building something with it, the hype is what it is, but the model is crazy useful.

            I'm building a woodworking app and I've managed to create an autopilot that can take a simple instruction ("get me 5 2x4s", "cut the middle 2x4 into 4 equal pieces", "move the 2x4 3 feet left") and the action instantly happens with next to no lag. There is already an llm but now it can share an intent, and the geometry system shows jev the various actions and jev chooses the action that gets it closer to the goal until it has found a state that matches the intent or gives up. The result is the llm can think "higher level" and let the cheap fast model grind out the options in a relative blink of the eye, without the 30s of reasoning the llm would have done about the various operations it could try.

      2. _menelaus · · focus · HN ↗
        If you're so inclined it would be easy, fast and cheap to distill Jev for your task.
    3. iamflimflam1 · · focus · HN ↗
      Nothing yet. Unfortunately it sometimes feels like our industry has been overrun by grifters and chancers.

      I’m sure this has been a gradual and long decline. Maybe it even started with the dot com boom and accelerated with crypto. With AI it seems to have got worse.

    4. jonmagic · · focus · HN ↗
      I've been following jevbench twice a day for the past week and that's been a lot of fun. Latest update:

      Rank System Score Public / sealed accuracy Evidence

      1 decider-4b v2 64.13 83.5% / 34.7% Evaluator-run, offline

      2 Jev 1.13 63.29 86.6% / 36.7% Evaluator-run API

      3 JevK5 v0.2 62.04 85.3% / 33.1% Evaluator-run

      4 Cygnet 12B 61.76 87.9% / 33.8% Evaluator-run, offline

      5 Hopper 59.43 82.3% / 34.1% Evaluator-run

      28 Kev 4B 36.14 66.2% / 22.4% Evaluator-run

      41 Laya 421M 30.25 58.4% / 30.8% Evaluator-run

      <a href="https:&#x2F;&#x2F;benchmarkheaven.com&#x2F;jev-models" rel="nofollow">https:&#x2F;&#x2F;benchmarkheaven.com&#x2F;jev-models

      1. philipodonnell · · focus · HN ↗
        What the best way to see how a homegrown version compares?
      2. Havoc · · focus · HN ↗
        Amazing - was looking for some benchmarks around this earlier
      3. rubymamis · · focus · HN ↗
        Did anyone else notice the huge gap between scores on private vs public for ALL Jev-like models compared to LLMs (such as GPT Luna)? Doesn&#x27;t it mean those models aren&#x27;t generalizing so not very useful on data they haven&#x27;t seen?
    5. verdverm · · focus · HN ↗
      one day, perhaps people will click through to the laya author&#x27;s arxiv paper content and the why may become clearer, you won&#x27;t have to read it, a skim will suffice
    6. jasonjmcghee · · focus · HN ↗
      In my experience it&#x27;s not close and the benchmarks I&#x27;ve seen don&#x27;t reflect my experience at all.

      But I&#x27;m guessing people will find the right training regime and data mix soon to close the gap.

      But big things I see are instability and inaccuracy - like pick a random problem.

    7. lgas · · focus · HN ↗
      This is just anecdotal and I might be doing it wrong but I made jev and laya versions of a simple semantic grep tool (<a href="https:&#x2F;&#x2F;github.com&#x2F;lgastako&#x2F;jevplay" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;lgastako&#x2F;jevplay) and played with them a bit, and at first it seemed like laya was comparable (eg on queries like &quot;this is a mans name&quot; or &quot;this is a womans name&quot; on names.txt) but the more I played with it, eg. &quot;this is a vegetable&quot; on foods.txt the further the gap widened in favor of jev. Then I started trying variations of the query eg simply &quot;mans name&quot; and for the most part laya just fell apart and didn&#x27;t return anything useful for a lot of stuff. I was hoping to find that laya was competitive because it&#x27;s much faster to have the model running locally but it&#x27;s just not, yet.
    8. cjonas · · focus · HN ↗
      I&#x27;ve been testing, for my use case laya didn&#x27;t come close to decider was just as good. My experience with the 3 models matches the result here

      <a href="https:&#x2F;&#x2F;benchmarkheaven.com&#x2F;jev-models" rel="nofollow">https:&#x2F;&#x2F;benchmarkheaven.com&#x2F;jev-models

    9. draginol · · focus · HN ↗
      That&#x27;s the thing I think some people keep missing.

      I mean, anyone here could have a decision maker just return a random number. Fast and easy.

      The question really is how GOOD is it?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.