‹ BackHN Continuity

Thread

Ollaya – Ollama for open-source, Jev-style decision models

618 points · 145 comments · Ardakilic

  1. pradn · · focus · HN ↗
    I'm not sure what this means for AI startups if their innovations can be copied by OSS so quickly (what, like 2 weeks?). There's "consumer surplus" for everyone, to borrow an economic concept. But we do ideally want some of the surplus to flow to the innovator, too. I know there were precursors, but that's fine - it's hard to have a totally novel idea in such a popular field. I don't know what the end game is for TypeSafe - they'd need to demonstrate perpetually better results, or compete in another axis: UX, support, custom solutions, etc. So much of the time, someone proving a concept, or it simply getting enough publicity, is enough for a "Cambrian explosion" of follow-ups and copies. Famously, that was true for "Attention is All You Need", and the general idea of "next-token prediction" being so powerful.

    We've stumbled into general differentiable models..

    1. redox99 · · focus · HN ↗
      Because what they did is kinda trivial. Its basically like the Dropbox comment really[0], except here you don't need petabytes of storage and infinite VC pockets.

      After chatgpt everything in AI mostly became LLMs and building wrappers around them. It's like people forgot how to do ML.

      To those of us who actually trained models back in the day, its kind of cute to see people wowed by a classifier. Yes, this is 0 shot and doesn't need training (most people wanting this would've used structured output, this is cool because it's cheaper and faster). But anyone with basic ML knowledge could've built this in a few hours.

      The question is mostly why wasn't this productized. And it's interesting indeed that it took this long to become a finished product.

      [0] <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=9224">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=9224

      1. andai · · focus · HN ↗
        &gt; The question is mostly why wasn&#x27;t this productized.

        Probably because doing it wrong (using an llm in place of a classifier) is more profitable? (For the people selling inference.)

      2. nickstinemates · · focus · HN ↗
        It wasn&#x27;t until recently that demand for classifiers at this scale existed. Jev exists because LLM&#x27;s exist. Without them it wouldn&#x27;t be (as) useful
        1. jmalicki · · focus · HN ↗
          That&#x27;s not at all true.

          There have been tons of applications for this. People were using earlier LLMs like BERT for classifiers long before LLMs became viable chatbots.

          1. janalsncm · · focus · HN ↗
            I think GP has a point, though. BERT models have existed for a while but OpenAI made classification via LLM convenient and accessible for regular developers.

            People didn’t know they wanted classifiers until OpenAI gave them a taste.

        2. calebkaiser · · focus · HN ↗
          I don&#x27;t know if that first part is true? Classifiers were&#x2F;are one of the dominant applications of classic ML and neural networks, especially in production. Even today, image classification, object recognition, language detection, segmentation models etc are still super common.

          I think the hype with Jev is just that, while structured generation is great, LLM judges tend to kind of suck for precise classification. And the more powerful the base model, the more accurate they can get, but they get increasingly expensive&#x2F;impossible to finetune. It was specifically the latency&#x2F;price point Jev offered vs. the general accuracy it claimed that generated all the excitement. Plus the promise of cheap calibration (tuning).

          &quot;Jev exists because LLMs exist&quot; is kind of a truism, as Jev apparently is literally a Transformer model.

          1. nickstinemates · · focus · HN ↗
            How many people were doing ML work pre LLMs? And how many are using LLMs now?
            1. calebkaiser · · focus · HN ↗
              I think the answers are &quot;a lot&quot; and &quot;a lot more&quot;? But what I&#x27;m saying is that Jev&#x27;s virality isn&#x27;t because people didn&#x27;t have access to classifiers before it. Jev&#x27;s general claims about capability and performance vs. cost would have been a very big deal 5 years ago too. In a vacuum, the idea that you can get a general classifier that is very accurate across any domain and on any modality with minimal latency and a very low price point is wild.

              In early 2016, Clarifai&#x27;s core product was basically just an image classifer exposed via an API. And at that point, they&#x27;d raised $40 million--the same amount as TypeSafe.ai&#x2F;Jev--and they were experiencing viral growth among developers + signing contracts with a bunch of flashy logos. The demand was so high that Amazon launched Rekognition and Google launched their similar APIs to compete.

              The AI hype cycle and the number of people thinking about using AI certainly puts more wind at Jev&#x27;s back, but even in an alternative universe where we don&#x27;t have contemporary LLMs, Jev&#x27;s core claims would be remarkable and there would be a big market for it.

      3. eadwu · · focus · HN ↗
        The answer for why it wasn&#x27;t productized might just be pretty straightforward.

        LLMs still are better than Jev at the task, just across the board slower.

        Anyone who had a reason to try this already tried it (ads&#x2F;recommendations) - back in 2023&#x2F;2024 during the first fine tuning wave and it was accurately determined that it was not worth the effort, the results were more bogus than just using CoT, so frankly parallelism meant nothing if bogus * parallel = bogus.

        So thrown into the dumpster and nobody really cared to revisit because it was already tried.

        Pretty much sometime between then and now it somehow became the state where the tradeoff makes sense now.

        1. wongarsu · · focus · HN ↗
          However in the last 2-3 years LLMs became a lot more efficient. Small models without CoT are still bad, but leagues ahead of where they used to be

          Maybe it&#x27;s just a case of the idea just now crossing the threshold into working just good enough to be worth it

      4. coolThingsFirst · · focus · HN ↗
        Can you explain how 0 shot classifiers work how come it isn’t trained on my data how does it know when issue is urgent lets say?
        1. dotancohen · · focus · HN ↗
          This is my burning question as well.

          I have 80,000 voice recordings to classify. Very few are in English, and the classes are not in English. Many of the classes are project names or other proper nouns. How could a system not trained on data specific to the problem possibly be expected to work?

          1. atroche · · focus · HN ↗
            you could help it along by providing clues about the domain in the prompt. including what various proper nouns might indicate (assuming they&#x27;re not chosen completely at random, and encode at least some meaning).

            and you might find the scores it gives back about its confidence are useful for escalating to a more expensive classifier.

            but also, there&#x27;s no requirement to use it as a zero-shot classifier, you can provide it as many examples as you like, and prioritise giving it examples it had previously gotten wrong. and with input caching it might be economical.

            I&#x27;d be surprised if they or others don&#x27;t start offering a fine tuning API for models like this, like openai does for some models (or used to, I haven&#x27;t checked in a long time).

            1. dotancohen · · focus · HN ↗
              If it&#x27;s not trained on my data, it&#x27;s going to be something on the order of a 500-shot prompt. This is a complex application with more corner cases than I&#x27;d like.

              As are, honestly, most projects I work on.

              There are already Jev-shaped open weight local models coming out that can be fine-tuned on specific data. We&#x27;ll see which of them the community settles on.

        2. danwitt · · focus · HN ↗
          It&#x27;s not magical and you won&#x27;t get good results if you ask a question that is that vague. Instead you will want to decompose it in to decidable questions that make up &quot;is urgent&quot; and if you find a case that corresponds to &quot;is urgent&quot; you add a new question to the panel that surfaces that.
      5. djfdat · · focus · HN ↗
        Theorizing, but I&#x27;m guessing that before LLMs, people weren&#x27;t using anything for situations like these. As people started using LLMs, people&#x27;s use cases grew, LLMs got slower and more expensive, and people lost the wow-factor and are now worrying about price. Great timing to launch a product like this, where certain use cases can be distilled down into something faster and cheaper.

        There&#x27;s probably other areas where people are using LLMs where a more tailored ML solution might work better.

      6. CactusOnFire · · focus · HN ↗
        I am looking forward to some hotshot implementing a jev-style model and then looking like a cost cutting, performance boosting wizard by replacing it with a logistic regression.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.