‹ BackHN Continuity

Thread

Ollaya – Ollama for open-source, Jev-style decision models

618 points · 145 comments · Ardakilic

  1. fooker · · focus · HN ↗
    For everyone dismissing Jev's innovation as being trivial, no it's not.

    It is definitely not the MNIST classifier you had trained in 2019.

    The difference is that you only train it once and the modern LLM machinery sort of takes care of that with large contexts.

    It's great that Jev proved this is a viable product. I'd expect a great many research innovations coming from making this work better/faster/cheaper, and around interfacing modern agents with it.

    1. syntaxing · · focus · HN ↗
      Hah you’re probably dating yourself. Keras came out in 2015 and that was one of the early examples with Theano backend. You could train MNIST since 2015 pretty straight forward. But comparing Jev to image classification is an unfaithful argument. Comparing it to ELmo or BERT is analogously better.
      1. fooker · · focus · HN ↗
        You missed the point - you had to train BERT or anything similar to get useful results out of it.

        Now all you need is to give it more context along with your query.

        1. c10o · · focus · HN ↗
          Exactly. And there’s a lot of situations where context change frequently and fine-tuning / re-training becomes impractical or even impossible. With a general classifier + context, you can change the context dynamically and get instant results. That capability opens up a whole lot of possibilities.
        2. k__ · · focus · HN ↗
          But aren't BERTs tiny compared to a LLM and can be trained cheaply with the help of an LLM?
          1. fooker · · focus · HN ↗
            Suppose you want it to make a decision based on ..say.. 300KBs of somewhat changing information per query.

            There's no scenario where you are training BERT online to give you an answer.

            1. k__ · · focus · HN ↗
              Can any model give you a reasonably correct answer on that amount of data?
              1. fooker · · focus · HN ↗
                Yes
    2. hodgehog11 · · focus · HN ↗
      Just because it is zero-shot does not mean that Jev is an architectural innovation, especially in the year 2026.

      Anyone fitting MNIST in 2019 was already outdated by several years at the very least. GPT-2 was 2019! We already had zero-shot classifiers then. In fact, the paper for GPT-3 was literally

      "Large Language Models are Zero-Shot Reasoners"

      These kinds of zero-shot classifiers were already developed and used in-house for many years. They just weren't commercialized as a separate product, because anyone who could use an LLM proper could build layers around it to fulfill any classification task like this.

      1. fooker · · focus · HN ↗
        To get a classifier that worked in 2019, you had to train a model. Either from scratch or from a starting point.

        GPT2 was absolutely unusable as a classifier. Using GPT3 as a classifier cost a few order of magnitude more than what this thing is priced at, and the context window was a few thousand tokens.

        > These kinds of zero-shot classifiers were already developed and used in-house for many years

        This is like Google's favorite coping mechanism for falling behind at AI. "We had everything inhouse for several years, we didn't release it for $reasons."

        > anyone who could use an LLM proper could build layers around it to fulfill any classification task like this

        You missed the part where it costs more than two orders of magnitude lower :)

        I'm not claiming there are major architectural innovations, but that's not the point. Once you prove there's a market, there's a cambrian explosion of innovations.

        1. hodgehog11 · · focus · HN ↗
          Of course it costs less and is more accurate. There have been several years of language model training, distillation, and reinforcement learning to get there. None of this is thanks to the Jev team.

          There already was a market for this, that's why people continued developing them in house for their own purposes. I trained people for this. It was just silently used at various companies without the loud marketing fanfare. The fact that other groups are able to replicate and beat Jev within days should tell you something. That's not "the new normal". That's an example of a product built on marketing hype alone.

          1. fooker · · focus · HN ↗
            And I was storing files on a server in 1985.

            Doesn't mean Dropbox was not innovative.

            LLMs better than chatgpt existed at Google two years before chatgpt. Doesn't mean chatgpt was not innovative.

            Not all innovation is about groundbreaking research. More often than not, it's figuring out what people will pay for. The research to make it work well often comes after that.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.