‹ BackHN Continuity

Thread

Ollaya – Ollama for open-source, Jev-style decision models

618 points · 145 comments · Ardakilic

  1. pradn · · focus · HN ↗
    I'm not sure what this means for AI startups if their innovations can be copied by OSS so quickly (what, like 2 weeks?). There's "consumer surplus" for everyone, to borrow an economic concept. But we do ideally want some of the surplus to flow to the innovator, too. I know there were precursors, but that's fine - it's hard to have a totally novel idea in such a popular field. I don't know what the end game is for TypeSafe - they'd need to demonstrate perpetually better results, or compete in another axis: UX, support, custom solutions, etc. So much of the time, someone proving a concept, or it simply getting enough publicity, is enough for a "Cambrian explosion" of follow-ups and copies. Famously, that was true for "Attention is All You Need", and the general idea of "next-token prediction" being so powerful.

    We've stumbled into general differentiable models..

    1. redox99 · · focus · HN ↗
      Because what they did is kinda trivial. Its basically like the Dropbox comment really[0], except here you don't need petabytes of storage and infinite VC pockets.

      After chatgpt everything in AI mostly became LLMs and building wrappers around them. It's like people forgot how to do ML.

      To those of us who actually trained models back in the day, its kind of cute to see people wowed by a classifier. Yes, this is 0 shot and doesn't need training (most people wanting this would've used structured output, this is cool because it's cheaper and faster). But anyone with basic ML knowledge could've built this in a few hours.

      The question is mostly why wasn't this productized. And it's interesting indeed that it took this long to become a finished product.

      [0] <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=9224">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=9224

      1. coolThingsFirst · · focus · HN ↗
        Can you explain how 0 shot classifiers work how come it isn’t trained on my data how does it know when issue is urgent lets say?
        1. dotancohen · · focus · HN ↗
          This is my burning question as well.

          I have 80,000 voice recordings to classify. Very few are in English, and the classes are not in English. Many of the classes are project names or other proper nouns. How could a system not trained on data specific to the problem possibly be expected to work?

          1. atroche · · focus · HN ↗
            you could help it along by providing clues about the domain in the prompt. including what various proper nouns might indicate (assuming they&#x27;re not chosen completely at random, and encode at least some meaning).

            and you might find the scores it gives back about its confidence are useful for escalating to a more expensive classifier.

            but also, there&#x27;s no requirement to use it as a zero-shot classifier, you can provide it as many examples as you like, and prioritise giving it examples it had previously gotten wrong. and with input caching it might be economical.

            I&#x27;d be surprised if they or others don&#x27;t start offering a fine tuning API for models like this, like openai does for some models (or used to, I haven&#x27;t checked in a long time).

            1. dotancohen · · focus · HN ↗
              If it&#x27;s not trained on my data, it&#x27;s going to be something on the order of a 500-shot prompt. This is a complex application with more corner cases than I&#x27;d like.

              As are, honestly, most projects I work on.

              There are already Jev-shaped open weight local models coming out that can be fine-tuned on specific data. We&#x27;ll see which of them the community settles on.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.