‹ BackHN Continuity

Thread

OpenAI is well positioned to fast-follow Jev

328 points · 233 comments · JohnBerryman

  1. orbital-decay · · focus · HN ↗
    Every major AI shop has a ton of in-house classifiers already, big, small, generalist, specialized. Some are used in inference pipelines (e.g. safeguards), some are used in data preparation, training, analysis and investigation, research, various one-off and intermediate tasks etc. Offering them on a public API doesn't always make business sense. I don't see much substance to this buzz, looks like people that are new to all this are discovering that classifiers exist, they are more efficient at classification, and many tasks commonly done with generative models are classification in disguise. Which is not bad at all, a fresh look at their use is great to have.
    1. bigmadshoe · · focus · HN ↗
      Correct me if I'm wrong, but a zero-shot classifier like Jev is fundamentally different to a classifier with a fixed task (e.g. for safeguards), unless they trained a general purpose system to complete the safeguard task, which seems unlikely.
      1. janalsncm · · focus · HN ↗
        Correct, but zero-shot classifiers are also not new.
        1. BoorishBears · · focus · HN ↗
          But zero-shot classifiers with this level of intelligence, world knowledge, ergonomics, cost profile, and ease of use are new.

          I feel like good engineering doesn't just ignore those things, or at least it didn't before recently. Now I guess social media has added a pressure to reduce everything to a hot take.

          1. hodgehog11 · · focus · HN ↗
            An LLM is a zero-shot classifier with a large number of classes. All you need to do is establish what the output means and you can fine-tune an LLM final layer for this task if you like (and others have done). A student of mine did this as an exercise two years ago, and it was cool, but not publishable.

            I agree with you on the "ease of use" business though. No one thought to make this sort of thing commercially available.

            But there is no hot take here. Jev is not some new paradigm; engineering-wise, it is a trivial modification to the existing pipeline. That doesn't mean it isn't commercially viable.

            1. ozgung · · focus · HN ↗
              > All you need to do is establish what the output means and you can fine-tune an LLM final layer for this task

              Yeah, that was the original idea with GPT, Generative Pre-trained Transformer, and earlier open pre-trained transformers.

              Today people use AI via APIs rather then fine-tuning models by themselves and when someone provides this as an API they got excited.

              1. hodgehog11 · · focus · HN ↗
                No it wasn't. Those were models trained from scratch, required large scale data, and the nontrivial parts involved training at scale and the autoregressive task which no one expected to work as well as it does. It is the difference between developing a foundation model, and using one. I believe Jev falls in the latter category, because the task itself is no different, only the output.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.