‹ BackHN Continuity

Thread

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

575 points · 225 comments · firelex

  1. adrithmetiqa · · focus · HN ↗
    Forgive my lack of understanding but how long before Jev type functionality is just built straight into all frontier models?
    1. bigyabai · · focus · HN ↗
      BeRT and FLAN-T5 were used as classifiers 5-7 years ago, they were technically "frontier" for their time.
      1. alanwreath · · focus · HN ↗
        This is the exact comment I’ve been waiting for, what is the difference between classifiers and jev?
        1. make3 · · focus · HN ↗
          FLAN-T5 generated text (Jev does not generate text), and BERT wasn't able to do tasks without fine-tuning.

          Jev is basically a kind of FLAN-BERT, if you want, where it has built-in multi-task ability, but doesn't generate text. It only generates 255 floats all at once, making it much faster, and what those floats mean (if anything) depends on the prompt.

          Eg, the following query is put in the encoder model:

          {"question": "Rank these 5 things by increasing order of how big they are", "choices": ["truck", "cow", "mouse", "ant", "building"] }

          The model returns [3., 2., 1., 0., 4.], and 249 other meaningless floats that are hidden from you by the UI.

          The UI stitches the first 5 floats with the choices and returns something like:

          {"rank": ["ant", "mouse", "cow", "truck", "tower"]}

          1. dannyw · · focus · HN ↗
            So it generates logits in a 255 token output space? ;)
            1. make3 · · focus · HN ↗
              logits assumed some form of softmax or logistic, which may not be the case
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.