‹ BackHN Continuity

Thread

Jeeves. Reasoning improves Jev-like decision models

242 points · 95 comments · nicowaltz

  1. sharih · · focus · HN ↗
    What is the point of this, if it is p90 17 seconds? Might as well use an LLM. The beauty of Jev is that it is dirt cheap and insanely fast.
    1. zihotki · · focus · HN ↗
      I would hold your horses to paint it as dirt cheap.. In my cases for spam detection Luna was 20% cheaper due to prompt caching, although not as fast.
      1. olgava · · focus · HN ↗

        [dead]

      2. atombender · · focus · HN ↗
        > hold your horses to paint it as dirt cheap

        For a moment I thought this was going to be a metaphor — maybe an ancient Chinese proverb about how paint brushes are made from horsehair and how you can't hold the horse to paint before you've turned the hair into a brush.

        1. idiotsecant · · focus · HN ↗
          Darmok and Jalad, at Tanagra
      3. tyre · · focus · HN ↗
        What are the costs compared to an ML model?
      4. nico · · focus · HN ↗
        For email you can use a classifier

        One way: separately embed sender, recipients, subject, body - then use the embedding vectors as input to a logistic classifier

        With that setup, I get 95% accuracy on email classification, training on 50-100 base examples. The model trains on CPU in under 1min, and it does inference in under 20ms (most of it is running the embeddings, so you can make it faster if you train your own embeddings model)

        Here’s a gist with some sample code: <a href="https:&#x2F;&#x2F;gist.github.com&#x2F;nicobrenner&#x2F;056a5aaff5d0119c0032ecdad5029557" rel="nofollow">https:&#x2F;&#x2F;gist.github.com&#x2F;nicobrenner&#x2F;056a5aaff5d0119c0032ecda...

        That code applies the embeddings + classifier setup on the Banking77 dataset. It gets 93-94% accuracy depending on the embeddings you use (SOTA for this is ~95%, with much bigger and slower models)

        1. zihotki · · focus · HN ↗
          I wonder what numbers you&#x27;d get using another system one model - Contrastive Language Model <a href="https:&#x2F;&#x2F;contrastive-lm.notion.site&#x2F;" rel="nofollow">https:&#x2F;&#x2F;contrastive-lm.notion.site&#x2F;

          That model scales very well with quantities of requests.

        2. janalsncm · · focus · HN ↗
          I can’t see your gist but spam classification is a textbook example of something you shouldn’t measure with accuracy. If 95% of your samples are not spam you can get 95% accuracy by always guessing not spam.

          You should use precision (when your model says “spam” how often is it spam?), recall (how many of the spam emails did it catch), or f1 (balanced between those two).

          1. nico · · focus · HN ↗
            That&#x27;s a great point. My case is not for spam, the classes are more balanced, but you are correct that precision, recall and f1 would be better measures for some of these tasks
      5. calebhwin · · focus · HN ↗
        How are you benefiting from prompt caching for simple classification?
        1. zihotki · · focus · HN ↗
          There are two parts in the data you supply to Jev for classification - the prompt describing your classification and the data. The data can be quite small - a simple chat message. And prompt part could be considerable since you need to describe your rubrics well.

          With Jev you each time pay for your prompt, you can&#x27;t cache it.

          1. sarkarghya · · focus · HN ↗
            I mean, it sounds like it&#x27;s only ideal for cases with significant system prompt overhead. I don&#x27;t think Jev was built to have a large well described prompt setup. To me its more like a happy go lucky small label classification tool with important decisions left to stronger agentic models or yk humans.
      6. jedberg · · focus · HN ↗
        Are you getting better performance from an LLM than a Bayesian classifier?
        1. StarlaAtNight · · focus · HN ↗
          BLASPHEMY! OUT WITH YOU!
      7. HawtAds · · focus · HN ↗
        How many requests per second do you have for spam that you are reliably hitting the Luna cache?
      8. simplisticelk · · focus · HN ↗
        Is that just because the Jev implementation is less mature? Couldn&#x27;t it also implement prompt caching?
      9. catlifeonmars · · focus · HN ↗
        Could you not just copycat jev and run a fast, small local model?
    2. esafak · · focus · HN ↗
      Jev ought to offer a flex mode that uses their spare capacity for a discount.
    3. amelius · · focus · HN ↗
      Next step: make it classify the next word.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.