For a moment I thought this was going to be a metaphor — maybe an ancient Chinese proverb about how paint brushes are made from horsehair and how you can't hold the horse to paint before you've turned the hair into a brush.
One way: separately embed sender, recipients, subject, body - then use the embedding vectors as input to a logistic classifier
With that setup, I get 95% accuracy on email classification, training on 50-100 base examples. The model trains on CPU in under 1min, and it does inference in under 20ms (most of it is running the embeddings, so you can make it faster if you train your own embeddings model)
Here’s a gist with some sample code: <a href="https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecdad5029557" rel="nofollow">https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...
That code applies the embeddings + classifier setup on the Banking77 dataset. It gets 93-94% accuracy depending on the embeddings you use (SOTA for this is ~95%, with much bigger and slower models)
I wonder what numbers you'd get using another system one model - Contrastive Language Model <a href="https://contrastive-lm.notion.site/" rel="nofollow">https://contrastive-lm.notion.site/
That model scales very well with quantities of requests.
I can’t see your gist but spam classification is a textbook example of something you shouldn’t measure with accuracy. If 95% of your samples are not spam you can get 95% accuracy by always guessing not spam.
You should use precision (when your model says “spam” how often is it spam?), recall (how many of the spam emails did it catch), or f1 (balanced between those two).
That's a great point. My case is not for spam, the classes are more balanced, but you are correct that precision, recall and f1 would be better measures for some of these tasks
There are two parts in the data you supply to Jev for classification - the prompt describing your classification and the data. The data can be quite small - a simple chat message. And prompt part could be considerable since you need to describe your rubrics well.
With Jev you each time pay for your prompt, you can't cache it.
I mean, it sounds like it's only ideal for cases with significant system prompt overhead. I don't think Jev was built to have a large well described prompt setup. To me its more like a happy go lucky small label classification tool with important decisions left to stronger agentic models or yk humans.
sharih · · focus · HN ↗
zihotki · · focus · HN ↗
olgava · · focus · HN ↗
[dead]
atombender · · focus · HN ↗
For a moment I thought this was going to be a metaphor — maybe an ancient Chinese proverb about how paint brushes are made from horsehair and how you can't hold the horse to paint before you've turned the hair into a brush.
idiotsecant · · focus · HN ↗
tyre · · focus · HN ↗
nico · · focus · HN ↗
One way: separately embed sender, recipients, subject, body - then use the embedding vectors as input to a logistic classifier
With that setup, I get 95% accuracy on email classification, training on 50-100 base examples. The model trains on CPU in under 1min, and it does inference in under 20ms (most of it is running the embeddings, so you can make it faster if you train your own embeddings model)
Here’s a gist with some sample code: <a href="https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecdad5029557" rel="nofollow">https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...
That code applies the embeddings + classifier setup on the Banking77 dataset. It gets 93-94% accuracy depending on the embeddings you use (SOTA for this is ~95%, with much bigger and slower models)
zihotki · · focus · HN ↗
That model scales very well with quantities of requests.
janalsncm · · focus · HN ↗
You should use precision (when your model says “spam” how often is it spam?), recall (how many of the spam emails did it catch), or f1 (balanced between those two).
nico · · focus · HN ↗
calebhwin · · focus · HN ↗
zihotki · · focus · HN ↗
With Jev you each time pay for your prompt, you can't cache it.
sarkarghya · · focus · HN ↗
jedberg · · focus · HN ↗
StarlaAtNight · · focus · HN ↗
HawtAds · · focus · HN ↗
simplisticelk · · focus · HN ↗
catlifeonmars · · focus · HN ↗