LLM Classification Is Feature Engineering
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
LLM Classification Is Feature Engineering
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
ltbarcly3 · · focus · HN ↗
It's very often (always?) the case that something general also solves particular problems.
It's true that LLM output can be used as an input to another classifier, this is also true of any classifier. The improvement on top of the straight LLM classification is relatively small, and I would argue that working on the prompt or just including in the prompt for the LLM what features might be useful to consider would likely work even better.Fundamentally I read this article as: We want to build a simpler, dumbed down clone of Mathematica, so we cobbled together the following pieces... We also needed a way to do arithmetic, so we also include a copy of Mathematica to do basic arithmetic.
michi883 · · focus · HN ↗
[dead]
Terr_ · · focus · HN ↗
ltbarcly3 · · focus · HN ↗
softwaredoug · · focus · HN ↗
<a href="https://softwaredoug.com/blog/2025/01/21/llm-judge-decision-tree" rel="nofollow">https://softwaredoug.com/blog/2025/01/21/llm-judge-decision-...
chrisweekly · · focus · HN ↗
FWIW I also liked this (very different) post of yours: <a href="https://softwaredoug.com/blog/2026/07/13/who-want-to-be" rel="nofollow">https://softwaredoug.com/blog/2026/07/13/who-want-to-be
softwaredoug · · focus · HN ↗
xerlait · · focus · HN ↗
twelfthnight · · focus · HN ↗
dist-epoch · · focus · HN ↗
The amount of thinking is relatively calibrated. Ask an obvious classification, you get an instant answer. Ask a tricky one, much more thinking.
drabbiticus · · focus · HN ↗
Maybe these are well understood terms in some field? Maybe I'm just lost?
probably_wrong · · focus · HN ↗
The point about beta (which I think the equation doesn't actually reflect) is just to indicate that this approach risks nothing because, worst case scenario, the weights you assign to the model can simply return the original LLM prediction. Keep in mind that 0*inf=0 and that LLM(x) only returns 0 or 1, so it's not accurate to say that the sigmoid would always return 1.
The function I() is the indicator function [1] which, in this case, returns 1 if the argument is true and 0 otherwise. It's only there to convert booleans to integers because summing booleans is not defined.
[1] <a href="https://en.wikipedia.org/wiki/Indicator_function" rel="nofollow">https://en.wikipedia.org/wiki/Indicator_function
drabbiticus · · focus · HN ↗
Having said that, if
then it seems that p(y=1|x) === LLM(x) already.Why bother with p(y=1|x) = σ( α + β*LLM(x) ) as in the article? You are right that I missed the case where LLM(x) = 0, but that just means that p(y=1|x) has exactly 2 values when defined as above.
This is substantially lossy and converts a continuous output LLM(x) of range [0,1] to a step function not even defined as a set {0,1} but instead the set {1/(1+e^-α),1} for unclear gain. It might make more sense if LLM(x) is not limited to [0,1] like you have claimed because then at least the logistic regression is clamping the output to [0,1]. If you were wrong about the range of LLM(x), this would come back to asking authors to actually define their terms.My point is that when math is used to justify something or communicate something, it should be explained or very apparently right. If when someone goes to try to understand the math it doesn't match the claims being made in the article ("recovering the LLM") then it throws the rest of the article into doubt.
elendilm · · focus · HN ↗
levocardia · · focus · HN ↗
aleksiy123 · · focus · HN ↗
It’s sort of like memoizing or distilling the knowledge. Works really well for certain type of problems.
vova_hn2 · · focus · HN ↗
1. Give "strong" LLM the task formulation and some labeled examples. Ask it to generate a prompt for the "weak" LLM.
2. Run "weak" LLM on the training set with generated prompt from 1, use replies as features for a smaller ML model (logreg, decision tree etc).
3. Pick examples from the training set that your small model is most wrong about and ask "strong" LLM to generate one more prompt (like in 1), except this time you are using the misclassified examples instead of random.
4. Run "weak" LLM on generated prompt from 3, add results as one more feature for your model.
5. Repeat 2 - 4 until your token budget for this task is exhausted or required precision on cross validation set it reached.
I was thinking about creating an open source library that implements this, but I'm not sure if anyone really needs it. I suspect that people who need something like this already made their own implementation.
danvayn · · focus · HN ↗
For this reason it’s why I feel the current approach where chatgpt mails Claude and refers to it by name is so ineffective for these reasons, IMO. They’re coerced to cooperate and naturally wouldn’t and naturally do not see eye to eye, which will naturally lead to issues within cooperation. It’s all natural.
iforgotmypasswo · · focus · HN ↗
This is a bit of an outdated take as of two days ago. Dear lord things move fast these last few years. Some of this is still relevant. Fine tuning Jev once available could address certain concerns.
(Very excited as I got an invite email for TypeSafe today! I don’t have time for all the little experiments I want to run with Jev and Astra combined!)
danvayn · · focus · HN ↗
Anyways, this article is plenty good. Thanks for the writeup OP.
aidos · · focus · HN ↗
I see a bunch of people saying that there were already similar solutions in this space. Maybe true but it definitely feels like the missing primitive for working with llms. Can immediately see how it can be deployed in real applications in a way that the autoregressive chat approach can’t.
suraj_phanindra · · focus · HN ↗
[dead]
suraj_phanindra · · focus · HN ↗
lhk931122 · · focus · HN ↗
Otterly99 · · focus · HN ↗
Use binary questions rather than multi-class. Then you can use the consistency as another signal of your pipeline working or not, on top of accuracy.
wodenokoto · · focus · HN ↗
MaxQuimby · · focus · HN ↗
[dead]