‹ BackHN Continuity

Thread

Contrastive Language Models

176 points · 59 comments · erichocean

  1. amluto · · focus · HN ↗
    I’m fascinated by this thing and by the way it’s interpreting Jev. It’s very cool, but is it actually a classifier?

    IIUC they took an already-trained “frozen” LLM and trained a little model on top that takes both a question and the hidden states after processing the input data and produces answer “probabilities”. (In contrast, the original LLM would have been run in AR mode to generate multiple output tokens representing its answer.) But then they used it for a purpose that isn’t really classification.

    IMO there is a rather large difference between “is this email spam” and “what character should I type in this agentic workload”. The former is classification: there is hopefully a ground truth (is the email spam?) and the model is trying to classify the email. You would score it with a proper scoring rule. The latter is a strategy: there usually isn’t a correct answer, now or in the future. The model is playing a game consisting of repeated rounds, and the only way to evaluate it is to see how well it plays. You can’t even usefully compare it to the optimal solution because you may not know the optimal solution and you don’t actually need the model to produce an optimal solution.

    I do think this approach is really cool, and it does suggest that one might be able to use a modern LLM to process an input and then extract the model’s next agentic step in a very fast, non-AR manner, with results comparably good to the usual AR decoding. And I think it’s very interesting to decouple the tokenized input representation from the model output representation, both because prefill tends to be faster and cheaper than AR output and because it’s never seemed particularly sensible to me that a model should be constrained to generate outputs at the cadence of one run through the model per output token. (AFAIK the main reason that models work on the same input and output token space is that this is how the pretraining process works.)

    I wonder how to fit “reasoning” into this framework. Maybe have the question be something like “do you need to think further and, if so, what is your first thinking token”. But maybe something more clever is possible.

    1. visarga · · focus · HN ↗
      > The latter is a strategy: there usually isn’t a correct answer, now or in the future. The model is playing a game consisting of repeated rounds, and the only way to evaluate it is to see how well it plays

      This is self contradictory. You can tell a correct answer as you said, by looking at the score. In my $DAYJOB I am making hundreds of RL environments that produce a score for each intermediate state.

      1. amluto · · focus · HN ↗
        Sure, you are scoring them. But I doubt that you are scoring them in the sense that you are minimizing loss where the loss is a proper scoring function that treats the model as a classifier where that treatment as a classifier actually means something. Unless you happen to be RLing a problem where there is either exactly one correct rollout or there is actually a bona fide natural distribution over rollouts independently from the model.
        1. visarga · · focus · HN ↗
          My environments don't score actions, they score state, so no matter how the agent chooses to solve a task it gets scored correctly. It's called Potential Based Reward Shaping and its main benefit is that it does not introduce reward hacking.
          1. paraschopra · · focus · HN ↗
            Super intriguing. Do you have more details on this? Example?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.