‹ BackHN Continuity

Thread

Jev Can't Be Calibrated

65 points · 61 comments · alexmolas

  1. abhgh · · focus · HN ↗
    I like this post. I haven't had time to dig into Jev (they aren't accepting new signups), but calibrated probabilities is one of their pitches that caught my attention. And I was wondering how does one offer them on user data. Standard calibration essentially ensures that if a score of 0.8 accompanies a positive prediction (assuming the simple case of binary classification), then if you gathered together all predictions with a score of 0.8, around 80% will be correct.

    If you have just one example you're sending to a model, how would they guarantee 80% over your data?

    FYI, for an overview, scikit's page on calibration is great [1], and my answer on Quora from a long time ago covers a specific type [2].

    [1] <a href="https:&#x2F;&#x2F;scikit-learn.org&#x2F;stable&#x2F;modules&#x2F;calibration.html" rel="nofollow">https:&#x2F;&#x2F;scikit-learn.org&#x2F;stable&#x2F;modules&#x2F;calibration.html

    [2] <a href="https:&#x2F;&#x2F;www.quora.com&#x2F;How-is-isotonic-regression-used-in-practice-for-calibration-in-machine-learning&#x2F;answer&#x2F;Abhishek-Ghose" rel="nofollow">https:&#x2F;&#x2F;www.quora.com&#x2F;How-is-isotonic-regression-used-in-pra...

    1. danielmarkbruce · · focus · HN ↗
      While I don&#x27;t believe they are doing the following: you can calibrate by inspecting the reasoning traces. That is the relevant distribution. If you ask someone to explain how&#x2F;why they are classifying something one way v another, you can get a reasonably good understanding of their confidence level.
      1. abhgh · · focus · HN ↗
        This tells me the confidence of the LLM&#x27;s belief about the response - which is different from the calibrated confidence score. The former also is useful (just not what I thought their advertisement sells - and from the article it seems like it tripped up others as well), and there are different techniques to extract such a value [1] [2], typically via &quot;response sampling&quot;, i.e., interrogate the LLM slightly differently to see if it changes its answer.

        [1] Semantic Entropy <a href="https:&#x2F;&#x2F;www.nature.com&#x2F;articles&#x2F;s41586-024-07421-0" rel="nofollow">https:&#x2F;&#x2F;www.nature.com&#x2F;articles&#x2F;s41586-024-07421-0

        [2] Kernel Language Entropy <a href="https:&#x2F;&#x2F;openreview.net&#x2F;pdf?id=j2wCrWmgMX" rel="nofollow">https:&#x2F;&#x2F;openreview.net&#x2F;pdf?id=j2wCrWmgMX

        1. danielmarkbruce · · focus · HN ↗
          I mean the model can learn from it during RL training. The confidence score is affected by the tokens prior to it it&#x27;s output. I was using the word &quot;you&quot; loosely.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.