‹ BackHN Continuity

Thread

Jev in 25 Lines of Python

691 points · 212 comments · bashbjorn

  1. armcat · · focus · HN ↗
    Looking at the logprobs on tokens works for the local models, but not on the frontier ones. It&#x27;s been more or less broken since GPT-4o for example. I wrote about it two years ago: <a href="https:&#x2F;&#x2F;medium.com&#x2F;data-science&#x2F;9-11-or-9-9-which-one-is-higher-6efbdbd6a025" rel="nofollow">https:&#x2F;&#x2F;medium.com&#x2F;data-science&#x2F;9-11-or-9-9-which-one-is-hig.... Also, I&#x27;ve done some work in estimating confidence and on rubric evals using the same method, and you actually get better correlation to &quot;real confidence&quot; by just getting the LLM to say it.
    1. hununu · · focus · HN ↗
      Interesting. Have you repeated these experiments with recent models? I&#x27;m thinking frontier models APIs have tools&#x2F;MCPs for math stuff but curious about recent Qwen models, etc.
      1. armcat · · focus · HN ↗
        Haven&#x27;t tried recent frontier ones, Astra for example doesn&#x27;t support logprobs emission on the API, and Sol and Luna supposedly support it with reasoning disabled. Haven&#x27;t tried local models like Qwen 3.8 27b (I&#x27;m actually exploring their thinking trace, it&#x27;s a lot of fun)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.