‹ BackHN Continuity

Thread

Jev in 25 Lines of Python

691 points · 212 comments · bashbjorn

  1. fzysingularity · · focus · HN ↗
    Am I missing something here:

    p(y = next thinking+decision token | x = question) != p(y = next decision token | x = question)

    The former is what LLMs are trained for, the latter is what Jev was likely trained on (likely used thinking alignment as an auxiliary loss, but not explicitly included in the probability calibration).

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.