‹ BackHN Continuity

Thread

Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data

158 points · 43 comments · Betelbuddy

  1. cyanydeez · · focus · HN ↗
    >The scaling laws hold that a language model grows more capable with more parameters and more training data.

    Which is a choice, not a "law":

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2510.13786" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2510.13786

    <a href="https:&#x2F;&#x2F;www.alphaxiv.org&#x2F;abs&#x2F;2512.20264" rel="nofollow">https:&#x2F;&#x2F;www.alphaxiv.org&#x2F;abs&#x2F;2512.20264

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2607.05155" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2607.05155

    1. largbae · · focus · HN ↗
      I think this is partially true: scaling parameter size will always go asymptotic to 100% accuracy because 100% is the ceiling of that metric.

      However 95% is still half the error rate of 90%, and 97.5% is half the error rate of that.

      And when test time compute like reasoning and looping harnesses stack many inference acts with many tokens each, those seemingly small accuracy gains stack tremendously.

      1. cyanydeez · · focus · HN ↗
        Error rates will never go to zero and as context grows ambiguities grow in reverse. So your arguement is great to some token...n but after that, it all unwinds.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.