‹ BackHN Continuity

Thread

PSSA: A non-transformer language model written from scratch in Rust

89 points · 38 comments · sparticle62

  1. JPLeRouzic · · focus · HN ↗
    I have read somewhere that Transformer architecture has a quadratic cost [0] (which explains the high costs associated with LLMs and the difficulty for constant improvement without state size pockets).

    For what I understand PSSA belongs to a line of research for LLMs with scalable architecture because you don't need to load the full KV in memory to generate a single token:

    [0] <a href="https:&#x2F;&#x2F;aclanthology.org&#x2F;2023.findings-emnlp.936&#x2F;" rel="nofollow">https:&#x2F;&#x2F;aclanthology.org&#x2F;2023.findings-emnlp.936&#x2F;

    <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2503.00392" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2503.00392

    <a href="https:&#x2F;&#x2F;papers.nips.cc&#x2F;paper_files&#x2F;paper&#x2F;2023&#x2F;hash&#x2F;6ceefa7b15572587b78ecfcebb2827f8-Abstract-Conference.html" rel="nofollow">https:&#x2F;&#x2F;papers.nips.cc&#x2F;paper_files&#x2F;paper&#x2F;2023&#x2F;hash&#x2F;6ceefa7b1...

    1. janalsncm · · focus · HN ↗
      Yes, and it’s not even research anymore! Qwen uses linear attention: <a href="https:&#x2F;&#x2F;sebastianraschka.com&#x2F;llms-from-scratch&#x2F;ch04&#x2F;08_deltanet&#x2F;" rel="nofollow">https:&#x2F;&#x2F;sebastianraschka.com&#x2F;llms-from-scratch&#x2F;ch04&#x2F;08_delta...
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.