‹ BackHN Continuity

Thread

PSSA: A non-transformer language model written from scratch in Rust

89 points · 38 comments · sparticle62

  1. janalsncm · · focus · HN ↗
    OP, you should not have written this in Rust. It should be in PyTorch, which is by far the most popular. We can’t tell if this architecture is good or whether there is a problem in your implementation.

    You can test the whole thing for free on a GPU with Google Colab. Test both the transformer and your new architecture on a larger dataset. Something that maxes out the GPU for an hour each run.

    Also, the readme mentions keeping the same optimizer schedule which sounds nice at first but they are completely different architectures. The loss is high on the transformer, did you try raising the learning rate on it?

    In general I’m interested in parameter efficient architectures. I don’t think transformers are optimal, and indeed many improvements have been made to vanilla transformers. But if you have an idea for something better you need to show it.

    1. yjftsjthsd-h · · focus · HN ↗
      I dunno, I could probably be convinced to try a new tool purely on the basis of not having to deal with installing pytorch
      1. intoXbox · · focus · HN ↗
        I’m curious, what’s the criticism for PyTorch?
        1. pseudocomposer · · focus · HN ↗
          The entire Python ecosystem is horrible and shouldn’t have been as falsely boosted by institutions as it was in the 2010s. Yes, it got less bad with 3.8 or whatever version added type annotations. But making so much of ML depend on Python has made it distasteful to a lot of devs who would otherwise have contributed more to it.

          We really need to move all AI/ML research off PyTorch to Candle or… just anything that isn’t Python or another old-gen, broken language like it.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.