Trying to be a little less negative than the guy that got flagged, I too don't understand the motivation behind these projects. I've seen a hundred of them at this point- and each of them is probably worse and has less learning value than the one Andrej Karpathy made to teach people the building blocks involved in a GPT
So, why do people keep making these? Asking genuinely
> So, why do people keep making these? Asking genuinely
I extensively used minGPT for home experiments on transformer architecture. It is great for learning!
However, if you want to scale the experiments up at home you need to go faster. Karpathy made optimized <a href="https://github.com/karpathy/nanoGPT" rel="nofollow">https://github.com/karpathy/nanoGPT, but it is tuned for "8XA100 40GB node in about 4 days of training".
13s is a bit overkill here (my machine builds that project in 30s). But it gives some space for experimentation with architectures that don't have optimized primitives.
_345 · · focus · HN ↗
So, why do people keep making these? Asking genuinely
lostmsu · · focus · HN ↗
I extensively used minGPT for home experiments on transformer architecture. It is great for learning!
However, if you want to scale the experiments up at home you need to go faster. Karpathy made optimized <a href="https://github.com/karpathy/nanoGPT" rel="nofollow">https://github.com/karpathy/nanoGPT, but it is tuned for "8XA100 40GB node in about 4 days of training".
13s is a bit overkill here (my machine builds that project in 30s). But it gives some space for experimentation with architectures that don't have optimized primitives.