The work by Valve's Timur Kristóf on improving old AMD GPUs on Linux
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
The work by Valve's Timur Kristóf on improving old AMD GPUs on Linux
Unofficial Hacker News client; not affiliated with Y Combinator.
segmondy · · focus · HN ↗
saghm · · focus · HN ↗
It would be awesome if someone manages to figure out how to get small enough models to fit on older cards to be viable, but I'm not optimistic that it will come without some sort of fundamental architectural innovation rather than incremental improvements, and it's not clear if and when that will happen.
BLKNSLVR · · focus · HN ↗
How long is this runway?
saghm · · focus · HN ↗
resistings-gend · · focus · HN ↗
Also you can use multiple GPUs at once, two of your GPUs could run qwen 27b very comfortably, and with great performance.
saghm · · focus · HN ↗
hparadiz · · focus · HN ↗
With AMD the best you can do in the consumer market right now is an RX 7900 XTX which is about 960 GB/s.
Tade0 · · focus · HN ↗
The other day I managed to get a context of 195k for Qwen3.5-9b Q4_K_M using a llama.cpp fork that supports TurboQuant:
<a href="https://github.com/TheTom/llama-cpp-turboquant" rel="nofollow">https://github.com/TheTom/llama-cpp-turboquant
I think you could replicate this with a larger model on your device.
Overall with the right quantisations for both the model and KV cache you can get a lot of mileage out of this old hardware. Speed remains the main limitation, as IIRC I was getting ~26-30tps on a 7700S.
saghm · · focus · HN ↗
I'm not saying there's no benefit to using local models. My point is still the same as before: you have to be willing to sacrifice both performance and quality even when just comparing to free models that are available today.
Tade0 · · focus · HN ↗
I spent some years in the pharma industry, where AI came really late, as there was (justified) concern that sensitive data might leak - even by accident. Eventually a solution was implemented - it was some kind of open model (they didn't say which) running on-premises.
You couldn't tell these people that this and that powerful model is free or inexpensive, because it's useless to them if it's running on someone else's computer i.e. the cloud.
saghm · · focus · HN ↗
Either way, my initial reply was responding to the idea of these 15 year old GPUs as potential "capable processing units" for LLMS. My argument is that they are not particularly capable at the current moment in time.
marknsikora · · focus · HN ↗
lp92 · · focus · HN ↗