AMD has acquired "AI stuff" for several billions at this point, yet somehow, doing "AI stuff" on their GPU/NPU stack still seems to be troublesome. Well, that is what I read from others anyway. Inference with llama.cpp on AMD GPUs pretty much just works.
I bet soon enough we can just ask AI to port stuff from CUDA to whatever language AMD is using. This is how nVidia will become a victim of their own success.
Keep going... extrapolate out further. There's a conclusion you beley will be true, but haven't fully thought through why that conclusion must be true if CUDA is no longer the moat it once was.
Who says that? I’m having a blast with rocm powering my R9700. Been able to run all models on day 0 at comparable speeds to Nvidia with similar memory bandwidth. Software is no longer the main bottleneck for AMD.
just use torch/vllm/sglang/llama.cpp and the rest of the ecosystem, most have decent support for AMD, it's not that different from CUDA once you get it working (and it is much easier now than it used to)
Really, it's just the Python frameworks on AMD consumer GPUs that suck. AMD datacenter GPUs seem to work with them well enough, and llama.cpp crushes it for the rest of us.
LarsDu88 · · focus · HN ↗
I was also shocked by how quickly AMD acquired Talaas. AMD may be preparing for the next way (ultra fast inference, and embodied AI inference)
ahartmetz · · focus · HN ↗
amelius · · focus · HN ↗
wmf · · focus · HN ↗
fragmede · · focus · HN ↗
nwah1 · · focus · HN ↗
fragmede · · focus · HN ↗
tybit · · focus · HN ↗
fragmede · · focus · HN ↗
homosapien97 · · focus · HN ↗
vlovich123 · · focus · HN ↗
alightsoul · · focus · HN ↗
sheepscreek · · focus · HN ↗
__rito__ · · focus · HN ↗
I don't know anything.
Things like training a Vision/RL/LM, downloading and running LMs, etc.
imjonse · · focus · HN ↗
whateverboat · · focus · HN ↗
ahartmetz · · focus · HN ↗
MrDrMcCoy · · focus · HN ↗