Okay, so, GPUs are taking one more step towards being general purpose massively parallel machines. That's cool.
What would be even cooler though would be for GPU vendors to start giving us the user manual. An I mean the real user manual, that explains how to use their piece of metal when all you have is that piece of metal. That means a precise description of the wire protocols, the data format of the buffers we send to & get from the GPU, the ISA of the cores we have access to, the relevant performance characteristics…
In other words, enough information to write a state-of-the-art driver for any OS. That would be cool.
They don't release it because exposing a stable instruction set would kill their ability to quickly iterate, to release silicon with bugs that can be papered over with software fixes, as fixing bugs in chips is very expensive in terms of time to market, and undoubtedly to charge more for what looks like a hardware feature but actually is a software feature.
It's been this way for 25 years and I don't see it changing.
I've been programming GPUs since the PlayStation1. The way they work under the hood has changed fundamentally maybe 4 times in that span.
I can't compare it to changes I've seen in CPU architecture since then. Maybe like: Compare the NES with its 6502 and per-cartridge mappers vs. a IBM 386 PC. Now repeat that shift 2 or 3 more times.
> The way they work under the hood has changed fundamentally maybe 4 times in that span.
How many times since we got shaders? And when exactly? I bet each change takes longer to come than the last. I mean, GPUs are increasingly general purpose nowadays. Sure they're optimised to specific kinds of embarrassingly parallel problems, but as more and more of their capabilities move out of the fixed pipeline to shaders, the need to change lessens.
When I started what you did was nothing but set up DMA streams to set registers. DMA sets registers, hardware reacts by rasterizing triangles.
In the GeForce3 era, the registers got complicated enough they resembled tiny "pixel shaders", but under the hood it was still a small struct held in registers. Vertex shaders were 1 to 128 asm instructions executed strictly linearly.
In the G80 era we got "general purpose shaders." But, they still depend heavily on the fix function pipeline for their dispatch/scheduling and I/O.
These days, everything is basically a dressed-up compute shader. The shared-memory SRAM is front-and-center in your attention. Dispatch and scheduling are manual and complicated. GPUs are transitioning into tensor evaluators.
So, maybe today we can start considering talking about planning committee meetings about stability. But, what I've observed is that this has been a request for a few decades now. And, in hindsight it would not have worked out in the past. Moving forward, maybe it would work out OK today for a while. But, I don't see the rate of change in GPUs slowing down any time soon. Wouldn't be surprised if we're racing towards some Cerebras + Tensor Cores + FPGA near future.
loup-vaillant · · focus · HN ↗
What would be even cooler though would be for GPU vendors to start giving us the user manual. An I mean the real user manual, that explains how to use their piece of metal when all you have is that piece of metal. That means a precise description of the wire protocols, the data format of the buffers we send to & get from the GPU, the ISA of the cores we have access to, the relevant performance characteristics…
In other words, enough information to write a state-of-the-art driver for any OS. That would be cool.
floil · · focus · HN ↗
It's been this way for 25 years and I don't see it changing.
VikingCoder · · focus · HN ↗
But hi, if I spent $10,000 on a piece of hardware, let me program the metal, thanks.
corysama · · focus · HN ↗
I can't compare it to changes I've seen in CPU architecture since then. Maybe like: Compare the NES with its 6502 and per-cartridge mappers vs. a IBM 386 PC. Now repeat that shift 2 or 3 more times.
loup-vaillant · · focus · HN ↗
How many times since we got shaders? And when exactly? I bet each change takes longer to come than the last. I mean, GPUs are increasingly general purpose nowadays. Sure they're optimised to specific kinds of embarrassingly parallel problems, but as more and more of their capabilities move out of the fixed pipeline to shaders, the need to change lessens.
corysama · · focus · HN ↗
In the GeForce3 era, the registers got complicated enough they resembled tiny "pixel shaders", but under the hood it was still a small struct held in registers. Vertex shaders were 1 to 128 asm instructions executed strictly linearly.
In the G80 era we got "general purpose shaders." But, they still depend heavily on the fix function pipeline for their dispatch/scheduling and I/O.
These days, everything is basically a dressed-up compute shader. The shared-memory SRAM is front-and-center in your attention. Dispatch and scheduling are manual and complicated. GPUs are transitioning into tensor evaluators.
So, maybe today we can start considering talking about planning committee meetings about stability. But, what I've observed is that this has been a request for a few decades now. And, in hindsight it would not have worked out in the past. Moving forward, maybe it would work out OK today for a while. But, I don't see the rate of change in GPUs slowing down any time soon. Wouldn't be surprised if we're racing towards some Cerebras + Tensor Cores + FPGA near future.