"Hardware that does arithmetic is cheap, so any CPU made this century has plenty of it. But you still only have one instruction decoding block and it is hard to get it to go fast, so the arithmetic hardware is vastly underutilized.
Warning, Nitpick. Saying "hardware [...] is cheap" and "instruction decoding is expensive" (implied) is a minor contradiction. You can actually duplicate instruction decoding just fine, it's the thing before it where things go to hell: Instruction fetching.
Nothing prevents you from building a computer that can fetch, decode and execute 16 instructions at once, assuming they don't all write to the same register.
But if you want to do the same add repeated 16 times you'll need 16 times more program memory and 16 wider read ports on your caches and so on. SRAM is really expensive so this strategy will waste a lot of area on memory that you probably didn't need in the first place. I say this as someone who had to design a chip in university and basically you couldn't even find the primitive CPU in-between the massive SRAM blocks. By reusing the same instruction you can now increase your compute to memory ratio in terms of area.
Just a heads up for people who want to know why SIMD is a thing. I'm not criticizing the article, I just want more people to realize the pain that SRAM represents to chip designers.
imtringued · · focus · HN ↗
Warning, Nitpick. Saying "hardware [...] is cheap" and "instruction decoding is expensive" (implied) is a minor contradiction. You can actually duplicate instruction decoding just fine, it's the thing before it where things go to hell: Instruction fetching.
Nothing prevents you from building a computer that can fetch, decode and execute 16 instructions at once, assuming they don't all write to the same register.
But if you want to do the same add repeated 16 times you'll need 16 times more program memory and 16 wider read ports on your caches and so on. SRAM is really expensive so this strategy will waste a lot of area on memory that you probably didn't need in the first place. I say this as someone who had to design a chip in university and basically you couldn't even find the primitive CPU in-between the massive SRAM blocks. By reusing the same instruction you can now increase your compute to memory ratio in terms of area.
Just a heads up for people who want to know why SIMD is a thing. I'm not criticizing the article, I just want more people to realize the pain that SRAM represents to chip designers.