Yes. there is real hardware architecture behind Apple’s “Unified Memory”, but the underlying idea is not uniquely Apple. What Apple did was design the CPU, GPU, memory controller, cache hierarchy, interconnect, package, OS, and graphics APIs together around the architecture. I mean, they can do shit like that because they own the entire product pipeline. They can fine tune the hardware in ways other OEMs can't.
In a Intel+Nvidia CPU+GPU pipeline you have the CPU, that load something in ram, then for the GPU to process it you need to move it from RAM to VRAM through a slow PCI connect then when the GPU is done you have to move it again to RAM for the CPU to handle it again. These copies cost bandwidth, latency, energy, extra memory and low level programing complexity. In the Apple unified architecture both de CPU and GPU use the same memory and avoid copying data.
You may say "ah but don't you can handle shared memory with just the DDR controller?" and that's what Intel integrated graphics do but on top of that Apple: build the CPU and GPU in the same SOC, gives the SOC a massive memory subsystem and a large system level cache.
Nvidia can do the same. Eg the Grace Hopper has the Grace CPU - nvlink C2C - Hopper/Blackwell GPU. But the trade-off is modularity. You can mix processors, ram, Nvidia GPU and all parts must work at their best capacity, but isn't even close to the fine tuning of an Apple system.
vivzkestrel · · focus · HN ↗
- how exactly does apple silicon s unified gpu + ram thingy work?
- how come intel and nvidia cannot do the same?
- is there an actual difference in terms of hardware architecture or something or is it pure apple marketing hype?
aeve890 · · focus · HN ↗
In a Intel+Nvidia CPU+GPU pipeline you have the CPU, that load something in ram, then for the GPU to process it you need to move it from RAM to VRAM through a slow PCI connect then when the GPU is done you have to move it again to RAM for the CPU to handle it again. These copies cost bandwidth, latency, energy, extra memory and low level programing complexity. In the Apple unified architecture both de CPU and GPU use the same memory and avoid copying data. You may say "ah but don't you can handle shared memory with just the DDR controller?" and that's what Intel integrated graphics do but on top of that Apple: build the CPU and GPU in the same SOC, gives the SOC a massive memory subsystem and a large system level cache.
Nvidia can do the same. Eg the Grace Hopper has the Grace CPU - nvlink C2C - Hopper/Blackwell GPU. But the trade-off is modularity. You can mix processors, ram, Nvidia GPU and all parts must work at their best capacity, but isn't even close to the fine tuning of an Apple system.
ranguna · · focus · HN ↗
Why not make a faster interconnect with an open communication standard?