Fujitsu launches made-in-Japan next-generation CPU FUJITSU-MONAKA
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Fujitsu launches made-in-Japan next-generation CPU FUJITSU-MONAKA
Unofficial Hacker News client; not affiliated with Y Combinator.
drob518 · · focus · HN ↗
swiftcoder · · focus · HN ↗
There is a fairly direct link between the two numbers. You can predict the latter from former reasonably well
michaelbuckbee · · focus · HN ↗
julius · · focus · HN ↗
hermitShell · · focus · HN ↗
You have to run inference on the GPU by reading and writing to VRAM. So TFLOPS of the compute matters, and bandwidth to the VRAM (Always integrated with the GPU, rarely a bottleneck), and this strongly affects tokens/s
If you're doing training workloads or offloading to system RAM, it gets more complicated. (And mostly bound up trying to feed compute on time)
(Edits for clarity.)
usrnm · · focus · HN ↗
On dedicated inference hardware I'd expect model weights to never leave the RAM, and you'd probably load them on startup before even starting to serve requests
chorylee · · focus · HN ↗
[dead]