‹ BackHN Continuity

Thread

Fujitsu launches made-in-Japan next-generation CPU FUJITSU-MONAKA

633 points · 254 comments · my123

  1. drob518 · · focus · HN ↗
    One problem we’re going to have with AI hardware is coming up with a standard set of specifications that are comparable. I don’t really care about the CPU GHz and the memory bandwidth, at least not directly. What I really want to know is how many tokens per second this will deliver, but that also depends on the model. We need a standard metric for that. Perhaps we agree on a specific open weight model (e.g. GLM 5.3 Flash or Qwen vWhatever) and then measure TPS on the hardware of interest.
    1. swiftcoder · · focus · HN ↗
      > I don’t really care about the CPU GHz and the memory bandwidth ... What I really want to know is how many tokens per second this will deliver

      There is a fairly direct link between the two numbers. You can predict the latter from former reasonably well

    2. michaelbuckbee · · focus · HN ↗
      I feel like even TPS is becoming less of a good metric as we're seeing certain models handle similar problems while burning far fewer tokens.
    3. julius · · focus · HN ↗
      I always heard of GB/s as most important number ... AI told me their 8800 MT/s on 8 Byte, 12 DDR5-Channels means 845 GB/s. A Nvidia RTX 4090 has 1008 GB/s. Nvidia B200 has 8000 GB/s. Is this the right way of looking at it?
      1. hermitShell · · focus · HN ↗
        You have to load the model weights into VRAM over PCI-E (from RAM). So the (PCI-E) bandwidth strongly affects time to first token.

        You have to run inference on the GPU by reading and writing to VRAM. So TFLOPS of the compute matters, and bandwidth to the VRAM (Always integrated with the GPU, rarely a bottleneck), and this strongly affects tokens/s

        If you're doing training workloads or offloading to system RAM, it gets more complicated. (And mostly bound up trying to feed compute on time)

        (Edits for clarity.)

        1. usrnm · · focus · HN ↗
          > So the (PCI-E) bandwidth strongly affects time to first token

          On dedicated inference hardware I'd expect model weights to never leave the RAM, and you'd probably load them on startup before even starting to serve requests

    4. chorylee · · focus · HN ↗

      [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.