‹ BackHN Continuity

Thread

Nvidia announces native GPU programming in Rust

970 points · 404 comments · nonmaskable

  1. jacobgorm · · focus · HN ↗
    I strongly dislike CUDA. Once you have allowed that proprietary cr*p into your C++ codebase, it is very hard to get rid, and you end up with code that is either tied to a single vendor or an #ifdef hell, probably both.

    The best way to program GPUs is face up to the reality that they are not the same machine as the CPU, write your kernels in separate files, and launch them manually, like in Metal, OpenCL, and D3D12, etc. These days we even have DSLs like Triton that make kernel writing much more ergonomic than anything you would hope to achieve in Rust.

    1. jacobgorm · · focus · HN ↗
      As it happens, I just got my employer&#x27;s permission to release as open source a Triton back-end for Metal and D3D12 GPUs here: <a href="https:&#x2F;&#x2F;github.com&#x2F;dropbox&#x2F;neso" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;dropbox&#x2F;neso .

      As an example of you how can use it to deploy real models there is this project doing ASR and TTS: <a href="https:&#x2F;&#x2F;github.com&#x2F;dropbox&#x2F;nspeech" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;dropbox&#x2F;nspeech .

      Finally, I am also going to be switching the inferencing part of Witchcraft from current Candle on MacOS and OpenVINO on Windows to just Candle with Neso; <a href="https:&#x2F;&#x2F;github.com&#x2F;dropbox&#x2F;witchcraft" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;dropbox&#x2F;witchcraft

      1. bbkane · · focus · HN ↗
        That sounds like it&#x27;ll be easier to maintain. Will it also be faster?
        1. jacobgorm · · focus · HN ↗
          It is currently faster than the stock Candle &#x2F; MPS shaders it replaces on MacOS&#x2F;ARM64, and IIRC a bit slower than OpenVINO&#x2F;CPU on my old Windows laptop, where I never got OpenVINO&#x2F;GPU to compute correctly. Candle didn&#x27;t have support for GPUs on MacOS&#x2F;Intel, and OpenVINO ceased to be supported there.

          Compared to OpenVINO (I tried ONNX runtime too, but never got it produce correct outputs with my quantized models) it is very nice to be able to build the exact kernels I need, at the quantization settings and precision that works for the models I have and with the custom operators required (speech models do a lot of non-standard stuff), run from a single set of sources, and not have to ship a hefty third-party DLL, and having to deal with their memory leaks and other stability issues.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.