‹ BackHN Continuity

Thread

Show HN: Janus – Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia

104 points · 19 comments · Maverick617

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. PcChip · · focus · HN ↗
    I didn't see any benchmarks against vllm, sglang, exllama, etc
    1. rancor · · focus · HN ↗
      Since this is basically a wrapper around libllama.so, I would assume that the performance is roughly the same as llama.cpp upstream.
      1. nullpoint420 · · focus · HN ↗
        Woof. Wonder if the creator knows that
  2. peddling-brink · · focus · HN ↗
    > llama.cpp via Vulkan (AMD / Intel / NVIDIA) or CPU fallback

    I got excited about someone paying attention to intel. Oh well.

    1. kamranjon · · focus · HN ↗
      llama.cpp sycl and vllm xmx work is pretty incredible right now - you just gotta build it with some extra flags
      1. peddling-brink · · focus · HN ↗
        Llama would be nice for the ggufs. Any specific flags or tutorials I should look at?
        1. gunalx · · focus · HN ↗
          The docs are a great start.

          <a href="https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;blob&#x2F;master&#x2F;docs&#x2F;backend&#x2F;SYCL.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;blob&#x2F;master&#x2F;docs&#x2F;backe...

    2. wronglebowski · · focus · HN ↗
      What hardware do you have? I’ve been playing with a 258V and OpenVINO has come a longggggg way.
      1. peddling-brink · · focus · HN ↗
        Two arc b60s. The intel vllm build is getting me ~15t&#x2F;s decode with heavy context using qwen3.8 27b.
        1. bitexploder · · focus · HN ↗
          Those cards should be able to do a lot more. They support int8 and they have very good processing speed. For reference I have a single V100 running around 900t&#x2F;s prefill and 100 t&#x2F;s decode during DeepSWE runs on 27B (4 bit quant). Those cards have half the bandwidth so a realistic decode is going to stop around 50 t&#x2F;s, but those cards can support NVFP4 more easily than the V100. Combined with a TP build you should be able to get 100 t&#x2F;s. And you should be able to 2x the prefill of a V100. I have had Claude optimizing llama.cpp for about a week and am at around 300 t&#x2F;s decode so far on two ancient V100s @ 64 GB RAM.

          Also for the 3 people that ever read this and are curious about local models still, Qwen 27B 3.8 matched Sonnet 5 in the 17 DeepSWE tasks I have run so far, solving the exact same 7 it has. Caveat: datacurve combined low&#x2F;medium&#x2F;high Sonnet 5 data.

          1. wingtw · · focus · HN ↗
            Very interesting! Will you be sharing the optimizations ?
            1. bitexploder · · focus · HN ↗
              I will make sure they get out there soon :)
  3. dlcarrier · · focus · HN ↗
    From what I&#x27;ve seen, Vulkan adds a lot of overhead on Intel hardware.
    1. gunalx · · focus · HN ↗
      Yes and no. I tested llamacpp on my intel gpu with both vulkan and sycl. I measured them to be in the same ballpark even if i have se en pepole claim marger differences than i observed.

      In the end i took the sligth slowdown of vulkan to have a more stable and higher development velocity backend While being able to use identical setups on both intel and and gpus.

  4. shayanjavadi · · focus · HN ↗

    [dead]

  5. aidiveyt · · focus · HN ↗
    claude code appends a role:&quot;system&quot; block after the user prompt, so a proxy rewriting the trailing user message is a no-op.
  6. DylanMerigaud · · focus · HN ↗
    Intel hardware can have Vulkan overhead, impacting performance.
  7. antonyragleap · · focus · HN ↗
    Curious about Vulkan overhead on Intel vs AMD&#x2F;Nvidia for long context. Any benchmarks vs vllm&#x2F;sglang?
  8. ContinuityLab · · focus · HN ↗

    [dead]

  9. kgeist · · focus · HN ↗
    This seems to be a wrapper around libllama, what&#x27;s the point? Llama.cpp already ships a web server. I don&#x27;t see anything in the README that llama.cpp doesn&#x27;t already support.

    Usually people advertise &quot;Go binary&quot; when it&#x27;s pure Go, not just a CGO wrapper. The project seems to be the result of 5 minutes of running Claude Code.

  10. colinsane · · focus · HN ↗
    &gt; Local inference — llama.cpp via Vulkan (AMD &#x2F; Intel &#x2F; NVIDIA) or CPU fallback

    so this is just a go binary that exec&#x27;s `llama-server --model $MODEL_PATH ...`? there&#x27;s room for llama wrappers, sure, but the readme only shows features that are already exposed directly from llama-server. it doesn&#x27;t seem to do anything besides rename the CLI arguments: i don&#x27;t get it.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.