‹ BackHN Continuity

Thread

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

575 points · 225 comments · firelex

  1. adrithmetiqa · · focus · HN ↗
    Forgive my lack of understanding but how long before Jev type functionality is just built straight into all frontier models?
    1. tbeseda · · focus · HN ↗
      If I had to guess, it's already built and is just waiting on Product's/Marketing's desk. How do you position this without looking like your roadmap is being determined by newcomers? Probably don't want to adopt the same verbiage+acronyms - but also can't be seen to be just sherlocking features.
      1. seizethecheese · · focus · HN ↗
        I think Apple has demonstrated that shipping second has essentially no negative impact if your product is seen as higher quality.
        1. bigyabai · · focus · HN ↗
          I think Nvidia has demonstrated that shipping first is a multi-trillion dollar opportunity if you don't shy away from a challenge.
          1. throwaway27448 · · focus · HN ↗
            Chip manufacturing intrinsically comes with one hell of a moat. There's not much parallel in software.
            1. slashdev · · focus · HN ↗
              That’s kind of funny because Nvidia’s biggest moat is arguably CUDA, the software ecosystem around their chips
              1. mcmcmc · · focus · HN ↗
                CUDA is a lock-in moat, the infrastructure needed for chip manufacturing is a barrier-to-entry moat. Two different things.
                1. angry_octet · · focus · HN ↗
                  There are many microarchitecture patents used in NVIDIA chips. I'm sure they have a team that rips apart AMD chips looking for infringement. The way CUDA works is tied to many GPU architecture decisions and it would be hard to decouple them efficiently. Obviously a huge effort was made to get PyTorch decoupled from CUDA.
                2. AtlasBarfed · · focus · HN ↗
                  If cuda is an API, and llms make apis effortless, then how big of a moat is cuda?
                  1. robflynn · · focus · HN ↗
                    I ran across this a few days ago: <a href="https:&#x2F;&#x2F;zluda.org&#x2F;" rel="nofollow">https:&#x2F;&#x2F;zluda.org&#x2F;
                    1. QuantumNomad_ · · focus · HN ↗
                      All of the buttons and links on that page redirect to spam pages. Most of the times I clicked, it brings me to some site that wants to sell me a VPN.
                      1. robflynn · · focus · HN ↗
                        Oh, yikes, I should&#x27;ve checked those links before posting it here. That&#x27;s certainly not a good look for them.

                        There was a github repo but I have not checked it.

                        edit I see, thats an unaffiliated site that latched onto that, my bad, here&#x27;s the GH that I should&#x27;ve linked: <a href="https:&#x2F;&#x2F;github.com&#x2F;vosen&#x2F;ZLUDA" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;vosen&#x2F;ZLUDA

                    2. [deleted] · · focus · HN ↗

                      [deleted]

                  2. bobmarleybiceps · · focus · HN ↗
                    IMO, yes a lot of the &quot;nvidia pays lots of people to make non-portable, tightly coupled backends to open source projects&quot; is potentially going to be less of a moat?

                    (Though it could turn into &quot;nvidia pay lots of people to use LLMs to make non-portable, tightly coupled backends to _even more_ open source projects&quot;)

                  3. trollbridge · · focus · HN ↗
                    It turns out the moat is “writing drivers that work”; Nvidia drivers simply work, and Intel’s are poor quality. So if I want stuff that works I need to buy Nvidia gear.
                  4. latentsea · · focus · HN ↗
                    It&#x27;s becoming less of one. Previously I would have shied away from buying an AMD card because of CUDA, but with local LLMs getting good enough to be usable and frontier models becoming as good as they have, I bit the bullet and got an R9700 for local inference. Dealing with working around CUDA used to be more of a manual process, but when you can point an agent at it and get stuff working, it&#x27;s dramatically less painful and scary than it used to be. Plus, at least in ComfyUI and local LLMs I&#x27;m finding support for AMD has gotten really good. Lately I&#x27;ve been witnessing a lot of people using agents to write custom kernels for RDNA4 and improving performance dramatically.
                3. flyinglizard · · focus · HN ↗
                  I’m totally guessing, but I can’t imagine CUDA has any significance at the frontier lab scale. The operational and capex costs are so massive that the convenience of the platform becomes a minuscule consideration.

                  It’s just that Nvidia’s stuff works, and available at scale, and includes the full stack with networking, cooling and such.

              2. angry_octet · · focus · HN ↗
                Historically there was a big patent moat in (Graphics) GPU design. This continues with CUDA, but obviously Intel and AMD could find ways to support eg PyTorch. What we don&#x27;t know is how much effort that cost them, or why they decided they couldn&#x27;t make a CUDA API compatible competitor.
                1. trollbridge · · focus · HN ↗
                  Intel could have beat the pants off Nvidia a long time ago with Arc if they’d bothered to ship usable drivers. But they refuse to, and simply can’t figure it out, so Arc cards remain cheap because they’re so #%£€ing hard to get working well, and everyone is nervous they’ll lay off the driver team again.
                  1. api · · focus · HN ↗
                    AMD has always had software problems too. Hardware companies often devalue software and suck at it.

                    If you can make them work Arc cards are a massive bargain. On raw compute the silicon is not bad.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.