‹ BackHN Continuity

Thread

Laya on Mac M4 CoreML Offline

174 points · 34 comments · putna

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. PaulRobinson · · focus · HN ↗
    Local LLMs are the future, and one of the reasons I think the data centre furore is just going to end in a market crash.

    LLMs that can reliably be used for control problems are the future, and I think classic/deep RL has generally been overlooked for years for a whole host of problems by wider industry because it felt inaccessible. The first thing I thought of when I saw Jev (and then Laya), was "this might move the needle in a really, really interesting way".

    Local LLMs that can reliably be used for control problems smash through a lot of barriers I'm interested in, and this intrigues me a lot. Guess I'm about to become a big Laya fan if it can run on this kind of hardware to this performance.

    1. frag · · focus · HN ↗
      that&#x27;s not a local LLM. If it&#x27;s local, it doesn&#x27;t matter in this case. Laya is a System 1 &quot;AI&quot;, namely works like a classifier, given a state and questions, it shoots probabilities for each. I publish an episode tomorrow about Laya and Typesafe AI on <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;@DataScienceatHome" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;@DataScienceatHome

      Stay tuned ;)

      1. putna · · focus · HN ↗
        cool, will check it
      2. EagnaIonat · · focus · HN ↗
        I’ve found they start to fail the more classifications you have, long before your typical ML classifier.

        To me it’s like a solution looking for a problem that is already solved.

    2. bigyabai · · focus · HN ↗
      It won&#x27;t. Laya is a finetuned version of Google&#x27;s BeRT model, which is almost 10 years old right now.

      If BeRT had any potential to disrupt the datacenter buildout, it already would have.

      1. viraptor · · focus · HN ↗
        Modernbert is from 2024. It&#x27;s also trained from scratch, not a fine tune.
    3. ipsi · · focus · HN ↗
      The future for whom? The general public? Not a chance, no way, not unless it&#x27;s able to run on a phone (anywhere from 20-40% of internet users, world-wide, are phone-only).

      For companies? I think that&#x27;s a lot more plausible, as that&#x27;s mostly just a question of money - is it cheaper to run and administrate our own models, or outsource that?

      For technically inclined users? I think that&#x27;s unlikely unless they&#x27;re able to operate on relatively cheap hardware while still being just as good as the hosted models. And by that I don&#x27;t mean &quot;a mac studio,&quot; that&#x27;s far more money than I think is reasonable. A single RTX 5080, maybe, once memory prices start to drop.

      1. Izmaki · · focus · HN ↗
        Compare the games your average high-end smartphone can run to the AAA titles of the 2010&#x27;s. It&#x27;s not a matter of &quot;unless it is able to&quot; but &quot;when it is able to&quot;.
        1. bigyabai · · focus · HN ↗
          That&#x27;s going from 150w 720p gaming to ~15w 720p gaming in about 10 years. A typical inference cluster will draw well over 1500w to deliver even a small-ish 500b model at reasonable speeds&#x2F;quantization.

          So, extrapolating from your gaming example, it will take smartphones only... *checks clipboard* ...100 years to achieve datacenter-level performance at the pace of 2010&#x27;s improvements.

          1. mynegation · · focus · HN ↗
            Uh, my clipboard says 20 years
            1. bigyabai · · focus · HN ↗
              Mea culpa, but it&#x27;s still a decent while.
          2. Izmaki · · focus · HN ↗
            We don&#x27;t need datacenter-level performance in a handheld device just like we didn&#x27;t need a storage room in our backyard for a mid-sized, tape-based NAS because technology did its thing and gave us something smarter, smaller and faster.

            Your argument is true if the number of parameters in an LLM is the only measurement for quality - a bit like number of bolts per aircraft or lines of code in software. I&#x27;d bet you that an 8b parameter LLM will in 5 years outperform the datacenter-level LLMs of today.

    4. itemize123 · · focus · HN ↗
      no way. one is functionally (slightly overhyped) magic. another is better tool.
  2. oezi · · focus · HN ↗

    [dead]

  3. tentacleuno · · focus · HN ↗
    This looks like a local AI model playing Snake -- is that correct? The article offers no explanation.
    1. ryuuseijin · · focus · HN ↗
      The linked github project [1] contains more information.

      [1]: <a href="https:&#x2F;&#x2F;github.com&#x2F;mizorewww&#x2F;laya-coreml" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;mizorewww&#x2F;laya-coreml

    2. putna · · focus · HN ↗
      Correct, the cli commands are just copy paste to run on your machine.
  4. frag · · focus · HN ↗
    Great job! Did you finetune your own Laya for the snake game or what?
    1. putna · · focus · HN ↗
      its just an example how to run it locally. it takes the laya from official repo. you can try running it in 2 minutes
  5. aidiveyt · · focus · HN ↗

    [dead]

  6. altano · · focus · HN ↗
    [delayed]
    1. WASDx · · focus · HN ↗
      The model is only 0.3B params so probably not much.
    2. putna · · focus · HN ↗
      Physical footprint: 560.4M Physical footprint (peak): 778.0M
    3. ImJasonH · · focus · HN ↗
      I had Fable one-shot the same demo from the same weights on an iPhone 15 Pro and it decides in ~40ms.
  7. imranq · · focus · HN ↗
    Based on my admittedly limited research, it seems like you should use Laya for much more deterministic tasks where you have some training data. It won&#x27;t be as good as Jev for zero shot cases.
    1. putna · · focus · HN ↗
      agree, too good to be true for one shot cases
    2. jwpapi · · focus · HN ↗
      I think it can make a lot of sense to make a set with jev and then train laya on that set and do the rest with laya for saving $
      1. adinb · · focus · HN ↗
        I did that on a little macbook m4 last night on my model of the innate immune system—fine tuning took 15m or so. Just wish it had a larger context window
        1. jwpapi · · focus · HN ↗
          How about runpod
          1. adinb · · focus · HN ↗
            No need—small and light enough was able to do everything offline, though runpod would work fine, though it‘d be be a quick job
  8. EgregiousCube · · focus · HN ↗
    Isn&#x27;t part of the Jev marketing that it has &quot;terra-class intelligence&quot;? I don&#x27;t know how much it actually achieves that, but unless that&#x27;s EXTREMELY wrong, it&#x27;s hard to see how a 0.3B model could claim to be an OS Jev.
    1. putna · · focus · HN ↗
      agree, maybe os jev is too strong of a wording
    2. physicallyIllfr · · focus · HN ↗
      Jevs marketing has a lot to do with claiming credit for this pardigm, when the devloper of Layla was actually the first to do it. Its quite annoying to see another member of the OpenAI mafia so clearly rip off someone elses IP.
  9. speedping · · focus · HN ↗
    So cool. I&#x27;ve fired up pumas (energy monitor) and it seems to run almost fully on the neural engine and not the GPU so it plays really nicely with CoreML
  10. brcmthrowaway · · focus · HN ↗
    Why is Laya being shilled here? It doesn&#x27;t have real intelligence backing it.
    1. putna · · focus · HN ↗
      all examples is to inspire. For example i am inspired and know that i can run simple tasks on local machine offline. its just cool.
  11. grejioh · · focus · HN ↗

    [dead]

  12. tc3oliver · · focus · HN ↗

    [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.