‹ BackHN Continuity

Thread

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

575 points · 225 comments · firelex

  1. hadlock · · focus · HN ↗
    If you&#x27;re looking for a more impressive doom example, I put together laya-duum which uses open source micropython implementation of doom (duum) and freeware freedoom1.wad. It uses the standard jev api and will play through the first two levels to completion: <a href="https:&#x2F;&#x2F;github.com&#x2F;Hadlock&#x2F;laya-duum" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;Hadlock&#x2F;laya-duum
    1. sunbum · · focus · HN ↗
      Why is this more impressive?
      1. hadlock · · focus · HN ↗
        It&#x27;s not just shooting at static monsters in an empty room, this will go after monsters, make decisions about prioritizing health, changing goals, and finish the level, etc. It&#x27;s not unique, it&#x27;s just more impressive than the demo they&#x27;re using.
        1. runeks · · focus · HN ↗
          Their demo could also change goals, e.g. avoid monsters instead of shooting them, by changing some variable. I&#x27;m sure that could also be used to change goals.
          1. mosselman · · focus · HN ↗
            But it doesn’t. That is why that other person is claiming their demo is more impressive.

            What you are saying is like me losing 13kg is as impressive as having the potential to do so if only I would stop eating ice cream and too many snacks every day.

            1. trencedamp · · focus · HN ↗
              You&#x27;re confusing demo with technology.

              The potential to lose weight, in this context, is the original post because there was potential but it wasn&#x27;t shown

              This commenter is the person who lost 13kg.

              The technology is not different but the demonstration is more impressive

    2. antman · · focus · HN ↗
      - Laya receives semantic snapshots, not the framebuffer.

      What does tjis mean? Another generative model?

      1. ovi256 · · focus · HN ↗
        Laya is the model this uses. Like Jev, it&#x27;s blind - it can&#x27;t take an image as an input. So how can it play Doom, or Atari games, or do other visual tasks, as people have shown it to do? They write a bit of app-specific adapter code that transforms the game state into a structured piece of data (here, the &quot;semantic snapshot&quot;). And that is what these models take as input
    3. prodigycorp · · focus · HN ↗
      laya sucks. it&#x27;s a dead end and for some treat it as a peer of jev or even the qwen jevs. base model doesnt have enough world knowledge to generalize.
      1. earino · · focus · HN ↗
        The base-model criticism is fair. Laya is weak zero-shot and doesn&#x27;t have the world knowledge of a much larger pretrained model. Their own site puts the base models at around 0.35 on the typed-decisions benchmark, which is close to random.

        I don&#x27;t think that makes it a dead end. Laya is designed to be fine-tuned for a specific task, and the site reports the fine-tuning gains. The 0.766 number comes from fine-tuning on the benchmark&#x27;s train split, not from the base checkpoint. They also report that fitting a single temperature scalar per question type cuts expected calibration error from 0.466 to 0.081. That&#x27;s a large gain, and it only shows up after you specialize the model.

        &quot;Peer&quot; is doing a lot of work in that comment. Jev can take on a new task without retraining because it starts with far more knowledge; Laya trades that away to stay small and trainable for a fixed task. So I wouldn&#x27;t compare base Laya to Jev and stop there. Compare Jev to fine-tuned Laya on the same task and test set, then look at accuracy, latency, cost, calibration, and robustness, depending on which of those matter for the deployment.

        1. prodigycorp · · focus · HN ↗
          btw your comment was immediately marked as flagged&#x2F;dead. Odd for a high quality comment. I vouched for it.

          My main point of disagreement would be that I fundamentally see a different use for a jev sort of model (generalism is applealing), but if you&#x27;re finetuning, a bert base is not bad.

          1. earino · · focus · HN ↗
            Thanks for the vouching. I don&#x27;t really understand what I did to anger the algorithms! But yeah, I think that that&#x27;s the real sticking point. If the value prop of jev is &quot;plug it in anywhere and it generally does smart things&quot; you&#x27;re absolutely right that Laya is Not A Thing. I&#x27;ve spent a bunch of time working on fine tuning models (i teach a class on it! <a href="https:&#x2F;&#x2F;earino.github.io&#x2F;applied-deep-learning&#x2F;" rel="nofollow">https:&#x2F;&#x2F;earino.github.io&#x2F;applied-deep-learning&#x2F; using the same ModernBERT stuff!) and so for me, I naturally went: &quot;wait a second, so how good is Jev compared to fixing the problem domain.&quot;

            Anyways, thanks for the vouching!

            1. prodigycorp · · focus · HN ↗
              Great course, thanks for sharing, going through your slides and papers right now.
            2. fwip · · focus · HN ↗
              My guess is, it&#x27;s because you sounded like an LLM in a few phrases there.
              1. earino · · focus · HN ↗
                I think that&#x27;s a pretty good guess. I move around different countries a lot so I also assume I can look like a bot farm from multiple source IPs.

                However at this point I talk to LLMs more than anyone except probably my wife. As a multiple times immigrant, I can absolutely believe I&#x27;m adjusting my speech patterns to its vernacular.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.