‹ BackHN Continuity

Thread

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

575 points · 225 comments · firelex

  1. hadlock · · focus · HN ↗
    If you&#x27;re looking for a more impressive doom example, I put together laya-duum which uses open source micropython implementation of doom (duum) and freeware freedoom1.wad. It uses the standard jev api and will play through the first two levels to completion: <a href="https:&#x2F;&#x2F;github.com&#x2F;Hadlock&#x2F;laya-duum" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;Hadlock&#x2F;laya-duum
    1. prodigycorp · · focus · HN ↗
      laya sucks. it&#x27;s a dead end and for some treat it as a peer of jev or even the qwen jevs. base model doesnt have enough world knowledge to generalize.
      1. earino · · focus · HN ↗
        The base-model criticism is fair. Laya is weak zero-shot and doesn&#x27;t have the world knowledge of a much larger pretrained model. Their own site puts the base models at around 0.35 on the typed-decisions benchmark, which is close to random.

        I don&#x27;t think that makes it a dead end. Laya is designed to be fine-tuned for a specific task, and the site reports the fine-tuning gains. The 0.766 number comes from fine-tuning on the benchmark&#x27;s train split, not from the base checkpoint. They also report that fitting a single temperature scalar per question type cuts expected calibration error from 0.466 to 0.081. That&#x27;s a large gain, and it only shows up after you specialize the model.

        &quot;Peer&quot; is doing a lot of work in that comment. Jev can take on a new task without retraining because it starts with far more knowledge; Laya trades that away to stay small and trainable for a fixed task. So I wouldn&#x27;t compare base Laya to Jev and stop there. Compare Jev to fine-tuned Laya on the same task and test set, then look at accuracy, latency, cost, calibration, and robustness, depending on which of those matter for the deployment.

        1. prodigycorp · · focus · HN ↗
          btw your comment was immediately marked as flagged&#x2F;dead. Odd for a high quality comment. I vouched for it.

          My main point of disagreement would be that I fundamentally see a different use for a jev sort of model (generalism is applealing), but if you&#x27;re finetuning, a bert base is not bad.

          1. earino · · focus · HN ↗
            Thanks for the vouching. I don&#x27;t really understand what I did to anger the algorithms! But yeah, I think that that&#x27;s the real sticking point. If the value prop of jev is &quot;plug it in anywhere and it generally does smart things&quot; you&#x27;re absolutely right that Laya is Not A Thing. I&#x27;ve spent a bunch of time working on fine tuning models (i teach a class on it! <a href="https:&#x2F;&#x2F;earino.github.io&#x2F;applied-deep-learning&#x2F;" rel="nofollow">https:&#x2F;&#x2F;earino.github.io&#x2F;applied-deep-learning&#x2F; using the same ModernBERT stuff!) and so for me, I naturally went: &quot;wait a second, so how good is Jev compared to fixing the problem domain.&quot;

            Anyways, thanks for the vouching!

            1. prodigycorp · · focus · HN ↗
              Great course, thanks for sharing, going through your slides and papers right now.
            2. fwip · · focus · HN ↗
              My guess is, it&#x27;s because you sounded like an LLM in a few phrases there.
              1. earino · · focus · HN ↗
                I think that&#x27;s a pretty good guess. I move around different countries a lot so I also assume I can look like a bot farm from multiple source IPs.

                However at this point I talk to LLMs more than anyone except probably my wife. As a multiple times immigrant, I can absolutely believe I&#x27;m adjusting my speech patterns to its vernacular.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.