‹ BackHN Continuity

Thread

Show HN: Jev Plays Pokémon Red

282 points · 124 comments · pancomplex

  1. stusmall · · focus · HN ↗
    This is so interesting to watch. For a couple minutes I was in awe of how quick and cheap it was. Then I saw just how bad the decision are and how it would get stuck in strange loops of going in and out of the same door to no end.

    This seems like a technology heading in the right direction but not quiet there yet. Excited for what they are cooking up but probably won't start building around it yet.

    1. binlog · · focus · HN ↗
      This entire conversation around Jev seems weird to me. Like... we started from neural nets that could do basic decision making and classifications pretty well, then trained larger and larger language models to get to where we are now. Now suddenly everyone is going crazy because someone trained a smaller model that is adequate at making decisions? We already went through the "look this AI can play pokemon terribly" phase like a decade ago.
      1. azan_ · · focus · HN ↗
        Making decisions quickly, cheaply and without having to train your own model.
        1. joshuat · · focus · HN ↗
          Math.random can make poor decisions quickly and cheaply
          1. azan_ · · focus · HN ↗
            Benchmark it against jev and you'll have your answer.
            1. zahlman · · focus · HN ↗
              I mean,

              > get stuck in strange loops of going in and out of the same door to no end

              Math.random is statistically unlikely to do this.

              1. ranguna · · focus · HN ↗
                Maybe math.random would've taken 10B years to get to that door in the first place.
                1. magicpin · · focus · HN ↗
                  Pathfinding is handled by the harness
            2. joshuat · · focus · HN ↗
              mathrandomplayspokemon.org
            3. catlifeonmars · · focus · HN ↗
              No you have to use Math.sqrt(Math.random()) to outperform jev
      2. ford · · focus · HN ↗
        I agree it's overhyped, but the transition to a general purpose classifier (vs a narrow scope classifier) is new and noteworthy.

        Ie the famous "Hotdog" clip from Silicon Valley [0]

        <a href="https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=ACmydtFDTGs" rel="nofollow">https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=ACmydtFDTGs

        1. janalsncm · · focus · HN ↗
          Maybe noteworthy but definitely not new. The category of zero-shot classification has been around for a while.

          Example (2022):

          <a href="https:&#x2F;&#x2F;developers.openai.com&#x2F;cookbook&#x2F;examples&#x2F;zero-shot_classification_with_embeddings" rel="nofollow">https:&#x2F;&#x2F;developers.openai.com&#x2F;cookbook&#x2F;examples&#x2F;zero-shot_cl...

      3. c7b · · focus · HN ↗
        A pre-trained universal classifier that can replace specifically-trained ones would have been considered just as much science fiction in the 2010&#x27;s as the capabilities of modern LLMs. I&#x27;m not sure Jev is actually there yet, but at least it sounds theoretically doable today.

        That being said, one thing having been unrealistic 10 years ago and just about possible today doesn&#x27;t mean that it&#x27;s going to change the world the same way another technically related, previously-impossible thing did. The Jev hype gives me a bit of the &quot;you&#x27;re still early to crypto&quot; vibes of some later altcoins. I really like the idea, I think it&#x27;s going to open up possibilities for using classifiers where we wouldn&#x27;t or couldn&#x27;t have trained one before. I&#x27;m crossing my fingers for an open weights version to drop. But it&#x27;s still just a classifier, people have built similar things before Jev, the one thing that really stands out about it is their ability to generate hype.

        1. someothherguyy · · focus · HN ↗
          &gt; but at least it sounds theoretically doable today

          why

          1. thornewolf · · focus · HN ↗
            We have a bad universal classifier now (via Jev). 0-&gt;1, one might say.

            A bad universal classifier does suggest a good one later. And that is exactly what I would call &quot;theoretically doable&quot;

            That said, I don&#x27;t think that Jev is a magic breakthrough or anything. I think it is just a particularly good narrative with an easy way to try it out.

            1. [deleted] · · focus · HN ↗

              [deleted]

            2. lukev · · focus · HN ↗
              Jev is interesting in that it&#x27;s much cheaper and faster than a frontier LLM.

              But I&#x27;ve seen nothing to indicate that the upper bound on classification tasks of a Jev-like model can exceed a frontier LLM with reasoning tokens. That seems nearly impossible even in principle (since Jev-style models are still based on LLM pretraining).

              So while they&#x27;re definitely on the Pareto frontier, which is valuable, they&#x27;re at the &quot;cheap&quot; end of the spectrum more than the &quot;good&quot; end and I don&#x27;t expect that to change.

          2. c7b · · focus · HN ↗
            LLMs are like lossy compression of ~all of written text ever produced, with useful recall. To the extent that the corpus contains labelled examples of the given classification task, it&#x27;s not unreasonable to think that we&#x27;ll be able to build a decoder for that, just like we already have a useful decoder for next-token prediction. Extend to image classification the same way we already have multimodal LLMs.
        2. daveguy · · focus · HN ↗
          &gt; A pre-trained universal classifier that can replace specifically-trained ones would have been considered just as much science fiction in the 2010&#x27;s as the capabilities of modern LLMs. I&#x27;m not sure Jev is actually there yet, but at least it sounds theoretically doable today.

          That&#x27;s the point. It can&#x27;t. And it&#x27;s not even close.

          1. rco8786 · · focus · HN ↗
            Sounds eerily familiar to how we talked about ChatGPT doing [various things that it now does easily] when it was released.
      4. osener · · focus · HN ↗
        It is impressive, but all the hype and fake demos are selling it as a model that is as smart as frontier reasoning LLMs in the decisions it makes yet much cheaper and much faster, which is not true.
        1. pancomplex · · focus · HN ↗
          Nothing fake here and fully open source if you wanna take a peek. It does make a bunch of mistakes, often. But it eventually recovers!

          <a href="https:&#x2F;&#x2F;github.com&#x2F;christianmat&#x2F;jev-pokemon" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;christianmat&#x2F;jev-pokemon

          1. tehsauce · · focus · HN ↗
            For those curious how it works, it’s essentially a script that plays the game but uses jev as a source of rng to make it stochastic
            1. Jach · · focus · HN ↗
              Yeah, there&#x27;s a lot hard-coded into the typescript files that make this much less impressive than the other &quot;AI plays&quot; versions that have come and gone that play with less help. A Claude one that only used screenshots would always get stuck in the rocket hideout...
            2. Leynos · · focus · HN ↗
              Does dropping a prng in place of Jev produce similar results?
          2. osener · · focus · HN ↗
            For the record I did not mean to say that about your work, nice job.
        2. gchamonlive · · focus · HN ↗
          I think that misses the point of Jev being ridiculously efficient while maintaining adequate intelligence for automation tasks. We have to train our minds to filter out branding and marketing.
          1. [deleted] · · focus · HN ↗

            [deleted]

      5. gchamonlive · · focus · HN ↗
        For what it does, it classifies, orchestrates, operates and delegates tasks exceedingly well for its size and weight. It&#x27;s ridiculously cheap and efficient, but if you can only see progress in terms of raw cognitive power then you&#x27;ll surely miss how interesting this is.
      6. solidasparagus · · focus · HN ↗
        The cheap, fast and smart-enough LLM space has been wildly neglected. Jev is one of the few players truly targeting that space. And for a lot of people it is the first time they are asking &quot;what could I build if llms were interaction-speed fast?&quot;. The answers are cool, the problem is that Jev is not, I think, smart-enough yet to have that many applications, but it&#x27;s smart enough that you can start to see what they will look like.
      7. michaelchisari · · focus · HN ↗
        Performance and efficiency have been neglected as everyone threw every available GPU and trillions of dollars trying (and failing) to create AGI. Though the results have been impressive.

        But good enough for pennies in an instant is very useful.

      8. irregularbowels · · focus · HN ↗

        [dead]

      9. aurareturn · · focus · HN ↗
        Yea but this model wasn’t trained to play Pokémon but it can. It’s general.
        1. daveguy · · focus · HN ↗
          So is a random number generator. That doesn&#x27;t mean it can play Pokemon.
      10. mtford · · focus · HN ↗
        This happens all the time in tech. A few years ago everybody got excited about static websites and server-side rendering as if we hadn&#x27;t been doing that with PHP long ago.
      11. Garlef · · focus · HN ↗
        The difference is that you don&#x27;t need training here; The decision graphs can be built on the fly by an LLM and contain instructions in plain language.

        The magic moment for me from the Jev release was not that there was some system playing doom: Rather it was the moment, they just changed a part of the prompt to &quot;don&#x27;t shoot, just dodge&quot; and the behavior changed immediately.

        This means you can have a system with fast decision-making but still interact with it via language.

        1. catlifeonmars · · focus · HN ↗
          You’re kind of screwed on the latency optimization side bc your bottleneck is a language model though. Why wouldn’t you eventually just use a specialized model that isn’t constrained by running a language model under the hood?
          1. Garlef · · focus · HN ↗
            True - in the long run. But that&#x27;s quite expensive. And you&#x27;d need a bigger org to even support this.

            They even named the model after &quot;Jevons Paradoxon&quot; - They anticipate that their model lowers the cost of adopting this kind of AI significantly, unlocking a lot of use cases.

      12. sky2224 · · focus · HN ↗
        The difference to me is that the models from a decade ago would need to be trained to play the game. As far as I&#x27;m aware, Jev has had zero specific training or fine-tuning to play the game. It&#x27;s simply told, &quot;Hey, here are the rules. Play.&quot; That&#x27;s it.

        I work in manufacturing. I think this will be fantastic for stuff like SPC.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.