‹ BackHN Continuity

Thread

Show HN: Jev Plays Pokémon Red

282 points · 124 comments · pancomplex

  1. stusmall · · focus · HN ↗
    This is so interesting to watch. For a couple minutes I was in awe of how quick and cheap it was. Then I saw just how bad the decision are and how it would get stuck in strange loops of going in and out of the same door to no end.

    This seems like a technology heading in the right direction but not quiet there yet. Excited for what they are cooking up but probably won't start building around it yet.

    1. ralusek · · focus · HN ↗
      The exact message I sent my friend this morning:

      > the most interesting thing about this jev stuff

      > is that people are seemingly like

      > completely disinterested in how smart it actually is

      > I haven't even heard it mentioned a single time how it actually compares to other LLMs coming up with their own classifications. Just: it's fast and cheap

      After watching a few minutes of this it makes me think that maybe we should be a little more interested in how smart it is.

      1. stusmall · · focus · HN ↗
        There is a lot of room for a lot of different models. For many use cases, intelligence beats out all.

        For me in my day job, having extremely fast low quality decision makers over noisy inputs is very valuable. I work in security and having something that can help triage alerts, classify items and group things together is extremely valuable. It doesn't need to be perfect. Just being able to take a set of inputs from deterministic tooling and to be make general priority classifications goes a long way on helping humans look at the most important items first.

        1. refactor_master · · focus · HN ↗
          Why not spend a few tokens or so on making a frontier LLM build the classifier from scratch? Then you also know exactly what the classifier is capable of, and whether it’s easily learnable/informative data.

          Why would you be confident in feeding garbage to a “cheap and fast” classifier with unknown domain-specific performance?

          I know we kind assume omniscience for frontier models, but at this point the evidence is kind of out there.

      2. johnsmith1840 · · focus · HN ↗
        They know it's not. The company actually has or had public statements that they didn't like public benchmarks for comparison.

        My main wonder is the difference between it and having a small llm no thinking output a single number only as a choice. Isn't that nearly the same here?

      3. raincole · · focus · HN ↗
        Because it's not that smart especially when you compare it to other LLMs. The top LLMs have completely change the baseline of being smart.
        1. hbrn · · focus · HN ↗
          You gotta admit that turning "our model is very dumb and can only make simple decisions" into "System One Model" is a pretty clever marketing move.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.