‹ BackHN Continuity

Thread

Introducing System One Models and Jev

1989 points · 520 comments · albelfio

  1. jacobgold · · focus · HN ↗
    First, congrats to the team on launching something genuinely interesting and new.

    Seems like a more accurate title would be "Jev: Trading general purpose generation for fast typed inference" or something like that.

    This is interesting, but the speed comparison seems misleading? A generative model that can output code in a Turing-complete language can do anything a computer can do.

    Jev can only generate structured output, right? This is probably super useful for classification/routing/scoring, but it's nothing like the code generating models we're all using today for code and automation.

    Also "can't hallucinate" seems wrong? Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value. You can enforce structured output from an LLM too, with an appropriate harness, etc.

    Assuming there's no funny business, the Doom demo is cool.

    1. dbbk · · focus · HN ↗
      When they say "can't hallucinate" they mean they produce a confidence value for every result, so you could see for example it has 0.1 confidence, and you can disregard the result - that'd be different from hallucinating where it believes it's correct
      1. 8note · · focus · HN ↗
        if it puts a high confidence value on a wrong answer, thats still hallucinating, no?

        llm hallucinations are high probability tokens that are incorrect vs the real world

        1. dozerly · · focus · HN ↗
          Yes, there is no magic sauce here that makes stochastic output binary if that’s what people are looking for.
          1. resonious · · focus · HN ↗
            Right and so maybe we should stop saying "can't hallucinate" when it can by definition.
          2. rpunkfu · · focus · HN ↗
            It’s not what people are looking for, but what they wrongly claim.
        2. jubilanti · · focus · HN ↗
          Correct, they have not made a universal all-knowing omniscient oracle, which is what would be required for "can't hallucinate".
          1. spencerflem · · focus · HN ↗
            Not to be tooo pedantic, but a bot that assigned 0 confidence to everything wouldn’t hallucinate.

            A calculator either gets the right answer or doesn’t answer.

            It wouldn’t have to be all knowing as long as it knew perfectly what it doesn’t know

            1. baq · · focus · HN ↗
              A quantum calculator answers in distributions.
              1. stpedgwdgfhgdd · · focus · HN ↗
                In one universe that is true, in another one not.
                1. baq · · focus · HN ↗
                  Copenhagen is a beautiful city.
          2. eru · · focus · HN ↗
            That seems like a weird standard.

            I would be happy enough with: only produces what it can verify with sources.

            If you eg try to remember a court case (ie produce the reference via LLM token generation only), it's easy enough to check with your data whether it really exists. Similar for following links and other references.

            If your data or sources are wrong, obviously your report about them will be wrong. But I wouldn't call that a hallucination.

            1. baq · · focus · HN ↗
              There isn’t a single human in this world and hasn’t ever been that meets your happy-enough standard. Make of it what you will.
              1. eru · · focus · HN ↗
                It's not a binary thing. You can get closer or further away from that standard.

                And humans also behave differently in different contexts. A conversation at the pub has more such hallucinations than a formal deposit in court. For the latter, a good lawyer will look at her shoes, when you ask him what colour her laces are.

              2. hdjrudni · · focus · HN ↗
                Why is that at all relevant?

                Humans are known to hallucinate a lot. Ask 10 different witnesses at a crime scene what they saw and they'll all report different things.

                A good, non-hallucinating LLM would only report things for which it has evidence. It would consult the facts every single time.

                It's a pain in the butt for humans to fact-check everything but LLMs can quickly look up all kinds of stuff. That's what makes them useful.

                1. eru · · focus · HN ↗
                  Yes, and for the LLM you can do it in multiple passes.

                  So you can bolt the fact-check / source-check pass onto whatever other system you have, without having to redesign the underlying system.

            2. tedbradley · · focus · HN ↗
              That's not true that incorrect sources means incorrect report. Often, LLMs have some sense of what is true, and due to that, they hallucinate plausible sources that appear to back that knowledge up.
        3. adastra22 · · focus · HN ↗
          No, I don't believe so. Hallucinations are not "high probability" in a real sense. They are an artifact of the random walk the inference algorithm takes, which causes it to latch on to and chase attractors in the noise. This random walk behavior is necessary for chat interfaces to be useful, but are less critical to typed output predictors. I'm guessing they found some optimization that is possible if you give up caring about chat.
        4. elil17 · · focus · HN ↗
          What we would want to see if a confidence value that is in line with the actual correctness. If the value is 0.9 for 1000 different answers, then approximately 900 of those answers should be correct.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.