‹ BackHN Continuity

Thread

Introducing System One Models and Jev

1989 points · 520 comments · albelfio

  1. jacobgold · · focus · HN ↗
    First, congrats to the team on launching something genuinely interesting and new.

    Seems like a more accurate title would be "Jev: Trading general purpose generation for fast typed inference" or something like that.

    This is interesting, but the speed comparison seems misleading? A generative model that can output code in a Turing-complete language can do anything a computer can do.

    Jev can only generate structured output, right? This is probably super useful for classification/routing/scoring, but it's nothing like the code generating models we're all using today for code and automation.

    Also "can't hallucinate" seems wrong? Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value. You can enforce structured output from an LLM too, with an appropriate harness, etc.

    Assuming there's no funny business, the Doom demo is cool.

    1. CompleteSkeptic · · focus · HN ↗
      I'm biased but I wouldn't call it misleading - generating text is super awesome and flexible, (we describe that in the blog post - and I personally use string models all the time) but it's true you pay a high tax for autoregressive generation

      > Also "can't hallucinate" seems wrong? Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value.

      that is likely true of all ML! perhaps we could debate semantics, but I don't think it's fair to say a random forest "hallucinates" in the way LLMs do

      1. WhitneyLand · · focus · HN ↗
        His claim was that the title is misleading, not sure how it's relevant to that claim that you use "string models" (full LLMs).

        The original title before it changed less than an hour ago was:

        "Jev: New frontier model 40-400x cheaper and 20-200x faster"

        I'm going to agree that was misleading.

        And on the second point:

        >>Also "can't hallucinate" seems wrong? Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value.

        >that is likely true of all ML! perhaps we could debate semantics, but I don't think it's fair to say a random forest "hallucinates" in the way LLMs do"

        Also going to disagree here, and I don't think it's semantics.

        Type safety is not factual correctness.

        1. CompleteSkeptic · · focus · HN ↗
          > Type safety is not factual correctness.

          I very much agree with this and want to hone in on where do actually disagree. Would you say a linear classifier hallucinates?

          1. thduabmd · · focus · HN ↗
            No. Your launch post puts “0%” on a hallucination chart, then explains that the number comes from guaranteed schema matching.

            You’ve already agreed that this doesn’t establish correctness. An approve for an unauthorized action still meets the schema guarantee.

            That’s why I find the messaging misleading. You’re acknowledging the limitations in these replies while defending the broader reliability pitch.

            Even granting that each answer is calibrated individually, that doesn’t establish calibration of the decision that combines them.

            Sure, I can threshold a composite score, but there may be many wrong answers with the same score. An unauthorized action doesn’t become acceptable because it scores highly on the other dimensions.

            I still have to define the constraints and test which wrong actions get through the complete workflow on my own data. That’s a substantial part of the work being pushed back onto the developer.

            1. agos · · focus · HN ↗
              hallucinations are not wrong answers, that's why we use a different term
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.