‹ BackHN Continuity

Thread

Introducing System One Models and Jev

1989 points · 520 comments · albelfio

  1. jacobgold · · focus · HN ↗
    First, congrats to the team on launching something genuinely interesting and new.

    Seems like a more accurate title would be "Jev: Trading general purpose generation for fast typed inference" or something like that.

    This is interesting, but the speed comparison seems misleading? A generative model that can output code in a Turing-complete language can do anything a computer can do.

    Jev can only generate structured output, right? This is probably super useful for classification/routing/scoring, but it's nothing like the code generating models we're all using today for code and automation.

    Also "can't hallucinate" seems wrong? Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value. You can enforce structured output from an LLM too, with an appropriate harness, etc.

    Assuming there's no funny business, the Doom demo is cool.

    1. dbbk · · focus · HN ↗
      When they say "can't hallucinate" they mean they produce a confidence value for every result, so you could see for example it has 0.1 confidence, and you can disregard the result - that'd be different from hallucinating where it believes it's correct
      1. orbital-decay · · focus · HN ↗
        Yeah but what stops it from producing confidently incorrect outputs...
        1. zenlikethat · · focus · HN ↗
          Nothing, but imagine using LLMs for a classification task

          People out there are so resigned to the models being unreliable that they are really doing things like hallucinating deliberately, and then matching the hallucinations to embeddings -

          <a href="https:&#x2F;&#x2F;softwaredoug.com&#x2F;blog&#x2F;2026&#x2F;08&#x2F;10&#x2F;hypothetical-classifications" rel="nofollow">https:&#x2F;&#x2F;softwaredoug.com&#x2F;blog&#x2F;2026&#x2F;08&#x2F;10&#x2F;hypothetical-classi...

          You could do that or you could just... use a model that will never produce unreliable outputs in the first place.

          1. ActivePattern · · focus · HN ↗
            It would be great to see benchmarks for Jev that demonstrate the value of calibrated uncertainty.

            For example, one could set a confidence threshold over which we trust the model decision, and otherwise reject. This provides a lever to trade-off accuracy and automation %.

            Then we can ask questions like &quot;What % of decisions can we automate to achieve 90% accuracy&quot;?

            1. suraj_phanindra · · focus · HN ↗
              i ran some interesting experiments on jev today and wanted to share the results. i compared cost and latency of support ticket triage with different arrival rates for jev vs mainstream LLMs: <a href="https:&#x2F;&#x2F;suraj-website-eta.vercel.app&#x2F;blog&#x2F;what-a-correct-decision-costs" rel="nofollow">https:&#x2F;&#x2F;suraj-website-eta.vercel.app&#x2F;blog&#x2F;what-a-correct-dec...
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.