‹ BackHN Continuity

Thread

Introducing System One Models and Jev

1989 points · 520 comments · albelfio

  1. big_toast · · focus · HN ↗
    It seems like the docs[0] are a better explanation? The comparison to llm tokens is kinda confusing.

    It looks like the model takes as input a state (structured text? not sure if multi-modal) and a question (as a "Choice", "Score", or "Noul") with some additional augmentations possible. Then outputs the question's answers as appropriate (e.g. a choice, accompanying probabilities, confidence).

    Edit: On the AI primer page, it looks like they do the RLCD on a pre-trained base model?

    [0]:<a href="https:&#x2F;&#x2F;docs.typesafe.ai&#x2F;concepts&#x2F;system-one" rel="nofollow">https:&#x2F;&#x2F;docs.typesafe.ai&#x2F;concepts&#x2F;system-one

    1. CompleteSkeptic · · focus · HN ↗
      CEO here - that is right!

      I do agree that the comparison to LLM tokens is hard to understand (also because output tokens are not comparable).

      But yes, text or structured state (like a JSON with multiple pieces of text in) -&gt; decisions out (e.g. choice maps to &quot;match&quot; statement, &quot;score&quot; maps to sorting, &quot;noul&quot; short for bernoulli maps to if-statements)

      1. mckngbrd · · focus · HN ↗
        here is how I attempted to explain it to my company&#x27;s AI group chat, is this roughly accurate?

        &quot;instead of autoregressive string output it instead outputs structured type-safe &#x27;decisions&#x27; with probabilities&#x2F;confidence scores, each generated in parallel

        so sort of more like a Large Classification Model than a Large Language Model? or, maybe better to think of it as a sort of &quot;shift left&quot; in the LLM&#x27;s transformer architecture, allowing you to replace the predefined token vocabulary of an LLM with a prescribed set of &#x27;decisions&#x27; that need to be made based off the input context; and exposing those probabilities directly so they can be integrated into the system logic, instead of just sampling from top-K.

        all of this while still being instruction-tuned (!!!)&quot;

        It&#x27;s always been possible to build classification pipelines using LLM embeddings as the input. seems like this is a much more sophisticated &#x2F; useful application of that concept

        1. CompleteSkeptic · · focus · HN ↗
          very accurate!

          the one nuance I&#x27;d get into is I&#x27;d call it &quot;zero-shot&quot; over &quot;instruction-tuned&quot; (the latter often implies a particular distribution), but very safe for sharing

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.