‹ BackHN Continuity

Thread

Introducing System One Models and Jev

1989 points · 520 comments · albelfio

  1. mortsnort · · focus · HN ↗
    I am confused why they say it is not an LLM and then in the documentation it is shown as being an LLM derivative. The documentation makes it sound like they're taking a pretrained LLM and then giving it their unique post-training. How is that not an LLM?

    FAQ: Is Jev just a smaller LLM?

    Jev is neither small nor an LLM, hence being off the intelligence Pareto curve.

    Image in documentation: <a href="https:&#x2F;&#x2F;mintcdn.com&#x2F;ts-docs&#x2F;aFVnpmCIX68NpsV1&#x2F;images&#x2F;ai-primer&#x2F;training-paths-dark.webp?fit=max&amp;auto=format&amp;n=aFVnpmCIX68NpsV1&amp;q=85&amp;s=2747633edb0e54fa3f14a8aba830f4fd" rel="nofollow">https:&#x2F;&#x2F;mintcdn.com&#x2F;ts-docs&#x2F;aFVnpmCIX68NpsV1&#x2F;images&#x2F;ai-prime...

    1. riknos314 · · focus · HN ↗
      LLM seems to have become synonymous with Generative Transformer architecture.

      While this model may share much with GPT-style models on the encoder side, it clearly has a different decoder architecture. So is a high-parameter count language model an LLM even when it doesn&#x27;t have a GPT-style decoder? The definitions are in flux.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.