‹ BackHN Continuity

Thread

MicroLLM Lab – Try 7 tiny LLM's in the browser

283 points · 113 comments · logicallee

  1. tolugenius · · focus · HN ↗
    I did the default arithmetic with PetitGPT research-v1

    >What is 2+2?

    Answer

    > To find 2 + 2, we need to add 2 to both sides of the equation.

    > 2 + 2 = 4

    > So, 2 + 2 = 4 + 2.

    Brilliant

    1. dotancohen · · focus · HN ↗
      That's not incorrect.

      LLMs produce semantically correct sentences, not factually correct statements. Have we forgotten this so soon?

      1. anyfoo · · focus · HN ↗
        It is, in every sense, incorrect. Which statement in this short snippet is "semantically correct"? (Better LLMs get this right, of course.)
      2. NicuCalcea · · focus · HN ↗
        It's not very good semantically either.

        > Give me a recipe for soup.

        > Here is a recipe for soup:

        > Saffa-Cake-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-

    2. logicallee · · focus · HN ↗
      I got the correct output for PetitGPT research-v1: <a href="https:&#x2F;&#x2F;ibb.co&#x2F;0pP9DS2T" rel="nofollow">https:&#x2F;&#x2F;ibb.co&#x2F;0pP9DS2T
    3. tecleandor · · focus · HN ↗
      GPT-2 124M is terrible...

        &gt; what is 2+2?
      
      Answer:

        &gt; 3+3? 4+4? 5+6? 7+8?
        &gt; Reply ~18000 0 ~10 min 2 By : 1-1: I&#x27;m a beginner. 3x2 is my best option, but if you&#x27;re not sure about the other options then just go for it and try again
      1. krackers · · focus · HN ↗
        It&#x27;s not instruct tuned looks like? It&#x27;s closer to a base model rather than a chatbot.
        1. logicallee · · focus · HN ↗
          That one is a 2019 model :) Years before the ChatGPT public preview.
      2. logicallee · · focus · HN ↗
        GPT-2 is an interesting one because it is a February 2019 model. (You can see some information about it below the card if you click on the card.)

        That was 2-3 years before the big &quot;ChatGPT moment&quot; (the highly coherent ChatGPT research preview was released in November 2022, I think it was ChatGPT 3.5). Back in 2019 the models really were not producing very coherent output. Now you can see it for yourself right in your browser :) Everything has come a really long way since then!

        1. anyfoo · · focus · HN ↗
          I&#x27;ve been playing with it for a bit, and I&#x27;m a tiny little bit surprised that they decided to continue pursuing that direction of research at all. What I&#x27;m getting from it looks like it could just be arbitrary sentence and paragraph fragments from the Internet, pasted together Markov-chain like.

          I&#x27;m not sure I would have ever believed that something useful would come out of it, yet here we are.

          1. wyrdcurt · · focus · HN ↗
            Not sure how you&#x27;re prompting it but remember that it&#x27;s not trained for chat or instruction following, it simply takes the text given to it and tries to continue it. Give it the right prompt structure, and it can (at least sometimes) output coherent completions, far more often than you&#x27;d see in a Markov-chain. Also, this version is more or less equivalent to the smallest version of GPT-2; the largest version was 1.5 billion parameters and was much more likely to generate impressive (at the time) output.

            The assumption that LLMs would always need sophisticated inputs to generate useful outputs is where the term &quot;prompt engineering&quot; came from. Now that idea is basically dead. Absolutely wild how far these models have come in less than a decade!

            1. anyfoo · · focus · HN ↗
              That is insightful, thanks. I didn’t start caring about models until very late, so I genuinely thought GPT-2 was only a single tiny model originally.

              And I now tried using it more as a “text completer”, and results are much better.

      3. kasumispencer2 · · focus · HN ↗
        Sounds about right about something that&#x27;s 124M. I trained one myself a few weeks ago and it&#x27;s about the same level of being terrible.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.