‹ BackHN Continuity

Thread

Language models for text classification: From bag-of-words to Jev

205 points · 10 comments · Anon84

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. nzoschke · · focus · HN ↗
    Great article. This matches my feelings:

    > Like ChatGPT in 2022 was exciting because it was a general-purpose chat model that could generate all kinds of texts, one of the reasons the tech community is excited about Jev is that it is the ChatGPT moment for classification, where it can cheaply classify all kinds of text inputs without having to fine-tune a custom classifier for each task.

    We&#x27;ve been comparing strategies for classifying email and Jev is looking promising. I compared some strategies here: <a href="https:&#x2F;&#x2F;housecat.com&#x2F;blog&#x2F;classifying-email" rel="nofollow">https:&#x2F;&#x2F;housecat.com&#x2F;blog&#x2F;classifying-email

  2. tomrod · · focus · HN ↗
    Dr. Raschka is wonderful. Great article.
  3. aidiscoverywire · · focus · HN ↗

    [dead]

    1. olgava · · focus · HN ↗

      [dead]

  4. Moon_Y · · focus · HN ↗

    [dead]

  5. jacek-123 · · focus · HN ↗
    Thats very nice, I always explain LLMs to ppl starting from old-school language models and then just replacing the predictor from bag-of-words to transformers to etc.) I think its a cool framing
    1. alansaber · · focus · HN ↗
      Agreed as well, this is my preferred framing, it only makes sense to get into semantic vectors once you realise the very real limits of discrete vectors and bag of words.
  6. alescalaios · · focus · HN ↗

    [dead]

  7. Topfi · · focus · HN ↗
    Solid assessment, very much what I assumed after announcement.

    The way quite a lot of brains fell out, some unreflectively quoting how this could get us to AGI, system 1, “no hallucinating”, etc, while others ignored the breadth and new data vs existing classifiers and saw no possible upside, was revealing. The game demos were especially harmful, was told repeatedly that Jev must have near instant visual input support, as few to none of the flashy Doom, Minecraft, etc. showcases explained this was using game state.

    Hype really is the worst aspect of this industry.

  8. frjj · · focus · HN ↗

    [dead]

  9. kevinwang · · focus · HN ↗
    Really nice explanation of how Jev is and isn&#x27;t special, something I was struggling to understand before this.
  10. malshe · · focus · HN ↗
    Sebastian is the author of two excellent books related to LLMs

    Build a Large Language Model (From Scratch): <a href="https:&#x2F;&#x2F;sebastianraschka.com&#x2F;llms-from-scratch&#x2F;" rel="nofollow">https:&#x2F;&#x2F;sebastianraschka.com&#x2F;llms-from-scratch&#x2F;

    Build a Reasoning Model (From Scratch): <a href="https:&#x2F;&#x2F;sebastianraschka.com&#x2F;reasoning-from-scratch&#x2F;" rel="nofollow">https:&#x2F;&#x2F;sebastianraschka.com&#x2F;reasoning-from-scratch&#x2F;

  11. aesthesia · · focus · HN ↗
    I like the way this builds up from simple models to transformers. There&#x27;s still a pretty big gap between bag-of-words and neural network models, though, and one step that helps bridge that gap is continuous bag-of-words models, where you create word embeddings and sum&#x2F;average them together for all the words in a document. You can use precomputed embeddings to improve performance for small training datasets in a way that&#x27;s analogous to fine-tuning a foundation model. This is more or less what libraries like fastText do.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.