The paper explains absolutely everything as if it was a tutorial "how to made your own modern agentic LLM". They even tell how they made their dataset. <a href="https://aleph-alpha.com/downloads/tech-report.pdf" rel="nofollow">https://aleph-alpha.com/downloads/tech-report.pdf ; It's the first time I see this level of openness.
By various forms of deduplication (exact, fuzzy, substring), heuristic filters and distilling quality classifiers that annotate our data. Synthetic rephrases can also be considered a form of cleaning/getting more out of existing noisy data.
We have a lot of details in the tech report if you want to go deeper.
Hey! I’m curious if you tried comparing luxical to model2vec classifiers for the pretraining.
I’m one of the authors of model2vec, and working on training classifiers for this. I think model2vec could be better, but I haven’t had the opportunity to try this at scale. So if you did, knowing about it would be helpful!
Are there plans to make much larger versions of this model? With 500B-1T params for general purpose knowledge tasks, similar to the current leading proprietary models?
miellaby · · focus · HN ↗
ivo-42 · · focus · HN ↗
brcmthrowaway · · focus · HN ↗
ivo-42 · · focus · HN ↗
We have a lot of details in the tech report if you want to go deeper.
stephantul · · focus · HN ↗
I’m one of the authors of model2vec, and working on training classifiers for this. I think model2vec could be better, but I haven’t had the opportunity to try this at scale. So if you did, knowing about it would be helpful!
xvfLJfx9 · · focus · HN ↗
idiotsecant · · focus · HN ↗
ivo-42 · · focus · HN ↗