The paper explains absolutely everything as if it was a tutorial "how to made your own modern agentic LLM". They even tell how they made their dataset. <a href="https://aleph-alpha.com/downloads/tech-report.pdf" rel="nofollow">https://aleph-alpha.com/downloads/tech-report.pdf ; It's the first time I see this level of openness.
By various forms of deduplication (exact, fuzzy, substring), heuristic filters and distilling quality classifiers that annotate our data. Synthetic rephrases can also be considered a form of cleaning/getting more out of existing noisy data.
We have a lot of details in the tech report if you want to go deeper.
Are there plans to make much larger versions of this model? With 500B-1T params for general purpose knowledge tasks, similar to the current leading proprietary models?
miellaby · · focus · HN ↗
ivo-42 · · focus · HN ↗
brcmthrowaway · · focus · HN ↗
ivo-42 · · focus · HN ↗
We have a lot of details in the tech report if you want to go deeper.
xvfLJfx9 · · focus · HN ↗
ivo-42 · · focus · HN ↗