The paper explains absolutely everything as if it was a tutorial "how to made your own modern agentic LLM". They even tell how they made their dataset. <a href="https://aleph-alpha.com/downloads/tech-report.pdf" rel="nofollow">https://aleph-alpha.com/downloads/tech-report.pdf ; It's the first time I see this level of openness.
In my opinion not being open about which data is ingested and trained on, and trying to make that a repeatable thing for a third party, is not worth being called "open". Glad they did that.
to be fair, this discussion has been had numerous times here and the industry has arrived on "open-weight" to describe the practice of releasing the post-training weights in an open manner but not releasing the data it was trained on.
That's what they call this and I think it's a pretty clear definition these days to people in the industry.
miellaby · · focus · HN ↗
zelphirkalt · · focus · HN ↗
davidjfelix · · focus · HN ↗
That's what they call this and I think it's a pretty clear definition these days to people in the industry.
gewetensleegte · · focus · HN ↗
to be a bit more exact