The paper explains absolutely everything as if it was a tutorial "how to made your own modern agentic LLM". They even tell how they made their dataset. <a href="https://aleph-alpha.com/downloads/tech-report.pdf" rel="nofollow">https://aleph-alpha.com/downloads/tech-report.pdf ; It's the first time I see this level of openness.
In my opinion not being open about which data is ingested and trained on, and trying to make that a repeatable thing for a third party, is not worth being called "open". Glad they did that.
miellaby · · focus · HN ↗
zelphirkalt · · focus · HN ↗
slow_typist · · focus · HN ↗