Did private frontier models use sparse embedding and ngram first? The article claims sparse attention was copied from open weight but we can't know that. We could just as easily argue that OAI and ANT had these improvements for years and decided to slash their margins only now to stay competitive with open weight neoclouds.
Second, sparse attention is an old area of active research. Offloaded N-gram tables are the next big open weight technological leap.
The O & A strategy has until very recently been to use brute force and just throw more money at the problem.
Deepseek was the company that invented some and improved some other ideas and got them to workreliably in production. Before that Sam and Dario were basically competing in who has the most expensive training.
samuelknight · · focus · HN ↗
Second, sparse attention is an old area of active research. Offloaded N-gram tables are the next big open weight technological leap.
throwa356262 · · focus · HN ↗
Deepseek was the company that invented some and improved some other ideas and got them to workreliably in production. Before that Sam and Dario were basically competing in who has the most expensive training.