> All this must mean the Western AI companies are now extremely inference-margin positive.
> So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.
That inference wasn't profitable is a widespread myth.
Analysis based on Kimi K3 suggests that OpenAI and Anthropic have margins well north of 95%: <a href="https://inferencex.semianalysis.com/run/kimi-k3-on-b200" rel="nofollow">https://inferencex.semianalysis.com/run/kimi-k3-on-b200
Over the last months I have seen news that OpenAI made breakthroughs in inference efficiency multiple times.
I have no reason to believe that the leading US labs don't have their own optimizations, or that they learned of this particular optimization from DeepSeek.
95% margin is really unlikely. Anthropic recently said they have 80% gross margin when using their adjusted ebidta (ie if they do not consider revenue sharing, training expenses, and a bunch of other costs). They wouldn’t be talking about non standard metrics if they had such high margin on inference
user43928 · · focus · HN ↗
> So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why.
That inference wasn't profitable is a widespread myth.
Analysis based on Kimi K3 suggests that OpenAI and Anthropic have margins well north of 95%: <a href="https://inferencex.semianalysis.com/run/kimi-k3-on-b200" rel="nofollow">https://inferencex.semianalysis.com/run/kimi-k3-on-b200
Over the last months I have seen news that OpenAI made breakthroughs in inference efficiency multiple times.
I have no reason to believe that the leading US labs don't have their own optimizations, or that they learned of this particular optimization from DeepSeek.
dgellow · · focus · HN ↗
user43928 · · focus · HN ↗
Cost per GPU hour versus API price of generated tokens assuming 100% utilization.
This could be a margin around 98.3% for 5.6 Sol.
If the utilization of the GPU was 25%, it would drop to 93.1%.
Revenue sharing or training expenses are not considered here in this "inference margin".