You can see the Pareto Frontier well in DeepSWE's chart here - <a href="https://deepswe.datacurve.ai/">https://deepswe.datacurve.ai/
ChatGPT 5.6 Luna on the right (cheaper) cover most of the frontier, with a point for Deepseek flash, and higher performance overlapping heavily between 5.6 Sol and Fable.
That DeepSeek point will probably move back towards Luna as deepseek announced a "significant" price increase coming to their API [1], which kind of demonstrates that beating the Pareto frontier is where the difficulty actually is).
I think Luna might be just small enough to provide some kind of stepwise improvement in how it is hosted.
Going from 81GB of weights to 79GB of weights can mean a 50% reduction in GPU capacity required.
If you can fit a model in just one GPU (or rack) as opposed to across an entire datacenter, the latency gains can be substantial too. If you can reduce token latency by half, that would double the amount of customers you could support.
gpt5 · · focus · HN ↗
ChatGPT 5.6 Luna on the right (cheaper) cover most of the frontier, with a point for Deepseek flash, and higher performance overlapping heavily between 5.6 Sol and Fable.
That DeepSeek point will probably move back towards Luna as deepseek announced a "significant" price increase coming to their API [1], which kind of demonstrates that beating the Pareto frontier is where the difficulty actually is).
[1] <a href="https://www.bloomberg.com/news/articles/2026-08-06/deepseek-plans-significant-price-increase-for-its-ai-services" rel="nofollow">https://www.bloomberg.com/news/articles/2026-08-06/deepseek-...
kooi · · focus · HN ↗
I think it's great and hope the price can stay the same.
bob1029 · · focus · HN ↗
Going from 81GB of weights to 79GB of weights can mean a 50% reduction in GPU capacity required.
If you can fit a model in just one GPU (or rack) as opposed to across an entire datacenter, the latency gains can be substantial too. If you can reduce token latency by half, that would double the amount of customers you could support.