Off topic:With sol pricing drop tbh kimi k3’s value prop has not been that great. For our internal use case/testing/benchmarks sol come out with way better quality and much cheaper costs.
Kimi really needs to drop their pricing (I heard it’s set by them across all the neoclouds)
Sol is at 2/10 vs kimi’s 3/15
Agreed, I think the only place where it’s still interesting is ui design. Visually kimi and muse feel much nicer than frontier models to me, but maybe it’s an artifact of everything terrible being Claude Design
I was surprised by that. I run my benchmark [1] every couple of days and was sure this model will be ath the pareto frontier, if not THE pareto frontier. But no:
Ember isn't picked yet. In planning, Opus 5.5 wins under the planning weights. In code, GPT-6 Sol dominates it: also 10/10, but with a higher quality score and a lower estimated cost. Ember has no intelligence index, so its starting score is only 0.73, which holds its 10/10 down to 0.954 against Sol's 0.975.
Sol pricing dropped but so did the quality few days ago. I wonder when these companies are sued for making the terms from their side to go downwards while taking the same subscription cost.
Sure they are, but do you think they are continuing training to improve a model after release without bumping the version number, presumably only to game the benchmarks?
If you're willing to cheat, isn't it just a matter of grepping for the benchmark's question and pasting the solution in the chain of thought?
They're impossible to objectively track by design and intent. Intelligence is such a nebulous target that one can find any number of metrics to support any premise: that the model is smarter, and that the model is dumber. We use benchmarks to attempt to standardise comparisons but these are quickly ingested into the training data and then become effectively useless. You might have heard the term "bench-maxed." Meaning that a benchmark has an effective lifespan in months.
They've been using the Pelican test in the /r/Codex subreddit with some success. One major finding is that OpenAI has been silently degrading the model while charging Astra prices. Another finding is that even when the model has not been silently degraded, pelican quality is significantly lower. Sometimes comically so. The general consensus right now is that the new GPT-6 Sol model is an updated Terra model. Many intelligence metrics are roughly similar. Meaning the most recent model updates were an attempt to rebalance compute rather than improve intelligence.
Ultimately I've never seen users as upset about GPT-6 Sol/Luna than I have right now. Even Astra has been noticeably degraded for me and everyone else I have asked. This is compounded by the fact that Opus 5.5 is a generational improvement at an affordable price. There is currently no competition.
I am glad I am not the only one to notice. I feel like I've gone back to Sonnet 4 levels of incompetence!
With Sol 6 I am back in a world where the model writes bad code because it is lazy ("You're absolutely right, I did not [do it properly] because I did not want to edit [a normal amount of files]").
- it actually failed to correctly understand a simple English grammar and logical implication of it, then when challenged it admitted its mistake but couldn't explain why it made it.
- for the code I am working on, I asked to create two PRs for the two small features (couple lines of code). It created one in upstream, as intended, and other one in my own fork. Just like that, out of nowhere, and called the job done.
- it said it would ask me to approve/ammend the suggested PR message, it never did and fired off right away
- it keeps forgetting the changes it did itself; no context compaction was used
- it said it tested the change visually, but it did not even try
On top of that, it ignores all of my AGENTS.md, which is short and concise. I mean I point it at ignoring it, it acknowledges and ignores again.
This is astonishingly bad and it is nowhere close to Sol 5.6, or even DeepSeek 4.1! I swear even Gemini 3.8 is slightly better.
To me, Sol6 is what Opus5 was for Claude.
I keep saying that with self-hosting, you at least know what to expect and don't have to trust they nerf their models as they go. I was a skeptic and considered nerfing a conspiracy theory, but at this point with enough experience, I have experienced enough to fully see this being a thing.
Probably only a matter of time before some class action happens.
100%, I wish for a legislation which would require the providers to give you at least a unique hash identifying the model (and infra running it, if it affects output) - such that the same hash must give the same output given the same seed. Right now it's all just vibes
netvarun · · focus · HN ↗
drob518 · · focus · HN ↗
segmondy · · focus · HN ↗
drob518 · · focus · HN ↗
nostrebored · · focus · HN ↗
copperx · · focus · HN ↗
pornel · · focus · HN ↗
alansaber · · focus · HN ↗
7777777phil · · focus · HN ↗
Ember isn't picked yet. In planning, Opus 5.5 wins under the planning weights. In code, GPT-6 Sol dominates it: also 10/10, but with a higher quality score and a lower estimated cost. Ember has no intelligence index, so its starting score is only 0.73, which holds its 10/10 down to 0.954 against Sol's 0.975.
[1] <a href="https://philippdubach.com/posts/jev-model-router-for-pi/" rel="nofollow">https://philippdubach.com/posts/jev-model-router-for-pi/
toasty228 · · focus · HN ↗
6 or 5.6? Because 6 is hot garbage
nicce · · focus · HN ↗
solarkraft · · focus · HN ↗
nicce · · focus · HN ↗
pupppet · · focus · HN ↗
bpavuk · · focus · HN ↗
—"Benchmarks!"
...I'll tell that they can be gamed so easily, and they are on a consistent basis.
solarkraft · · focus · HN ↗
nananana9 · · focus · HN ↗
conorcleary · · focus · HN ↗
copperx · · focus · HN ↗
hn8726 · · focus · HN ↗
copperx · · focus · HN ↗
Gareth321 · · focus · HN ↗
They've been using the Pelican test in the /r/Codex subreddit with some success. One major finding is that OpenAI has been silently degrading the model while charging Astra prices. Another finding is that even when the model has not been silently degraded, pelican quality is significantly lower. Sometimes comically so. The general consensus right now is that the new GPT-6 Sol model is an updated Terra model. Many intelligence metrics are roughly similar. Meaning the most recent model updates were an attempt to rebalance compute rather than improve intelligence.
Ultimately I've never seen users as upset about GPT-6 Sol/Luna than I have right now. Even Astra has been noticeably degraded for me and everyone else I have asked. This is compounded by the fact that Opus 5.5 is a generational improvement at an affordable price. There is currently no competition.
vlyan · · focus · HN ↗
<a href="https://arxiv.org/pdf/2307.09009" rel="nofollow">https://arxiv.org/pdf/2307.09009
the accusations are quite a few because people notice.
[deleted] · · focus · HN ↗
[deleted]
koyote · · focus · HN ↗
With Sol 6 I am back in a world where the model writes bad code because it is lazy ("You're absolutely right, I did not [do it properly] because I did not want to edit [a normal amount of files]").
cromka · · focus · HN ↗
- it actually failed to correctly understand a simple English grammar and logical implication of it, then when challenged it admitted its mistake but couldn't explain why it made it.
- for the code I am working on, I asked to create two PRs for the two small features (couple lines of code). It created one in upstream, as intended, and other one in my own fork. Just like that, out of nowhere, and called the job done.
- it said it would ask me to approve/ammend the suggested PR message, it never did and fired off right away
- it keeps forgetting the changes it did itself; no context compaction was used
- it said it tested the change visually, but it did not even try
On top of that, it ignores all of my AGENTS.md, which is short and concise. I mean I point it at ignoring it, it acknowledges and ignores again.
This is astonishingly bad and it is nowhere close to Sol 5.6, or even DeepSeek 4.1! I swear even Gemini 3.8 is slightly better.
To me, Sol6 is what Opus5 was for Claude.
I keep saying that with self-hosting, you at least know what to expect and don't have to trust they nerf their models as they go. I was a skeptic and considered nerfing a conspiracy theory, but at this point with enough experience, I have experienced enough to fully see this being a thing.
Probably only a matter of time before some class action happens.
hn8726 · · focus · HN ↗
k__ · · focus · HN ↗
conception · · focus · HN ↗