Off topic:With sol pricing drop tbh kimi k3’s value prop has not been that great. For our internal use case/testing/benchmarks sol come out with way better quality and much cheaper costs.
Kimi really needs to drop their pricing (I heard it’s set by them across all the neoclouds)
Sol is at 2/10 vs kimi’s 3/15
Sol pricing dropped but so did the quality few days ago. I wonder when these companies are sued for making the terms from their side to go downwards while taking the same subscription cost.
Sure they are, but do you think they are continuing training to improve a model after release without bumping the version number, presumably only to game the benchmarks?
If you're willing to cheat, isn't it just a matter of grepping for the benchmark's question and pasting the solution in the chain of thought?
They're impossible to objectively track by design and intent. Intelligence is such a nebulous target that one can find any number of metrics to support any premise: that the model is smarter, and that the model is dumber. We use benchmarks to attempt to standardise comparisons but these are quickly ingested into the training data and then become effectively useless. You might have heard the term "bench-maxed." Meaning that a benchmark has an effective lifespan in months.
They've been using the Pelican test in the /r/Codex subreddit with some success. One major finding is that OpenAI has been silently degrading the model while charging Astra prices. Another finding is that even when the model has not been silently degraded, pelican quality is significantly lower. Sometimes comically so. The general consensus right now is that the new GPT-6 Sol model is an updated Terra model. Many intelligence metrics are roughly similar. Meaning the most recent model updates were an attempt to rebalance compute rather than improve intelligence.
Ultimately I've never seen users as upset about GPT-6 Sol/Luna than I have right now. Even Astra has been noticeably degraded for me and everyone else I have asked. This is compounded by the fact that Opus 5.5 is a generational improvement at an affordable price. There is currently no competition.
netvarun · · focus · HN ↗
nicce · · focus · HN ↗
solarkraft · · focus · HN ↗
nicce · · focus · HN ↗
pupppet · · focus · HN ↗
bpavuk · · focus · HN ↗
—"Benchmarks!"
...I'll tell that they can be gamed so easily, and they are on a consistent basis.
solarkraft · · focus · HN ↗
nananana9 · · focus · HN ↗
conorcleary · · focus · HN ↗
copperx · · focus · HN ↗
hn8726 · · focus · HN ↗
copperx · · focus · HN ↗
Gareth321 · · focus · HN ↗
They've been using the Pelican test in the /r/Codex subreddit with some success. One major finding is that OpenAI has been silently degrading the model while charging Astra prices. Another finding is that even when the model has not been silently degraded, pelican quality is significantly lower. Sometimes comically so. The general consensus right now is that the new GPT-6 Sol model is an updated Terra model. Many intelligence metrics are roughly similar. Meaning the most recent model updates were an attempt to rebalance compute rather than improve intelligence.
Ultimately I've never seen users as upset about GPT-6 Sol/Luna than I have right now. Even Astra has been noticeably degraded for me and everyone else I have asked. This is compounded by the fact that Opus 5.5 is a generational improvement at an affordable price. There is currently no competition.
vlyan · · focus · HN ↗
<a href="https://arxiv.org/pdf/2307.09009" rel="nofollow">https://arxiv.org/pdf/2307.09009
the accusations are quite a few because people notice.