Notwithstanding the duplication of these posts across social media, OpenAI, and Anthropic ("it's all the model, they just tell it to keep going" if you beat anything with RL enough ... it's going to do the thing)
Here's my issue with this post:
> True to the spirit of the challenge, they didn’t use millions of dollars in computer power. They used Fable 5.1, working within Claude Science, a platform scientists can pay to use.
Okay, billions of dollars have been poured into these agentic LMs, right? Each training run to get the next increment is costing millions of dollars?
This feels like an obvious jab at Navier-Stokes, but where we get to shift the numbers around to hide where the compute actually is being spent ... compute is being spent. It's either being spent in amortization to make the search smarter ahead of time, during training, or its being spent after.
Also love: scientists get to pay Anthropic to work within their special science harness to do science. That's exactly what I dreamed of doing when I pursued physics in undergrad, one or two companies holding the keys to "progress" for a monthly subscription price.
The training costs get rolled into the usage prices. It's not hiding anything to talk about the cost of usage without the cost of training, any more than I'd be hiding something by talking about the price of a $1,000 CPU without mentioning the tens of billions of dollars in R&D and infrastructure needed to create it.
> Also love: scientists get to pay Anthropic to work within their special science harness to do science. That's exactly what I dreamed of doing when I pursued physics in undergrad, one or two companies holding the keys to "progress" for a monthly subscription price.
I get this anxiety, and am largely an AI skeptic, but at one point there were only a handful of computers in the world too (same for batteries, or engines, or crucibles, or stills -- it goes way back), and the organizations that had them had a stranglehold on progress in the field, as did the small number of companies who knew how to make them. It got better as they got cheaper and more plentiful.
I guess my point is that there are more important anxieties to feed when it comes to LLMs and the current state of the world.
> This feels like an obvious jab at Navier-Stokes, but where we get to shift the numbers around to hide where the compute actually is being spent ... compute is being spent. It's either being spent in amortization to make the search smarter ahead of time, during training, or its being spent after.
I think that argument is recursive? These posts aren't very complicated for either of us, but they're written on devices that are fabricated with billions of dollars of semiconductor equipment. At what point do we just acknowledge that we stand on the shoulders of giants?
To me, the distinguishing factor is that the expense not special-purpose but upfront. The model here is trained without foreknowledge of what problems it will solve. Solutions like nine loops are genuine expressions of a pre-existing model capability, even if that capability has not pre-existed for very long.
Oh I agree with you! I just don't see "we're standing on the shoulders of giants" in most of these marketing blog posts?
If I'm wrong here, I'd love reference links. I think of these companies as trying to inspire the idea that Claude (or GPT) are these special alien entities, in a sense?
What a weird thing to be butthurt about. You don't have to pay them. Do it the old fashioned way, with elbow grease and pots of coffee. Or invest in local AI.
Or embrace the future and realize that you couldn't imagine everything that would unfold, when you pursued your undergrad.
There was a time when a generation of hackers got (rightfully) worked up about the Microsoft tax for every PC sold. It's not hard to see that somebody gets mad if they believe that an Anthropic/OpenAI tax is about to become the standard when doing any serious Physics/Math/... research or writing code.
The term "hackers" didn't just apply to those of us who wiped the thing off first chance they got. Folks working in IT of large corps which had fleets of a 5-digit numbers of PCs to administer and for each of them, Microsoft took their cut.
Just because a description doesn't apply to you personally doesn't mean it's misrepresenting things.
Divide the compute of your argument (pre-training, training, post-training) through everything this model now can do and it will not look that bad at all.
> Also love: scientists get to pay Anthropic to work within their special science harness to do science. That's exactly what I dreamed of doing when I pursued physics in undergrad, one or two companies holding the keys to "progress" for a monthly subscription price.
The moat is very limited. Harnesses aren't crazy hard to engineer. Open models are quite capable.
One of the most interesting aspects of MiMo 2.6 is that they shared their RL costs, sub $5M (yes, million)
This stuff is going to get a lot cheaper, just like Stable Diffusion
I for one find the incessant side questing the moment something goes awry to be very annoying. Please stop and ask the human for clarification or guidence
mccoyb · · focus · HN ↗
Here's my issue with this post:
> True to the spirit of the challenge, they didn’t use millions of dollars in computer power. They used Fable 5.1, working within Claude Science, a platform scientists can pay to use.
Okay, billions of dollars have been poured into these agentic LMs, right? Each training run to get the next increment is costing millions of dollars?
This feels like an obvious jab at Navier-Stokes, but where we get to shift the numbers around to hide where the compute actually is being spent ... compute is being spent. It's either being spent in amortization to make the search smarter ahead of time, during training, or its being spent after.
Also love: scientists get to pay Anthropic to work within their special science harness to do science. That's exactly what I dreamed of doing when I pursued physics in undergrad, one or two companies holding the keys to "progress" for a monthly subscription price.
wat10000 · · focus · HN ↗
teiferer · · focus · HN ↗
[dead]
ElevenLathe · · focus · HN ↗
I get this anxiety, and am largely an AI skeptic, but at one point there were only a handful of computers in the world too (same for batteries, or engines, or crucibles, or stills -- it goes way back), and the organizations that had them had a stranglehold on progress in the field, as did the small number of companies who knew how to make them. It got better as they got cheaper and more plentiful.
I guess my point is that there are more important anxieties to feed when it comes to LLMs and the current state of the world.
teiferer · · focus · HN ↗
I don't see how that makes the state of affairs any better.
badrequest · · focus · HN ↗
teiferer · · focus · HN ↗
How many search engines and video platforms and browsers are people using exactly? I can count them on one hand.
Majromax · · focus · HN ↗
I think that argument is recursive? These posts aren't very complicated for either of us, but they're written on devices that are fabricated with billions of dollars of semiconductor equipment. At what point do we just acknowledge that we stand on the shoulders of giants?
To me, the distinguishing factor is that the expense not special-purpose but upfront. The model here is trained without foreknowledge of what problems it will solve. Solutions like nine loops are genuine expressions of a pre-existing model capability, even if that capability has not pre-existed for very long.
mccoyb · · focus · HN ↗
If I'm wrong here, I'd love reference links. I think of these companies as trying to inspire the idea that Claude (or GPT) are these special alien entities, in a sense?
sejje · · focus · HN ↗
Or embrace the future and realize that you couldn't imagine everything that would unfold, when you pursued your undergrad.
How you gonna get mad about all this?
teiferer · · focus · HN ↗
sejje · · focus · HN ↗
I was mad that I had to pay for windows just to wipe that shit and put linux on. If I used it, I wouldn't have been mad about it at all.
I've never been mad about paying for my software.
teiferer · · focus · HN ↗
Just because a description doesn't apply to you personally doesn't mean it's misrepresenting things.
sejje · · focus · HN ↗
teiferer · · focus · HN ↗
sejje · · focus · HN ↗
I guess you're mad about the quasi-monopoly of Tesla, too, since they're the only electric car selling worth a hoot.
Glemmlko · · focus · HN ↗
tomrod · · focus · HN ↗
The moat is very limited. Harnesses aren't crazy hard to engineer. Open models are quite capable.
mccoyb · · focus · HN ↗
verdverm · · focus · HN ↗
This stuff is going to get a lot cheaper, just like Stable Diffusion
I for one find the incessant side questing the moment something goes awry to be very annoying. Please stop and ask the human for clarification or guidence