The Economics of Open-Weight Inference
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
The Economics of Open-Weight Inference
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
jcmontx · · focus · HN ↗
Silagi · · focus · HN ↗
I've taken to using them as micro-review subagents at development milestones, where a "frontier - 1" model like Opus or Sol launches 10-15 of them on small review tasks that each run for ~10 minutes. Costs about $1 per cycle, and they usually catch something Astra or Fable didn't. Then the orchestrator validates each claim before passing it back to the planning session so we can fold the findings in.
physicsguy · · focus · HN ↗
datsci_est_2015 · · focus · HN ↗
And once frontier models can architect then I guess no one needs a job because that’s the digital singularity.
lowbloodsugar · · focus · HN ↗
orsorna · · focus · HN ↗
dig1 · · focus · HN ↗
Perhaps they simply don’t advertise it and investors (currently) love companies that spend heavily on frontier models. That said, OW models, especially when combined with RAG, work quite well, and given the current state of the industry, they may be the only sensible way to keep inference costs under control.
gyyu · · focus · HN ↗
Why do you think OAI etc are all up in arms? They hate what’s going on. Most here are delusional.
I work in a very large market cap firm and i’m telling you - more and more managers are pushed to use open source and squeeze employees to get the max out of them.
codybontecou · · focus · HN ↗
barbazoo · · focus · HN ↗
whatshisface · · focus · HN ↗
perrygeo · · focus · HN ↗
Perhaps because they are dumber, they produce better results? IMO an excellent well-tuned harness combined with a "frontier minus x" model produces the highest quality result. DS4.1 and Qwen3.8, far from being a compromise, legit give me better results. For my personal definition of "better".
gyyu · · focus · HN ↗
augment_me · · focus · HN ↗
This is clearly not true, and we are starting to see compute markets, but only for rental prices/H, not for the hardware itself. I feel like there is some artificial moat being built here to stimulate sales of new hardware, because an H100 at 1/16th the price will have comparable dollar/FLOP as Vera Rubin.
linuxftw · · focus · HN ↗
kurthr · · focus · HN ↗
The claims of many GW of installed training/inference have come under scrutiny lately. The first VeraRubins aren't even there yet, so it's all GB300 NVL72s as the peak performers and probably <<1GW of those so far. Even xAI Colossus is mostly H200s and B200s.
csmoak · · focus · HN ↗
whatshisface · · focus · HN ↗
mathisfun123 · · focus · HN ↗
Which banks have analysts that understand the difference between H100 and A100? Do you have actual experience with being denied a loan based on this or are you just making things up?
simianwords · · focus · HN ↗
gyyu · · focus · HN ↗
hungryhobbit · · focus · HN ↗
HarHarVeryFunny · · focus · HN ↗
fred_is_fred · · focus · HN ↗
itsmeduncan · · focus · HN ↗
[dead]
nylonstrung · · focus · HN ↗
But I'd still rather use them since it's inevitable the unsustainable margins of the closed weight providers lead to enshittification once they stop subsidizing demand
I'd rather spend more today with a workflow whose underlying unit economics are sustainable and don't force me to inevitably look for other solutions once the closed weight providers start optimizing for profit and extracting value
cmiles8 · · focus · HN ↗