I don't see how anyone can be using Claude with prices like this, it's pretty incredible what the OpenAI team is doing, w.r.t model quality and pricing.
Prices per 1M tokens Claude Opus 5.5 Claude Opus 5
Cache reads $0.20 $0.50
Input tokens $4 $5
Output tokens $20 $25
Cache writes $5 $6.25
You're building your livelihood/workflows on a set of inputs that you have no idea what they actually cost or how reliable they'll be when the VC cash stops flowing. If you're OK with that, do your thing but it seems a little foolish to me.
I wouldn't bet on hardware flooding the market. I bet the machines running in the data centers don't use traditional PCIe connectors and cards. Maybe somebody could pull the chips and put them on standardized PCIe cards, but that is not a given.
V100s are three generations behind current and missing many of the features that modern inference benefits from, but they are the cheapest way to get a 32GB gpu.
It seems silly to say we have no idea when we actually do, though. We know how much hardware costs, we know how to reliably run a webservice that hits an API hosted on a machine with a GPU, we know how to operate these things at scale outside of OpenAI and Anthropic (not Nvidia). VC money can be patient, Uber's profitable, yeah $1 Uber rides got us hooked and they're running the same playbook. Unfortunately the convenience is worth paying for, so it seems dumb to think we can control the beast or ignore it, or get everyone to agree to hold back.
Is there a world where OpenAI starts charging $2,000/month for what we previously were paying $20 for? What are we going to do? AWS could totally jack up the prices for EC2 instances as well, but we've come to rely on that as well.
Push comes to shove, OpenAI could go out of business tomorrow and I could pick up roughly where I left off for $25k, which is the cost to serve GLM 5.3 Flash on four Nvidia GB10s. Granted, if OpenAI et al go kaput all at the same time, I could probably get a whole lot more compute for a whole lot less money.
huh? i use the plans because they're cheap and i get strong models, but i could go back to deepseek flash on commodity api pricing and be just fine
I'm fairly sure most open weight model providers are serving them at a sustainable price - and I've used them enough to know that I could live with them if the big boys did a rug pull.
pookieinc · · focus · HN ↗
Input
Output
Price reduction
GPT‑6 Sol vs. GPT‑5.6 Sol
$4 → $2
$20 → $10
50% cheaper
GPT‑6 Luna vs. GPT‑5.6 Luna
$0.20 → $0.10
$1.20 → $0.50
50% cheaper
an0malous · · focus · HN ↗
solenoid0937 · · focus · HN ↗
blovescoffee · · focus · HN ↗
selectodude · · focus · HN ↗
If they’re subsidizing my usage, that’s great.
infinitezest · · focus · HN ↗
derac · · focus · HN ↗
ssl-3 · · focus · HN ↗
foepys · · focus · HN ↗
Leynos · · focus · HN ↗
External example: <a href="https://ebay.io/m/lV8UsD" rel="nofollow">https://ebay.io/m/lV8UsD
Internal example: <a href="https://ebay.io/m/z1ygRU" rel="nofollow">https://ebay.io/m/z1ygRU
V100s are three generations behind current and missing many of the features that modern inference benefits from, but they are the cheapest way to get a 32GB gpu.
fragmede · · focus · HN ↗
Is there a world where OpenAI starts charging $2,000/month for what we previously were paying $20 for? What are we going to do? AWS could totally jack up the prices for EC2 instances as well, but we've come to rely on that as well.
selectodude · · focus · HN ↗
slopinthebag · · focus · HN ↗
goosejuice · · focus · HN ↗
andybak · · focus · HN ↗
minimaxir · · focus · HN ↗