I don't see how anyone can be using Claude with prices like this, it's pretty incredible what the OpenAI team is doing, w.r.t model quality and pricing.
Prices per 1M tokens Claude Opus 5.5 Claude Opus 5
Cache reads $0.20 $0.50
Input tokens $4 $5
Output tokens $20 $25
Cache writes $5 $6.25
You're building your livelihood/workflows on a set of inputs that you have no idea what they actually cost or how reliable they'll be when the VC cash stops flowing. If you're OK with that, do your thing but it seems a little foolish to me.
I wouldn't bet on hardware flooding the market. I bet the machines running in the data centers don't use traditional PCIe connectors and cards. Maybe somebody could pull the chips and put them on standardized PCIe cards, but that is not a given.
V100s are three generations behind current and missing many of the features that modern inference benefits from, but they are the cheapest way to get a 32GB gpu.
pookieinc · · focus · HN ↗
Input
Output
Price reduction
GPT‑6 Sol vs. GPT‑5.6 Sol
$4 → $2
$20 → $10
50% cheaper
GPT‑6 Luna vs. GPT‑5.6 Luna
$0.20 → $0.10
$1.20 → $0.50
50% cheaper
an0malous · · focus · HN ↗
selectodude · · focus · HN ↗
If they’re subsidizing my usage, that’s great.
infinitezest · · focus · HN ↗
derac · · focus · HN ↗
ssl-3 · · focus · HN ↗
foepys · · focus · HN ↗
Leynos · · focus · HN ↗
External example: <a href="https://ebay.io/m/lV8UsD" rel="nofollow">https://ebay.io/m/lV8UsD
Internal example: <a href="https://ebay.io/m/z1ygRU" rel="nofollow">https://ebay.io/m/z1ygRU
V100s are three generations behind current and missing many of the features that modern inference benefits from, but they are the cheapest way to get a 32GB gpu.