The finding that flips the usual story is the sparse-MoE reversal: on gpt-oss-120b the A100 comes in cheaper per output token than the H100, because the workload is latency-tolerant and the MXFP4 experts fit in 80GB. Chip age stops predicting cost once you decouple "newest silicon" from "cheapest inference" and let price-elastic work route to whatever's efficient.
The point about no salvage value: the paper is really pricing a rental service, not the devices, and even there the 5-year A100 term holds 80% of the one-month mark. The reason there's no clean resale market yet isn't that old GPUs stop earning. It's that the workloads keeping them earning (batch, RL rollouts, long-horizon agents) are the ones that tolerate latency, and the people running those tend to rent rather than buy secondhand.
The part the paper understates: rented A100s aren't the biggest pool of already-paid-for, latency-tolerant capacity. That would be the Apple Silicon and flagship phones sitting idle, where the capex is fully sunk and there's no hourly meter at all. Same economics, one step further. The deciding factor is who gets to make the deployment decision, and open weights are what move that decision to the operator.
itsmeduncan · · focus · HN ↗
The point about no salvage value: the paper is really pricing a rental service, not the devices, and even there the 5-year A100 term holds 80% of the one-month mark. The reason there's no clean resale market yet isn't that old GPUs stop earning. It's that the workloads keeping them earning (batch, RL rollouts, long-horizon agents) are the ones that tolerate latency, and the people running those tend to rent rather than buy secondhand.
The part the paper understates: rented A100s aren't the biggest pool of already-paid-for, latency-tolerant capacity. That would be the Apple Silicon and flagship phones sitting idle, where the capex is fully sunk and there's no hourly meter at all. Same economics, one step further. The deciding factor is who gets to make the deployment decision, and open weights are what move that decision to the operator.