> It's not "cheap". 6 months of S3 is at around price of outright buying 4TB SSD at retail price. That before you do any IOPS to it
This is badly misunderstanding the problem: S3 is a highly available service with geographic redundancy and a huge range of integrated features. Your comparison would need to be updated to include multiple running servers in addition to redundant storage, and the software stack implementing things like immutability, not to mention all of the security features.
That’s not to say that you can’t build equivalents for the parts you use but you either need massive scale or giving up features to do that. For example, if you can tolerate bitrot or long access times if hardware fails, you can definitely get a lower cost per terabyte.
The reason why most people don’t do that, even when they have scale, is that it adds cost and risk everywhere else if you need engineering/ops people working on storage. If you’re, say, the internet archive that might make sense—it’s quite literally why your organization exists—but most other places are going to see all of the integrated features that they don’t have to build and operate paying for the difference between S3 and physical media pricing. For example, if I want to process files as soon as they’re uploaded or have immutability, I can just turn that on rather than having to build more services.
> This is badly misunderstanding the problem: S3 is a highly available service with geographic redundancy
First off, for geo-redundancy you need to pay extra, you only get region, and with recent strait crisis amazon already showed that can be hit and be unrecoverable.
And you're badly misunderstanding the point; that it is fucking expensive from every angle.
You can get geo-redundancy cheaper (their competitions prove it all the time); you can get IOPS cheaper; and with massive difference in end of the month bill.
Especially if you need a tons of bandwidth, S3 is very, very, very expensive way to get there.
It might be acceptable if you just run a website, but once you start running data-heavy operations on it, a big SSD attached to server can be order of magnitude cheaper than running it on S3 (and you can still use it to backup the data in case something fails)
PunchyHamster · · focus · HN ↗
It's not "cheap". 6 months of S3 is at around price of outright buying 4TB SSD at retail price. That before you do any IOPS to it
S3 is terrible deal on any front. it's just easy
acdha · · focus · HN ↗
This is badly misunderstanding the problem: S3 is a highly available service with geographic redundancy and a huge range of integrated features. Your comparison would need to be updated to include multiple running servers in addition to redundant storage, and the software stack implementing things like immutability, not to mention all of the security features.
That’s not to say that you can’t build equivalents for the parts you use but you either need massive scale or giving up features to do that. For example, if you can tolerate bitrot or long access times if hardware fails, you can definitely get a lower cost per terabyte.
The reason why most people don’t do that, even when they have scale, is that it adds cost and risk everywhere else if you need engineering/ops people working on storage. If you’re, say, the internet archive that might make sense—it’s quite literally why your organization exists—but most other places are going to see all of the integrated features that they don’t have to build and operate paying for the difference between S3 and physical media pricing. For example, if I want to process files as soon as they’re uploaded or have immutability, I can just turn that on rather than having to build more services.
PunchyHamster · · focus · HN ↗
First off, for geo-redundancy you need to pay extra, you only get region, and with recent strait crisis amazon already showed that can be hit and be unrecoverable.
And you're badly misunderstanding the point; that it is fucking expensive from every angle.
You can get geo-redundancy cheaper (their competitions prove it all the time); you can get IOPS cheaper; and with massive difference in end of the month bill.
Especially if you need a tons of bandwidth, S3 is very, very, very expensive way to get there.
It might be acceptable if you just run a website, but once you start running data-heavy operations on it, a big SSD attached to server can be order of magnitude cheaper than running it on S3 (and you can still use it to backup the data in case something fails)