> It's not "cheap". 6 months of S3 is at around price of outright buying 4TB SSD at retail price. That before you do any IOPS to it
This is badly misunderstanding the problem: S3 is a highly available service with geographic redundancy and a huge range of integrated features. Your comparison would need to be updated to include multiple running servers in addition to redundant storage, and the software stack implementing things like immutability, not to mention all of the security features.
That’s not to say that you can’t build equivalents for the parts you use but you either need massive scale or giving up features to do that. For example, if you can tolerate bitrot or long access times if hardware fails, you can definitely get a lower cost per terabyte.
The reason why most people don’t do that, even when they have scale, is that it adds cost and risk everywhere else if you need engineering/ops people working on storage. If you’re, say, the internet archive that might make sense—it’s quite literally why your organization exists—but most other places are going to see all of the integrated features that they don’t have to build and operate paying for the difference between S3 and physical media pricing. For example, if I want to process files as soon as they’re uploaded or have immutability, I can just turn that on rather than having to build more services.
Try doing the math and ask why your time costs. You need to buy a lot of storage to pay for the engineering work and most places would prefer to spend that time and attention on the product rather than shaving a few percent off of the storage bill (especially since the savings will be negative for quite some time).
Again, this is completely misunderstanding the problem. Your single SSD with no redundancy, integrity, or security, or running servers competes with S3 in the same way that my backyard garden competes with local restaurants.
If you don't pay the premium for redundancy, your S3 is the same sitting duck as the single SSD in your office, albeit an expensive one. Thinking AWS and therefore integrity and security is magical thinking.
Ask those who lost their data in Bahrain AWS.
But you are entitled to your expensive opinion. Some of us just don't buy it is all.
No, that was literally the first paragraph of my original comment.
> If you don't pay the premium for redundancy, your S3 is the same sitting duck as the single SSD in your office, albeit an expensive one.
Again, this is misunderstanding what the service provides relative to raw storage. I suggest reading the documentation about reliability and thinking about how that SSD in your office fails to provide either geographic redundancy (which all but the one-zone flavor of S3 has) or bitrot protection. Similarly, while it is true that some security concerns still require you to configure things correctly, you should think about the infrastructure you would need to build to match the core service: none of that is something you can’t match but building and operating that has costs.
> Ask those who lost their data in Bahrain AWS.
Let’s think about this slightly longer: how is the SSD in your Bahrain office doing? Iran has to hit multiple AWS data centers to lose your S3 data so to match that you need to build a system which syncs changes at multiple separate campuses in the region, which is significantly more expensive, and when you realize that you probably want a copy somewhere far away, the work to build that will be more than checking a box — and don’t forget that you have to regularly test all of this to make sure you can trust it to work in a disaster.
Again, as readers of the original comment know, I’m not saying just you can’t do this but that it’s a category error to think the comparison is the raw hardware and not the combination of redundant hardware, software, and staffing you need to match capabilities. You can beat that with enough scale but that’s a very different equation.
The original comment by PunchyHamster was about cost, i.e how AWS S3 base tier costs cannot be justified if you purchase your own SSD.
You shifted the conversation to redundancy and how AWS provides bells and whistles. You are consistent in your position, just not with me and not with PunchyHamster because we both are comparing costs by self hosting.
And for those of you who are willing to pay the premium and don't care about costs, like you, would obviously find AWS attractive.
For a startup like ours and many others, cost is everything. And spending a day or two for bitrot protection and having a couple of SSDs for redundancy and scalability is a no brainer.
You don't need big buildings for resilience. I suggest building things that require real scaling and resilience and you would know that it is about being distributed, having low cost, and owning your systems.
Your argument that it is expensive and tedious to maintain compared to AWS S3 is false. No it isn't.
Fear mongering only works against those who haven't built reliable systems. Building an S3 equivalent with cheap providers for redundancy or having one's own small resilient distributed hardware with staffing most often dwarfs the horribly insane premium paid for AWS S3 services.
But this is only if you want the redundancy which is a different topic altogether.
If all you want is a single replica, which was what PunchyHamster originally commented on, the base price of what AWS S3 provides, which does not include redundancy, makes even lesser sense compared to having your own SSD.
PunchyHamster · · focus · HN ↗
It's not "cheap". 6 months of S3 is at around price of outright buying 4TB SSD at retail price. That before you do any IOPS to it
S3 is terrible deal on any front. it's just easy
acdha · · focus · HN ↗
This is badly misunderstanding the problem: S3 is a highly available service with geographic redundancy and a huge range of integrated features. Your comparison would need to be updated to include multiple running servers in addition to redundant storage, and the software stack implementing things like immutability, not to mention all of the security features.
That’s not to say that you can’t build equivalents for the parts you use but you either need massive scale or giving up features to do that. For example, if you can tolerate bitrot or long access times if hardware fails, you can definitely get a lower cost per terabyte.
The reason why most people don’t do that, even when they have scale, is that it adds cost and risk everywhere else if you need engineering/ops people working on storage. If you’re, say, the internet archive that might make sense—it’s quite literally why your organization exists—but most other places are going to see all of the integrated features that they don’t have to build and operate paying for the difference between S3 and physical media pricing. For example, if I want to process files as soon as they’re uploaded or have immutability, I can just turn that on rather than having to build more services.
ehe78qhe · · focus · HN ↗
acdha · · focus · HN ↗
elendilm · · focus · HN ↗
You just plug it in and basically it runs. Thats pretty much it.
Having to explain to an HN audience how installing an SSD is trivial is weird.
acdha · · focus · HN ↗
elendilm · · focus · HN ↗
If you don't pay the premium for redundancy, your S3 is the same sitting duck as the single SSD in your office, albeit an expensive one. Thinking AWS and therefore integrity and security is magical thinking.
Ask those who lost their data in Bahrain AWS.
But you are entitled to your expensive opinion. Some of us just don't buy it is all.
acdha · · focus · HN ↗
No, that was literally the first paragraph of my original comment.
> If you don't pay the premium for redundancy, your S3 is the same sitting duck as the single SSD in your office, albeit an expensive one.
Again, this is misunderstanding what the service provides relative to raw storage. I suggest reading the documentation about reliability and thinking about how that SSD in your office fails to provide either geographic redundancy (which all but the one-zone flavor of S3 has) or bitrot protection. Similarly, while it is true that some security concerns still require you to configure things correctly, you should think about the infrastructure you would need to build to match the core service: none of that is something you can’t match but building and operating that has costs.
> Ask those who lost their data in Bahrain AWS.
Let’s think about this slightly longer: how is the SSD in your Bahrain office doing? Iran has to hit multiple AWS data centers to lose your S3 data so to match that you need to build a system which syncs changes at multiple separate campuses in the region, which is significantly more expensive, and when you realize that you probably want a copy somewhere far away, the work to build that will be more than checking a box — and don’t forget that you have to regularly test all of this to make sure you can trust it to work in a disaster.
Again, as readers of the original comment know, I’m not saying just you can’t do this but that it’s a category error to think the comparison is the raw hardware and not the combination of redundant hardware, software, and staffing you need to match capabilities. You can beat that with enough scale but that’s a very different equation.
elendilm · · focus · HN ↗
You shifted the conversation to redundancy and how AWS provides bells and whistles. You are consistent in your position, just not with me and not with PunchyHamster because we both are comparing costs by self hosting.
And for those of you who are willing to pay the premium and don't care about costs, like you, would obviously find AWS attractive.
For a startup like ours and many others, cost is everything. And spending a day or two for bitrot protection and having a couple of SSDs for redundancy and scalability is a no brainer.
You don't need big buildings for resilience. I suggest building things that require real scaling and resilience and you would know that it is about being distributed, having low cost, and owning your systems.
Your argument that it is expensive and tedious to maintain compared to AWS S3 is false. No it isn't.
Fear mongering only works against those who haven't built reliable systems. Building an S3 equivalent with cheap providers for redundancy or having one's own small resilient distributed hardware with staffing most often dwarfs the horribly insane premium paid for AWS S3 services.
But this is only if you want the redundancy which is a different topic altogether.
If all you want is a single replica, which was what PunchyHamster originally commented on, the base price of what AWS S3 provides, which does not include redundancy, makes even lesser sense compared to having your own SSD.