Ever tried Ceph? It's a bit painful to set up and operate, but it seems to work pretty well.
(Don't bother with Ceph Object Gateway unless you really really need it - since you're programming your own application, access the storage pool directly with librados)
As someone with TBs of SSDs and HDDs laying around I am someone who agrees with your point, but this is not a good way to compare costs.
For one, this has no redundancy and doesn't factor the cost of the machine providing the storage. But also it's an upfront cost for storing 4TB. It would be implied you wouldn't literally store 4TB on the SSD because now you're out of capacity for more data. If you only have 2GB of data in a bucket then S3 is still magnitudes cheaper, faster, and more redundant than anything you could put together yourself because the cost is spread across all the users of the service.
Reality is there are so many times that S3 makes the most sense that it makes people short-circuit and always pick S3 despite the huge hidden IOPS and bandwidth costs everyone rightfully tries to point out.
If you have only 2GB of data it's trivial, it hardly matters how you store it. You might as well hold a full copy in RAM on each server and synchronize it with your developer laptop every half hour in case they all go offline at once, for all it matters.
To reliably hold a bit over 2GB of data in RAM on AWS, you'll probably need something like a t4g.medium, about $0.0336/hr on-demand. 730 hours in a month, so $24.528, but ideally you'll have like three or so instances so like $74/mo. That's before thinking about stuff like IP address costs, the EBS volumes underpinning those boxes, etc. And then managing all the syncing and what not.
Of course, that's list prices, you can get savings plans and RIs and discounts.
Once again proving that AWS is expensive any way you slice it. Even in today's RAM crisis, an extra 2GB is what, $50? and you probably already have 2GB free since they don't make RAM modules that small.
Just buying a few sticks of RAM isn't hosting data in that RAM indefinitely. Its not the same as what S3 (or similar services) gets you.
What is your datacenter costs for multiple different buildings (assuming colocation/office space)? How much are you going to pay for the electricity? How much are you paying for networking? How much is rent for each U of server? How much do you spend on payroll to have knowledgeable people to manage it 24 hours a day, 7 days a week? Suddenly its not just a one-time $50 payment.
Maybe you don't actually need all of that, just some slices of some of that whole stack of things. Probably true in many situations. But just looking at it as a couple of sticks of RAM is not a like-for-like comparison.
I'm not arguing the cloud is cheap. You can definitely go cheaper and choose your own cost optimizations when managing the hardware yourself, deciding what is actually a necessary cost and what isn't. I totally agree AWS overchages for a lot of what they do; they price their stuff at a premium because they know businesses will pay it. But hosting 2GB of stuff in RAM reliably and indefinitely to anywhere near the guarantees of S3 is going to be a lot more than just a single $50 payment.
Server? What server? Suddenly now I need to also buy a server? Sounds like quite a bit more than $50!
Things always seem free when you've already spent the money. It doesn't change the fact you spent quite a bit of money, you're just failing to account for it properly.
The standard here was I've just got 2GB of data I need to reliably store someplace. Its an expensive assumption I've got multiple servers in multiple physically separate buildings already lying around.
> It's not "cheap". 6 months of S3 is at around price of outright buying 4TB SSD at retail price. That before you do any IOPS to it
This is badly misunderstanding the problem: S3 is a highly available service with geographic redundancy and a huge range of integrated features. Your comparison would need to be updated to include multiple running servers in addition to redundant storage, and the software stack implementing things like immutability, not to mention all of the security features.
That’s not to say that you can’t build equivalents for the parts you use but you either need massive scale or giving up features to do that. For example, if you can tolerate bitrot or long access times if hardware fails, you can definitely get a lower cost per terabyte.
The reason why most people don’t do that, even when they have scale, is that it adds cost and risk everywhere else if you need engineering/ops people working on storage. If you’re, say, the internet archive that might make sense—it’s quite literally why your organization exists—but most other places are going to see all of the integrated features that they don’t have to build and operate paying for the difference between S3 and physical media pricing. For example, if I want to process files as soon as they’re uploaded or have immutability, I can just turn that on rather than having to build more services.
Try doing the math and ask why your time costs. You need to buy a lot of storage to pay for the engineering work and most places would prefer to spend that time and attention on the product rather than shaving a few percent off of the storage bill (especially since the savings will be negative for quite some time).
Wait, I'm confused, are you arguing for or against S3? Did you mean factoring the time needed to obtain the multiple PhD levels of information in tracking the cost, usage, and security of those interoperable services plugged into S3?
I’m saying that once you make the fully-loaded comparison, the gap closes a lot and most places will choose to spend their time and attention elsewhere unless they have a truly massive storage need.
Again, this is completely misunderstanding the problem. Your single SSD with no redundancy, integrity, or security, or running servers competes with S3 in the same way that my backyard garden competes with local restaurants.
If you don't pay the premium for redundancy, your S3 is the same sitting duck as the single SSD in your office, albeit an expensive one. Thinking AWS and therefore integrity and security is magical thinking.
Ask those who lost their data in Bahrain AWS.
But you are entitled to your expensive opinion. Some of us just don't buy it is all.
No, that was literally the first paragraph of my original comment.
> If you don't pay the premium for redundancy, your S3 is the same sitting duck as the single SSD in your office, albeit an expensive one.
Again, this is misunderstanding what the service provides relative to raw storage. I suggest reading the documentation about reliability and thinking about how that SSD in your office fails to provide either geographic redundancy (which all but the one-zone flavor of S3 has) or bitrot protection. Similarly, while it is true that some security concerns still require you to configure things correctly, you should think about the infrastructure you would need to build to match the core service: none of that is something you can’t match but building and operating that has costs.
> Ask those who lost their data in Bahrain AWS.
Let’s think about this slightly longer: how is the SSD in your Bahrain office doing? Iran has to hit multiple AWS data centers to lose your S3 data so to match that you need to build a system which syncs changes at multiple separate campuses in the region, which is significantly more expensive, and when you realize that you probably want a copy somewhere far away, the work to build that will be more than checking a box — and don’t forget that you have to regularly test all of this to make sure you can trust it to work in a disaster.
Again, as readers of the original comment know, I’m not saying just you can’t do this but that it’s a category error to think the comparison is the raw hardware and not the combination of redundant hardware, software, and staffing you need to match capabilities. You can beat that with enough scale but that’s a very different equation.
The original comment by PunchyHamster was about cost, i.e how AWS S3 base tier costs cannot be justified if you purchase your own SSD.
You shifted the conversation to redundancy and how AWS provides bells and whistles. You are consistent in your position, just not with me and not with PunchyHamster because we both are comparing costs by self hosting.
And for those of you who are willing to pay the premium and don't care about costs, like you, would obviously find AWS attractive.
For a startup like ours and many others, cost is everything. And spending a day or two for bitrot protection and having a couple of SSDs for redundancy and scalability is a no brainer.
You don't need big buildings for resilience. I suggest building things that require real scaling and resilience and you would know that it is about being distributed, having low cost, and owning your systems.
Your argument that it is expensive and tedious to maintain compared to AWS S3 is false. No it isn't.
Fear mongering only works against those who haven't built reliable systems. Building an S3 equivalent with cheap providers for redundancy or having one's own small resilient distributed hardware with staffing most often dwarfs the horribly insane premium paid for AWS S3 services.
But this is only if you want the redundancy which is a different topic altogether.
If all you want is a single replica, which was what PunchyHamster originally commented on, the base price of what AWS S3 provides, which does not include redundancy, makes even lesser sense compared to having your own SSD.
I have done the math as part of my job and often found that yes, it was literally worth the money to do our own storage in cases. Running a highly available, durable storage system is not particularly difficult. There was an era when it was a common skill for a sysadmin. In the current era there are off the shelf solutions for it, both open and proprietary.
The biggest reason people pay for S3 is because of data transfer cost to other cloud services they are using.
Yes, I’ve also built those services. As I said before, it is totally possible to beat S3 on pricing but the margin narrows considerably when you factor in the full costs — e.g. think about how many places had replicated storage but not bit-rot protection or versioning and built that at the application layer instead — and the savings is often not big enough to make it your top engineering pick for the money & attention required.
Fundamentally you’re playing a commodity game and that means you need a consistent edge to be worth the fixed costs. If your fully-loaded advantage is less than, say, 30% you’re better off negotiating with your AWS rep and investing your engineering resources in something your customers want.
Again, spec out a multi-data center server + storage + staffing deployment and calculate the full cost per gigabyte. As I said earlier, if you buy enough storage you will definitely hit a level where you can beat S3 by a fair margin but many places never reach that point and many of the ones who do are going to have other projects which make more sense because they offer a high rate of return or provide something which isn’t a commodity.
Recent Iranian attacks on AWS locations in Bahrain resulted in loss of data that simply could not be recovered. <a href="https://www.infoq.com/news/2026/09/aws-middle-east-data-loss/" rel="nofollow">https://www.infoq.com/news/2026/09/aws-middle-east-data-loss...
Of course, if you had paid extra for the backups and redundancy, your data would survive.
So to address your original point, unless you are paying extra for redundancy, it is cheaper if you use a 4TB SSD at retail price as the original commenter discussed.
In this case AWS did not save you. AWS is supposed to be resilient against AZ failures within a region. Iran destroyed all AZs at once, so all the data was lost. If you had a hard drive in your office, either directly serving your project or as an off-site backup, you still have your data. In the end, was it worth paying several times as much for AWS to provide durability for you, only to have them lose the data anyway?
it is advertised as eleven nines durability, however. That is better than anything you can build yourself, so if true, the best backup for your S3 bucket is an independent S3 bucket
Do you need that? If you want your own replicated storage cluster at home, you can use Ceph. I've seen Ceph deployed to utilize the spare hard drive ports in server clusters that otherwise mostly do compute (transcoding), alongside memcached to utilize the RAM, etc.
I'd bet the majority of projects, but perhaps not the majority of traffic, can run adequately on a single server and tolerate a few hour maintenance window one weekend every month. Don't waste money on overkill.
> This is badly misunderstanding the problem: S3 is a highly available service with geographic redundancy
First off, for geo-redundancy you need to pay extra, you only get region, and with recent strait crisis amazon already showed that can be hit and be unrecoverable.
And you're badly misunderstanding the point; that it is fucking expensive from every angle.
You can get geo-redundancy cheaper (their competitions prove it all the time); you can get IOPS cheaper; and with massive difference in end of the month bill.
Especially if you need a tons of bandwidth, S3 is very, very, very expensive way to get there.
It might be acceptable if you just run a website, but once you start running data-heavy operations on it, a big SSD attached to server can be order of magnitude cheaper than running it on S3 (and you can still use it to backup the data in case something fails)
If you play the game correctly it can be a good deal. The ~1 USD per TB per month of glacier deep archive is hard to compete with if you (probably) don’t need to read the data back.
S3 cost growth is so gradual most businesses will just absorb the cost. We have some physical backups, but everyone is more than happy to not be shuffling around physical hard drives, and pay Amazon to deal with that and securely store it.
That's AWS in general for you. It has its place, but it's vastly overused in our industry. A lot of businesses would find their TCO would go down if they ditched cloud services, but that's not trendy enough for execs to do it.
PunchyHamster · · focus · HN ↗
It's not "cheap". 6 months of S3 is at around price of outright buying 4TB SSD at retail price. That before you do any IOPS to it
S3 is terrible deal on any front. it's just easy
CodesInChaos · · focus · HN ↗
someonebaggy · · focus · HN ↗
(Don't bother with Ceph Object Gateway unless you really really need it - since you're programming your own application, access the storage pool directly with librados)
0xCMP · · focus · HN ↗
For one, this has no redundancy and doesn't factor the cost of the machine providing the storage. But also it's an upfront cost for storing 4TB. It would be implied you wouldn't literally store 4TB on the SSD because now you're out of capacity for more data. If you only have 2GB of data in a bucket then S3 is still magnitudes cheaper, faster, and more redundant than anything you could put together yourself because the cost is spread across all the users of the service.
Reality is there are so many times that S3 makes the most sense that it makes people short-circuit and always pick S3 despite the huge hidden IOPS and bandwidth costs everyone rightfully tries to point out.
someonebaggy · · focus · HN ↗
vel0city · · focus · HN ↗
Of course, that's list prices, you can get savings plans and RIs and discounts.
The cost of RAM per hour in AWS isn't cheap.
someonebaggy · · focus · HN ↗
vel0city · · focus · HN ↗
What is your datacenter costs for multiple different buildings (assuming colocation/office space)? How much are you going to pay for the electricity? How much are you paying for networking? How much is rent for each U of server? How much do you spend on payroll to have knowledgeable people to manage it 24 hours a day, 7 days a week? Suddenly its not just a one-time $50 payment.
Maybe you don't actually need all of that, just some slices of some of that whole stack of things. Probably true in many situations. But just looking at it as a couple of sticks of RAM is not a like-for-like comparison.
I'm not arguing the cloud is cheap. You can definitely go cheaper and choose your own cost optimizations when managing the hardware yourself, deciding what is actually a necessary cost and what isn't. I totally agree AWS overchages for a lot of what they do; they price their stuff at a premium because they know businesses will pay it. But hosting 2GB of stuff in RAM reliably and indefinitely to anywhere near the guarantees of S3 is going to be a lot more than just a single $50 payment.
someonebaggy · · focus · HN ↗
vel0city · · focus · HN ↗
Things always seem free when you've already spent the money. It doesn't change the fact you spent quite a bit of money, you're just failing to account for it properly.
The standard here was I've just got 2GB of data I need to reliably store someplace. Its an expensive assumption I've got multiple servers in multiple physically separate buildings already lying around.
acdha · · focus · HN ↗
This is badly misunderstanding the problem: S3 is a highly available service with geographic redundancy and a huge range of integrated features. Your comparison would need to be updated to include multiple running servers in addition to redundant storage, and the software stack implementing things like immutability, not to mention all of the security features.
That’s not to say that you can’t build equivalents for the parts you use but you either need massive scale or giving up features to do that. For example, if you can tolerate bitrot or long access times if hardware fails, you can definitely get a lower cost per terabyte.
The reason why most people don’t do that, even when they have scale, is that it adds cost and risk everywhere else if you need engineering/ops people working on storage. If you’re, say, the internet archive that might make sense—it’s quite literally why your organization exists—but most other places are going to see all of the integrated features that they don’t have to build and operate paying for the difference between S3 and physical media pricing. For example, if I want to process files as soon as they’re uploaded or have immutability, I can just turn that on rather than having to build more services.
ehe78qhe · · focus · HN ↗
acdha · · focus · HN ↗
aeonik · · focus · HN ↗
acdha · · focus · HN ↗
elendilm · · focus · HN ↗
You just plug it in and basically it runs. Thats pretty much it.
Having to explain to an HN audience how installing an SSD is trivial is weird.
acdha · · focus · HN ↗
elendilm · · focus · HN ↗
If you don't pay the premium for redundancy, your S3 is the same sitting duck as the single SSD in your office, albeit an expensive one. Thinking AWS and therefore integrity and security is magical thinking.
Ask those who lost their data in Bahrain AWS.
But you are entitled to your expensive opinion. Some of us just don't buy it is all.
acdha · · focus · HN ↗
No, that was literally the first paragraph of my original comment.
> If you don't pay the premium for redundancy, your S3 is the same sitting duck as the single SSD in your office, albeit an expensive one.
Again, this is misunderstanding what the service provides relative to raw storage. I suggest reading the documentation about reliability and thinking about how that SSD in your office fails to provide either geographic redundancy (which all but the one-zone flavor of S3 has) or bitrot protection. Similarly, while it is true that some security concerns still require you to configure things correctly, you should think about the infrastructure you would need to build to match the core service: none of that is something you can’t match but building and operating that has costs.
> Ask those who lost their data in Bahrain AWS.
Let’s think about this slightly longer: how is the SSD in your Bahrain office doing? Iran has to hit multiple AWS data centers to lose your S3 data so to match that you need to build a system which syncs changes at multiple separate campuses in the region, which is significantly more expensive, and when you realize that you probably want a copy somewhere far away, the work to build that will be more than checking a box — and don’t forget that you have to regularly test all of this to make sure you can trust it to work in a disaster.
Again, as readers of the original comment know, I’m not saying just you can’t do this but that it’s a category error to think the comparison is the raw hardware and not the combination of redundant hardware, software, and staffing you need to match capabilities. You can beat that with enough scale but that’s a very different equation.
elendilm · · focus · HN ↗
You shifted the conversation to redundancy and how AWS provides bells and whistles. You are consistent in your position, just not with me and not with PunchyHamster because we both are comparing costs by self hosting.
And for those of you who are willing to pay the premium and don't care about costs, like you, would obviously find AWS attractive.
For a startup like ours and many others, cost is everything. And spending a day or two for bitrot protection and having a couple of SSDs for redundancy and scalability is a no brainer.
You don't need big buildings for resilience. I suggest building things that require real scaling and resilience and you would know that it is about being distributed, having low cost, and owning your systems.
Your argument that it is expensive and tedious to maintain compared to AWS S3 is false. No it isn't.
Fear mongering only works against those who haven't built reliable systems. Building an S3 equivalent with cheap providers for redundancy or having one's own small resilient distributed hardware with staffing most often dwarfs the horribly insane premium paid for AWS S3 services.
But this is only if you want the redundancy which is a different topic altogether.
If all you want is a single replica, which was what PunchyHamster originally commented on, the base price of what AWS S3 provides, which does not include redundancy, makes even lesser sense compared to having your own SSD.
ehe78qhe · · focus · HN ↗
The biggest reason people pay for S3 is because of data transfer cost to other cloud services they are using.
acdha · · focus · HN ↗
Fundamentally you’re playing a commodity game and that means you need a consistent edge to be worth the fixed costs. If your fully-loaded advantage is less than, say, 30% you’re better off negotiating with your AWS rep and investing your engineering resources in something your customers want.
everfrustrated · · focus · HN ↗
Only someone who had never done this would say this.
ehe78qhe · · focus · HN ↗
Dylan16807 · · focus · HN ↗
Really? No, it's not a few percent.
acdha · · focus · HN ↗
Dylan16807 · · focus · HN ↗
This is very different from the idea that the savings are a few percent. In almost all cases the savings are nothing or a large fraction.
elendilm · · focus · HN ↗
Of course, if you had paid extra for the backups and redundancy, your data would survive.
So to address your original point, unless you are paying extra for redundancy, it is cheaper if you use a 4TB SSD at retail price as the original commenter discussed.
someonebaggy · · focus · HN ↗
elendilm · · focus · HN ↗
The parent commentator is under the illusion that AWS automatically means security and scale and reliability.
I was merely pointing out to him that Iranian attacks must serve as a wake up call for him.
everfrustrated · · focus · HN ↗
elendilm · · focus · HN ↗
someonebaggy · · focus · HN ↗
inkyoto · · focus · HN ↗
[dead]
someonebaggy · · focus · HN ↗
I'd bet the majority of projects, but perhaps not the majority of traffic, can run adequately on a single server and tolerate a few hour maintenance window one weekend every month. Don't waste money on overkill.
PunchyHamster · · focus · HN ↗
First off, for geo-redundancy you need to pay extra, you only get region, and with recent strait crisis amazon already showed that can be hit and be unrecoverable.
And you're badly misunderstanding the point; that it is fucking expensive from every angle.
You can get geo-redundancy cheaper (their competitions prove it all the time); you can get IOPS cheaper; and with massive difference in end of the month bill.
Especially if you need a tons of bandwidth, S3 is very, very, very expensive way to get there.
It might be acceptable if you just run a website, but once you start running data-heavy operations on it, a big SSD attached to server can be order of magnitude cheaper than running it on S3 (and you can still use it to backup the data in case something fails)
tempay · · focus · HN ↗
elendilm · · focus · HN ↗
Yes you can make the argument for that specific use case.
For everyday use case, using s3 as your primary storage is costly and is not at all ideal.
someonebaggy · · focus · HN ↗
hadlock · · focus · HN ↗
bigstrat2003 · · focus · HN ↗