‹ BackHN Continuity

Thread

S3 Is the Future, S3 Is the Past

79 points · 120 comments · tkhattra

  1. PunchyHamster · · focus · HN ↗
    > S3’s dominance is due to its many advantages: effectively infinite capacity, high durability, and low per-gigabyte capacity cost.

    It's not "cheap". 6 months of S3 is at around price of outright buying 4TB SSD at retail price. That before you do any IOPS to it

    S3 is terrible deal on any front. it's just easy

    1. acdha · · focus · HN ↗
      > It's not "cheap". 6 months of S3 is at around price of outright buying 4TB SSD at retail price. That before you do any IOPS to it

      This is badly misunderstanding the problem: S3 is a highly available service with geographic redundancy and a huge range of integrated features. Your comparison would need to be updated to include multiple running servers in addition to redundant storage, and the software stack implementing things like immutability, not to mention all of the security features.

      That’s not to say that you can’t build equivalents for the parts you use but you either need massive scale or giving up features to do that. For example, if you can tolerate bitrot or long access times if hardware fails, you can definitely get a lower cost per terabyte.

      The reason why most people don’t do that, even when they have scale, is that it adds cost and risk everywhere else if you need engineering/ops people working on storage. If you’re, say, the internet archive that might make sense—it’s quite literally why your organization exists—but most other places are going to see all of the integrated features that they don’t have to build and operate paying for the difference between S3 and physical media pricing. For example, if I want to process files as soon as they’re uploaded or have immutability, I can just turn that on rather than having to build more services.

      1. someonebaggy · · focus · HN ↗
        Do you need that? If you want your own replicated storage cluster at home, you can use Ceph. I've seen Ceph deployed to utilize the spare hard drive ports in server clusters that otherwise mostly do compute (transcoding), alongside memcached to utilize the RAM, etc.

        I'd bet the majority of projects, but perhaps not the majority of traffic, can run adequately on a single server and tolerate a few hour maintenance window one weekend every month. Don't waste money on overkill.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.