‹ BackHN Continuity

Thread

Backblaze drive stats for Q2 2026

307 points · 113 comments · HieronymusBosch

  1. Gigachad · · focus · HN ↗
    Something that's interested me is how hard drive capacities have continuously gone up, but read and write speeds have not really changed. To the point where it would take multiple days to read these new high capacity drives.

    I wonder what the use case for them is, they aren't ideal for RAID as they would take far too long to rebuild the array in a drive failure, but they also aren't great for long term cold storage. Perhaps they are ideal for periodic backups where you want to take a daily 1TB backup of a server or CCTV where you don't really care about backing up the data as it's all deleted on schedule anyway.

    1. ssl-3 · · focus · HN ↗
      What's "far too long"? Why is it too long?

      I keep my important stuff with ZFS raidz2.

      It doesn't matter much to me if the overall process of replacing a dead drive can be done in 40 minutes or if it drags out to 40 hours. The storage system is still alive, still available for reading and writing, and still redundant for that entire time. It can take as long as it takes.

      My break-glass-in-case-of-emergency offsite backups are done with restic over the network, so I only had to send 1 initial copy. That first copy took quite a long time. But restic is pretty clever about how it handles things; all subsequent daily backups are what might best be described as incremental-ish. They contain only recent changes and are pretty small with the usual churn rate on this particular pile of data.

      (This matters because it keeps the cost of keeping it running quite low, which means that backups actually happen instead of being just a good idea. :) )

      ---

      I'd probably pick different methods if I had a real budget for this stuff, but even then: Why would I be motivated to care about how long it takes to replace a hard drive?

      1. justincormack · · focus · HN ↗
        Because thats a reduced redundancy period, and the longer it takes the more likely another drive will fail during that time.
        1. ssl-3 · · focus · HN ↗
          > Because thats a reduced redundancy period

          So move the bar up to raidz3 or the equivalent? This adds cost, of course, but enhancing redundancy isn't free.

          It was something I kept in-mind when I made the decision to use raidz2 for my own purposes: It tolerates two concurrent drive failures before local data loss happens, which I feel is acceptable for my purposes. raidz3 moves that tolerance to three concurrent drive failures.

          I mean: What's the alternative for storing X amount of data? A greater number of relatively small drives? That's expensive in its own ways, too.

          And adding complexity in that way tends to nosedive metrics like MTBF: If a system relies upon one widget and this system has an MTBF of n, then a system that relies upon two such widgets has an MTBF of n/2.

          If the drives all have the same failure rate, then: A system with raidz2 and 10x 10TB drives will tend experience more disk failures than a system with raidz2 and 6x 20TB drives. Both systems can store ~80TB of data, but the one with 10TB drives has a worse ratio of redundancy and includes more parts that will eventually fail.

          KISS, etc. :)

          > and the longer it takes the more likely another drive will fail during that time.

          You mean the old light-weight RAID-5-esque fear? A drive fails and redundancy drops to zero. So the failed drive is replaced or the hot spare is rotated in (or whatever), and the rebuild process begins. Another drive crashes during the rebuild because of the stress of actually-being-read and now the data on the array is presumed to be trash.

          That process seems to actually work more like this: Old data that hasn't been read for years finally gets accessed (because RAID rebuild), and that's when some other drive starts reporting errors. But those errors were already there; they just weren't detected until it was way too late.

          It doesn't really seem to work that way with filesystems like ZFS. The contents, whether old and stale or brand new, can (and should) be read routinely and checksummed to confirm integrity as part of a regularly-performed scrub.

          The weak drives that can't stand to be read will be discovered during this read-only exercise, but they were already dead. We just didn't know it until we ran a scrub. So we scrub early, and often.

          That's the same as with RAID, except: With RAID, the same disk still dies in the same way. We just don't know it was dead until a recovery is in-process. RAID, as commonly practiced, can therefore function as a problem-multiplier in ways that other systems can avoid.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.