Something that's interested me is how hard drive capacities have continuously gone up, but read and write speeds have not really changed. To the point where it would take multiple days to read these new high capacity drives.
I wonder what the use case for them is, they aren't ideal for RAID as they would take far too long to rebuild the array in a drive failure, but they also aren't great for long term cold storage. Perhaps they are ideal for periodic backups where you want to take a daily 1TB backup of a server or CCTV where you don't really care about backing up the data as it's all deleted on schedule anyway.
You use them for backups, data archiving, or object storage with triple parity RAID or erasure coding. The combined bandwidth of dozens to hundreds of magnetic hard drives is pretty decent.
I imagine that there is a ton of data in the "would be nice to retain but not so critical as to warrant a constant backup" category. Alternatively stuff that requires large storage, but only occasional small writes, so incremental backups are feasible.
For one, home servers can go on maintenance mode (read only?) for a day or two, it will be fine most of the time.
In bigger applications, you have enough redundancy to allow one or two nodes to go down for a while. If a company had terabytes upon terabytes of critical data, having a multi-tiered storage would make sense, even of a full backup restore would take a days.on some tiers.
Spinning rust is still the only cost effective answer to redundant 10+ tb storage. If you have 4+ disks, use raid 6 or raid 10 and you get decent fault tolerance too. Raid 10 also increases write performance but it's a bit less robust imo and more expensive per gb.
As for speed it's mainly tied to rpm and the market decided no one wants higher failure rates / lower mtbf for marginal speed gains.
If you need speed and capacity, throw an ssd or two in front of a raid disk array for caching with primocache or whatever.
Speeds do go up with density. Take an empty 4 TB drive and an empty 16 TB drive, and then write and read 2 TB of data to both of them. Since the platters on the 16 TB one are that much denser, you go over a lot more blocks in a single spin than the 4 TB one.
That's a very logical theory, but it doesn't really seem to hold up in real-world data transfer speeds.
For instance: 7200 RPM SATA drives have been stuck in the peak-performance realm of between about 250 and 300 megabytes per second for a decade, even though maximum capacity has more than tripled over that time.
The 16 TB drive won't be anywhere close to 4x faster than the 4 TB drive, though. Maybe 20% (~200 MB/s vs ~230-250 MB/s). That's because most of the capacity increase comes from more tracks and more platters, not linear density, which determines throughput.
I would say 250-260MB/s for a 16TB drive, and when 4TB drives were newer they were more like 160-180MB/s. That's about a 50% increase.
Doing the math, you're probably going from 5 to 8-9 platters, so 2.2-2.5 areal density, and a naive square root suggests it should be about 50% faster. So maybe track width is taking a fair share after all?
Though when I look at more sizes, it seems like 12TB drives were also hitting 260MB/s and speed has almost entirely stalled since, with 24TB being only a hair faster.
What's "far too long"? Why is it too long?
I keep my important stuff with ZFS raidz2.
It doesn't matter much to me if the overall process of replacing a dead drive can be done in 40 minutes or if it drags out to 40 hours. The storage system is still alive, still available for reading and writing, and still redundant for that entire time. It can take as long as it takes.
My break-glass-in-case-of-emergency offsite backups are done with restic over the network, so I only had to send 1 initial copy. That first copy took quite a long time. But restic is pretty clever about how it handles things; all subsequent daily backups are what might best be described as incremental-ish. They contain only recent changes and are pretty small with the usual churn rate on this particular pile of data.
(This matters because it keeps the cost of keeping it running quite low, which means that backups actually happen instead of being just a good idea. :) )
---
I'd probably pick different methods if I had a real budget for this stuff, but even then: Why would I be motivated to care about how long it takes to replace a hard drive?
So move the bar up to raidz3 or the equivalent? This adds cost, of course, but enhancing redundancy isn't free.
It was something I kept in-mind when I made the decision to use raidz2 for my own purposes: It tolerates two concurrent drive failures before local data loss happens, which I feel is acceptable for my purposes. raidz3 moves that tolerance to three concurrent drive failures.
I mean: What's the alternative for storing X amount of data? A greater number of relatively small drives? That's expensive in its own ways, too.
And adding complexity in that way tends to nosedive metrics like MTBF: If a system relies upon one widget and this system has an MTBF of n, then a system that relies upon two such widgets has an MTBF of n/2.
If the drives all have the same failure rate, then: A system with raidz2 and 10x 10TB drives will tend experience more disk failures than a system with raidz2 and 6x 20TB drives. Both systems can store ~80TB of data, but the one with 10TB drives has a worse ratio of redundancy and includes more parts that will eventually fail.
KISS, etc. :)
> and the longer it takes the more likely another drive will fail during that time.
You mean the old light-weight RAID-5-esque fear? A drive fails and redundancy drops to zero. So the failed drive is replaced or the hot spare is rotated in (or whatever), and the rebuild process begins. Another drive crashes during the rebuild because of the stress of actually-being-read and now the data on the array is presumed to be trash.
That process seems to actually work more like this: Old data that hasn't been read for years finally gets accessed (because RAID rebuild), and that's when some other drive starts reporting errors. But those errors were already there; they just weren't detected until it was way too late.
It doesn't really seem to work that way with filesystems like ZFS. The contents, whether old and stale or brand new, can (and should) be read routinely and checksummed to confirm integrity as part of a regularly-performed scrub.
The weak drives that can't stand to be read will be discovered during this read-only exercise, but they were already dead. We just didn't know it until we ran a scrub. So we scrub early, and often.
That's the same as with RAID, except: With RAID, the same disk still dies in the same way. We just don't know it was dead until a recovery is in-process. RAID, as commonly practiced, can therefore function as a problem-multiplier in ways that other systems can avoid.
When you're regularly working with largely sequential datasets with sizes ranging from 100GB to >1TB that you're not processing at speeds in excess of 100MB/sec, hard drives are very much good enough.
Especially now that the same 4TB NVMe SSDs I was buying for $300 are now selling for over $1,000.
In my case, the source of the data is S3, so there's no need for redundancy, as the hard drive is effectively a local cache for cloud storage and the work isn't nearly so time-critical that an occasional failure would be more than a mild annoyance to a customer, if that.
Same for archival backups. Faster recovery time than Glacier Deep Archive, less cost over a period of years, simpler and less expensive than tape for modest volume.
I'm not doubting the usefulness of hard drives in general, just the ones above 20TB where it would take you most of a week to write to the drive. Hard drive fill times have been steadily going up to the point it becomes almost impractical to use these 30tb drives.
I guess if you have it all backed up in S3 it's fine, but if you put this thing in a NAS and one drive dies, your NAS is going to spend the whole week rebuilding where it's running at 100%, and another drive dying during this means data loss.
It depends? In something like a 800TB Ceph cluster a failure of a single 20TB drive is not felt at all if configured redundantly enough. Then it does not matter how long replacing it takes, only how many bucks each TB costs and/or how much electricity each drive needs.
If the cluster as whole can stomach the drive failures without losing data and is overall fast enough, there is no need to run many small drives where running fewer big drives is cheaper.
Can't you just not run a Raid 6 or whatever ? If it takes a week to rebuild but is just making a new redundant copy of the old failed drive then it doesn't seem like a huge deal ?
RAID 5 being too risky doesn't invalidate the whole concept. You can use RAID 6 with two parity drives instead of one, or something higher level like ZFS or Ceph. If you have a trustworthy backup and can tolerate an amount of downtime risk, you can even run RAID 0 or mergerfs on your main array. Many ways to arrange storage.
FWIW WD is working on high bandwidth drives. We'll see when it hits the market though.
<a href="https://blog.westerndigital.com/performance-optimized-hdd-4x-throughput-hbdt-dual-pivot/" rel="nofollow">https://blog.westerndigital.com/performance-optimized-hdd-4x...
Gigachad · · focus · HN ↗
I wonder what the use case for them is, they aren't ideal for RAID as they would take far too long to rebuild the array in a drive failure, but they also aren't great for long term cold storage. Perhaps they are ideal for periodic backups where you want to take a daily 1TB backup of a server or CCTV where you don't really care about backing up the data as it's all deleted on schedule anyway.
UltraSane · · focus · HN ↗
hgoel · · focus · HN ↗
Most of the data on my NAS is of that form.
makeitdouble · · focus · HN ↗
In bigger applications, you have enough redundancy to allow one or two nodes to go down for a while. If a company had terabytes upon terabytes of critical data, having a multi-tiered storage would make sense, even of a full backup restore would take a days.on some tiers.
eyegor · · focus · HN ↗
As for speed it's mainly tied to rpm and the market decided no one wants higher failure rates / lower mtbf for marginal speed gains.
If you need speed and capacity, throw an ssd or two in front of a raid disk array for caching with primocache or whatever.
Hamuko · · focus · HN ↗
ssl-3 · · focus · HN ↗
For instance: 7200 RPM SATA drives have been stuck in the peak-performance realm of between about 250 and 300 megabytes per second for a decade, even though maximum capacity has more than tripled over that time.
formerly_proven · · focus · HN ↗
Dylan16807 · · focus · HN ↗
Doing the math, you're probably going from 5 to 8-9 platters, so 2.2-2.5 areal density, and a naive square root suggests it should be about 50% faster. So maybe track width is taking a fair share after all?
Though when I look at more sizes, it seems like 12TB drives were also hitting 260MB/s and speed has almost entirely stalled since, with 24TB being only a hair faster.
killingtime74 · · focus · HN ↗
ssl-3 · · focus · HN ↗
I keep my important stuff with ZFS raidz2.
It doesn't matter much to me if the overall process of replacing a dead drive can be done in 40 minutes or if it drags out to 40 hours. The storage system is still alive, still available for reading and writing, and still redundant for that entire time. It can take as long as it takes.
My break-glass-in-case-of-emergency offsite backups are done with restic over the network, so I only had to send 1 initial copy. That first copy took quite a long time. But restic is pretty clever about how it handles things; all subsequent daily backups are what might best be described as incremental-ish. They contain only recent changes and are pretty small with the usual churn rate on this particular pile of data.
(This matters because it keeps the cost of keeping it running quite low, which means that backups actually happen instead of being just a good idea. :) )
---
I'd probably pick different methods if I had a real budget for this stuff, but even then: Why would I be motivated to care about how long it takes to replace a hard drive?
justincormack · · focus · HN ↗
ssl-3 · · focus · HN ↗
So move the bar up to raidz3 or the equivalent? This adds cost, of course, but enhancing redundancy isn't free.
It was something I kept in-mind when I made the decision to use raidz2 for my own purposes: It tolerates two concurrent drive failures before local data loss happens, which I feel is acceptable for my purposes. raidz3 moves that tolerance to three concurrent drive failures.
I mean: What's the alternative for storing X amount of data? A greater number of relatively small drives? That's expensive in its own ways, too.
And adding complexity in that way tends to nosedive metrics like MTBF: If a system relies upon one widget and this system has an MTBF of n, then a system that relies upon two such widgets has an MTBF of n/2.
If the drives all have the same failure rate, then: A system with raidz2 and 10x 10TB drives will tend experience more disk failures than a system with raidz2 and 6x 20TB drives. Both systems can store ~80TB of data, but the one with 10TB drives has a worse ratio of redundancy and includes more parts that will eventually fail.
KISS, etc. :)
> and the longer it takes the more likely another drive will fail during that time.
You mean the old light-weight RAID-5-esque fear? A drive fails and redundancy drops to zero. So the failed drive is replaced or the hot spare is rotated in (or whatever), and the rebuild process begins. Another drive crashes during the rebuild because of the stress of actually-being-read and now the data on the array is presumed to be trash.
That process seems to actually work more like this: Old data that hasn't been read for years finally gets accessed (because RAID rebuild), and that's when some other drive starts reporting errors. But those errors were already there; they just weren't detected until it was way too late.
It doesn't really seem to work that way with filesystems like ZFS. The contents, whether old and stale or brand new, can (and should) be read routinely and checksummed to confirm integrity as part of a regularly-performed scrub.
The weak drives that can't stand to be read will be discovered during this read-only exercise, but they were already dead. We just didn't know it until we ran a scrub. So we scrub early, and often.
That's the same as with RAID, except: With RAID, the same disk still dies in the same way. We just don't know it was dead until a recovery is in-process. RAID, as commonly practiced, can therefore function as a problem-multiplier in ways that other systems can avoid.
jasomill · · focus · HN ↗
Especially now that the same 4TB NVMe SSDs I was buying for $300 are now selling for over $1,000.
In my case, the source of the data is S3, so there's no need for redundancy, as the hard drive is effectively a local cache for cloud storage and the work isn't nearly so time-critical that an occasional failure would be more than a mild annoyance to a customer, if that.
Same for archival backups. Faster recovery time than Glacier Deep Archive, less cost over a period of years, simpler and less expensive than tape for modest volume.
Gigachad · · focus · HN ↗
I guess if you have it all backed up in S3 it's fine, but if you put this thing in a NAS and one drive dies, your NAS is going to spend the whole week rebuilding where it's running at 100%, and another drive dying during this means data loss.
atoav · · focus · HN ↗
If the cluster as whole can stomach the drive failures without losing data and is overall fast enough, there is no need to run many small drives where running fewer big drives is cheaper.
Every application will have different trade offs
lostlogin · · focus · HN ↗
I’ve got a Synology that’s mostly full of 20TB drives. As you say, adding a drive takes about a week. Expanding the array is even longer.
Melatonic · · focus · HN ↗
someonebaggy · · focus · HN ↗
7bbfea · · focus · HN ↗