‹ BackHN Continuity

Thread

US halts flights at busy East Coast airports, says fiber line cut

239 points · 150 comments · allanbreyes

  1. cube00 · · focus · HN ↗
    > When it went to flip into the backup, we discovered that the backup fiber had a break

    Pretty grim that a life critical system wasn't designed to report that the backup fibre was unserviceable until they attempted to switch over to it.

    I wonder how long it was down? Days, weeks, months?

    1. readthenotes1 · · focus · HN ↗
      A lot of people don't check their backups until they need to restore.

      A lot of people are incompetent.

      1. thewebguyd · · focus · HN ↗
        A backup without a restore test isn't a backup at all
        1. bonoboTP · · focus · HN ↗
          You don't always have a copy of your hardware to restore onto. And the test's entire purpose is that you're not yet sure whether your restore will truly work. So you can't just run a backup and restore on your true prod system, because you're not sure it won't wreck it. So you need extra money to have a second system onto which you try to restore. If you don't have a lot of money, you will want to actually use your disks for storage, not to put them into a second testing server. Of course I'm not talking about very professional companies with super critical data. Just simpler smaller scale places or consumers.
          1. lovich · · focus · HN ↗
            No, even professional companies with critical data balk at this.

            I worked at a company worth a few billion and the leadership balked when they told engineering they wanted a near instantaneous failover system and our department informed them that would require paying for a second environment that could be rolled over to.

            It is rare to find leaders who can accept the cost of redundant infrastructure that is there for emergency backup.

            What puzzles me is why they can’t accept it when they are perfectly fine with insurance costs and I can’t see much of a difference between the two when looking at a spreadsheet of costs other than possibly tax differences between the type of expenditure.

            1. bonoboTP · · focus · HN ↗
              Sure, that's another category and different considerations. I've worked at an academic lab with a limited budget where we set up file servers, but couldn't afford to do anything approximating 3-2-1. We did a nightly backup of a tiny part (most important) of the data onto another server in a different building, but like 90% of the data was just YOLO (well, RAID, but that's not a backup), and that's just how it is. Disks are pretty good though, they don't die often nowadays, and when people accidentally deleted their data, it was just gone. Would have been cool to have a backup of everything, but even just pulling out a nightly backup from a dense server with dozens of terabytes isn't simple and you don't want to slow down the server by constantly reading just for constructing the backups. It's a tradeoff.

              Redundancy has costs and those costs can be spent elsewhere like having higher quality or bigger disks, or a faster network switch or better CPUs etc.

              In theory, it would also be better to own two cars instead of one, because what if the first one gets in an accident or just breaks down. Yet, not everyone can afford that. Should you just buy two half-as-expensive cars than what you can buy one of, so you can say you have a "backup"? Likely the two half-price ones would be so much crappier that the one good car would cause you less trouble in expectation than driving a shitty one and then having another shitty spare one, both of which will constantly have issues.

              1. hubbahubbahubba · · focus · HN ↗
                Resilvering operations with large capacity data drives take a ridiculously long time to complete. A double disk failure in a volume can result in a loss of the LUN. Populating a volume with drives from the same supplier with matching batch numbers can result in correlated failures not independent from each other. You have to ask yourself - do you feel lucky?
                1. bonoboTP · · focus · HN ↗
                  We ran with RAID 6 + a hot spare and a mix of disks, not from the same batch. Again, the alternative is sometimes not even having this. Compared to that it's a pretty good deal, and losing the entire data would be quite inconvenient but not world ending in an academic lab.

                  And in other labs I've already seen data loss because they messed with their setup in incompetent ways.

                  You can't always have everything. There is a budget. There are storage needs. You can try to reduce your storage and instead beef up the backups and stuff, but then you can barely do anything and people have to constantly delete their data, and worry about space and that slows things down, generates hostility and conflict between colleagues, and just causes pain in general. You can't see these things from just a technical desirability point of view or what setup would get the most upvotes on r/DataHoarder or whatnot.

            2. 486sx33 · · focus · HN ↗

              [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.