‹ BackHN Continuity

Thread

US halts flights at busy East Coast airports, says fiber line cut

239 points · 150 comments · allanbreyes

  1. cube00 · · focus · HN ↗
    > When it went to flip into the backup, we discovered that the backup fiber had a break

    Pretty grim that a life critical system wasn't designed to report that the backup fibre was unserviceable until they attempted to switch over to it.

    I wonder how long it was down? Days, weeks, months?

    1. readthenotes1 · · focus · HN ↗
      A lot of people don't check their backups until they need to restore.

      A lot of people are incompetent.

      1. jimt1234 · · focus · HN ↗
        I've worked on backup/failure systems since the mid-90s, and I've found there's one universal truth: If you don't fully test your backup/failure system, you don't have a backup/failure system.

        There's generally two wrong responses: (1) We spent a lot of money on 'blah blah blah', a lot of other companies use it, so yeah, we've got a backup/failure system. And, (2) inadequate testing - either, we tested 1 of 50 services, and it worked, so the whole system can be restored; or, we gracefully tested, and it worked, so it will obviously work during not-graceful incidents.

        And the root cause of this is generally that no one gets promoted for implementing an adequate backup/failure system, or it's extremely rare.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.