‹ BackHN Continuity

Thread

US halts flights at busy East Coast airports, says fiber line cut

239 points · 150 comments · allanbreyes

  1. cube00 · · focus · HN ↗
    > When it went to flip into the backup, we discovered that the backup fiber had a break

    Pretty grim that a life critical system wasn't designed to report that the backup fibre was unserviceable until they attempted to switch over to it.

    I wonder how long it was down? Days, weeks, months?

    1. toast0 · · focus · HN ↗
      Sometimes your multiple fiber paths end up in the same bundle, severed by the same backhoe. It's always a fun day when that becomes apparent.
      1. unethical_ban · · focus · HN ↗
        That is negligence, practical if not contractual.

        Over ten years ago, my employer was spinning up a new DC across the state and had three links between it and the primary DC. Two were pretty direct, but we needed a third because at one point in the 300 mile path, the two main links went within 400 meters of each other.

        So how national-security-adjacent critical systems like an airport system doesn't have a larger set of backup links, and validates that they are geographically separate up until the connections, and have realtime alerting on the connection status, is surprising to me.

        1. Jtsummers · · focus · HN ↗
          > validates that they are geographically separate up until the connections

          This isn't relevant in this case. The problems occurred at different places. The backup fiber was, separately from the issues with the primary system, cut by a construction crew. It's not two fiber lines both cut at the same spot.

          1. hdgvhicv · · focus · HN ↗
            Yes the probe here is a lack of monitoring.

            You can get two separate failures in short succession, where the first one hasn’t been fixed. For critical devices that’s why you have tertiary backups, ideally on different technology compel you. For one major event I had diverse fibre and satelite and microwave. Lost microwave and satelite to electronic warfare but the fibres were ok. Ofcom were really pissed about it, but nothing they could do.

            1. justsomehnguy · · focus · HN ↗
              > Yes the probe here is a lack of monitoring.

              No?

              Like, okay, monitoring reports what there is no link.

              Now what? Somebody gets a hi vis vest, a hard hat and goes along the cable to investigate the reason?

              There is nothing monitoring could help, especially if the cable was cut recent enough.

              1. benzible · · focus · HN ↗
                <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Optical_time-domain_reflectometer" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Optical_time-domain_reflectome...

                &gt; An RFTS enables fiber to be automatically tested from a central location. A central computer is used to control the operation of OTDR-like test components located at key points in the fiber network. The test components scan the fiber to locate problems. If a problem is found, its location is noted and the appropriate personnel are notified to begin the repair process.

                1. justsomehnguy · · focus · HN ↗
                  There are a lot of more thing than just making sure the media is working. And you would be quite surprised (considering you are throwing Wiki links for a reflectometry) but the physical state of the line is the least important thing in making sure what something can communicate with something else. A sibling comment can provide a glimpse on how it can be done.
              2. iso1631 · · focus · HN ↗
                If I lose a link anywhere in the world then BGP reroutes within a second or so, and we get a notification in seconds or maybe minutes on slack depending on how busy slack is being.

                If it&#x27;s down for more than 15 minutes in country or 2 hours internationally it gets flagged for manual attention. Shorter outages are logged but only looked at monthly as part of the operational report process.

                Now sure you can have two independent faults on the same day, but the reports are that they didn&#x27;t know the backup line was down.

                1. justsomehnguy · · focus · HN ↗
                  Meh, I knew people would just clutch to the monitoring ignoring everything else.

                  In your case you already have way more than monitoring. You have the infrastructure designed for resilience. You have that design implemented. You have the protocols to do if the things go north. You have automation to disregard minor events and to bring to the attention more serious things. You have way more than monitoring alone.

        2. mikeweiss · · focus · HN ↗
          For all we know the contract says the FAA is responsible to notify Verizon if the line becomes dead. We don&#x27;t know what the agreement is.
        3. bonestamp2 · · focus · HN ↗
          That is fascinating. Out of curiosity, how far apart would the paths have had to remain in order to be ok with just two links? Like, I&#x27;m thinking a plane crash between those two links in the 400m zone could destory both links so a third link makes sense, but if it were 800m that would probably be fine? Or what magnitude of disaster are they trying to mitigate and how much separation would that require for just two links?
          1. unethical_ban · · focus · HN ↗
            IIRC it also had to do with them passing near a railroad track, which they identified as an increased risk. So I&#x27;m guessing it is about perceived chance of risk - backhoes and groundwork are more likely near certain areas than there is of a plane crash.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.