‹ BackHN Continuity

Thread

US halts flights at busy East Coast airports, says fiber line cut

239 points · 150 comments · allanbreyes

  1. cube00 · · focus · HN ↗
    > When it went to flip into the backup, we discovered that the backup fiber had a break

    Pretty grim that a life critical system wasn't designed to report that the backup fibre was unserviceable until they attempted to switch over to it.

    I wonder how long it was down? Days, weeks, months?

    1. toast0 · · focus · HN ↗
      Sometimes your multiple fiber paths end up in the same bundle, severed by the same backhoe. It's always a fun day when that becomes apparent.
      1. koolba · · focus · HN ↗
        Put all your backups in one basket, and then pray that nobody crushes the basket.
        1. oasisbob · · focus · HN ↗
          This is a pretty classic network operator story: go to great lengths to provide for physically diverse paths, then not notice when your provider refactors and grooms them onto the same bundle.
          1. jacquesm · · focus · HN ↗
            In NL this often happens because of waterways. All starts off with good intentions, multiple providers, totally different fibers exiting the building in different directions. And then they have to cross a canal or a river. It's a 50/50 at that point whether they converge on the same conduit across the water.
          2. hdgvhicv · · focus · HN ↗
            I find intonation it’s better to select two competing provides as they are less likely to use the same ducts.

            Had a few issues in some lamdlpcked counties on occasion - Kabul had all traffic routing via the same route into Pakistan for a few years, Kathmandu had an earthquake knock out both ISPs within 30 seconds of each other.

            Backups via starlink global roaming work well now in particularly challenging environments.

      2. konfusinomicon · · focus · HN ↗
        last time this happened in my area a local farmer was burying a cow and took out the whole towns connection
        1. burningChrome · · focus · HN ↗
          Nest door neighbor did this when he just decided on a whim to put in a new driveway and started digging with a bobcat. 10 mins in and BLAM severed the main Comcast coax serving the entire neighborhood that was running under his driveway.

          Luckily we were on Century Link so weren't affected by his stupidity. lol

          1. hubbahubbahubba · · focus · HN ↗
            How much was the fine for not using the call before you dig service?
      3. unethical_ban · · focus · HN ↗
        That is negligence, practical if not contractual.

        Over ten years ago, my employer was spinning up a new DC across the state and had three links between it and the primary DC. Two were pretty direct, but we needed a third because at one point in the 300 mile path, the two main links went within 400 meters of each other.

        So how national-security-adjacent critical systems like an airport system doesn't have a larger set of backup links, and validates that they are geographically separate up until the connections, and have realtime alerting on the connection status, is surprising to me.

        1. Jtsummers · · focus · HN ↗
          > validates that they are geographically separate up until the connections

          This isn't relevant in this case. The problems occurred at different places. The backup fiber was, separately from the issues with the primary system, cut by a construction crew. It's not two fiber lines both cut at the same spot.

          1. hdgvhicv · · focus · HN ↗
            Yes the probe here is a lack of monitoring.

            You can get two separate failures in short succession, where the first one hasn’t been fixed. For critical devices that’s why you have tertiary backups, ideally on different technology compel you. For one major event I had diverse fibre and satelite and microwave. Lost microwave and satelite to electronic warfare but the fibres were ok. Ofcom were really pissed about it, but nothing they could do.

            1. justsomehnguy · · focus · HN ↗
              > Yes the probe here is a lack of monitoring.

              No?

              Like, okay, monitoring reports what there is no link.

              Now what? Somebody gets a hi vis vest, a hard hat and goes along the cable to investigate the reason?

              There is nothing monitoring could help, especially if the cable was cut recent enough.

              1. benzible · · focus · HN ↗
                <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Optical_time-domain_reflectometer" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Optical_time-domain_reflectome...

                &gt; An RFTS enables fiber to be automatically tested from a central location. A central computer is used to control the operation of OTDR-like test components located at key points in the fiber network. The test components scan the fiber to locate problems. If a problem is found, its location is noted and the appropriate personnel are notified to begin the repair process.

                1. justsomehnguy · · focus · HN ↗
                  There are a lot of more thing than just making sure the media is working. And you would be quite surprised (considering you are throwing Wiki links for a reflectometry) but the physical state of the line is the least important thing in making sure what something can communicate with something else. A sibling comment can provide a glimpse on how it can be done.
              2. iso1631 · · focus · HN ↗
                If I lose a link anywhere in the world then BGP reroutes within a second or so, and we get a notification in seconds or maybe minutes on slack depending on how busy slack is being.

                If it&#x27;s down for more than 15 minutes in country or 2 hours internationally it gets flagged for manual attention. Shorter outages are logged but only looked at monthly as part of the operational report process.

                Now sure you can have two independent faults on the same day, but the reports are that they didn&#x27;t know the backup line was down.

                1. justsomehnguy · · focus · HN ↗
                  Meh, I knew people would just clutch to the monitoring ignoring everything else.

                  In your case you already have way more than monitoring. You have the infrastructure designed for resilience. You have that design implemented. You have the protocols to do if the things go north. You have automation to disregard minor events and to bring to the attention more serious things. You have way more than monitoring alone.

        2. mikeweiss · · focus · HN ↗
          For all we know the contract says the FAA is responsible to notify Verizon if the line becomes dead. We don&#x27;t know what the agreement is.
        3. bonestamp2 · · focus · HN ↗
          That is fascinating. Out of curiosity, how far apart would the paths have had to remain in order to be ok with just two links? Like, I&#x27;m thinking a plane crash between those two links in the 400m zone could destory both links so a third link makes sense, but if it were 800m that would probably be fine? Or what magnitude of disaster are they trying to mitigate and how much separation would that require for just two links?
          1. unethical_ban · · focus · HN ↗
            IIRC it also had to do with them passing near a railroad track, which they identified as an increased risk. So I&#x27;m guessing it is about perceived chance of risk - backhoes and groundwork are more likely near certain areas than there is of a plane crash.
      4. jonah · · focus · HN ↗
        Years ago I was touring a POP in a small city. They were very proud to point out one fiber coming in one side of the building going South and another on the opposite side of the building going North.
      5. to11mtm · · focus · HN ↗
        I actually got to see one of those once.

        For a major trans-oceanic backbone provider, at least 15ish years ago they had a mile or two between Detroit and Chicago where both ends were on the same side of the interstate highway.

        But it&#x27;s more frequent on DAS (Distributed Antenna Systems, AKA small-cell or micro-cell) networks.

        Also the challenge of when fibers are leased (if that&#x27;s still a thing, based on the networks I helped design I&#x27;d say &#x27;probably&#x27;).

        They really are analogous to Lamport&#x27;s &quot;Distributed System&quot; quip; A damaged fiber owned by a company you have never heard of can wreck your day.

        1. hdgvhicv · · focus · HN ↗
          In route one circuit from Washington to London via ashburn and down to Texas before going on subsea circuits through mainland Europe to ensure path diversity. The other heads north east towards Newark and and surfaces in Bude. That took two years of contractural arguments.

          Fibres on the Texas link die all the time (yearly) but that’s fine.

          My circuits in the far east are often carried one via dues and one via the USA.

          It is possible to guarentee diversity, you just need to get specific detail and have good contractual protections.

          1. to11mtm · · focus · HN ↗
            &gt; It is possible to guarentee diversity, you just need to get specific detail and have good contractual protections.

            Yeah, FWIW the non-redundant DAS links we designed when I was in that industry, it was typically a case of the client telling us &#x27;That is too expensive, we accept the risk&#x27;.

            The Detroit-chicago thing, I don&#x27;t know the story on that, it was something that fiber provider had already built.

      6. Havoc · · focus · HN ↗
        &gt;Sometimes your multiple fiber paths end up in the same bundle

        For mission critical stuff like airports I would like to think they&#x27;re go for a more rigorous methodology than hope for the best on paths

        1. [deleted] · · focus · HN ↗

          [deleted]

        2. happyPersonR · · focus · HN ↗
          lol there’s supposed to be …. But then greed gets in the way lol.

          Usually it’s a combo of a tier 1 that decides that a resell agreement is “good enough” and “well just eat the loss” and then by the time stuff like this rolls around “we’ll get back to you” + a bunch of silly explanations that boil down to “you’re not gonna sue us though” start coming out lol

        3. iso1631 · · focus · HN ↗
          We have a faction at work who are all about the outsourcing. &quot;We contract with supplier A to prevent single points of failure&quot;

          Then there&#x27;s another faction that doesn&#x27;t trust a word suppliers say, who will insist on two different suppliers, as while in theory they could be sharing the same routes, it&#x27;s far less likely.

          Comes down to &quot;do you trust contracts&quot;.

          Of course this is for third tier sites like branch offices. Major sites require far more oversight, including visibility of cable routes and any hops on the internal network

      7. hdgvhicv · · focus · HN ↗
        If you design your network badly sure. Spend 3 months with BR once trying to get diverse routing out of one location to their two nearby exchanges, took a long time for them to work out what they needed to do as they had the two paths crossing in a single location for quite some time, but eventually they sorted it.
        1. lazide · · focus · HN ↗
          Bad design is the default - it takes hard work to have real redundancy, and if no one checks&#x2F;demands it? Luck of the draw, or worse.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.