‹ BackHN Continuity

Thread

US halts flights at busy East Coast airports, says fiber line cut

239 points · 150 comments · allanbreyes

  1. atonse · · focus · HN ↗
    I don't quite understand this sort of thing happening. Wasn't the whole point of the internet to be a self-healing network where we route around severed cables, etc?

    Or is it that these ATC networks are their own air-gapped network with less redundancy? That just doesn't add up. Or maybe there was only one line going to the ATC, with no multiple "ISPs" like a datacenter would have?

    1. bri3d · · focus · HN ↗
      * Yes, TRACON have their own dedicated links. You can look into ASTERIX and STARS to learn some of the cursed ways the data processing and dataflow work.

      * There is supposed to be a primary and a secondary link, in this case the primary failed and the fail-over also failed. It's unclear from the reporting if they were damaged in the same incident or if the failover was not tested or monitored adequately.

      The Philadelphia TRACON site has been notoriously unreliable and was supposedly improved in 2025, it's also unclear if these issue actually could stem from that implementation.

      1. wrs · · focus · HN ↗
        I’ve had enough double-redundant links fail - and I mean pretty thoroughly double-redundant, different provider, different media, different physical path - that I’m quite surprised something like the air traffic control system only has two links.
        1. pixl97 · · focus · HN ↗
          One thing that needs to be reviewed is if they are dual links or redundant links.

          In any system where both sides of redundancy always carry traffic then loss of one link can cause congestion failure if any link fails. This is a very common means of failure in electrical networks that requires load shedding. Well, you can't load shed air traffic.

          A more complex system that I like, but comes with it's own set of constraints and implementation issues is a system where both lines carry all the traffic at all times. This way the default state of the system is always working and your first failure isn't invisibly critical.

          But this is very hard as we see in TCP when things get out of order and high latency creeps in. You have to manage a lot more state at the data level.

        2. Polizeiposaune · · focus · HN ↗
          Contractually dual-path, or actually dual-path? My understanding is that there's enough infrastructure horse-trading going on behind the scenes that it's very difficulty to be certain that two circuits between points A and B don't share the same infrastructure somewhere in between.

          First big oops of this form that I remember:

          "In December 1986, the ARPANET had 7 dedicated trunk lines between NY and Boston, except that they all went through the same conduit -- which was accidentally cut by a backhoe. "

          <a href="https:&#x2F;&#x2F;www.csl.sri.com&#x2F;~neumann&#x2F;insiderisks06.html" rel="nofollow">https:&#x2F;&#x2F;www.csl.sri.com&#x2F;~neumann&#x2F;insiderisks06.html

          1. wrs · · focus · HN ↗
            I&#x27;m talking about data not able to leave the building, not even data getting lost on the way.

            One example was a site that had fiber and coax, from different companies. They might have shared a pipe at some point along their length, hard to say. But both connections went down at the same time for digital reasons, not physical. The providers had simultaneous unrelated backend router problems and the site lost contact to both gateways.

            This was with a nice SD-WAN system that routed everything dynamically across both pipes to address latency and errors, but if the packets get dropped at the first hop in both networks, there&#x27;s not much you can do!

            The saving grace in that case was the third redundant connection, a 4G cell modem, with a lot less bandwidth but able to keep the critical transactions going. These days I&#x27;d certainly want Starlink on the roof as well.

          2. FireBeyond · · focus · HN ↗
            Yup. Seven different paths comes with seven different sets of equipment, and seven different rights-of-way that have to be negotiated, purchased, leased, whatever.
      2. cmiles8 · · focus · HN ↗
        Reporting is saying that the backup was cut a while ago and they only discovered it when the primary failed. Apparently nobody was ping testing the backup link.
    2. readthenotes1 · · focus · HN ↗
      Do you want air traffic control to be on the general internet?

      That seems like an extremely foolhardy thing to do.

      1. ssl-3 · · focus · HN ↗
        As a tertiary backup: Yeah, maybe I do want that.

        It seems like it would present a less-chaotic solution than that provided by having no data communications at all.

      2. kqgnkqgn · · focus · HN ↗
        You certainly don’t need to run this over public Internet to get resiliency over multiple paths.
      3. bell-cot · · focus · HN ↗
        Why not? Assuming it&#x27;s competently set up - serious encryption, details not blabbered about (to attracted DDoS or whatever), etc.
        1. topspin · · focus · HN ↗
          Competency appears to be the limited resource here.
          1. bell-cot · · focus · HN ↗
            True. But what if there are more people competent to set up &amp; maintain secure tunnels through the internet than there are people competent to set up &amp; maintain dedicated fiber links?
            1. topspin · · focus · HN ↗
              But what if the actual bottleneck isn&#x27;t the size of the various cohorts of people, but a bureaucracy that inflicts some irrational set of conditions that impedes competent operation? If so, then changing media won&#x27;t help; whatever scheme you envision will fail.

              I&#x27;m rather certain the FAA does indeed either employ or contract people that are capable of providing the solution. I&#x27;m also quite certain those people live in a perpetual blizzard of responsibility diffusion bullshit: the FAA is a venerable and highly risk-adverse operation. If someone told me it takes months of meetings and sign-offs to so much as log into a router and perform a benign check of some setting, I wouldn&#x27;t blink.

              1. bell-cot · · focus · HN ↗
                Yes, but low-functioning bureaucracies are also the most prone to outsourcing. They can&#x27;t manage to get anything done internally, so they hire outside orgs (hopefully higher functioning) to perform necessary basic services. Often late in the game, when (figuratively) a VIP visit could reveal that their bureaucracy is too broken to keep its office bathrooms clean.

                In this case, that outsourcing could easily look like &quot;call local ISP&#x27;s, get connections, set up tunnel&quot;.

                1. topspin · · focus · HN ↗
                  &gt; outsourcing

                  One of the more popular means of diffusing responsibility. A mere symptom, in other words. Not a cause. There is a abundance of qualified contractors in our world that could provide a functioning fail-over solution using whatever medium one might imagine. Somehow, this does not happen. The problem is elsewhere.

    3. Hikikomori · · focus · HN ↗
      It is, but you need to set it up correctly.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.