‹ BackHN Continuity

Thread

Ask HN: Who's still keeping a DOS machine up because the business depends on it?

297 points · 304 comments · mlaux

  1. freeli · · focus · HN ↗
    A certain nuclear power plant had a Windows NT 4.0 machine running as late as 2007. The reason is interesting.

    The machine's purpose was to report status of the control rods that mitigate nuclear reactions. Basically, "are the rods inserted, and if so, how many / how far?". I want to emphasize that this was reporting only, NOT control.

    The original software was written back in the 80's, when the plant was originally commissioned, for AmigaOS. Of course, it's hard to buy Amigas anymore, and the original one died long ago (nobody remembers when).

    So in the mid '90s, the utility purchased an AmigaOS emulator that ran on Windows NT 4.0, which was current at the time. The emulator (IIRC) was developed by a firm in the UK. The firm went out of business sometime in the late '90s. The control rod monitoring software ran under this emulator on top of NT4.

    Windows NT 4.0 was the last OS to allow the emulation software direct access to the physical hardware that produced the status signal. Later versions of Windows abstracted the hardware access away, and the monitoring software broke. Because the emulation company had gone belly up, there was no way to fix the incompatibility.

    So the utility had a choice: get new hardware/software certified (by NRC?), or keep doing what they were doing with the software (and hardware) that they had. They chose the latter.

    So this is how, in 2007, during a tour of the facility, I stumbled across a Pentium 1 system running an AmigaOS emulator on Windows NT 4.0 that was responsible for displaying the status of the control rods of a nuclear power plant.

    Spare hardware for this setup was purchased off of eBay and stocked on an adjacent shelf.

    1. sajithdilshan · · focus · HN ↗
      This is crazy. Why on earth wouldn’t the respective government spend money to modernize a critical infrastructure like a nuclear power plant?

      This is how we would end up with nuclear disasters, not because the technology is bad, but purely because of mismanagement and human negligence.

      1. onion2k · · focus · HN ↗
        This is crazy. Why on earth wouldn’t the respective government spend money to modernize a critical infrastructure like a nuclear power plant?

        Safety and reliability come from understanding the system. Something that's old but that you know everything about is far safer than something new. This risk is that the people who know move on, or the sources of replacement parts stop working, or that other things around the system change. Then the risk curve inverts and you find the old system more of a liability than a asset.

        Almost always people choose to update things either too early or too late. Knowing when to do something is hard.

        1. sajithdilshan · · focus · HN ↗
          That’s pure negligence. It’s not like the people that have the know how disappear suddenly . It’s the responsibility of the management to make sure the knowledge transferred to new generation.

          I guess that kind of thinking is what led to the collapse of ancient civilisations like why even bother to improve anything and let everything decay and die out

          1. jodrellblank · · focus · HN ↗
            Basically everything collapses as far as it is allowed to, and the collapse is slowed only by:

            - profit, if a system generates it, people will keep the system running.

            - dilligence, usually driven by personal idealism, and unrewarded.

            - legal or financial penalties, or insurance costs.

            - it has collapsed to a temporarily stable state for now.

            For another comment I was looking at the Grenfell Tower fire in the UK around 2017 and the people who lived there had been raising fire risks for years and the landlord (management organization) and the council had been ignoring them, and ignoring the fire department. The companies which replaced the cladding on the tower with flammable cladding were choosing the cheaper flammable option and pointing fingers as well. And during the fire, the fire service had never dealt with such a big fire and didn't have a truck with a long ladder, didn't have extended-breathing gear, had radio problems, water pressure problems from the local water company (who deny that).

            This kind of backstory is typical for disasters, and for IT outages - and I've taken to believing that if something "should work" but wasn't tested recently then you should expect that it doesn't work. This is a common saying in backups (you need to test restoring), but it seems to apply everywhere. Backup internet connections that don't have enough bandwidth for the company to keep running. Disaster recovery sites that share resources with the production site. 'Disaster recovery' that recovered into a remote site, but they couldn't do CAD work remotely over their slow connection so it was still disastrous. Multiple power feeds but they were cross-wired so one failing still functionall took down everything.

            If knowledge transfer "should be happening" but it's not critical to a job, or it's not profitable, or it's not legally required and audited, then it isn't happening. Subject to the dilligent idealist employee mentioned above who are temporary in the long view. Management's involvement seems not to arrange the larger system to effectively shoulder responsibility, but rather find ways to avoid personal blame while cutting costs beyond the point where everything that needs to happen keeps happening.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.