Ask HN: Who's still keeping a DOS machine up because the business depends on it?
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Ask HN: Who's still keeping a DOS machine up because the business depends on it?
Unofficial Hacker News client; not affiliated with Y Combinator.
freeli · · focus · HN ↗
The machine's purpose was to report status of the control rods that mitigate nuclear reactions. Basically, "are the rods inserted, and if so, how many / how far?". I want to emphasize that this was reporting only, NOT control.
The original software was written back in the 80's, when the plant was originally commissioned, for AmigaOS. Of course, it's hard to buy Amigas anymore, and the original one died long ago (nobody remembers when).
So in the mid '90s, the utility purchased an AmigaOS emulator that ran on Windows NT 4.0, which was current at the time. The emulator (IIRC) was developed by a firm in the UK. The firm went out of business sometime in the late '90s. The control rod monitoring software ran under this emulator on top of NT4.
Windows NT 4.0 was the last OS to allow the emulation software direct access to the physical hardware that produced the status signal. Later versions of Windows abstracted the hardware access away, and the monitoring software broke. Because the emulation company had gone belly up, there was no way to fix the incompatibility.
So the utility had a choice: get new hardware/software certified (by NRC?), or keep doing what they were doing with the software (and hardware) that they had. They chose the latter.
So this is how, in 2007, during a tour of the facility, I stumbled across a Pentium 1 system running an AmigaOS emulator on Windows NT 4.0 that was responsible for displaying the status of the control rods of a nuclear power plant.
Spare hardware for this setup was purchased off of eBay and stocked on an adjacent shelf.
sajithdilshan · · focus · HN ↗
This is how we would end up with nuclear disasters, not because the technology is bad, but purely because of mismanagement and human negligence.
onion2k · · focus · HN ↗
Safety and reliability come from understanding the system. Something that's old but that you know everything about is far safer than something new. This risk is that the people who know move on, or the sources of replacement parts stop working, or that other things around the system change. Then the risk curve inverts and you find the old system more of a liability than a asset.
Almost always people choose to update things either too early or too late. Knowing when to do something is hard.
sajithdilshan · · focus · HN ↗
I guess that kind of thinking is what led to the collapse of ancient civilisations like why even bother to improve anything and let everything decay and die out
somenameforme · · focus · HN ↗
I think almost nobody from that era would have believed you if you told them that more than 50 years later not a single human would have traveled beyond low Earth orbit, and that we'd now be struggling to recreate what we did in 1969. Furthermore a significant chunk of the world doesn't even believe we landed on the Moon anymore, no doubt in part because of this apparent anachronism.
It's not like there was any sort of collective or even singular decision to just let things die out. I mean sure Nixon did his thing intentionally, but even as a President in the golden years of the US, he was still just one man.
ac29 · · focus · HN ↗
Four people went around the moon earlier this year (leaving earth's orbit for the first time since '72)
onion2k · · focus · HN ↗
That's just lazy thinking, and probably a sign you've never been a manager. For all the responsibility to lie with management you'd have to be leading people who enthusiastically do what they're asked to do, raise problems, document things, follow processes, and make sure everything is handed over to the next person when they leave.
That's not how people are. They'll have times when they're unmotivated, underpaid, grumpy, over-worked so they miss things, etc. That's when the organisational tech debt starts piling up, and there's nothing a manager can do to stop it except manage it as best they can even though they're under the same pressure, with the same lack of motivation and crap pay as everyone else.
It's so easy to think you can fix it just by keeping on top of it and micromanaging where it starts to show up. It doesn't work that way.
jagged-chisel · · focus · HN ↗
These two (and many others) are well within management’s purview.
> … there's nothing a manager can do … [emphasis added]
Right, but Management (capitalized and plural) can. A manager should report, and management should fix - if not, it’s management’s fault.
onion2k · · focus · HN ↗
The fact that you're willing to absolve an individual manager but still believe Management can fix the problems is a big part of the problem. The Management usually includes the last person in the chain who has the ultimate decision making power. If you assume that when I said 'a manager' I meant that person, it should be fairly obvious that even they don't have unlimited budget or unlimited time, and they can only try to fix problems they hear about. Within those constraints Management are still going to find fixing systemic problems hard.
It's very easy to look at something with hindsight and think the outcome should have been predicted and corrected. If you look at the real world you'll see that just isn't the case. Personally, I wish we had a lot more people applying some systems thinking approaches to at least consider the second- and third-order impacts of their decisions but I imagine that won't ever happen. It makes me sad because it's actually really interesting.
jodrellblank · · focus · HN ↗
- profit, if a system generates it, people will keep the system running.
- dilligence, usually driven by personal idealism, and unrewarded.
- legal or financial penalties, or insurance costs.
- it has collapsed to a temporarily stable state for now.
For another comment I was looking at the Grenfell Tower fire in the UK around 2017 and the people who lived there had been raising fire risks for years and the landlord (management organization) and the council had been ignoring them, and ignoring the fire department. The companies which replaced the cladding on the tower with flammable cladding were choosing the cheaper flammable option and pointing fingers as well. And during the fire, the fire service had never dealt with such a big fire and didn't have a truck with a long ladder, didn't have extended-breathing gear, had radio problems, water pressure problems from the local water company (who deny that).
This kind of backstory is typical for disasters, and for IT outages - and I've taken to believing that if something "should work" but wasn't tested recently then you should expect that it doesn't work. This is a common saying in backups (you need to test restoring), but it seems to apply everywhere. Backup internet connections that don't have enough bandwidth for the company to keep running. Disaster recovery sites that share resources with the production site. 'Disaster recovery' that recovered into a remote site, but they couldn't do CAD work remotely over their slow connection so it was still disastrous. Multiple power feeds but they were cross-wired so one failing still functionall took down everything.
If knowledge transfer "should be happening" but it's not critical to a job, or it's not profitable, or it's not legally required and audited, then it isn't happening. Subject to the dilligent idealist employee mentioned above who are temporary in the long view. Management's involvement seems not to arrange the larger system to effectively shoulder responsibility, but rather find ways to avoid personal blame while cutting costs beyond the point where everything that needs to happen keeps happening.