‹ BackHN Continuity

Thread

Booted up in 1993, this server still runs – but not for much longer (2017)

193 points · 110 comments · doener

  1. colonwqbang · · focus · HN ↗
    When you think about it, 24 years is not such a long time for a machine to keep doing what it's supposed to be doing. It says something about our profession that we find a 24-year service life surprising.

    There are airliners that have been in service for over 50 years.

    1. [deleted] · · focus · HN ↗

      [deleted]

    2. afavour · · focus · HN ↗
      I think it just speaks to how quickly the tech has developed. Airliners work more or less the same as they did 50 years ago, computers are vastly different (and vastly more powerful).
      1. Leonard_of_Q · · focus · HN ↗
        Airliners may look the same but they don't work as they did 50 years ago when the was no fly by wire in these planes, no high-bypass engines, etc.
        1. yCombLinks · · focus · HN ↗
          Computers are many thousands of times faster and more efficient than they were 50 years ago. Airplanes still fly at the same speeds, and with only marginal safety improvements. Fuel efficiency has doubled however. Overall significantly less changes.
          1. Leonard_of_Q · · focus · HN ↗
            Faster but for the rest basically the same. More megahertz, more megabytes, more transistors per cm2 of silicon, more processors but that's basically it. Von Neumann architecture, processing separated from memory, block/file-based storage, input and output through keyboards and displays.

            Yes, this drastically understates the change in computing but the actual technology hasn't changed that much, it just got denser and faster.

        2. zrail · · focus · HN ↗
          Eh, sorta. Fifty years ago was 1976 (I know, right?!). The 747 debuted in 1970 with high bypass engines. The A320 came out in 1988 with fly-by-wire and all Airbus aircraft since have been the same.

          The biggest change from then to now is probably the amount of solid state electronics onboard. That's ridden roughly the same curve as other industrial and commercial applications.

    3. vanviegen · · focus · HN ↗
      It's not so much surprising that the machine hasn't broken, but that someone apparently considered it worthwhile to keep it around, having about the compute of a throwaway vape.
    4. streetfighter64 · · focus · HN ↗
      Airliners are up against physical limitations that don't incentivize upgrading. If there was an airliner that could fit millions of people and cost less than the 1980s version, I don't think you'd see many 80s aircraft around anymore, except in niece applications such as this computer. Perhaps now that we're at the end of Moore's law, you might see a larger amount of long-lived computers. But on the other hand, we'll probably find new ways to innovate in computing power.
      1. StilesCrisis · · focus · HN ↗
        Computers are still evolving rapidly. Instead of raw GHz increases, it will be things like tensor cores or other extremely-wide/GPU-type processing units. And of course, more RAM and more caches to feed them.
    5. asah · · focus · HN ↗
      yabut... airliners require constant maintenance and periodic rebuilds.
      1. andrew_lettuce · · focus · HN ↗
        Not to mention service disruptions and unplanned downtime.
      2. [deleted] · · focus · HN ↗

        [deleted]

      3. bitwize · · focus · HN ↗
        One of the major differences between like, your PC, and a high-availability server is that high-availability servers require those things as well—and get them, piecewise, while still running. It's like doing a complete overhaul while the plane is flying.

        IBM mainframes, famously, phone home if they detect a faulty component. Service personnel will be on site, same day, to replace it before you even knew it was there, let alone had the opportunity to ask.

        1. organsnyder · · focus · HN ↗
          I wonder how many datacenters bother with those sorts of things anymore, given how many workloads are no longer tied to individual nodes. Seems like it would just be easier to accept that a certain percentage of nodes will be down at any one time, and go through and repair/replace them as needed.
          1. bitwize · · focus · HN ↗
            Cloudslop doesn't provide high availability the way high-availability platforms provide high availability. IBM is still doing brisk mainframe business. You won't find mainframes in many data centers because you need nearly a data center's worth of hardware to provide the volume and availability of one mainframe, and that can fit in one cabinet. (Mainframes also benefited from VLSI; new ones aren't dinosaurs that require three-phase power, their own floor of a building, and their own aircon infra.)
    6. Transformanshen · · focus · HN ↗
      That may well be the case, but such a long period still seems extraordinary
    7. wazoox · · focus · HN ↗
      There are industrial machines such as steam hammers that have been in constant use for more than a century.
      1. Aardwolf · · focus · HN ↗
        This is the most dwarven thing I've read today
    8. tyrabound · · focus · HN ↗
      You raised in my mind that we are looking at this all wrong, the server is not actually one component, it’s really a system of systems, and just like an airliner is a system of systems too, the conflict arises in that the comparison is at the wrong level.

      It seems when people think of a server today, they think of a single computer, if not some software server somewhere in the cloud. These subject kinds of servers are far closer to complex systems of separate components, more like a network system, it’s why components of what are really separate networked computers can and need to be replaced.

      An airliner is of course also made up of components that are also systems, however at least in my mind, the difference is the actual expected use case. An airliner is not a good comparison because it is never expected to remain in continuous operation, e.g., that at least one engine is always running even when it is being overhauled or repaired, to satisfy a requirement of continuous operation.

      1. pitched · · focus · HN ↗
        FTA, this is a fault-tolerant server where all hardware is redundant and can be hot swapped while running. The servers you’re thinking about are redundant at the software level so rebooting one server won’t cause the service to go down.

        What is remarkable is that apparently that VOS thing has never crashed. I wonder if they also do software redundancy under the hood to keep uptime going during reboots. If the CPU is hotswappable, it must have something.

        1. wildzzz · · focus · HN ↗
          Probably uses formally verified code along with plenty of housekeeping processes to keep any failures from shitting the whole bed. In critical system design, you build in redundancies that work in parallel such that any one failure will not interrupt the system.

          The main computer system in the Space Shuttle is an excellent example of this. It had 5 identical IBM System/4 Pi machines. Three of them ran identical code and handled the same work. The fourth ran a completely different codebase to handle the same work, preventing a bug in the main code from killing the whole system. A fifth computer handled other tasks but could be swapped over to the critical role if needed. You could lose 2/5 computers and still have insurance against a cosmic ray flipping a bit.

        2. serf · · focus · HN ↗
          VOS is a parallel lockstep OS. You drop nodes and replace them to keep the whole operational.
        3. theamk · · focus · HN ↗
          It's really not that hard to keep computers non-crashing, as long as you have good hardware and run a limited subset of software.

          Many servers I've owned had multiple years of uptime, and the only reason they'd go down is because they will get decommissioned or because of power outage.

          The article says:

          > disk drives, power supplies and some other components have been replaced but Hogan estimates that close to 80% of the system is original.

          so I am guessing there was no reboots, nor CPU replacements.

    9. lp92 · · focus · HN ↗
      Those aircraft have maintainance windows/refurbishment/refits and aren't flying 24/7 though.
      1. debo_ · · focus · HN ↗
        This server also had planned downtime maintenance. The article specifically mentions "no unplanned downtime."
    10. VectorLock · · focus · HN ↗
      >There are airliners that have been in service for over 50 years.

      Are there any airliners that have flown for 24 years non-stop?

      1. russianGuy83829 · · focus · HN ↗
        That would not be a useful airliner.
        1. literalAardvark · · focus · HN ↗
          Mr Bones wild ride

          But really didn't the NSA have something that was roughly an airliner in perpetual flight?

          1. quietsegfault · · focus · HN ↗
            Spy blimp!
      2. serf · · focus · HN ↗
        this is why I find comparing mostly-solid-state stuff with big mechanical contraptions a frustrating exercise.

        it's more impressive to me that the electric fans within the Stratus server still operate than it is that the chips still work, and if we're going with the plane metaphor it's the silicon that puts the thing into the air, not the simple fans.

        So really it just boils down to "well, the planes work is harder." as to why it fails, which is of course abstract and unsatisfying; comparing the comparative lifetime workloads of a chip that has shifted trillions of bits versus some jet engines that have moved millions of kilograms around the world.

        and then another layer comes and makes it even weirder : the vast majority of hardware around the world that has been EoLd and replaced has been in that situation due to software, not the hardware itself; a paradigm that really doesn't exist in the physical engineering world in the same was as it does CS.

        1. semi-extrinsic · · focus · HN ↗
          Mechanical integrity in industrial processes is often divided into two categories: static equipment (tanks, pipes, heat exchangers...) and rotating equipment.

          These two have very distinct failure modes. The failure mechanisms for solid state electronics are much more similar to those for static equipment (corrosion, thermal fatigue, creep, ...)

        2. nextaccountic · · focus · HN ↗
          The chips themselves aren't surprising, but there are other electronic components that may fail earlier like capacitors
      3. dlisboa · · focus · HN ↗
        Yeah, it’s not a good comparison. Give me endless parts and a maintenance window of 6 months every so often and I can make any server last a century.
      4. giancarlostoro · · focus · HN ↗
        Was going to say, there's servers that last several decades, they receive routine maintenance, most people buy a computer, and rarely replace parts, and if enough years pass, you opt to buy a new one instead.
      5. jrootabega · · focus · HN ↗
        I'm not sure, but I think there are a few submarines that have remained underwater for that long.
        1. loloquwowndueo · · focus · HN ↗
          Name one.
          1. jrootabega · · focus · HN ↗
            I think most of the time the navy does that.
            1. loloquwowndueo · · focus · HN ↗
              You could have just said “I can’t”
        2. literalAardvark · · focus · HN ↗
          Absolutely not. They come up for shift changes at the very least, and more.
        3. ok_dad · · focus · HN ↗
          Naval vessels are some of the most maintained machines in the world. They get hauled out for hull repairs routinely, or put into long term maintenance facilities for months while they are repaired. Probably second only to airplanes and space craft.
      6. rurban · · focus · HN ↗
        You&#x27;ll find for sure an old DC3 somewhere, and a Cessna 172 also. And according to <a href="https:&#x2F;&#x2F;www.oldest.org&#x2F;technology&#x2F;planes-still-flying&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.oldest.org&#x2F;technology&#x2F;planes-still-flying&#x2F; there are still 40 year old 737&#x27;s in the air
    11. glitchc · · focus · HN ↗
      &gt; There are airliners that have been in service for over 50 years.

      ...with a very aggressive maintenance cycle (typical part life is 1000-10000 hrs of service). Most airliners of that vintage resemble the &quot;Ship of Theseus&quot; in real life.

    12. verzali · · focus · HN ↗
      Isn&#x27;t mostly capacitors that limit the lifespan? With a proper maintenance plan (as airliners follow) you could probably keep servers running for a long time.
    13. PunchyHamster · · focus · HN ↗
      Those are getting regular maintenance and overhauls while not flying, it&#x27;s not really comparable.

      On other side only moving element would be hard drives and only wear element really would be the electrolytic caps.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.