‹ BackHN Continuity

Thread

We broke an Over-The-Air update on the ESP32 on purpose

14 points · 4 comments · adunk

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. dole · · focus · HN ↗
    tl;dr: Not broken. "Our experiment show that we are able to both survive and recover our interrupted OTA attempts. This is what we expected: we were expecting the raw OTA mechanism in the Espressif IDF to be solid. Across every configuration we interrupted, the device found its own way back through ESP-IDF's own rollback mechanism."
  2. leoedin · · focus · HN ↗
    What an empty article. Testing an OTA update can survive a power interruption during the downloading is such a basic test. It doesn't need an article that long. I guess AI wrote it from a few sentence prompt?

    They didn't do any interesting tests. What about brownouts, voltage ripples, timing the reset signal to see if there's any critical moments in the update process, corrupted update files, high EMC environment etc etc. That would be interesting.

    I think the article is an advert for a testing platform that enables automating this sort of test. But weirdly they don't show how their platform does the automating, so it just looks like they're really pleased with themselves for doing really basic engineering.

    1. rckoepke · · focus · HN ↗
      Indeed. Would be nice to have a rock-solid “correct” example to critique and learn from, which goes into detail about all the traps for new players, what failure states / eFuse settings could be truly unrecoverable even if you have remote control of the MCU’s PWR, COM, and RST, etc.

      Years ago I had to design some robust OTA systems for both ATmega’s and ESP32’s. I was very confident that my ATmega solution could recover from absolutely any failure state that wasn’t a true “not my fault” hardware failure. At the time though, I was never quite sure if I properly covered every last edge case of the ESP32, which has significantly higher complexity of things that can go wrong from a bad OTA firmware update.

    2. pseudohadamard · · focus · HN ↗
      > Testing an OTA update can survive a power interruption during the downloading is such a basic test. It doesn't need an article that long.

      Tell that to the developers of containerloads of IoS (Internet of Shit) devices that can barely manage an OTA update under perfect conditions without bricking themselves.

      While you're at it, also let a well-known company that I'll leave unnamed know for their Windows Update service.

  3. ec109685 · · focus · HN ↗
    I was expecting something much more rigorous than a few trials resetting the controller at random points in its update process.
  4. rurban · · focus · HN ↗
    You just enough flash ram for 2 firmwares and a proper bootloader, which knows which partition is active/verified. Simple as that. Biggest problem is getting enough flash. writing the bootloader is trivial.
    1. pseudohadamard · · focus · HN ↗
      +1 to the above. And while a works-under-every-failure-case bootloader isn't quite that trivial, it's still absolutely nothing compared to telling the hardware guys that you want twice as much flash as the SoC they're using can manage.
  5. mikewarot · · focus · HN ↗
    It was circa 1985 that I encountered a similar problem. I hand a hand held barcode scanner that had to upload the collected data for reporting to a PC.

    The customer asked "what happens if I disconnect this right now?" during the upload. It was a corner case I hadn't considered. It only took a few days to make it bulletproof.

    I'd be worried about corruption in the download, and check for that. I'd also make very sure the watchdog hardware was on and the code respected it.

    It seems to me you need 3 buffers for the OTA code, not 2. There should be a way to keep the last version that runs for X seconds in addition to any new updates, where X is quite large.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.