‹ BackHN Continuity

Thread

A 40ms Go garbage collector pause caused by swap

92 points · 85 comments · shellpipe

  1. jacobgold · · focus · HN ↗
    "It hurts when I do this"

    "Stop doing that"

    If you care about latency, disable swap. System wide or for the specific the cgroup.

    1. delamon · · focus · HN ↗
      This is not entirely correct. If you care about latency, then it doesn't matter on which major fault your application gets paused. Disabling swap protects you from data pages being evicted, but code pages can still be paged out.

      If you care about latency, mlock() your memory, do not disable swap. Swap is good and gives the kernel an equal opportunity to evict data and code pages.

      1. xxs · · focus · HN ↗
        For GC enabled languages swap is universally bad. Some gc-pauses are indistinguishable from a system crash. It's a side effect on not having tightly specified memory limits.

        I'd rather have applications be oom_killed than having them swap out, the former is rather obvious and demands action.

        1. delamon · · focus · HN ↗
          If you want you application to stay in memory, then make it explicitly with mlock()/mlockall().

          Disabling swap will just moves pressere elsewhere: to code pages. And evicted code page is no better: full stall while kernel loads that page from disk.

          1. yvdriess · · focus · HN ↗
            Disabling swap disables loading a process' binary to memory? Or do you mean that evicted code pages will need to get loaded from disk every time.
            1. MobiusHorizons · · focus · HN ↗
              Under memory pressure the system will free the ram used for code pages because it can always load them back from the executable on disk. It’s the same virtual memory mechanism as swap but without needing dedicated swap space. The op is saying disabling swap doesn’t prevent long pauses during memory pressure, because the system just swaps code out instead of dynamically allocated memory.
            2. delamon · · focus · HN ↗
              There are two types of pages: anonymous and mapped from files. The code pages are mapped from a binary; they are not very special.

              Under memory pressure, the kernel evicts less popular pages from memory. If a page has been mapped from a file, it is dropped (if dirty, then it is written out first). If it is needed later, the kernel can read it back from the file. If a page is anonymous (read: heap page), then there is no backing file and the kernel copies it to swap before dropping it. This is swapping.

              So, what happens if you disable swap and the kernel is low on memory? What can it evict? Anonymous pages cannot be evicted: there is no swap to put a copy in. The only choice the kernel has is to evict pages that are mapped from files. Those include pages mapped from the executable. You don't eliminate stalls by disabling swap, you just move them elsewhere: the kernel will page out code and your app gets paused whenever the execution flow hits such a page.

              1. mrob · · focus · HN ↗
                >So, what happens if you disable swap and the kernel is low on memory? What can it evict?

                Your userspace early OOM killer triggers and lets you know you're trying to run more than will fit in memory so you don't do it again. (In my experience, the kernel OOM killer can't be trusted to kill processes soon enough.)

              2. yvdriess · · focus · HN ↗
                Thanks for the clarification, learning everyday.

                I haven't had to analyze the performance of no-swap processes before. My assumption is that code is hot enough to avoid eviction and that evicted code pages are rather the exception. To strong-man the argument, I can imagine long running complex (bloated) services could have parts that are not touched unless a specific request comes in.

            3. dwattttt · · focus · HN ↗
              There are good sibling replies answering your question, but to give a specific example: if your application writes something but then doesn't need it anymore. Since we're talking about go, perhaps a data structure you need also holds a reference to data you don't need.

              When the kernel needs memory, it goes hunting for a page it can discard. But since that's transparent, the kernel can only discard a page if it knows it can get it back (after all, it's still got valid data on it).

              If there's swap, a page full of stale/unneeded data can be written out to swap. But if there's no swap, your page of "dangling data that you'll never use, but is still valid & referenced" can't be discarded; the kernel doesn't know you won't want it later, and it can't recreate the page if it throws it away.

              So like sibling said, at that point it has to find other pages it can evict from memory, ones that _do_ have somewhere persistent they can be written out to. Pages loaded from binaries on disk satisfy that, so those will get dropped instead.

              1. xxs · · focus · HN ↗
                For JIT languages most of the code is dynamically generated, not available to be loaded back from the disk. Morealso, the allocations do happen at large chucks where the oom_killer also happens, so it's significantly less of an issue.
                1. dwattttt · · focus · HN ↗
                  That doesn't necessarily improve matters. Now any code that gets JIT'd is also going to be permanently pinned in memory too, no matter whether it'll get used again, if you don't have swap.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.