‹ BackHN Continuity

Thread

A 40ms Go garbage collector pause caused by swap

48 points · 11 comments · shellpipe

Loading the complete thread in the background. This saved snapshot is available now. Refresh

  1. octoberfranklin · · focus · HN ↗
    Yet another reason why "no runtime" is such a huge advantage.
    1. my-next-account · · focus · HN ↗
      You're probably coding for a 'runtime' in the broad sense: libc and your kernel.
  2. pizlonator · · focus · HN ↗
    Is there a reason why Go isn’t using on the fly GC, where there’s no STW at all?
    1. klodolph · · focus · HN ↗
      What I found when I searched for &quot;on the fly&quot; is this: <a href="https:&#x2F;&#x2F;github.com&#x2F;mthom&#x2F;on-the-fly-gc" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;mthom&#x2F;on-the-fly-gc — it looks like it has pauses, from the docs.

      In general when you tune knobs for GC, you pay for benefits in one area with sacrifices in another. Two big knobs to turn are pause latency and throughput. You probably wouldn’t want to go full “optimize for latency” because you’d end up with poor throughput. Also vice versa. Java’s reputation for poor GC performance is partly due to historical defaults that tune it for throughput.

      1. pizlonator · · focus · HN ↗
        The classic on the fly GC algorithm is DLG, hilariously published in two papers, because the first one had a bug. Here&#x27;s the second paper: <a href="https:&#x2F;&#x2F;caml.inria.fr&#x2F;pub&#x2F;papers&#x2F;doligez_gonthier-gc-popl94.pdf" rel="nofollow">https:&#x2F;&#x2F;caml.inria.fr&#x2F;pub&#x2F;papers&#x2F;doligez_gonthier-gc-popl94....

        It&#x27;s a well known algorithm. Folks who do GCs for a living know about it. The folks who work on Go are surely aware of it. I&#x27;m assuming that they do not use it for a good reason, hence my question!

        Fil-C&#x27;s GC (Fil&#x27;s Unbelievable Garbage Collector) uses an alternative on-the-fly algorithm, which I call Phil&#x27;s Concurrent Marking.

        I&#x27;ve documented it here: <a href="https:&#x2F;&#x2F;fil-c.org&#x2F;fugc" rel="nofollow">https:&#x2F;&#x2F;fil-c.org&#x2F;fugc

        Here&#x27;s the source: <a href="https:&#x2F;&#x2F;github.com&#x2F;pizlonator&#x2F;fil-c&#x2F;blob&#x2F;deluge&#x2F;libpas&#x2F;src&#x2F;libpas&#x2F;fugc.c" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;pizlonator&#x2F;fil-c&#x2F;blob&#x2F;deluge&#x2F;libpas&#x2F;src&#x2F;l...

        Phil&#x27;s Concurrent Marking differs from DLG in that it only requires a Djikstra barrier and uses a permagrey stack (something that Go used to do).

        However, FUGC does clever things for coroutines (as in ucontexts, which Fil-C supports) - they are not permagrey; they only become grey if they execute. That&#x27;s relevant to Go because Go moved away from permagrey stacks because of coroutine scan overheads, which the FUGC coroutine strategy might avoid.

        But even if Go could not go back to permagrey, then the answer would be to use DLG, which would involve using the combined Yuasa+Dijstra barrier, which Go uses today anyway

  3. truth_seeker · · focus · HN ↗
    What is stopping the author to use latest version of Go and Linux Kernel ?
  4. soltanov · · focus · HN ↗
    A GC latency SLO should include operating-system memory pressure. Otherwise, a page-fault problem will look like a collector problem and lead to the wrong fix.
  5. jacobgold · · focus · HN ↗
    &quot;It hurts when I do this&quot;

    &quot;Stop doing that&quot;

    If you care about latency, disable swap. System wide or for the specific the cgroup.

    1. delamon · · focus · HN ↗
      This is not entirely correct. If you care about latency, then it doesn&#x27;t matter on which major fault your application gets paused. Disabling swap protects you from data pages being evicted, but code pages can still be paged out.

      If you care about latency, mlock() your memory, do not disable swap. Swap is good and gives the kernel an equal opportunity to evict data and code pages.

  6. ahmedmostafa16 · · focus · HN ↗
    The nasty bit is that swap doesn&#x27;t just make the allocation slower; if GC metadata gets paged out, you have turned memory pressure into a stop-the-world latency spike.
  7. raverbashing · · focus · HN ↗
    I don&#x27;t get why people do not prefer reference counting, it has more predictable runtime performance

    (though of course a swap is a swap - but you can &quot;trigger&quot; it depending on your memory or file access pattern)

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.