‹ BackHN Continuity

Thread

Jemalloc 5.4.0

339 points · 94 comments · gkfasdfasdf

  1. skavi · · focus · HN ↗
    does anyone familiar with the art have thoughts on why only tcmalloc switched from thread caches to cpu caches? would it make linux behavior diverge too much from other platforms?
    1. rwmj · · focus · HN ↗
      (Not an expert but ...) unless you pin threads to cores, which is not the default and somewhat awkward in Linux for user applications, having a per-thread cache doesn't really make sense as your thread could be moved to another core and then your cache will no longer be local to the physical cache.
      1. skavi · · focus · HN ↗
        i think we agree that per cpu caching seems superior. i’m looking for the other side of this. most allocators seem to have stuck with per thread.
        1. Someone · · focus · HN ↗
          If you use a thread-local data structure, your allocator can pretend that it is running on a single-core, single-task system.

          If you use a CPU-local data structure, you must handle the case where, mid-way through a call to your allocator, the CPU runs a second thread that makes another call to your allocator (and that, too, can get interrupted by another thread that allocates memory, etc.)

          That makes thread-local easier to implement and likely faster (it doesn’t require any memory barriers in the fast path)

          Also, good schedulers try to avoid moving threads between CPUs. The better they manage to do that, the lower the cost of having per thread data structures (there likely still is a price, as there most of the time are more threads than CPUs on a system)

          1. skavi · · focus · HN ↗
            <a href="https:&#x2F;&#x2F;google.github.io&#x2F;tcmalloc&#x2F;rseq.html" rel="nofollow">https:&#x2F;&#x2F;google.github.io&#x2F;tcmalloc&#x2F;rseq.html

            i don’t believe rseq based cpu local caches require memory barriers on the fast path.

            1. Someone · · focus · HN ↗
              <a href="https:&#x2F;&#x2F;lwn.net&#x2F;Articles&#x2F;1033957&#x2F;" rel="nofollow">https:&#x2F;&#x2F;lwn.net&#x2F;Articles&#x2F;1033957&#x2F;:

                rseq_cs
              
                 The rseq_cs field is a pointer to a struct rseq_cs. Is is NULL when no
                 rseq assembly block critical section is active for the registered
                 thread. Setting it to point to a critical section descriptor (struct
                 rseq_cs) marks the beginning of the critical section.
              
              I’m not sure I fully understand that man page (it never seems to say callers have to clear that field at the end of a critical section, for example), but doesn’t that mean the caller has to guarantee setting rseq_cs happens_before any code in the critical section? That’s a memory barrier.
              1. Veserv · · focus · HN ↗
                That is because you do not need to clear the field at the end of a critical section. It contains the contiguous instruction range where it fires so there is no problem with leaving it active forever unless you have another critical section where you want to use it.

                No explicit memory barrier is required anywhere as the value is only read in supervisor mode and a privilege switch implicitly issues a LS-LS barrier on all major architectures. Even if you did not want to rely on that, you would only need a single S-LS barrier when you store the control structure the very first time.

                1. Someone · · focus · HN ↗
                  &gt; That is because you do not need to clear the field at the end of a critical section. It contains the contiguous instruction range where it fires so there is no problem with leaving it active forever unless you have another critical section where you want to use it.

                  Aha! So, to take advantage of that, a memory allocator uses the same abort handler for all operations?

                  1. ckennelly · · focus · HN ↗
                    TCMalloc has different abort handlers for each function: If preempted, we need to know where to restart.

                    If `rseq_cs` is no longer describing a relevant address, that is, the program counter has moved past it, the kernel just ignores it.

              2. fweimer · · focus · HN ↗
                It&#x27;s just a compiler barrier (signal fence), not a memory barrier that concerns the CPU. The CPU is free to reorder loads and stores.
              3. ckennelly · · focus · HN ↗
                TCMalloc TL here :)

                The code is executed by a single thread. Everything retires in program order. There is no need for a memory barrier between starting the critical section and its body.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.