does anyone familiar with the art have thoughts on why only tcmalloc switched from thread caches to cpu caches? would it make linux behavior diverge too much from other platforms?
(Not an expert but ...) unless you pin threads to cores, which is not the default and somewhat awkward in Linux for user applications, having a per-thread cache doesn't really make sense as your thread could be moved to another core and then your cache will no longer be local to the physical cache.
If you use a thread-local data structure, your allocator can pretend that it is running on a single-core, single-task system.
If you use a CPU-local data structure, you must handle the case where, mid-way through a call to your allocator, the CPU runs a second thread that makes another call to your allocator (and that, too, can get interrupted by another thread that allocates memory, etc.)
That makes thread-local easier to implement and likely faster (it doesn’t require any memory barriers in the fast path)
Also, good schedulers try to avoid moving threads between CPUs. The better they manage to do that, the lower the cost of having per thread data structures (there likely still is a price, as there most of the time are more threads than CPUs on a system)
rseq_cs
The rseq_cs field is a pointer to a struct rseq_cs. Is is NULL when no
rseq assembly block critical section is active for the registered
thread. Setting it to point to a critical section descriptor (struct
rseq_cs) marks the beginning of the critical section.
I’m not sure I fully understand that man page (it never seems to say callers have to clear that field at the end of a critical section, for example), but doesn’t that mean the caller has to guarantee setting rseq_cs happens_before any code in the critical section? That’s a memory barrier.
skavi · · focus · HN ↗
rwmj · · focus · HN ↗
skavi · · focus · HN ↗
Someone · · focus · HN ↗
If you use a CPU-local data structure, you must handle the case where, mid-way through a call to your allocator, the CPU runs a second thread that makes another call to your allocator (and that, too, can get interrupted by another thread that allocates memory, etc.)
That makes thread-local easier to implement and likely faster (it doesn’t require any memory barriers in the fast path)
Also, good schedulers try to avoid moving threads between CPUs. The better they manage to do that, the lower the cost of having per thread data structures (there likely still is a price, as there most of the time are more threads than CPUs on a system)
skavi · · focus · HN ↗
i don’t believe rseq based cpu local caches require memory barriers on the fast path.
Someone · · focus · HN ↗
fweimer · · focus · HN ↗