does anyone familiar with the art have thoughts on why only tcmalloc switched from thread caches to cpu caches? would it make linux behavior diverge too much from other platforms?
(Not an expert but ...) unless you pin threads to cores, which is not the default and somewhat awkward in Linux for user applications, having a per-thread cache doesn't really make sense as your thread could be moved to another core and then your cache will no longer be local to the physical cache.
Exactly. You want memory arenas that are hot in this CPU's caches. If your thread moves, its per-thread caches are now elsewhere. Original TCMalloc was developed in the days of 2-4 core servers. Current TCMalloc was an evolution in the context of 32+ core servers.
my feeling is that the space efficiency gains are probably more significant than the reduction in core migration costs. many applications have far more threads than the system has cores.
Sure, also true that the per-CPU scheme co-evolved with the proliferation of services with thread-per-request architectures having way more TIDs than cores.
skavi · · focus · HN ↗
rwmj · · focus · HN ↗
jeffbee · · focus · HN ↗
skavi · · focus · HN ↗
jeffbee · · focus · HN ↗