FTA: “When multiple threads simultaneously allocate or deallocate memory from the allocator, the allocator will serialize them. Programs making intensive use of the allocator actually slow down as the number of processors increases.”
The article does later retract on that, but that’s no reason to lead with such a blatantly false (with current allocators) statement.
Also FTA “In 2006, a third pool was introduced (after operating system memory pool and library-based memory pool) called the “arena”. Arena is a jemalloc-term”
Jemalloc is from around 2005 (<a href="http://jemalloc.net/" rel="nofollow">http://jemalloc.net/), the idea of arenas is from the 1960s, and Wikipedia claims the term was coined in 1990 (<a href="https://en.wikipedia.org/wiki/Region-based_memory_management#History_and_concepts" rel="nofollow">https://en.wikipedia.org/wiki/Region-based_memory_management...), and the linked paper (<a href="https://www.cs.princeton.edu/techreports/1988/191.pdf" rel="nofollow">https://www.cs.princeton.edu/techreports/1988/191.pdf) is from 1988.
Then, a typo: “as well as memory tied to specific to each of the multiple CPU core or even CPU infinity.”
yup and the characterization of each allocator is so fuzzy, with zero methodology provided.
allocators are so simple to just swap into your program. if you can put together a few representative workloads, you should just try out a few allocators and profile whatever metrics you care about.
no article, sorry, grabbed the numbers from an old PR.
to confirm, you were using tcmalloc from <a href="https://github.com/google/tcmalloc" rel="nofollow">https://github.com/google/tcmalloc and not from <a href="https://github.com/gperftools/gperftools" rel="nofollow">https://github.com/gperftools/gperftools, right?
Good point. It was from the <a href="https://packages.debian.org/trixie/google-perftools" rel="nofollow">https://packages.debian.org/trixie/google-perftools Debian package, which points to the gperftools project - so it was the worse one I guess.
yeah that’s a common mistake when evaluating tcmalloc. gperftools tcmalloc diverged quite a while ago. doesn’t have a lot of the fancier features of modern tcmalloc [0].
Someone · · focus · HN ↗
FTA: “When multiple threads simultaneously allocate or deallocate memory from the allocator, the allocator will serialize them. Programs making intensive use of the allocator actually slow down as the number of processors increases.”
The article does later retract on that, but that’s no reason to lead with such a blatantly false (with current allocators) statement.
Also FTA “In 2006, a third pool was introduced (after operating system memory pool and library-based memory pool) called the “arena”. Arena is a jemalloc-term”
Jemalloc is from around 2005 (<a href="http://jemalloc.net/" rel="nofollow">http://jemalloc.net/), the idea of arenas is from the 1960s, and Wikipedia claims the term was coined in 1990 (<a href="https://en.wikipedia.org/wiki/Region-based_memory_management#History_and_concepts" rel="nofollow">https://en.wikipedia.org/wiki/Region-based_memory_management...), and the linked paper (<a href="https://www.cs.princeton.edu/techreports/1988/191.pdf" rel="nofollow">https://www.cs.princeton.edu/techreports/1988/191.pdf) is from 1988.
Then, a typo: “as well as memory tied to specific to each of the multiple CPU core or even CPU infinity.”
“Infinity” should be “affinity” there.
skavi · · focus · HN ↗
allocators are so simple to just swap into your program. if you can put together a few representative workloads, you should just try out a few allocators and profile whatever metrics you care about.
imp0cat · · focus · HN ↗
skavi · · focus · HN ↗
for us, tc was among the fastest in runtime while being very space efficient [0]. large rust application using far too many threads.
we’ve since also had great success with tc’s built in profiling tools.
[0]: <a href="https://news.ycombinator.com/item?id=47403847">https://news.ycombinator.com/item?id=47403847
imp0cat · · focus · HN ↗
We've tried tcmalloc, too. I don't remember the exact details, but we basically ended using jemalloc because it was using way less memory.
Same story with mimalloc - it usually provided a tiny bit more speed, but required more cpu and memory.
skavi · · focus · HN ↗
to confirm, you were using tcmalloc from <a href="https://github.com/google/tcmalloc" rel="nofollow">https://github.com/google/tcmalloc and not from <a href="https://github.com/gperftools/gperftools" rel="nofollow">https://github.com/gperftools/gperftools, right?
the latter is a lot worse iiuc.
imp0cat · · focus · HN ↗
skavi · · focus · HN ↗
[0]: <a href="https://github.com/google/tcmalloc/blob/master/docs/gperftools.md#differences" rel="nofollow">https://github.com/google/tcmalloc/blob/master/docs/gperftoo...