Jemalloc 5.4.0
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Jemalloc 5.4.0
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
ZenoArrow · · focus · HN ↗
aktenlage · · focus · HN ↗
charcircuit · · focus · HN ↗
ck45 · · focus · HN ↗
defrost · · focus · HN ↗
- <a href="https://news.ycombinator.com/item?id=49715318">https://news.ycombinator.com/item?id=49715318
baq · · focus · HN ↗
rfgplk · · focus · HN ↗
groomlake · · focus · HN ↗
HackerThemAll · · focus · HN ↗
Writing allocators for domain-specific access patterns is easy. Writing a general-purpose high performing, stable allocator with bounded P99 latency is hard.
Give your friend, Dunning–Kruger, some better pills to keep him from speaking through you.
Pannoniae · · focus · HN ↗
matheusmoreira · · focus · HN ↗
cv5005 · · focus · HN ↗
Should one even want a global, general purpose heap allocator for that? Seems like a crazy idea to even consider.
Mikhail_Edoshin · · focus · HN ↗
And that's rather hard, because a general purpose allocator makes all decisions based only on the requested size. This is a very simple interface and such a tool is worth having. But a custom allocator can both bake in a specific scenario and provide more nuanced interaction.
jonkerz · · focus · HN ↗
My first thought would be to use per thread pool allocators.
stackghost · · focus · HN ↗
stackghost · · focus · HN ↗
Why would the scads of people writing JavaScript, Java, python, go, rails, etc need to be aware of jemalloc?
iam-da-author · · focus · HN ↗
I work for the runtime team of JPG @ Oracle. We use malloc in Hotspot, quite a lot actually! Providing your JVM with a good malloc can improve the performance of the runtime, both in terms of CPU and memory, by quite a bit.
I don't think you need the details, but it's good to be aware that some mallocs are better than others, and there are multiple of them. Being aware of jemalloc is a good way of being aware of the facts I just mentioned :-).
stackghost · · focus · HN ↗
iam-da-author · · focus · HN ↗
xxs · · focus · HN ↗
and even then recently it costed (us) quite a few months to blame JVM and later the default glibc memory allocator for running out native (not java heap memory) - had to exclude all possible native libs (zlib, zstd via jna), direct buffers, sockets, thread stacks and so on. Changing the malloc to jemalloc solved the issue, even though initially it was done for its debugging capabilities.
It's just a great memory allocator.
fc417fc802 · · focus · HN ↗
javier2 · · focus · HN ↗
nh2 · · focus · HN ↗
In our Python program, a bit of numpy processing of large pictures led to 100 GB not being returned to the OS by glibc's default allocator and the machine running out of memory shortly after. With jemalloc's reliable memory return settings, those problems disappear.
majora2007 · · focus · HN ↗
nh2 · · focus · HN ↗
This is controlled by jemalloc settings `dirty_decay_ms`, `muzzy_decay_ms`, and their interaction with `background_thread`.
`dirty_decay_ms` currently defaults to 10 seconds, so it's not that instant.
That is important e.g. for single-threaded programs that start other programs, such as my Python example: If it starts a subprocess before the 10 seconds elapse after `free()`, Python (and jemalloc) do not run, and get no chance to return memory to the OS.
In such cases, either enable `background_thread`, or set the `_decay_` values to `0` to ensure immediate return to the OS upon `free()`. (This costs some performance.)
See e.g. <a href="https://github.com/jemalloc/jemalloc/issues/2688" rel="nofollow">https://github.com/jemalloc/jemalloc/issues/2688
stackghost · · focus · HN ↗
The number of programmers who are in positions to care about jemalloc vs other malloc is minuscule
yxhuvud · · focus · HN ↗
I wish that wasn't the case, but it is.
imhoguy · · focus · HN ↗
adityapatadia · · focus · HN ↗
smartmic · · focus · HN ↗
gjvc · · focus · HN ↗
kreco · · focus · HN ↗
This question was not necessary. You know the answer, because people upvoted this.
vocx2tx · · focus · HN ↗
[0] <a href="https://news.ycombinator.com/item?id=44264958">https://news.ycombinator.com/item?id=44264958
ezst · · focus · HN ↗
Far from the only case. Trillion dollar companies having sudden interest in your open source project is not necessarily a long term benefit.
titanomachy · · focus · HN ↗
beanjuiceII · · focus · HN ↗
eatonphil · · focus · HN ↗
<a href="https://theconsensus.dev/p/2026/04/16/who-even-uses-jemalloc-anyway.html" rel="nofollow">https://theconsensus.dev/p/2026/04/16/who-even-uses-jemalloc...
nextaccountic · · focus · HN ↗
> Development of Jemalloc will transition back to the original open source project: <a href="https://github.com/jemalloc/jemalloc" rel="nofollow">https://github.com/jemalloc/jemalloc. Meta’s fork of the Jemalloc project will be fully archived as part of the transition. We look forward to continuing the collaboration with the community to drive jemalloc.
> See blog post for details: <a href="https://engineering.fb.com/2026/03/02/data-infrastructure/investing-in-infrastructure-metas-renewed-commitment-to-jemalloc/" rel="nofollow">https://engineering.fb.com/2026/03/02/data-infrastructure/in...
pstoll · · focus · HN ↗
lordnacho · · focus · HN ↗
sam_lowry_ · · focus · HN ↗
P.S. I also wondered whether systemd had any cultural reference to Système D aka Système Débrouillard, but Pottering does not seem to engage in word play on other occasions, so probably not.
tefkah · · focus · HN ↗
eliaspro · · focus · HN ↗
Source: <a href="https://brand.systemd.io/" rel="nofollow">https://brand.systemd.io/
NewJazz · · focus · HN ↗
fizzbuzzbarbazz · · focus · HN ↗
justincormack · · focus · HN ↗
> Systemd was launched in April 2010; it adopts various ideas from previous init systems and combines them with a uniform configuration and administration interface. Systemd operates as a background service (daemon) and controls important system configuration tasks including hardware initialisation and the starting of server processes. The developers thought that its name is suitably reminiscent of the French term "système D", an expression that relates to "thinking on your feet" and describes high-speed technical problem-solving abilities such as those displayed by TV action hero MacGyver.
<a href="https://web.archive.org/web/20121014173559/http://www.h-online.com/open/features/Control-Centre-The-systemd-Linux-init-system-1565543.html" rel="nofollow">https://web.archive.org/web/20121014173559/http://www.h-onli...
simonask · · focus · HN ↗
u8080 · · focus · HN ↗
jibcage · · focus · HN ↗
thrance · · focus · HN ↗
progval · · focus · HN ↗
ButlerianJihad · · focus · HN ↗
hobo123 · · focus · HN ↗
wpollock · · focus · HN ↗
The more you know!
echoangle · · focus · HN ↗
> It is disputed whether the name Wi-Fi is short-form for 'Wireless Fidelity',[34] although the Wi-Fi Alliance did use the advertising slogan "The Standard for Wireless Fidelity" for a short time after the brand name was created,[31][35] referenced "the Wi-Fi (Wireless Fidelity) logo" in a white paper[33] and the Wi-Fi Alliance was also called the "Wireless Fidelity Alliance Inc." in some publications.[36] IEEE, a separate but related organization, has stated "WiFi is a short name for Wireless Fidelity" on their website.[37][38] The name Wi-Fi was partly chosen because it sounds similar to Hi-Fi, which consumers take to mean high fidelity or high quality. Interbrand hoped consumers would find the name catchy, and that they would assume this wireless protocol has high fidelity because of its name.[39]
So it’s not really safe to say it’s not meant to mean wireless fidelity, in fact it sounds pretty likely.
throw0101a · · focus · HN ↗
Jason Evans, JEmalloc:
* <a href="https://jasone.github.io/2025/06/12/jemalloc-postmortem/" rel="nofollow">https://jasone.github.io/2025/06/12/jemalloc-postmortem/
ksec · · focus · HN ↗
Is Meta still using it and developing it? If not who are the driving force behind it now? I just checked there wasn't a release since 2022 and then we have this now. Something changed?
Just wish we have a little bit of context. But it is also great it is continue being maintained. It makes a huge difference for Ruby on Rails Apps.
aktenlage · · focus · HN ↗
See the comment of vocx2tx
throw0101a · · focus · HN ↗
<a href="https://news.ycombinator.com/item?id=49750698">https://news.ycombinator.com/item?id=49750698
To:
* <a href="https://news.ycombinator.com/item?id=44264958">https://news.ycombinator.com/item?id=44264958
* <a href="https://jasone.github.io/2025/06/12/jemalloc-postmortem/" rel="nofollow">https://jasone.github.io/2025/06/12/jemalloc-postmortem/
feldrim · · focus · HN ↗
loeg · · focus · HN ↗
skavi · · focus · HN ↗
enduku · · focus · HN ↗
jeffbee · · focus · HN ↗
tcmalloc only works on Linux.
rwmj · · focus · HN ↗
skavi · · focus · HN ↗
Someone · · focus · HN ↗
skavi · · focus · HN ↗
i don’t believe rseq based cpu local caches require memory barriers on the fast path.
fc417fc802 · · focus · HN ↗
Meanwhile the better the scheduler performs the more competitive the thread local approach becomes.
skavi · · focus · HN ↗
cost for interruption in an rseq critical section is that the PC gets overwritten to the rseq abort entry point before the task is rescheduled. no management thread necessary.
should be fairly minimal cost, especially assuming interruptions in the critical section are rare.
fc417fc802 · · focus · HN ↗
Giving it some more thought, I expect caches will typically be wiped out by a context switch. So the only place rseq is likely to benefit an allocator is on systems with multiple NUMA nodes where you'd like to make sure any management code isn't paying a penalty by hitting the wrong address range.
IIUC the primary thing rseq (and CPU local data) is good for is obviating the need for atomics (specifically the resultant cache line ping-pong) but thread local data already accomplishes that.
skavi · · focus · HN ↗
maybe we weigh things differently, because:
> massive oversubscription of physical CPU cores (ie tens of thousands of threads) where TLS becomes utterly wasteful while also thrashing the cache
seems worth solving to me
fc417fc802 · · focus · HN ↗
Outside of that, TLS has zero overhead and doesn't suffer from contention while rseq involves a small dance and carries a penalty if preempted. I'm certainly open to benchmarks but to me it very much looks like a mixed bag that only comes up when you're already in questionable territory to begin with. If it does come up I'd guess that a handful of threads with contention is a much more common scenario than thousands of threads per physical core and only minimal preemption.
I also expect something like a web server servicing thousands of requests in parallel to use an event loop instead of spawning an equivalent number of threads. I'm struggling to come up with a scenario where you haven't already fatally shot yourself in the foot and this remains a useful optimization to make. It's certainly relevant if you're using fibers (green threads, whatever you want to call them) but at that point you aren't in c calling malloc and your language runtime will (one hopes) already be taking care of all this for you.
skavi · · focus · HN ↗
Totally fair, and I agree that this is what you should do. But I’m only grinding this axe because I’ve personally had to deal with a system that went against most of that guidance.
We ran far too many threads in a memory-constrained environment. Thread count was many multiples of core count. And we weren’t even doing too much worse than a typical Rust tokio application (where all fs ops are dispatched to a thread pool of effectively unbounded size).
So I totally agree that this is “questionable territory”, but honestly, any application that’s outgrown the basic glibc malloc has made a few mistakes.
Someone · · focus · HN ↗
Veserv · · focus · HN ↗
No explicit memory barrier is required anywhere as the value is only read in supervisor mode and a privilege switch implicitly issues a LS-LS barrier on all major architectures.
Someone · · focus · HN ↗
ckennelly · · focus · HN ↗
If `rseq_cs` is no longer describing a relevant address, that is, the program counter has moved past it, the kernel just ignores it.
fweimer · · focus · HN ↗
ckennelly · · focus · HN ↗
The code is executed by a single thread. Everything retires in program order. There is no need for a memory barrier between starting the critical section and its body.
jeffbee · · focus · HN ↗
skavi · · focus · HN ↗
jeffbee · · focus · HN ↗
loeg · · focus · HN ↗
The main issue is thread preemption while you're holding a per-core cache mutex. Some other thread can't do meaningful work using the cache while the holder is sleeping.
ckennelly · · focus · HN ↗
loeg · · focus · HN ↗
albertgoeswoof · · focus · HN ↗
There’s a slow memory leak somewhere in my code but with jemalloc it no longer actually matters.
Thanks to jemalloc team for this!
nvartolomei · · focus · HN ↗
paper2d · · focus · HN ↗
cduzz · · focus · HN ↗
<a href="https://www-users.cse.umn.edu/~arnold/disasters/patriot.html" rel="nofollow">https://www-users.cse.umn.edu/~arnold/disasters/patriot.html
albertgoeswoof · · focus · HN ↗
throw0101a · · focus · HN ↗
* <a href="https://news.ycombinator.com/item?id=49715318">https://news.ycombinator.com/item?id=49715318
atombender · · focus · HN ↗
ggg011012 · · focus · HN ↗
collimarco · · focus · HN ↗