The intro of this article repeats the common assertion that
> ARM is the most relaxed, allowing significant hardware optimizations; and x86 is the most strict, enforcing a very strong coherency model that doesn’t allow a lot of room for optimization
but I've seen some compelling arguments that a relaxed model doesn't necessarily have much of a benefit, <a href="https://fgiesen.wordpress.com/2026/08/25/memory-ordering-in-cpus/#320ecacb-31dd-4a74-bd34-de2dbf46b1e0" rel="nofollow">https://fgiesen.wordpress.com/2026/08/25/memory-ordering-in-...
[Disclaimer: I wrote Rosetta 2 and determined the spec for Apple's TSO mode, so I am obviously biased.]
Giesen's article comes off as well-meaning cope from an x86 fan. A relaxed memory model really does give you some performance. Another memory model flaw here in x86 is more architectural, which is that every instruction with the LOCK prefix is essentially a full barrier (of course, x86 could have provided different instructions while still being under TSO). In programs that make heavy usage of atomic reference counting, this actually helps quite a bit.
I would probably put that performance benefit in the single digit percentage range like my sibling comment, which may not seem like much to a SW engineer but is actually pretty serious in CPU microarchitecture. It also helps to be stacked with other architectural advantages over x86, e.g. fixed-length instructions, 32 GPRs (which Intel copied in APX), LDP/STP (which Intel also copied in APX), etc.
One of the old arguments from TSO enjoyers was that TSO helps avoid concurrency bugs that people would accidentally introduce, but this was before the C++ memory model propagated throughout the programming world. Nowadays, I think people generally conceptualize memory consistency in terms of acquire/release anyways, so why not use a CPU architecture that uses the same model?
Do you think there would be any worthwhile gains from relaxing address-dependent load ordering, like on Alpha/AXP? Or was that just a lot of extra pain for little reward?
Fun fact: ARM actually has relaxed address dependencies for non-temporal loads, although I don't know if too many implementations of ARM take advantage of this relaxation.
I think it is an interesting question. During the Alpha's lifetime as a non-hobbyist architecture, this decision was pretty much universally derided, but this was in the prehistoric eras of concurrent memory models, where people were just trying their best with a mix of C code, intrinsics, uses of `volatile` sprinkled around to hopefully disable optimizations, and inline assembly. When the C++11 memory model came around, they tried to integrate dependency ordering with `memory_order_consume`, and this famously failed, along with every attempt to fix it. I believe the plan is now for C++ (and later C?) to add special-cased RCU primitives.
The relaxation makes sense in the abstract. As evidenced from the `memory_order_consume` saga, compiler optimizations regularly violate dependency ordering anyways, so in the strictest sense you can't really rely on it. Of course, that doesn't stop people from "knowing" what their compilers will do in such a situation, but that strategy has become a worse one as the years have passed. I would feel better about the whole situation if there was a good greenfield design for a low-level PL that incorporates explicit dependency ordering.
I think it's less clear where the potential HW benefit is in a contemporary CPU. The obvious answer is value prediction, but any CPU performing value prediction has to deal with so many other microarchitectural conditions that can invalidate its speculation decisions that it's not clear this minor one is a huge burden. Of course, people who love TSO might make the same argument, i.e. that it's not a huge burden to snoop cache traffic and invalidate loads (although this is only the load half of TSO, not the store half). I know an Alpha architect who argued that Alpha was right for this decision. I even knew a Transmeta architect who argued for implementing sequential consistency in HW (as Transmeta and derived CPUs actually did).
In practice, microarchitectural structures have capacity/throughput limitations, and there are implementations and complexities that only come up in a real design, so everything needs to be evaluated in the context of a real project. I personally think the sweet spot falls to the weaker side of TSO, which also happens to be near the memory consistency model of the low-level languages we're using anyways.
> and this famously failed, along with every attempt to fix it.
What was the cause of failure? I see that it is now deprecated which surprises me. It was my understanding that large portions of this were taken from the Java memory model and in the JVM, it seems to have been pretty successful.
Is it because the compilers themselves refused to respect it?
As my sibling comment from jcranmer points out, it was only `memory_order_consume` that failed outright. The other aspects of the C++ memory model mostly worked out, at least after revision.
I believe the last attempt to come up with a comprehensive new proposal was P0190 (<a href="https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2017/p0190r4.pdf" rel="nofollow">https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2017/p01...), which reduced the ordering-imposing dependencies in the model to essentially just be pointer dependencies (which does eliminate some legitimately useful cases, e.g. stealing bits from pointers). The paper mentions the unresolved problem of control dependencies from compiler knowledge of pointer equality, which is intertwined with the pointer provenance discussions from around the same time.
I would like to see the basic concept realized in another language some day.
pdw · · focus · HN ↗
> ARM is the most relaxed, allowing significant hardware optimizations; and x86 is the most strict, enforcing a very strong coherency model that doesn’t allow a lot of room for optimization
but I've seen some compelling arguments that a relaxed model doesn't necessarily have much of a benefit, <a href="https://fgiesen.wordpress.com/2026/08/25/memory-ordering-in-cpus/#320ecacb-31dd-4a74-bd34-de2dbf46b1e0" rel="nofollow">https://fgiesen.wordpress.com/2026/08/25/memory-ordering-in-...
cwzwarich · · focus · HN ↗
Giesen's article comes off as well-meaning cope from an x86 fan. A relaxed memory model really does give you some performance. Another memory model flaw here in x86 is more architectural, which is that every instruction with the LOCK prefix is essentially a full barrier (of course, x86 could have provided different instructions while still being under TSO). In programs that make heavy usage of atomic reference counting, this actually helps quite a bit.
I would probably put that performance benefit in the single digit percentage range like my sibling comment, which may not seem like much to a SW engineer but is actually pretty serious in CPU microarchitecture. It also helps to be stacked with other architectural advantages over x86, e.g. fixed-length instructions, 32 GPRs (which Intel copied in APX), LDP/STP (which Intel also copied in APX), etc.
One of the old arguments from TSO enjoyers was that TSO helps avoid concurrency bugs that people would accidentally introduce, but this was before the C++ memory model propagated throughout the programming world. Nowadays, I think people generally conceptualize memory consistency in terms of acquire/release anyways, so why not use a CPU architecture that uses the same model?
spijdar · · focus · HN ↗
cwzwarich · · focus · HN ↗
I think it is an interesting question. During the Alpha's lifetime as a non-hobbyist architecture, this decision was pretty much universally derided, but this was in the prehistoric eras of concurrent memory models, where people were just trying their best with a mix of C code, intrinsics, uses of `volatile` sprinkled around to hopefully disable optimizations, and inline assembly. When the C++11 memory model came around, they tried to integrate dependency ordering with `memory_order_consume`, and this famously failed, along with every attempt to fix it. I believe the plan is now for C++ (and later C?) to add special-cased RCU primitives.
The relaxation makes sense in the abstract. As evidenced from the `memory_order_consume` saga, compiler optimizations regularly violate dependency ordering anyways, so in the strictest sense you can't really rely on it. Of course, that doesn't stop people from "knowing" what their compilers will do in such a situation, but that strategy has become a worse one as the years have passed. I would feel better about the whole situation if there was a good greenfield design for a low-level PL that incorporates explicit dependency ordering.
I think it's less clear where the potential HW benefit is in a contemporary CPU. The obvious answer is value prediction, but any CPU performing value prediction has to deal with so many other microarchitectural conditions that can invalidate its speculation decisions that it's not clear this minor one is a huge burden. Of course, people who love TSO might make the same argument, i.e. that it's not a huge burden to snoop cache traffic and invalidate loads (although this is only the load half of TSO, not the store half). I know an Alpha architect who argued that Alpha was right for this decision. I even knew a Transmeta architect who argued for implementing sequential consistency in HW (as Transmeta and derived CPUs actually did).
In practice, microarchitectural structures have capacity/throughput limitations, and there are implementations and complexities that only come up in a real design, so everything needs to be evaluated in the context of a real project. I personally think the sweet spot falls to the weaker side of TSO, which also happens to be near the memory consistency model of the low-level languages we're using anyways.
cogman10 · · focus · HN ↗
What was the cause of failure? I see that it is now deprecated which surprises me. It was my understanding that large portions of this were taken from the Java memory model and in the JVM, it seems to have been pretty successful.
Is it because the compilers themselves refused to respect it?
cwzwarich · · focus · HN ↗
I believe the last attempt to come up with a comprehensive new proposal was P0190 (<a href="https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2017/p0190r4.pdf" rel="nofollow">https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2017/p01...), which reduced the ordering-imposing dependencies in the model to essentially just be pointer dependencies (which does eliminate some legitimately useful cases, e.g. stealing bits from pointers). The paper mentions the unresolved problem of control dependencies from compiler knowledge of pointer equality, which is intertwined with the pointer provenance discussions from around the same time.
I would like to see the basic concept realized in another language some day.