‹ BackHN Continuity

Thread

The scourge of x86 emulation

295 points · 100 comments · dagmx

  1. pdw · · focus · HN ↗
    The intro of this article repeats the common assertion that

    > ARM is the most relaxed, allowing significant hardware optimizations; and x86 is the most strict, enforcing a very strong coherency model that doesn’t allow a lot of room for optimization

    but I&#x27;ve seen some compelling arguments that a relaxed model doesn&#x27;t necessarily have much of a benefit, <a href="https:&#x2F;&#x2F;fgiesen.wordpress.com&#x2F;2026&#x2F;08&#x2F;25&#x2F;memory-ordering-in-cpus&#x2F;#320ecacb-31dd-4a74-bd34-de2dbf46b1e0" rel="nofollow">https:&#x2F;&#x2F;fgiesen.wordpress.com&#x2F;2026&#x2F;08&#x2F;25&#x2F;memory-ordering-in-...

    1. Aissen · · focus · HN ↗
      A recent study seemed to support Fabian&#x27;s well written article: <a href="https:&#x2F;&#x2F;dl.acm.org&#x2F;doi&#x2F;epdf&#x2F;10.1145&#x2F;3779212.3790129" rel="nofollow">https:&#x2F;&#x2F;dl.acm.org&#x2F;doi&#x2F;epdf&#x2F;10.1145&#x2F;3779212.3790129
      1. wat10000 · · focus · HN ↗
        Is this just a question of what one considers to be “significant”? I’d consider 3% to be significant but the authors apparently don’t.
        1. deater · · focus · HN ↗
          well it depends on your error bars. On many modern system you can have +&#x2F;- 5% variation or more run to run just due to the non-determinism present in modern CPU architectures and operating systems (even things like room temperature, time of day, the number of environment variables, etc, can affect this). While you could maybe run a set of careful experiments to characterize and remove this, in my experience most researchers don&#x27;t bother. So something as small as 3% would need a lot of convincing to me to make the argument that it is significant.
          1. wat10000 · · focus · HN ↗
            Multiple separate questions here.

            First, is a measured improvement actually real or just an artifact of noise?

            Second, if it is real, is 3% anywhere close to the true value?

            Third, if 3% is real, is it an important difference?

            I&#x27;m just commenting on the third one. If 3% is real, it&#x27;s important.

            I&#x27;m inclined to believe there&#x27;s a real improvement. They made a lot of different measurements. If the measured improvement was a result of noise, you&#x27;d expect a lot of variation an a lot of measurements where TSA was actually faster, and then 3% was the average of that variation. There was a lot of variation (expected, because they were measuring different things) but nearly all of them had TSO being either neutral or slower. Looking at their benchmark graphs, I see two (out of dozens) where TSO was faster.

            As far as being close to the true value, these results suggest there is no single true value, as it depends on the workload. No surprise there.

      2. kccqzy · · focus · HN ↗
        It’s a question of how much resources to allocate to the hardware team, and how much resources to be distributed diffusely to the software engineers but especially to the compiler team.

        Even your linked paper contends that the actual observed slowdown is as much as 22% in the Geekbench example, but the thesis is that the slowdown is not inherent to TSO, but merely to the specific hardware implementation. Is it worthwhile for a company to optimize its TSO to chase the final gains, or is it better not to have this feature in the first place and just change the compiler?

        Indeed my instinct is that it is better to do this in software, where the programmer clearly communicates which stores are ordered, and which may happen in arbitrary order.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.