‹ BackHN Continuity

Thread

Writing Efficient C++ Code (2013)

176 points · 161 comments · ibobev

  1. hn_submit · · focus · HN ↗
    I write in C++ almost every day but never have the need to optimize for speed. Even when you write straightforward code it's already blazingly fast.
    1. cjbgkagh · · focus · HN ↗
      I rarely use C++ but when I do it is for speed. It’s not uncommon that carefully crafted intrinsics can 10x the straightforward naive implementation.
    2. gbin · · focus · HN ↗
      It is probably very domain specific. In robotics for example everything is a zero sum game: CPU, memory bandwidth, GPU, battery life etc ... So it is really a topic, probably true for anything embedded actually. Some other offline applications: HFT, Telco etc.. I wish the GUI apps devs respect more the laptop resources they are running on, don't get me started on the 4 instances of chrome I need to run just for discord, signal etc ...
      1. [deleted] · · focus · HN ↗

        [deleted]

      2. serbuvlad · · focus · HN ↗
        GUI engine developers need to trade EVERYTHING for execution time, otherwise JavaScript would simply not be fast enough to handle modern applications.

        If your device has enough resources to power V8, modern GUIs are certainly very pleasant and snappier than a more minimal GUI like HN. Otherwise they are horrendous and very laggy.

        1. gbin · · focus · HN ↗
          I don't know if you got my point. Starting a multi gigabyte machinery for the web just for a chatting app is pure insanity sorry. It will be slow to start, slow to react and a battery hog vs a comparable quality QT app. The worse part is usually people use the web stack for desktop app because they don't want to bother giving a good experience to the people on their own native platform.
          1. serbuvlad · · focus · HN ↗
            I do wish more things were written in Qt, but Qt Widgets is very unpleasant to work in for a modern dynamic app.

            QtQuick/QML/JS is very pleasant and I do wish more people would use it, but from I've seen it's 25-50% the resource use of electron, not some multi-order-of-maginute improvement, do I understand why many people still prefer electron for portability in this case.

          2. rubymamis · · focus · HN ↗
            True, I wrote my own chat app in Qt, and it is substantially faster and more efficient (lower battery usage) than anything made with web technologies.[1]

            [1] <a href="https:&#x2F;&#x2F;www.get-vox.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.get-vox.com&#x2F;

    3. glouwbug · · focus · HN ↗
      True, but moving from a list of unique polymorphic pointers to a std::variant gains you at least a 2-3x speed up in terms of TLB and cacheline locality. From there, swapping to SOA will net you another 4-8x, so you&#x27;re looking at nearly 25x improvement by going data first. That may not matter in the unique case of say, games, where rendering a million entities will dwarf the cost of SIMD processing a million entities, but in something like numerical simulations (fluids) or quant it will be warmly welcomed
      1. cenamus · · focus · HN ↗
        Is the improvement from using std:variant vs polymorphism just due to the indirection you save on?
        1. glouwbug · · focus · HN ↗
          That, and it frees the compiler from reasoning about virtual inlining, and that the std::variant approach can pack potentially more than one object into a single cacheline. TLBs also work with 4096 byte pages, so 32 polymorphic 128 byte entities may (at the absolute worst case) use 32 distinct pages which requires 32 TLB virtual translations, while the std::variant one uses 1.

          The next step of going SOA benefits from all of the above, it just further unlocks you packed quad and oct instructions (AVX256 and 512 depending if you buy AMD or not).

      2. tcfhgj · · focus · HN ↗
        are you sure?

        <a href="https:&#x2F;&#x2F;stackoverflow.com&#x2F;questions&#x2F;69444641&#x2F;c17-stdvariant-is-slower-than-dynamic-polymorphism" rel="nofollow">https:&#x2F;&#x2F;stackoverflow.com&#x2F;questions&#x2F;69444641&#x2F;c17-stdvariant-...

        1. creata · · focus · HN ↗
          If you take the linked benchmark and use the latest compiler version, the std::variant version is faster. The annoying thing about std::variant (and with some other features of modern C++) is that it generates a bunch of code that the compiler has to optimize away.
          1. vlovich123 · · focus · HN ↗
            Somehow and for some reason rust enums don’t have this problem and are far more ergonomic and easier to work with (not to mention compile times are amazing). I don’t know exactly why it’s better to have it as a first class language primitive and why the compiler has such a problem with std::visit, but clearly c++ meta programming slows down things in a super linear way such that the compiler has problems both from code gen and then optimization.
        2. jasode · · focus · HN ↗
          To add to sibling comment...

          GCC 11 (2021) std::visit was slower than virtual dispatch.

          GCC 12 (2022) optimized std::visit so it can be faster than virtual dispatch.

          <a href="https:&#x2F;&#x2F;shubhankar-gambhir.github.io&#x2F;posts&#x2F;your-stdlib-implementation-matters-more-than-the-dispatch-pattern&#x2F;" rel="nofollow">https:&#x2F;&#x2F;shubhankar-gambhir.github.io&#x2F;posts&#x2F;your-stdlib-imple...

          1. someonebaggy · · focus · HN ↗
            It&#x27;s odd reading an AI generated article about something that happened in 2022. Anachronistic.
      3. someonebaggy · · focus · HN ↗
        You don&#x27;t render a million entities, only the ones that are actually on-screen. But when you do render them, you also want to render them in SoA style, culling them by comparing all the bounding boxes against the view frustum, generating a single big list of draw commands and passing it to draw-indirect one time.
      4. jplusequalt · · focus · HN ↗
        &gt;That may not matter in the unique case of say, games, where rendering a million entities will dwarf the cost of SIMD processing a million entities

        Modern graphics APIs allow you to render as many objects as you want with a single draw indirect call.

    4. flowerbreeze · · focus · HN ↗
      When writing code for end-user applications, I think it&#x27;s mostly true. When it&#x27;s writing code for a database engine, a game engine, a 3d renderer, or anything else that involves heavy data processing, optimization is the core &quot;thing&quot; often and it might not even be a good enough solution without it. Although, a lot of time even then C++ is good enough even then when picking reasonable data structures to represent the data.
      1. hn_submit · · focus · HN ↗
        These are what I like to call &quot;infinity applications&quot; where the need for speed is essentially infinite.

        Even if you write them in hand-optimized assembly they would still clamor for more speed.

        1. flowerbreeze · · focus · HN ↗
          Oh, that&#x27;s a great name for it! I&#x27;ll need to start using it.
    5. hackrmn · · focus · HN ↗
      I started writing a [CPU-only] 3-D rendering library in C++ recently, after having written the equivalent in C as a proof-of-concept and an experiment. The reason I decided to write it in C++ after C, is not only because I wanted to tap into meta-programming which is facilitated much better with C++, or that I wanted niceties like procedure overloading, but because some things with C or C++ aren&#x27;t automagically optimised -- like if you want to leverage struct-of-array (SoA) memory layouts because it allows fewer SIMD (AVX in my case) instructions in the rendering pipeline. You do _not_ get that &quot;for free&quot; just writing a single procedure in C++, much less with C. Both languages are layout-sensitive, I mean this is in part what gives you the speed -- optimising with memory layout for cache locality etc. But you have to do it yourself. Meaning that if you need array-of-struct (AoS) or in fact don&#x27;t know which path the CPU would prefer, there&#x27;s no other way than roll up your sleeves and one way or another implement both.

      The kicker is, in my case I chose C++ because templates allow me to reuse most of the code in the rendering pipeline _regardless_ of whether I go for AoS or SoA layout. I leverage operator overloading to do vector by matrix multiplication which is implemented in both variants. I do have to specify the desired variant during building, but I&#x27;ve profiled and for Intel x86 AVX in my case SoA is something like twice as efficient because I process (&quot;shade&quot;) 8 vertices with 4-5 instructions instead of 1 vertex at a time (still shaded with vectorisation -- just &quot;rotated&quot;, i.e in the pipeline axis and not vertex buffer axis).

      TL;DR; C++ gives you plenty fast by default, but it&#x27;s not always enough. The difference between 5 and 15 frames per second, well, makes all the difference -- our eyes are only fooled once the frames-per-second rate goes sufficiently up, anything below an acceptable threshold and it&#x27;s completely different experience. You then either sacrifice resolution or level of detail etc, or decide to squeeze more from the language by helping the compiler.

      1. smallstepforman · · focus · HN ↗
        To do something like this, you really need to look at both cache locality and CPU core access patterns:

        <a href="https:&#x2F;&#x2F;youtu.be&#x2F;jsdwRf3JvZM?si=0tysoCsvaWl0R_PZ" rel="nofollow">https:&#x2F;&#x2F;youtu.be&#x2F;jsdwRf3JvZM?si=0tysoCsvaWl0R_PZ

        40’000 NPC in game with collision avoidance, steering, on 8yo hardware.

        1. hackrmn · · focus · HN ↗
          I am sorry but what does cache locality and CPU core access pattern(s) have to do with the video you linked? Indeed it shows a lot of instances in animation, but the only technicalities it mentions are bare mention of Vulkan (which I assume implies GPU), and I am using the CPU and without any libraries.
    6. mathisfun123 · · focus · HN ↗
      Then you don&#x27;t work on a product that has any scale &lt;shrug&gt;.
      1. [deleted] · · focus · HN ↗

        [deleted]

    7. 8n4vidtmkvmk · · focus · HN ↗
      Definitely need to optimize a bit for games and huge scale web apps. We&#x27;ve been finding big optimizations in our app recently. App works without them because we can scale horizontally but cutting CPU usage by 30% by eliminating redundant work and reducing copies of big objects? Why wouldn&#x27;t we want to do that? This isn&#x27;t even fancy algorithm stuff, mostly just shoddy initial implementations by 100s of eng working on a codebase over 7 years (not even that old). Stuff like that creeps in.
    8. loeg · · focus · HN ↗
      This is a sign you might be more productive in a higher-level language.
      1. creata · · focus · HN ↗
        And the code might be faster, too, like in that 2005 series of articles by Raymond Chen and Rico Mariani (in which one of them wrote a program in C++ and the other wrote the same program in C#).

        <a href="https:&#x2F;&#x2F;devblogs.microsoft.com&#x2F;oldnewthing&#x2F;20060731-15&#x2F;?p=30293&#x2F;" rel="nofollow">https:&#x2F;&#x2F;devblogs.microsoft.com&#x2F;oldnewthing&#x2F;20060731-15&#x2F;?p=30...

        <a href="https:&#x2F;&#x2F;learn.microsoft.com&#x2F;en-us&#x2F;archive&#x2F;blogs&#x2F;ricom&#x2F;performance-quiz-6-chineseenglish-dictionary-reader" rel="nofollow">https:&#x2F;&#x2F;learn.microsoft.com&#x2F;en-us&#x2F;archive&#x2F;blogs&#x2F;ricom&#x2F;perfor...

        1. someonebaggy · · focus · HN ↗
          Titles for future searchers:

          &quot;Just because I don&#x27;t write about .NET doesn&#x27;t mean that I don&#x27;t like it&quot;

          &quot;Performance Quiz #6 -- Chinese&#x2F;English Dictionary reader&quot;

    9. oso2k · · focus · HN ↗
      This follows Rob Pike’s 5 Rules for Programming

      <a href="https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20250201145327&#x2F;https:&#x2F;&#x2F;users.ece.utexas.edu&#x2F;~adnan&#x2F;pike.html" rel="nofollow">https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20250201145327&#x2F;https:&#x2F;&#x2F;users.ece...

    10. wat10000 · · focus · HN ↗
      There’s code where there exists a concept of “fast enough,” and code where there is no such thing.
    11. senderista · · focus · HN ↗
      Then why are you using C++? Java&#x2F;C#&#x2F;Go are already fast enough for general application development. Why would you accept the footguns if not for performance?
      1. bluGill · · focus · HN ↗
        In my case, about 5% of our code needs the power of C++. Mixing C++ with any other language is a huge pain. Even if we were using C, mixing C with anything else is a pain, and that&#x27;s despite being the most supported FFI.

        Note that we started our project before Rust was an option. These days I would certainly look at rust to see if that would cover our 5% of the needs but now we have a lot of C++ and mixing rust with C++ is a pain.

        1. einpoklum · · focus · HN ↗
          &gt; Mixing C++ with any other language is a huge pain.

          It is less pain than for most other languages, except for C. The pain is in exposing a C API for your C++ code. Then you build a library and you&#x27;re set - because basically every language has the ability to call C code. Python, Rust, Java, etc. etc.

          1. palata · · focus · HN ↗
            Well once you have a C API, it is relatively easy to call it from any other language.

            The painful part is to have to go through a C API (modern languages can express much richer APIs and of course there are different constraints on the different runtimes, e.g. GC).

            The annoying part is that each language adds overhead (its runtime). I wouldn&#x27;t call it painful (I don&#x27;t have much to do about it), I say &quot;annoying&quot; just because I would rather minimise the amount of code I ship.

            1. bluGill · · focus · HN ↗
              I disagree. While it is possible, it is never easy. You are limited to exactly the features the C provides.

              Pointers - better null check them even if you guarantee they won&#x27;t be null. (assuming they are supported at all in the target, if not copy all the data in out)

              Want to use a string - your language probably has a better string type the null terminated C string - are you going to pay the overhead to copy the string (remember to copy it back and forth for each call); or are you going to use C strings in your non-C code? Don&#x27;t forget to remember the length of the buffer is different from the length of the string if you modify that string.

              What to use a list - C supports arrays (which are just syntactic pointers). Nearly ever language has a better list type, which for starters encodes length.

              Many languages are garbage collected, which is going to require a lot of work to make work well across the C API.

              And so on.

              For the above, as a C++ programmer I&#x27;m reaching for std::string and std::vector - which are both very good. I have a much better type than the C API allows, but I can&#x27;t use it. Odds are your language has an equivalent that is just as good (possibly with different trade offs - those details are not the point so lets not argue it here), but since the memory layout isn&#x27;t 100% the same you can&#x27;t use my string&#x2F;vectors directly in code as if they are the native types.

              On the surface it seems easy. If you are only doing it a few times it isn&#x27;t hard. However the details are hard and important and the more you mix the harder it gets since the above details (any many many more on the same lines) start to matter.

              1. palata · · focus · HN ↗
                I guess we disagree on the meaning of &quot;relatively easy&quot;. I do it a lot, I consider it boilerplate. It&#x27;s not like passing 10 strings is exponentially harder than passing 1 string, is it?

                I agree that the more you mix, the harder it gets. Maybe I&#x27;m just never working on &quot;serious&quot; projects, but I have never been in a situation where I had to mix 10 languages.

        2. palata · · focus · HN ↗
          I agree that mixing languages adds complexity. And I say that as someone who routinely does it, because many times it&#x27;s better to reuse a mature&#x2F;audited component than rewrite it from scratch.
      2. palata · · focus · HN ↗
        Sometimes it&#x27;s about the libraries. E.g. writing Computer Vision is nicer in C++ right now (IMHO) because most CV libraries are in C++.

        Similarly I like to do video stuff in C just because I call gstreamer&#x2F;ffmpeg directly in C, rather than having to bridge everything.

    12. aldanor · · focus · HN ↗
      Depends on your field. When one microsecond is considered &quot;hellishly slow&quot;, you might reconsider
    13. [deleted] · · focus · HN ↗

      [deleted]

    14. blacklion · · focus · HN ↗
      &quot;blazingly fast&quot; as in how much trading rules could you apply to 10Gbit&#x2F;s stream of stock market data on one core? On one socket?

      How 100Gbit NICs could your filter through your stateful firewall at line speed? And with 64 byte packets?

    15. spacechild1 · · focus · HN ↗
      It really depends a lot on the domain.

      In (soft) realtime audio programming, your audio callback might only have a time budget of 1.3 milliseconds. Everytime you exceed that limit, you&#x27;ll hear a dropout. That&#x27;s when you&#x27;ll start to optimize the hell out of your program :)

    16. Agentlien · · focus · HN ↗
      I work in game development with graphics, porting, and performance. The stuff mentioned in this article is absolutely essential to ship games with sufficient performance
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.