‹ BackHN Continuity

Thread

Writing Efficient C++ Code (2013)

176 points · 161 comments · ibobev

  1. asveikau · · focus · HN ↗
    This article reminds me of performance advice I was starting to see in the 2000s decade. Basically it was to not introduce a bunch of pointer heavy data structures to get lower algorithmic complexity. Stuff it all into a vector. You will use some algorithms that the computer science textbook will say it's slower, but if it fits all in cache it doesn't matter. The cache misses following pointers all over town hurts you more.
    1. stackghost · · focus · HN ↗
      I too am in the "premature optimization bad" camp.

      Beyond the low-hanging fruit like ensuring you aren't creating O(n^2) complexity by accident, I think C++ is fast enough/has mature-enough compilers that by the time you're worrying about cache hits materially affecting performance, you're probably also sufficiently staffed and capitalized to pay people to A/B test that performance.

      1. Pannoniae · · focus · HN ↗
        1. Compilers barely do even basic optimisations such as interprocedural register allocation when faced with non-trivial code. You often also need the most aggressive optimisation settings, LTO or even PGO enabled for many of these.

        2. Virtuals are, with the exception of PGO, mostly a black box i.e. you get a hard optimisation boundary, no inlining at all.

        3. The C++ standard library is usually comically slow (yes, even compared to Java/C#/the likes) so if your project uses std::vector and the such instead of specialised libraries, you've already lost at the beginning.

        4. If you don't pay attention to performance from the get-go, the approximate amount of autovectorisation you'll get is close to zero. Some compilers are better than others (Clang>MSVC for example) but I've seen codebases with 8 figures of LoC where the number of vectorised divides/multiplys was like less than ten when you dumped the object listing. In the whole program.

        5. Since aliasing and other optimisation barriers (you didn't use restrict or manually hoist, did ya?), it's not uncommon for large C++ programs to spend a third of their runtime doing atomic increments because shared_ptr is supposedly cheap and who cares about lifetimes anyway.

        6. If you're targeting Windows, the default new operator / malloc is also comically slow. Luckily that one is fairly easy to fix with installing mimalloc and deploying the hijack dll, but the negative effects on cache by the fragmented allocations is also significant.

        1. stinos · · focus · HN ↗
          Regarding 3: depends on what you're doing with it? Take std::vector: we have a codebase where pretty much everything is allocated once but we still want bounds checking on that memory. So we have a lot of std::vector in those places. What would a specialised library change there?
          1. Pannoniae · · focus · HN ↗
            1. I'm not familiar with the hardened stdlib stuff except for the msvc debug runtime but if you have a solution for this, skip this one. You presumably want boundschecking (and throwing/failing hard) or at the very least, logging out of bounds accesses.

            2. A non-inlined grow. If you have large collections you modify often, you want a vector implementation where the reallocation is out of line and the rare case. All the STLs treat it as a normal method and have inlined by codegen.

            3. Trivial relocation support so you don't need to destruct objects where there are no pointers inside or external objects pointing to them.

            In your case it's probably not as relevant/important, yes

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.