‹ BackHN Continuity

Thread

C's Flexible Integer Sizes Were Not a Design Mistake

113 points · 166 comments · ibobev

  1. stkdump · · focus · HN ↗
    The problem begins when you start mixing the traditional types and (u)intN_t, because the latter are merely aliases for the internal types, and it messes up overload resolution. All relevant platforms have pretty much agreed the size of char, short (int), int and long long (int). They have different opinions about long (int) and thus an int64_t might use either long (int) or long long (int).

    So the best solution for nowadays is to use just char, short, int and long long (and make strong assumptions that these are exactly 8, 16, 32 and 64 bits wide respectively), never use long or long double. Never use (u)intNN_t. Then you are good.

    Those caveats of the past (but int might be 16 or 36 bits), are exactly that. An artifact of the past. A historical curiosity. Not relevant for today or the future. No, I don't believe for a second that any future platform will change their size.

    Platforms also still disagree on the signedness of char, so when an 8 bit numeric type (as opposed to an ascii character type) is needed, one should always explicitly specify signed char or unsigned char, both of which are separate types from char.

    Further things of note: platforms also have agreed on little endian (so called "network byte order" is dead and should never be used in new protocols, because it forces everyone to convert) and on IEEE memory representation of float and double. Contrary to popular belief the main floating point operations (+,-,*,/,==,<,>,<=,>=) are also precisely defined and always behave exactly the same (leaving out strange edge cases such as denormals). And yes, of course platforms have very long agreed on twos-complement for negative integers. This even made it into the standard at some point, I believe. Same happened with the memory layout of a vector<>, which in the past wasn't standardized, but because everyone of course did the obvious (and made it the same as a normal C array), it was added to the standard later.

    What I am saying, what the C++ standard guarantees isn't everything. There are much more guarantees modern C++ code can (and should) rely on.

    1. Joker_vD · · focus · HN ↗
      > No, I don't believe for a second that any future platform will change their size.

      ILP64 (wherein int is 64 bits) exists. It's not very popular, but it exists; e.g. ICC supports it. So it happened in the past once already; it may again happen in the future. In any case, predicting the future is very hard, you really shouldn't be doing this.

      > IEEE memory representation of float and double

      Wait, what? I'm fairly certain that a) IEEE does not mandate the in-memory representation, and b) ARM actually uses big-endian byte order for floats/doubles when storing them in memory.

      > always behave exactly the same (leaving out strange edge cases such as denormals)

      So not always, but please pretend so? Yeah, no, thank you.

      > platforms have very long agreed on twos-complement for negative integers. This even made it into the standard at some point, I believe.

      Only in C23. It was explicitly rejected for C++ 23 (and C++ 26 too, I believe).

      > but because everyone of course did the obvious

      No, not everyone did the obvious. That's why it took so long to standardize because divergent implementations existed.

      > There are much more guarantees modern C++ code can (and should) rely on.

      As long as you only use only GCC (or Clang) exclusively, yes, you can. Otherwise, no, you can't and shan't.

      1. dgrunwald · · focus · HN ↗
        > ILP64 (wherein int is 64 bits) exists. It's not very popular, but it exists; e.g. ICC supports it.

        ILP64 is problematic for existing code: there is lots of stuff like hashcode computations using uint32_t with multiplications, relying on the C standard guaranteeing wraparound for unsigned overflows. But with 64-bit int, uint32_t will promote to a signed int, and overflows will thus be undefined behavior. This problem already exists with uint16_t multiplications on current architectures, but moving the problem to uint32_t will cause trouble for a lot of existing code that thought using fixed-size types like uint32_t would be safe.

        1. Joker_vD · · focus · HN ↗
          > stuff like hashcode computations using uint32_t with multiplications, relying on the C standard guaranteeing wraparound for unsigned overflows. But with 64-bit int, uint32_t will promote to a signed int, and overflows will thus be undefined behavior.

          Yeah, except that multiplying two 32-bit values, recast as 64-bit signed integers, will not overflow. Even adding another 32-bit value to this product will not overflow. Throw in the final cast to uint32_t to throw away the upper sign bits, and you get the identical result.

          1. nayuki · · focus · HN ↗
            > multiplying two 32-bit values, recast as 64-bit signed integers, will not overflow

            Factually wrong. Consider: (int64_t)0xFFFFFFFF * (int64_t)0xFFFFFFFF. It definitely overflows.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.