‹ BackHN Continuity

Thread

C's Flexible Integer Sizes Were Not a Design Mistake

113 points · 166 comments · ibobev

  1. stkdump · · focus · HN ↗
    The problem begins when you start mixing the traditional types and (u)intN_t, because the latter are merely aliases for the internal types, and it messes up overload resolution. All relevant platforms have pretty much agreed the size of char, short (int), int and long long (int). They have different opinions about long (int) and thus an int64_t might use either long (int) or long long (int).

    So the best solution for nowadays is to use just char, short, int and long long (and make strong assumptions that these are exactly 8, 16, 32 and 64 bits wide respectively), never use long or long double. Never use (u)intNN_t. Then you are good.

    Those caveats of the past (but int might be 16 or 36 bits), are exactly that. An artifact of the past. A historical curiosity. Not relevant for today or the future. No, I don't believe for a second that any future platform will change their size.

    Platforms also still disagree on the signedness of char, so when an 8 bit numeric type (as opposed to an ascii character type) is needed, one should always explicitly specify signed char or unsigned char, both of which are separate types from char.

    Further things of note: platforms also have agreed on little endian (so called "network byte order" is dead and should never be used in new protocols, because it forces everyone to convert) and on IEEE memory representation of float and double. Contrary to popular belief the main floating point operations (+,-,*,/,==,<,>,<=,>=) are also precisely defined and always behave exactly the same (leaving out strange edge cases such as denormals). And yes, of course platforms have very long agreed on twos-complement for negative integers. This even made it into the standard at some point, I believe. Same happened with the memory layout of a vector<>, which in the past wasn't standardized, but because everyone of course did the obvious (and made it the same as a normal C array), it was added to the standard later.

    What I am saying, what the C++ standard guarantees isn't everything. There are much more guarantees modern C++ code can (and should) rely on.

    1. Joker_vD · · focus · HN ↗
      > No, I don't believe for a second that any future platform will change their size.

      ILP64 (wherein int is 64 bits) exists. It's not very popular, but it exists; e.g. ICC supports it. So it happened in the past once already; it may again happen in the future. In any case, predicting the future is very hard, you really shouldn't be doing this.

      > IEEE memory representation of float and double

      Wait, what? I'm fairly certain that a) IEEE does not mandate the in-memory representation, and b) ARM actually uses big-endian byte order for floats/doubles when storing them in memory.

      > always behave exactly the same (leaving out strange edge cases such as denormals)

      So not always, but please pretend so? Yeah, no, thank you.

      > platforms have very long agreed on twos-complement for negative integers. This even made it into the standard at some point, I believe.

      Only in C23. It was explicitly rejected for C++ 23 (and C++ 26 too, I believe).

      > but because everyone of course did the obvious

      No, not everyone did the obvious. That's why it took so long to standardize because divergent implementations existed.

      > There are much more guarantees modern C++ code can (and should) rely on.

      As long as you only use only GCC (or Clang) exclusively, yes, you can. Otherwise, no, you can't and shan't.

      1. dgrunwald · · focus · HN ↗
        > ILP64 (wherein int is 64 bits) exists. It's not very popular, but it exists; e.g. ICC supports it.

        ILP64 is problematic for existing code: there is lots of stuff like hashcode computations using uint32_t with multiplications, relying on the C standard guaranteeing wraparound for unsigned overflows. But with 64-bit int, uint32_t will promote to a signed int, and overflows will thus be undefined behavior. This problem already exists with uint16_t multiplications on current architectures, but moving the problem to uint32_t will cause trouble for a lot of existing code that thought using fixed-size types like uint32_t would be safe.

        1. nayuki · · focus · HN ↗
          Thank you for being one of the few people who understands that in C/C++, `unsigned OP unsigned` can have each operand be promoted to a signed integer and then have the operation overflow and cause undefined behavior.

          I chose to deal with this problem by doing a "pointless" operation to force a promotion to at least unsigned int. For example:

              uint16_t x = 0xFFFF;
              uint16_t y = 0xFFFF;
              uint16_t z = (uint16_t)((x + 0U) * (y + 0U));
          
          This piece of code will work on any machine, such as: (uint16_t = unsigned short = 16 bits, uint32_t = unsigned int = 32 bits); (uint16_t = unsigned short = unsigned int = 16 bits, uint32_t = unsigned long = 32 bits).
          1. Joker_vD · · focus · HN ↗
            But the result is 1, whether you calculate it as 16-by-16 unsigned multiplication (you get 0xFFFE0001 truncated down to 1), or 32-by-32 signed (you multiply -1 by -1 and get 1, with no overflow).
            1. nayuki · · focus · HN ↗
              > 32-by-32 signed (you multiply -1 by -1 and get 1, with no overflow)

              Wrong. You mentally casted each operand to int16_t before subsequently casting to int32_t. The first step is unjustified.

              The correct calculation according to the C standard is: (int32_t)0xFFFF * (int32_t)0xFFFF, which definitely overflows.

              1. Joker_vD · · focus · HN ↗
                > You mentally casted each operand to int16_t before subsequently casting to int32_t. The first step is unjustified.

                That's horrifying. Why were unsigned shorts made to convert to signed ints by zero-extension, again?

                1. nayuki · · focus · HN ↗
                  Because the language rules say so, and I&#x27;m not the one who made the rules. <a href="https:&#x2F;&#x2F;en.cppreference.com&#x2F;c&#x2F;language&#x2F;conversion#Integer_promotions" rel="nofollow">https:&#x2F;&#x2F;en.cppreference.com&#x2F;c&#x2F;language&#x2F;conversion#Integer_pr...

                  &gt; Integer promotion is the implicit conversion of a value of any integer type with rank less or equal to rank of int or of {a bit-field of type _Bool(until C23)&#x2F;bool(since C23), int, signed int, unsigned int}, to the value of type int or unsigned int.

                  &gt; If int can represent the entire range of values of the original type (or the range of values of the original bit-field), the value is converted to type int. Otherwise the value is converted to unsigned int.

                  The fact that the C&#x2F;C++ integer conversion rules have these unintuitive footguns is why I made it a talking point.

                2. someonebaggy · · focus · HN ↗
                  One reason is that the language didn&#x27;t originally intend for UB to be so broad. They probably expected that each value would be loaded into an int-sized register, and the calculation would then be performed on two int-sized registers with overflow handled in a platform-dependent but generally sane manner (not time travel).
                  1. Joker_vD · · focus · HN ↗
                    Yeah, reading the rationale in the standard itself, it&#x27;s is kinda obvious that they didn&#x27;t consider potential UB at all: when talking about &quot;questionable signedness&quot; they only mention division&#x2F;remainder and comparisons as being possibly affected, but not addition&#x2F;multiplication (due to overflow), and comparisons behaving mostly expectedly (for some value of &quot;expectedly&quot;) is the main selling point. But comparing signed and unsigned values without actual thought and explicit cast is almost always a bug anyway! And we had the warnings for that for, like, forever.

                    Makes me wonder how ergonomic would a language with actually correct signatures for the arithmetic operators be.

                        +&lt;S, M, N&gt;: int&lt;S, M&gt; × int&lt;S, N&gt; → int&lt;S, max(M, N)+1&gt; &#x2F;&#x2F; alternatively int&lt;S, max(M, N)&gt; × bool
                        -&lt;S, M, N&gt;: int&lt;S, M&gt; × int&lt;S, N&gt; → int&lt;signed, max(M, N)+1&gt; &#x2F;&#x2F; alternatively int&lt;signed, max(M, N)&gt; × bool
                        *&lt;S, M, N&gt;: int&lt;S, M&gt; × int&lt;S, N&gt; → int&lt;S, M+N&gt;
                        &#x2F;&lt;S, M, N&gt;: int&lt;S, M&gt; × int&lt;S, N&gt; → int&lt;S, M&gt;
                        %&lt;S, M, N&gt;: int&lt;S, M&gt; × int&lt;S, N&gt; → int&lt;S, N&gt;
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.