‹ BackHN Continuity

Thread

C's Flexible Integer Sizes Were Not a Design Mistake

113 points · 166 comments · ibobev

  1. stkdump · · focus · HN ↗
    The problem begins when you start mixing the traditional types and (u)intN_t, because the latter are merely aliases for the internal types, and it messes up overload resolution. All relevant platforms have pretty much agreed the size of char, short (int), int and long long (int). They have different opinions about long (int) and thus an int64_t might use either long (int) or long long (int).

    So the best solution for nowadays is to use just char, short, int and long long (and make strong assumptions that these are exactly 8, 16, 32 and 64 bits wide respectively), never use long or long double. Never use (u)intNN_t. Then you are good.

    Those caveats of the past (but int might be 16 or 36 bits), are exactly that. An artifact of the past. A historical curiosity. Not relevant for today or the future. No, I don't believe for a second that any future platform will change their size.

    Platforms also still disagree on the signedness of char, so when an 8 bit numeric type (as opposed to an ascii character type) is needed, one should always explicitly specify signed char or unsigned char, both of which are separate types from char.

    Further things of note: platforms also have agreed on little endian (so called "network byte order" is dead and should never be used in new protocols, because it forces everyone to convert) and on IEEE memory representation of float and double. Contrary to popular belief the main floating point operations (+,-,*,/,==,<,>,<=,>=) are also precisely defined and always behave exactly the same (leaving out strange edge cases such as denormals). And yes, of course platforms have very long agreed on twos-complement for negative integers. This even made it into the standard at some point, I believe. Same happened with the memory layout of a vector<>, which in the past wasn't standardized, but because everyone of course did the obvious (and made it the same as a normal C array), it was added to the standard later.

    What I am saying, what the C++ standard guarantees isn't everything. There are much more guarantees modern C++ code can (and should) rely on.

    1. dnautics · · focus · HN ↗
      That is way too much to remember
    2. Joker_vD · · focus · HN ↗
      > No, I don't believe for a second that any future platform will change their size.

      ILP64 (wherein int is 64 bits) exists. It's not very popular, but it exists; e.g. ICC supports it. So it happened in the past once already; it may again happen in the future. In any case, predicting the future is very hard, you really shouldn't be doing this.

      > IEEE memory representation of float and double

      Wait, what? I'm fairly certain that a) IEEE does not mandate the in-memory representation, and b) ARM actually uses big-endian byte order for floats/doubles when storing them in memory.

      > always behave exactly the same (leaving out strange edge cases such as denormals)

      So not always, but please pretend so? Yeah, no, thank you.

      > platforms have very long agreed on twos-complement for negative integers. This even made it into the standard at some point, I believe.

      Only in C23. It was explicitly rejected for C++ 23 (and C++ 26 too, I believe).

      > but because everyone of course did the obvious

      No, not everyone did the obvious. That's why it took so long to standardize because divergent implementations existed.

      > There are much more guarantees modern C++ code can (and should) rely on.

      As long as you only use only GCC (or Clang) exclusively, yes, you can. Otherwise, no, you can't and shan't.

      1. magicalhippo · · focus · HN ↗
        > I'm fairly certain that a) IEEE does not mandate the in-memory representation,

        That's not how I interpret section 3.2 in the standard[1]. Figure 1 seems quite explicit in how a single and a double should be encoded. The section on extended values specify they can be encoded in an implementation-depended manner, which makes the case stronger IMO.

        edit: I note that in the 2008 revision[2], it's more explicitly mentioned that the specified encoding is a binary interchange format. So that's a lot more specific than the original.

        [1]: <a href="https:&#x2F;&#x2F;pub.sergev.org&#x2F;doc&#x2F;ieee754-1985.pdf" rel="nofollow">https:&#x2F;&#x2F;pub.sergev.org&#x2F;doc&#x2F;ieee754-1985.pdf

        [2]: <a href="https:&#x2F;&#x2F;pub.sergev.org&#x2F;doc&#x2F;ieee754-2008.pdf" rel="nofollow">https:&#x2F;&#x2F;pub.sergev.org&#x2F;doc&#x2F;ieee754-2008.pdf

        1. Joker_vD · · focus · HN ↗
          It only talks about MSBs and LSBs. It does not specify whether the LSB of the value as the whole resides in the first byte of the memory representation or in the fourth&#x2F;eighth.

          And of course, if you accept the network byte order as the one intended for the interchange, then IEEE-754 mandates big-endian encoding.

          1. sparkie · · focus · HN ↗
            Maybe you are talking past grandparent, but I think they meant &quot;it&#x27;s safe to assume float and double are IEEE-754,&quot; which is not mandated by the C standard.
          2. magicalhippo · · focus · HN ↗
            I forgot about the PDP-11, so yeah fair point. As I noted the 2008 revision is a lot more specific, which I presume is for good reason.
      2. dgrunwald · · focus · HN ↗
        &gt; ILP64 (wherein int is 64 bits) exists. It&#x27;s not very popular, but it exists; e.g. ICC supports it.

        ILP64 is problematic for existing code: there is lots of stuff like hashcode computations using uint32_t with multiplications, relying on the C standard guaranteeing wraparound for unsigned overflows. But with 64-bit int, uint32_t will promote to a signed int, and overflows will thus be undefined behavior. This problem already exists with uint16_t multiplications on current architectures, but moving the problem to uint32_t will cause trouble for a lot of existing code that thought using fixed-size types like uint32_t would be safe.

        1. nayuki · · focus · HN ↗
          Thank you for being one of the few people who understands that in C&#x2F;C++, `unsigned OP unsigned` can have each operand be promoted to a signed integer and then have the operation overflow and cause undefined behavior.

          I chose to deal with this problem by doing a &quot;pointless&quot; operation to force a promotion to at least unsigned int. For example:

              uint16_t x = 0xFFFF;
              uint16_t y = 0xFFFF;
              uint16_t z = (uint16_t)((x + 0U) * (y + 0U));
          
          This piece of code will work on any machine, such as: (uint16_t = unsigned short = 16 bits, uint32_t = unsigned int = 32 bits); (uint16_t = unsigned short = unsigned int = 16 bits, uint32_t = unsigned long = 32 bits).
          1. Joker_vD · · focus · HN ↗
            But the result is 1, whether you calculate it as 16-by-16 unsigned multiplication (you get 0xFFFE0001 truncated down to 1), or 32-by-32 signed (you multiply -1 by -1 and get 1, with no overflow).
            1. nayuki · · focus · HN ↗
              &gt; 32-by-32 signed (you multiply -1 by -1 and get 1, with no overflow)

              Wrong. You mentally casted each operand to int16_t before subsequently casting to int32_t. The first step is unjustified.

              The correct calculation according to the C standard is: (int32_t)0xFFFF * (int32_t)0xFFFF, which definitely overflows.

              1. Joker_vD · · focus · HN ↗
                &gt; You mentally casted each operand to int16_t before subsequently casting to int32_t. The first step is unjustified.

                That&#x27;s horrifying. Why were unsigned shorts made to convert to signed ints by zero-extension, again?

                1. nayuki · · focus · HN ↗
                  Because the language rules say so, and I&#x27;m not the one who made the rules. <a href="https:&#x2F;&#x2F;en.cppreference.com&#x2F;c&#x2F;language&#x2F;conversion#Integer_promotions" rel="nofollow">https:&#x2F;&#x2F;en.cppreference.com&#x2F;c&#x2F;language&#x2F;conversion#Integer_pr...

                  &gt; Integer promotion is the implicit conversion of a value of any integer type with rank less or equal to rank of int or of {a bit-field of type _Bool(until C23)&#x2F;bool(since C23), int, signed int, unsigned int}, to the value of type int or unsigned int.

                  &gt; If int can represent the entire range of values of the original type (or the range of values of the original bit-field), the value is converted to type int. Otherwise the value is converted to unsigned int.

                  The fact that the C&#x2F;C++ integer conversion rules have these unintuitive footguns is why I made it a talking point.

                2. someonebaggy · · focus · HN ↗
                  One reason is that the language didn&#x27;t originally intend for UB to be so broad. They probably expected that each value would be loaded into an int-sized register, and the calculation would then be performed on two int-sized registers with overflow handled in a platform-dependent but generally sane manner (not time travel).
                  1. Joker_vD · · focus · HN ↗
                    Yeah, reading the rationale in the standard itself, it&#x27;s is kinda obvious that they didn&#x27;t consider potential UB at all: when talking about &quot;questionable signedness&quot; they only mention division&#x2F;remainder and comparisons as being possibly affected, but not addition&#x2F;multiplication (due to overflow), and comparisons behaving mostly expectedly (for some value of &quot;expectedly&quot;) is the main selling point. But comparing signed and unsigned values without actual thought and explicit cast is almost always a bug anyway! And we had the warnings for that for, like, forever.

                    Makes me wonder how ergonomic would a language with actually correct signatures for the arithmetic operators be.

                        +&lt;S, M, N&gt;: int&lt;S, M&gt; × int&lt;S, N&gt; → int&lt;S, max(M, N)+1&gt; &#x2F;&#x2F; alternatively int&lt;S, max(M, N)&gt; × bool
                        -&lt;S, M, N&gt;: int&lt;S, M&gt; × int&lt;S, N&gt; → int&lt;signed, max(M, N)+1&gt; &#x2F;&#x2F; alternatively int&lt;signed, max(M, N)&gt; × bool
                        *&lt;S, M, N&gt;: int&lt;S, M&gt; × int&lt;S, N&gt; → int&lt;S, M+N&gt;
                        &#x2F;&lt;S, M, N&gt;: int&lt;S, M&gt; × int&lt;S, N&gt; → int&lt;S, M&gt;
                        %&lt;S, M, N&gt;: int&lt;S, M&gt; × int&lt;S, N&gt; → int&lt;S, N&gt;
        2. Joker_vD · · focus · HN ↗
          &gt; stuff like hashcode computations using uint32_t with multiplications, relying on the C standard guaranteeing wraparound for unsigned overflows. But with 64-bit int, uint32_t will promote to a signed int, and overflows will thus be undefined behavior.

          Yeah, except that multiplying two 32-bit values, recast as 64-bit signed integers, will not overflow. Even adding another 32-bit value to this product will not overflow. Throw in the final cast to uint32_t to throw away the upper sign bits, and you get the identical result.

          1. nayuki · · focus · HN ↗
            &gt; multiplying two 32-bit values, recast as 64-bit signed integers, will not overflow

            Factually wrong. Consider: (int64_t)0xFFFFFFFF * (int64_t)0xFFFFFFFF. It definitely overflows.

      3. drysine · · focus · HN ↗
        &gt;&gt; platforms have very long agreed on twos-complement for negative integers. This even made it into the standard at some point, I believe.

        &gt;Only in C23. It was explicitly rejected for C++ 23 (and C++ 26 too, I believe).

        It was added in C++20[0], see the note[1] &quot;This is also known as two&#x27;s complement representation&quot;.

        [0] <a href="https:&#x2F;&#x2F;timsong-cpp.github.io&#x2F;cppwp&#x2F;n4868&#x2F;basic.fundamental#3" rel="nofollow">https:&#x2F;&#x2F;timsong-cpp.github.io&#x2F;cppwp&#x2F;n4868&#x2F;basic.fundamental#...

        [1] <a href="https:&#x2F;&#x2F;timsong-cpp.github.io&#x2F;cppwp&#x2F;n4868&#x2F;basic.fundamental#footnote-44" rel="nofollow">https:&#x2F;&#x2F;timsong-cpp.github.io&#x2F;cppwp&#x2F;n4868&#x2F;basic.fundamental#...

      4. stkdump · · focus · HN ↗
        I think you missed the little word &quot;relevant&quot; on top. I am of course aware that there are exceptions. And yes, if you write C++ for embedded or other minor or outdated platforms you will know and understand that these things don&#x27;t apply to you. But there are tons of developers around writing code that will never touch any of these specialized systems that spend way too much effort trying to be clever, because of &quot;the standard&quot;, and &quot;that platform over there&quot;.
    3. bobmcnamara · · focus · HN ↗
      &gt; All relevant platforms

      Don&#x27;t make me tap the sign: the majority of processors running C are weird little dirtbag chips of 16 bits or less sprinkled by the dozen.

      1. pornel · · focus · HN ↗
        but the C for them is its own little world playing by its own rules, separate from C used everywhere else.

        And I bet that when you have only 16 bits of address space, you care how many bits every integer has.

        1. bobmcnamara · · focus · HN ↗
          &gt; but the C for them is its own little world playing by its own rules, separate from C used everywhere else.

          Sadly, this isn&#x27;t unique to micros...

    4. someonebaggy · · focus · HN ↗
      Interestingly the world agreed on little endian at a point in time where endian no longer matters for performance. Little-endian is the only good way to do things in an 8-bit processor, or one with an 8-bit data bus, because the bytes of a multi-byte number are read in the same order they are processed (for addition or subtraction). Modern processors load up to 64 memory bytes in one go, so this constraint does not exist.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.