‹ BackHN Continuity

Thread

C's Flexible Integer Sizes Were Not a Design Mistake

113 points · 166 comments · ibobev

  1. adrian_b · · focus · HN ↗
    I agree that the C flexible integer sizes were still necessary at the time of its creation, when some important computers still had word sizes that were not powers of two.

    Nonetheless, I started to use C for programming only in 1990, when I got access to the Microsoft C and Borland Turbo C compilers.

    At that time, 36 years ago, the C flexible integer sizes were already obsolete.

    Since that time until now, while using C on a great variety of computers, from servers and workstations to the smallest microcontrollers, I have seen plenty of portability problems created by the existence of the flexible integer sizes.

    The only programs that had no portability problems were those that never used the flexible integer sizes, but only integers with a definite size, e.g. 8-bit, 16-bit, 32-bit or 64-bit.

    While sizeof solves the problems of memory allocation or copying, it does not help in preventing unexpected integer overflows, because even the size of "char" may be unknown, and even if the size of "char" is known, writing code with multiple paths that would check or prevent overflow for different integer sizes is very cumbersome.

    Flexible integer sizes would work well only on the old computers, where integer overflow generated a hardware exception, so installing an overflow handler would have been sufficient to make the C code work correctly regardless of the size of the native integers.

    1. locknitpicker · · focus · HN ↗
      > At that time, 36 years ago, the C flexible integer sizes were already obsolete.

      This is a highly ignorant comment. You're confusing the fact that you only had to work with a single target architecture with the whole concept of multiple processor architectures being somehow obsolete, as if there was a sudden law of nature that forced every single computer, being full blown HPC stuff or small microcontrollers used in embedded applications.

      Take a look at arduino. They still have 16-bit models out there. Also noteworthy, it seems some DSPs also have ints larger than 32 bits.

      1. flohofwoe · · focus · HN ↗
        The parent is completely right in the sense that for actually portable C code it was always better to use fixed-width integer types which were chosen for the problem to solve instead of target hardware capabilities.

        For instance if your integer arithmetic needs to happen with 32 bit precision (no matter if the code runs on a 16- or 32-bit CPU), there is no scenario where using 'int' makes sense. Instead you'd use a fixed-width 32-bit integer type and accept that math operations are compiled into two instructions on a 16-bit CPU.

        And OTH if you only require 16 bits integer width, there's not much point in picking a 32 bit integer type. Since two's-complement integer encoding has been standard since at least the 70s, the CPU can do narrow operations in the native register width. Any overflow/wraparound is still correct when only looking at the lowest 16-bits of the result.

        1. jabl · · focus · HN ↗
          > And OTH if you only require 16 bits integer width, there's not much point in picking a 32 bit integer type.

          Some common architectures like x86 can suffer from an issue called partial register stalls. So from a performance perspective choosing a 32 bit integer can be better.

          1. sparkie · · focus · HN ↗
            Partial register stalls occurred when you assigned something to say `eax`, then read `ax`, or vice versa - if you wrote `ax` then read `eax`. If you wrote to `ax` then read `ax`, there was no stall. This was due to the register renaming implementation.

            AFAIK, this was only an issue in some older CPUs and isn't a problem today.

            32-bit `int` is still cheaper than 16-bits though, because 16-bit instructions require an operand size override prefix (0x66), or address size override prefix (0x67), or both. Technically, these aren't "16 bit prefixes" - if the machine was running in 16-bit protected mode, then you would need those prefixes to use 32-bit instructions and the non-prefixed ones would be 16-bit and thus cheaper, so `int` would be better as 16-bits in 16-bit protected mode - though this mode is essentially unused today, so for all intents and purposes the prefixes are used to issue 16-bit instructions and 16-bits is more expensive (in code size, i-cache usage, which may impact performance).

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.