The key problem with C's type sizes is that ANSI C had no way to get fixed integer types, so if the basic types char, short, int, and long didn't have the size you wanted, it might as well not have existed.
For this reason, platform designers were motivated to provide some 32 bit type, which by elimination needs to be int (it can't be char, which must be 8 bits, it can't be short or else you have no 16 bit type, so it must be int or long, but a 16 bit int is silly). Initial 64 bit ABIs were defined with 64 bit int types (ILP64 platforms, such as the first SPARC64 Solaris port) and the problems quickly became apparent when random C software packages failed to configure, compile, or work. Switching to LP64/LLP64 (where int is 32 bits and long is 32 resp. 64 bits) fixed this, at the cost of having to reeducate programmers to use size_t for array indexing. The decision between LP64 and LLP64 was then probably motivated by time_t, which traditionally is a long. If you do LLP64, long and thus time_t is 32 bits and you'll fall prey to the year 2036 problem. Or you use long long for time_t, breaking compatibility with ANSI C and a lot of old software. Most UNIX platforms decided to use LP64 thus, while non-UNIX platforms without this issue frequently went with LLP64 instead, reaping the benefit of staying compatible with software that assumes long is 32 bits.
These days, you probably need only three platform-dependently-sized types (as Go does):
- int/uint for array indexing; native word size, signed and unsigned
- uintptr for round-tripping data pointers
uintptr needs to be different as there are a few platforms where pointers are larger than indices, for example:
- on 8086 and i286, pointers are 32 bits (segment+offset) in some code models, while indices are only 16 bits
- on IBM iSeries, pointers are 128 bits, while indices are 32 or 64 bits
- on capability architectures like CHERI, capabilities (pointers) are 128 bits and need special treatment, while integers are 32 or 64 bits
- on MSP430X, pointers are 20 bits (!) while integers are 16 bits
Conceivably one might also need a separate type for text pointers, as these are some times of different size than data pointers (e.g. on PIC), but as arithmetic on text pointers does not generally make sense, a corresponding integer type can perhaps be omitted (use a generic function type). This also helps with platforms like PowerPC, where text pointers are pairs of addresses (entry+toc).
C's ptrdiff_t can be omitted if you follow C semantics, where differences between pointers are only meaningful within the same object, as that's just int. intptr_t does not really make a lot of sense, pointers are unsigned except on weirdo platforms (such as balanced ternary machines).
C's maxalign_t and (u)intmax_t are problematic, as they fail when the architecture is expanded with new features. For example, maxalign_t might have an 8 byte alignment on x86, but then you add SSE which requires a 16 byte alignment, AVX with a 32 byte alignment, AVX-512 with a 64 byte alignment and so on. Not good news. Same for (u)intmax_t, which end up stopping you from ever supporting larger integer types. It is not entirely clear what the solution is. For maxalign_t, perhaps the maximum alignment can be queried at runtime. For (u)intmax_t, perhaps use generic programming to instantiate the relevant code paths for the actual type used.
clausecker · · focus · HN ↗
For this reason, platform designers were motivated to provide some 32 bit type, which by elimination needs to be int (it can't be char, which must be 8 bits, it can't be short or else you have no 16 bit type, so it must be int or long, but a 16 bit int is silly). Initial 64 bit ABIs were defined with 64 bit int types (ILP64 platforms, such as the first SPARC64 Solaris port) and the problems quickly became apparent when random C software packages failed to configure, compile, or work. Switching to LP64/LLP64 (where int is 32 bits and long is 32 resp. 64 bits) fixed this, at the cost of having to reeducate programmers to use size_t for array indexing. The decision between LP64 and LLP64 was then probably motivated by time_t, which traditionally is a long. If you do LLP64, long and thus time_t is 32 bits and you'll fall prey to the year 2036 problem. Or you use long long for time_t, breaking compatibility with ANSI C and a lot of old software. Most UNIX platforms decided to use LP64 thus, while non-UNIX platforms without this issue frequently went with LLP64 instead, reaping the benefit of staying compatible with software that assumes long is 32 bits.
These days, you probably need only three platform-dependently-sized types (as Go does):
- int/uint for array indexing; native word size, signed and unsigned - uintptr for round-tripping data pointers
uintptr needs to be different as there are a few platforms where pointers are larger than indices, for example:
- on 8086 and i286, pointers are 32 bits (segment+offset) in some code models, while indices are only 16 bits - on IBM iSeries, pointers are 128 bits, while indices are 32 or 64 bits - on capability architectures like CHERI, capabilities (pointers) are 128 bits and need special treatment, while integers are 32 or 64 bits - on MSP430X, pointers are 20 bits (!) while integers are 16 bits
Conceivably one might also need a separate type for text pointers, as these are some times of different size than data pointers (e.g. on PIC), but as arithmetic on text pointers does not generally make sense, a corresponding integer type can perhaps be omitted (use a generic function type). This also helps with platforms like PowerPC, where text pointers are pairs of addresses (entry+toc).
C's ptrdiff_t can be omitted if you follow C semantics, where differences between pointers are only meaningful within the same object, as that's just int. intptr_t does not really make a lot of sense, pointers are unsigned except on weirdo platforms (such as balanced ternary machines).
C's maxalign_t and (u)intmax_t are problematic, as they fail when the architecture is expanded with new features. For example, maxalign_t might have an 8 byte alignment on x86, but then you add SSE which requires a 16 byte alignment, AVX with a 32 byte alignment, AVX-512 with a 64 byte alignment and so on. Not good news. Same for (u)intmax_t, which end up stopping you from ever supporting larger integer types. It is not entirely clear what the solution is. For maxalign_t, perhaps the maximum alignment can be queried at runtime. For (u)intmax_t, perhaps use generic programming to instantiate the relevant code paths for the actual type used.