UTF-8 originally supported up to six-byte encodings (see eg. RFC 2279), but it was restricted to four bytes in 2003 in order to match UTF-16 constraints :(
We still have about 85% of codepoint space unused. Hopefully, by the time it becomes a problem, UTF-16 will be long dead
But by then, the 4-byte limit of UTF-8 will itself have ossified. Even today, reverting back to the 6-byte limit is nigh impossible.
By then we will have quaternary quantum computers and FTL circuits where the information appears request it before you
Sharlin · · focus · HN ↗
delamon · · focus · HN ↗
colejohnson66 · · focus · HN ↗
Razengan · · focus · HN ↗