UTF-8 originally supported up to six-byte encodings (see eg. RFC 2279), but it was restricted to four bytes in 2003 in order to match UTF-16 constraints :(
We still have about 85% of codepoint space unused. Hopefully, by the time it becomes a problem, UTF-16 will be long dead
i hope so too, but UTF-16 being used by languages such as java and javascript makes me fear it might be here to stay.... i hope im wrong
Sharlin · · focus · HN ↗
delamon · · focus · HN ↗
nasso_dev · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]