UTF-8 originally supported up to six-byte encodings (see eg. RFC 2279), but it was restricted to four bytes in 2003 in order to match UTF-16 constraints :(
The number of glyphs available by adding additional bytes drops exponentially because each subsequent byte has one less bit available.
So I think if we ever were in a situation where > 1 million code points isn’t enough, then we should look at an entirely new way to serialise those code points.
I don't quite get it. 5-byte utf-8 encoding gets extra 5 bits compared to 4 byte, and 6-byte gets extra 10 bits. If you were thinking about bits in leading byte, then yes, you are losing one bit for every extra trailing byte, but you also get 6 bits from it. So adding a byte gives you extra 5 bits.
And UTF-8 isn't even fully compatible with windows UTF-16 - UTF8 can't encode a lot of truncated windows UTF-16 filenames.. You need WTF-8 for that.
It seems when designing Unicode most energy went into emoji. And there was nothing left for fancy things like fixed-length string buffers. The only explaination why UTF8 Buffers aren't compatible with UTF16 Buffers... is a really strong emoji...
Thankfully, this is changing. Win32's 'A' ANSI/ASCII APIs now support UTF-8 if your app declares such a wish. <a href="https://stackoverflow.com/a/69181417/1350209" rel="nofollow">https://stackoverflow.com/a/69181417/1350209
The internal string encoding of a programming language doesn't matter as long as it supports UTF-8 at the boundaries. E.g. the text encoding standard on the web is clearly UTF-8, even though JS strings may be internally stored as UTF-16 (or any other encoding).
Same on macOS/iOS btw: AFAIK NSString is internally UTF-16, but I've never seen a UTF-16 text file on macOS, it's all UTF-8 (unless the file originated on Windows of course).
Sharlin · · focus · HN ↗
delamon · · focus · HN ↗
nasso_dev · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
hnlmorg · · focus · HN ↗
So I think if we ever were in a situation where > 1 million code points isn’t enough, then we should look at an entirely new way to serialise those code points.
delamon · · focus · HN ↗
hnlmorg · · focus · HN ↗
7bit · · focus · HN ↗
adornKey · · focus · HN ↗
<a href="https://artoria2e5.github.io/XB18030/" rel="nofollow">https://artoria2e5.github.io/XB18030/
It seems when designing Unicode most energy went into emoji. And there was nothing left for fancy things like fixed-length string buffers. The only explaination why UTF8 Buffers aren't compatible with UTF16 Buffers... is a really strong emoji...
colejohnson66 · · focus · HN ↗
flohofwoe · · focus · HN ↗
Same on macOS/iOS btw: AFAIK NSString is internally UTF-16, but I've never seen a UTF-16 text file on macOS, it's all UTF-8 (unless the file originated on Windows of course).
account42 · · focus · HN ↗
bjoli · · focus · HN ↗