‹ BackHN Continuity

Thread

UTF-8000: Unlimited UTF-8

134 points · 122 comments · vismit2000

  1. Sharlin · · focus · HN ↗
    UTF-8 originally supported up to six-byte encodings (see eg. RFC 2279), but it was restricted to four bytes in 2003 in order to match UTF-16 constraints :(
    1. delamon · · focus · HN ↗
      We still have about 85% of codepoint space unused. Hopefully, by the time it becomes a problem, UTF-16 will be long dead
      1. nasso_dev · · focus · HN ↗
        i hope so too, but UTF-16 being used by languages such as java and javascript makes me fear it might be here to stay.... i hope im wrong
        1. [deleted] · · focus · HN ↗

          [deleted]

        2. hnlmorg · · focus · HN ↗
          The number of glyphs available by adding additional bytes drops exponentially because each subsequent byte has one less bit available.

          So I think if we ever were in a situation where > 1 million code points isn’t enough, then we should look at an entirely new way to serialise those code points.

          1. delamon · · focus · HN ↗
            I don't quite get it. 5-byte utf-8 encoding gets extra 5 bits compared to 4 byte, and 6-byte gets extra 10 bits. If you were thinking about bits in leading byte, then yes, you are losing one bit for every extra trailing byte, but you also get 6 bits from it. So adding a byte gives you extra 5 bits.
            1. hnlmorg · · focus · HN ↗
              Yeah, you’re right. I might have attempted to do mental arithmetic before coffee…
        3. 7bit · · focus · HN ↗
          Utf-16 is famously used by Windows for everything important as well.
          1. adornKey · · focus · HN ↗
            And UTF-8 isn't even fully compatible with windows UTF-16 - UTF8 can't encode a lot of truncated windows UTF-16 filenames.. You need WTF-8 for that.

            <a href="https:&#x2F;&#x2F;artoria2e5.github.io&#x2F;XB18030&#x2F;" rel="nofollow">https:&#x2F;&#x2F;artoria2e5.github.io&#x2F;XB18030&#x2F;

            It seems when designing Unicode most energy went into emoji. And there was nothing left for fancy things like fixed-length string buffers. The only explaination why UTF8 Buffers aren&#x27;t compatible with UTF16 Buffers... is a really strong emoji...

          2. colejohnson66 · · focus · HN ↗
            Thankfully, this is changing. Win32&#x27;s &#x27;A&#x27; ANSI&#x2F;ASCII APIs now support UTF-8 if your app declares such a wish. <a href="https:&#x2F;&#x2F;stackoverflow.com&#x2F;a&#x2F;69181417&#x2F;1350209" rel="nofollow">https:&#x2F;&#x2F;stackoverflow.com&#x2F;a&#x2F;69181417&#x2F;1350209
        4. flohofwoe · · focus · HN ↗
          The internal string encoding of a programming language doesn&#x27;t matter as long as it supports UTF-8 at the boundaries. E.g. the text encoding standard on the web is clearly UTF-8, even though JS strings may be internally stored as UTF-16 (or any other encoding).

          Same on macOS&#x2F;iOS btw: AFAIK NSString is internally UTF-16, but I&#x27;ve never seen a UTF-16 text file on macOS, it&#x27;s all UTF-8 (unless the file originated on Windows of course).

          1. account42 · · focus · HN ↗
            The internal string encoding and its limitations does leak into the APIs.
          2. bjoli · · focus · HN ↗
            The problem is the boundary with, say, windows. In rust they have WTF-8 more or less just to deal with windows filenames.
      2. colejohnson66 · · focus · HN ↗
        But by then, the 4-byte limit of UTF-8 will itself have ossified. Even today, reverting back to the 6-byte limit is nigh impossible.
        1. Razengan · · focus · HN ↗
          By then we will have quaternary quantum computers and FTL circuits where the information appears request it before you
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.