‹ BackHN Continuity

Thread

UTF-8000: Unlimited UTF-8

134 points · 122 comments · vismit2000

  1. Sharlin · · focus · HN ↗
    UTF-8 originally supported up to six-byte encodings (see eg. RFC 2279), but it was restricted to four bytes in 2003 in order to match UTF-16 constraints :(
    1. delamon · · focus · HN ↗
      We still have about 85% of codepoint space unused. Hopefully, by the time it becomes a problem, UTF-16 will be long dead
      1. nasso_dev · · focus · HN ↗
        i hope so too, but UTF-16 being used by languages such as java and javascript makes me fear it might be here to stay.... i hope im wrong
        1. hnlmorg · · focus · HN ↗
          The number of glyphs available by adding additional bytes drops exponentially because each subsequent byte has one less bit available.

          So I think if we ever were in a situation where > 1 million code points isn’t enough, then we should look at an entirely new way to serialise those code points.

          1. delamon · · focus · HN ↗
            I don't quite get it. 5-byte utf-8 encoding gets extra 5 bits compared to 4 byte, and 6-byte gets extra 10 bits. If you were thinking about bits in leading byte, then yes, you are losing one bit for every extra trailing byte, but you also get 6 bits from it. So adding a byte gives you extra 5 bits.
            1. hnlmorg · · focus · HN ↗
              Yeah, you’re right. I might have attempted to do mental arithmetic before coffee…
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.