‹ BackHN Continuity

Thread

UTF-8000: Unlimited UTF-8

134 points · 122 comments · vismit2000

  1. sph · · focus · HN ↗
    > UTF-8000 is in no way endorsed by or representative of the Unicode Consortium.

    Not until they decide to expand the emoji range, allocate space for all past and future fictional languages, as well as birdsong and dog barks.

    Someone at the consortium is rubbing their hands with glee with all the newfound space.

    But honestly, cool hack! If you invent a method to encode large numbers into bytes, why limit yourself to 24-bit numbers?

    1. throw0101a · · focus · HN ↗
      > […] why limit yourself to 24-bit numbers?

      For compatibility with UTF-16:

          o  Restricted the range of characters to 0000-10FFFF (the UTF-16
             accessible range).
      
      * <a href="https:&#x2F;&#x2F;datatracker.ietf.org&#x2F;doc&#x2F;html&#x2F;rfc3629#section-12" rel="nofollow">https:&#x2F;&#x2F;datatracker.ietf.org&#x2F;doc&#x2F;html&#x2F;rfc3629#section-12

      * <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;UTF-16" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;UTF-16

      The original spec had 31 bits (the UTF-32&#x2F;UCS-4 range):

      * <a href="https:&#x2F;&#x2F;datatracker.ietf.org&#x2F;doc&#x2F;html&#x2F;rfc2279" rel="nofollow">https:&#x2F;&#x2F;datatracker.ietf.org&#x2F;doc&#x2F;html&#x2F;rfc2279

      * <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;UTF-32" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;UTF-32

      1. orangeboats · · focus · HN ↗
        We really ought to deprecate UTF-16 someday. The fact that it pretends to be a fixed-length encoding has caused all sorts of bugs over the years, with many people assuming n(UTF-16 codepoints) == n(characters) which breaks when the string contains non-BMP characters.

        And also, for personal aesthetic reasons I hate that it limits the Unicode codepoint range to an awkward non-power-of-two number (now there are 0x110000 codepoints in total). UTF-8 and UTF-32&#x27;s 2^31 feels much more natural.

        1. bjoli · · focus · HN ↗
          I wonder what it would take. For things like Java, JSA, and c# there would have to be things like a parallel utf8 api (yes please), but the really hard parts is the things that are as old as time (windows).

          I don&#x27;t think it will ever happen, but one can dream.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.