‹ BackHN Continuity

Thread

UTF-8000: Unlimited UTF-8

134 points · 122 comments · vismit2000

  1. 2shortplanks · · focus · HN ↗
    On a practical matter, it seems like a bad idea to have codepoints that can take up to an arbitrary number of bytes - this just screams buffer overflow problems.

    So in practicality, you’re going to want an arbitrary limit on this (the article suggests as much). But if you place a limit on it then you’ve got one implementation of the standard that can decode certain characters and another that can’t. Better to have one standard that puts a hard limit on the number of bytes and another standard that uses more bytes and so on.

    1. flohofwoe · · focus · HN ↗
      OTH UTF-8 is just one variable-length stream encoding among many others (RLE, LBE128, etc...).
      1. DmitryOlshansky · · focus · HN ↗
        The bonus is synchonizing at arbitrary point in stream and that ASCII is UTF-8
      2. saghm · · focus · HN ↗
        Unless I'm misremembering, even UTF-16 is variable. You need to bump up to UTF-32 to get fixed-width.
        1. cyphar · · focus · HN ↗
          Even better, it's arguably both -- surrogate characters are valid codepoint values so technically UTF-16 is fixed-width but programs need to have special handling for surrogate pairs meaning it is practically variable-width.

          Truly the worst of all worlds.

        2. mafuy · · focus · HN ↗
          Correct me if I'm wrong, but I think all kinds of UTF, including 16 and 32, support arbitrary length for a single effective character. This would be because you can stack modifications as long as you like.
          1. ElectricalUnion · · focus · HN ↗
            What you meant by "single effective character" is grapheme clusters. This whole discussion is about variable sized code points.
            1. Dylan16807 · · focus · HN ↗
              The first comment was kind of iffy when it was also talking about buffers and characters, and focusing on code points is mostly a bad focus. It's worth bringing up so nobody thinks fixed width at a single layer is particularly useful, because other layers will still be variable.
        3. flohofwoe · · focus · HN ↗
          Yes, UTF-16 is the worst of all alternatives and should be abolished rather sooner than later.

          UTF-32 is fixed-width for UNICODE code points, but a single visual character (e.g. a "grapheme cluster") can be built from multiple code points. This is separate from the encoding algorithm though, grapheme clusters are mostly a problem for the high level code working with already decoded text data (text rendering, comparison, sorting etc...).

          1. nickyvdicarlo · · focus · HN ↗
            I was working on optimizing a UUID scanner recently, and was baffled when I discovered that Java represents all char primitives in UTF-16, regardless of the original encoding.

            So even given an ascii String, if you want to do something like

            char c = input.charAt(index)

            the JVM is going to jump to that index, check whether the character at that index isLatin(), then cast that single byte into to a 2 byte char... every single time.

            in the naive solution, omething like 15% of the CPU cycles were spent checking isLatin() over and over again

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.