‹ BackHN Continuity

Thread

UTF-8000: Unlimited UTF-8

134 points · 122 comments · vismit2000

  1. lukasgelbmann · · focus · HN ↗
    Self-synchronization in UTF-8 is intuitively a great thing to have, yet I don’t remember actively relying on it ever. Does anyone have a good example of when it‘s useful?

    Another related nice property that UTF-8 has: substring search reduces to bytestring substring search. I.e. given two Unicode strings in UTF-8 encoding, you can check if one is a substring of the other by just treating them as bytestrings and checking if one bytestring is a substring of the other bytestring. This is a stronger property than self-synchronization: UTF-8 has it, but UTF-8000 doesn’t.

    1. yencabulator · · focus · HN ↗
      I can think of 3 things:

      1. You can partition an input file at any offsets, parallelize, and adjust partition boundaries to a valid offset independently.

      Without the property, parallelization is hard.

      This is how mapreduce has been used to process large text files, except at line boundaries.

      Now, for this that might not be a useful enough property, given that we already do similar things for newlines, and UTF-8 guarantees ASCII is always recognizable and hence newlines are always recognizable.

      2. It might have been more useful in the era of dial-up where we still had occasional corrupted bytes in the transmission.

      3. It helps regain sanity if e.g. a background process outputs bytes that get interleaved at the tty. For example, cat a large text file, the write boundaries won't always align at UTF-8 boundaries, then have a background process output get interleaved in an unfortunate way. If it self-synchronizes, it'll knock itself back into sync after a small amount of garbage.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.