‹ BackHN Continuity

Thread

UTF-8000: Unlimited UTF-8

134 points · 122 comments · vismit2000

  1. Dwedit · · focus · HN ↗
    FF bytes are an easy way to identify an invalid UTF-8 file. This idea doesn't have that property.
    1. gzitscrux · · focus · HN ↗
      C0 and C1 can do so though according to the author.

      > because all 4 of 2-byte UTF-8's mandatory content bits lie in the first-and-final start byte, we can explicitly rule out 11000000 (0xC0) and 11000001 (0xC1) as permanently invalid bytes. They will never ever appear anywhere in a valid UTF-8000 code unit!

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.