‹ BackHN Continuity

Thread

UTF-8000: Unlimited UTF-8

134 points · 122 comments · vismit2000

  1. yyyk · · focus · HN ↗
    Just limit it to 8 bytes at which point you always do 'know the number of follow on bytes' from the first byte.

    Nobody needs more than 4.47 trillion characters. (famous last words)

    1. mitxela · · focus · HN ↗
      Important to recognize that characters have individuality, that's why there can only be a limited number of them. Unicode is enumerating a finite set of things, not encoding an infinite set. Aenything without this property - any generic form of encoding - is not characters, it's something else like images. If it's not in any alphabet it shouldn't be in unicode, you should use an escape tag for image data instead. (Emojis probably shouldn't, but they do behave like an alphabet)

      There cannot be 4 trillion characters because humans would need to know all of them and humans cannot know that many things.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.