‹ BackHN Continuity

Thread

UTF-8000: Unlimited UTF-8

134 points · 122 comments · vismit2000

  1. Sharlin · · focus · HN ↗
    UTF-8 originally supported up to six-byte encodings (see eg. RFC 2279), but it was restricted to four bytes in 2003 in order to match UTF-16 constraints :(
    1. delamon · · focus · HN ↗
      We still have about 85% of codepoint space unused. Hopefully, by the time it becomes a problem, UTF-16 will be long dead
      1. nasso_dev · · focus · HN ↗
        i hope so too, but UTF-16 being used by languages such as java and javascript makes me fear it might be here to stay.... i hope im wrong
        1. flohofwoe · · focus · HN ↗
          The internal string encoding of a programming language doesn't matter as long as it supports UTF-8 at the boundaries. E.g. the text encoding standard on the web is clearly UTF-8, even though JS strings may be internally stored as UTF-16 (or any other encoding).

          Same on macOS/iOS btw: AFAIK NSString is internally UTF-16, but I've never seen a UTF-16 text file on macOS, it's all UTF-8 (unless the file originated on Windows of course).

          1. bjoli · · focus · HN ↗
            The problem is the boundary with, say, windows. In rust they have WTF-8 more or less just to deal with windows filenames.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.