‹ BackHN Continuity

Thread

UTF-8000: Unlimited UTF-8

134 points · 122 comments · vismit2000

  1. Dwedit · · focus · HN ↗
    FF bytes are an easy way to identify an invalid UTF-8 file. This idea doesn't have that property.
    1. sph · · focus · HN ↗
      True, but not all non-UTF8 bytestrings contain 0xFF bytes, so it’s not very useful in practice.
      1. da_chicken · · focus · HN ↗
        Yes, I agree.

        It's more common for programs that say they support UTF-8 to not really do so at all. It wasn't that long ago that "UTF-8" support was often just single byte, so it was little more than ASCII. Even now it's common for programs to choke on the optional BOM. Yes, it is redundant, congratulations. The spec still explicitly allows it. Three and four byte character support is still not the best, too.

        1. flohofwoe · · focus · HN ↗
          > "UTF-8" support was often just single byte, so it was little more than ASCII

          "Single byte UTF-8" is ASCII. That's one of its most important properties.

          > Even now it's common for programs to choke on the optional BOM

          And they should... BOMs (and especially the hilarious UTF-8 BOM) are strictly a legacy Microsoft/Windows thing and should be abolished along with "extended" 8-bit ASCII encodings and UCS-2/UTF-16 (only UTF-32 makes sense, but should only be used at runtime to allow random access on UNICODE code points, but not for data exchange.

          1. da_chicken · · focus · HN ↗
            Your opinion on the BOM isn't wrong, but it's also not germaine to whether or not you're actually following the spec. The spec is the spec. If you don't like it you can get the spec changed. You don't get to ignore the spec and then claim support. That's not how standards work. "I don't like it," isn't a good explanation.

            Otherwise I'd be inclined to fix the spelling error in the HTTP referrer.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.