Why encode the length at all in the extension? Either the next byte is a continuation byte or not. Are there real world use cases where the length encoding is used in UTF-8 where individual bytes don’t have to be read anyway (for example, codepoint counting, but in practice what good is that
without knowing if the code points are combining)? Even in the case where you just want a byte count, an invalid UTF-8 stream won’t respect the first count marker. The arbitrarily sized count is no better than an arbitrarily sized code point - you’d still have to handle all the same resource and sizing issues for safety. And then what, are you going to allocate a buffer that can fit a length that needs more than 64-bits to represent?
Having the length of encoded in each codepoint means that you know if a stream ends at codepoint boundary or if it has a truncated final codepoint. You could probably do away with the length and maintain this property if you had distinct forms for start, continuer-mid and continuer-final bytes.
jonhohle · · focus · HN ↗
faithful_droog · · focus · HN ↗