On a practical matter, it seems like a bad idea to have codepoints that can take up to an arbitrary number of bytes - this just screams buffer overflow problems.
So in practicality, you’re going to want an arbitrary limit on this (the article suggests as much). But if you place a limit on it then you’ve got one implementation of the standard that can decode certain characters and another that can’t. Better to have one standard that puts a hard limit on the number of bytes and another standard that uses more bytes and so on.
Even better, it's arguably both -- surrogate characters are valid codepoint values so technically UTF-16 is fixed-width but programs need to have special handling for surrogate pairs meaning it is practically variable-width.
Correct me if I'm wrong, but I think all kinds of UTF, including 16 and 32, support arbitrary length for a single effective character.
This would be because you can stack modifications as long as you like.
The first comment was kind of iffy when it was also talking about buffers and characters, and focusing on code points is mostly a bad focus. It's worth bringing up so nobody thinks fixed width at a single layer is particularly useful, because other layers will still be variable.
Yes, UTF-16 is the worst of all alternatives and should be abolished rather sooner than later.
UTF-32 is fixed-width for UNICODE code points, but a single visual character (e.g. a "grapheme cluster") can be built from multiple code points. This is separate from the encoding algorithm though, grapheme clusters are mostly a problem for the high level code working with already decoded text data (text rendering, comparison, sorting etc...).
I was working on optimizing a UUID scanner recently, and was baffled when I discovered that Java represents all char primitives in UTF-16, regardless of the original encoding.
So even given an ascii String, if you want to do something like
char c = input.charAt(index)
the JVM is going to jump to that index, check whether the character at that index isLatin(), then cast that single byte into to a 2 byte char...
every single time.
in the naive solution, omething like 15% of the CPU cycles were spent checking isLatin() over and over again
2shortplanks · · focus · HN ↗
So in practicality, you’re going to want an arbitrary limit on this (the article suggests as much). But if you place a limit on it then you’ve got one implementation of the standard that can decode certain characters and another that can’t. Better to have one standard that puts a hard limit on the number of bytes and another standard that uses more bytes and so on.
flohofwoe · · focus · HN ↗
DmitryOlshansky · · focus · HN ↗
saghm · · focus · HN ↗
cyphar · · focus · HN ↗
Truly the worst of all worlds.
mafuy · · focus · HN ↗
ElectricalUnion · · focus · HN ↗
Dylan16807 · · focus · HN ↗
flohofwoe · · focus · HN ↗
UTF-32 is fixed-width for UNICODE code points, but a single visual character (e.g. a "grapheme cluster") can be built from multiple code points. This is separate from the encoding algorithm though, grapheme clusters are mostly a problem for the high level code working with already decoded text data (text rendering, comparison, sorting etc...).
nickyvdicarlo · · focus · HN ↗
So even given an ascii String, if you want to do something like
char c = input.charAt(index)
the JVM is going to jump to that index, check whether the character at that index isLatin(), then cast that single byte into to a 2 byte char... every single time.
in the naive solution, omething like 15% of the CPU cycles were spent checking isLatin() over and over again