On a practical matter, it seems like a bad idea to have codepoints that can take up to an arbitrary number of bytes - this just screams buffer overflow problems.
So in practicality, you’re going to want an arbitrary limit on this (the article suggests as much). But if you place a limit on it then you’ve got one implementation of the standard that can decode certain characters and another that can’t. Better to have one standard that puts a hard limit on the number of bytes and another standard that uses more bytes and so on.
Yes, UTF-16 is the worst of all alternatives and should be abolished rather sooner than later.
UTF-32 is fixed-width for UNICODE code points, but a single visual character (e.g. a "grapheme cluster") can be built from multiple code points. This is separate from the encoding algorithm though, grapheme clusters are mostly a problem for the high level code working with already decoded text data (text rendering, comparison, sorting etc...).
I was working on optimizing a UUID scanner recently, and was baffled when I discovered that Java represents all char primitives in UTF-16, regardless of the original encoding.
So even given an ascii String, if you want to do something like
char c = input.charAt(index)
the JVM is going to jump to that index, check whether the character at that index isLatin(), then cast that single byte into to a 2 byte char...
every single time.
in the naive solution, omething like 15% of the CPU cycles were spent checking isLatin() over and over again
2shortplanks · · focus · HN ↗
So in practicality, you’re going to want an arbitrary limit on this (the article suggests as much). But if you place a limit on it then you’ve got one implementation of the standard that can decode certain characters and another that can’t. Better to have one standard that puts a hard limit on the number of bytes and another standard that uses more bytes and so on.
flohofwoe · · focus · HN ↗
saghm · · focus · HN ↗
flohofwoe · · focus · HN ↗
UTF-32 is fixed-width for UNICODE code points, but a single visual character (e.g. a "grapheme cluster") can be built from multiple code points. This is separate from the encoding algorithm though, grapheme clusters are mostly a problem for the high level code working with already decoded text data (text rendering, comparison, sorting etc...).
nickyvdicarlo · · focus · HN ↗
So even given an ascii String, if you want to do something like
char c = input.charAt(index)
the JVM is going to jump to that index, check whether the character at that index isLatin(), then cast that single byte into to a 2 byte char... every single time.
in the naive solution, omething like 15% of the CPU cycles were spent checking isLatin() over and over again