On a practical matter, it seems like a bad idea to have codepoints that can take up to an arbitrary number of bytes - this just screams buffer overflow problems.
So in practicality, you’re going to want an arbitrary limit on this (the article suggests as much). But if you place a limit on it then you’ve got one implementation of the standard that can decode certain characters and another that can’t. Better to have one standard that puts a hard limit on the number of bytes and another standard that uses more bytes and so on.
You don't have a buffer overflow problem if you read it in a memory-safe way i.e. read it in chunks and realloc when you reach the size of your allocation.
What you will have is a potential denial-of-service attack - although this one isn't particularly great because there's zero amplification (they might as well just send garbage into your firewall)
I think the DoS has pretty common amplification vectors in the form of APIs that split or otherwise copy (e.g. materializing code points for Unicode regex searching).
Additionally, OOM inside a low level routine can be a troublesome attack, since OOM handling in many applications does questionable (nee vulnerable) things when crashes occur in not-known-to-be-memory-intensive code. Sure, that’s sloppy engineering, but it’s common.
2shortplanks · · focus · HN ↗
So in practicality, you’re going to want an arbitrary limit on this (the article suggests as much). But if you place a limit on it then you’ve got one implementation of the standard that can decode certain characters and another that can’t. Better to have one standard that puts a hard limit on the number of bytes and another standard that uses more bytes and so on.
Pannoniae · · focus · HN ↗
What you will have is a potential denial-of-service attack - although this one isn't particularly great because there's zero amplification (they might as well just send garbage into your firewall)
zbentley · · focus · HN ↗
Additionally, OOM inside a low level routine can be a troublesome attack, since OOM handling in many applications does questionable (nee vulnerable) things when crashes occur in not-known-to-be-memory-intensive code. Sure, that’s sloppy engineering, but it’s common.