‹ BackHN Continuity

Thread

UTF-8000: Unlimited UTF-8

134 points · 122 comments · vismit2000

  1. mqus · · focus · HN ↗
    Some ideas of what to do with this space:

    - fully-customizable emojis (think of a RPG-like character customization screen)

    - heck, why not full jpegs/gifs?

    - some unicode programming script (running Doom)

    - ?

    That said, some very minor (HN-style) nitpick:

    > Otherwise for an n byte code unit this is (5n+1) / 8n, that is 5n+1 content bits out of a total of 8n bits from n bytes. We can rewrite this as (5/8) + 1/(8n) which moderately quickly approaches 5/8 = 62.5%. It is nice that this limit is nonzero and does not depend on n.

    Isn't a limit by definition no longer dependent on n?

    1. jeroenhd · · focus · HN ↗
      U+E000–U+F8FF, U+F0000–U+FFFFD, and U+100000–U+10FFFD can already provide you with your own emoji, as that range has been reserved for private use. Extending the range further might make sense if you need even more space in your program, but that's a lot of space already.
      1. mqus · · focus · HN ↗
        2-3 bytes are not much space for anything. Sure, you could use multiple successive ones of these code points and define your own "continuation" encoding in these ranges, but that doesn't seem right to me somehow
        1. flohofwoe · · focus · HN ↗
          That combination is how it already works. You can build combined "characters" (grapheme clusters) from multiple code points, e.g. you could have a "base emoji" followed by a "modifier" emoji, and AFAIK that's how emojis with different skin colors work (one code point for the base emoji (e.g. 'thumbs up'), and a number of skin color modification code points which can be applied to all emojis that involve skin color.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.