‹ BackHN Continuity

Thread

Generate fonts where every LLM token is the same width

95 points · 23 comments · z-mach9

  1. kittikitti · · focus · HN ↗
    I wonder how this would look in Chinese Mandarin.
    1. fyredge · · focus · HN ↗
      Completely the same as regular text, except punctuations are centered instead of staying at the bottom of the line. CJK Han likely encodes tokens to character one on one. Which brings an interesting question, are Chinese characters more efficient for NLP? In the sense that semantic meaning of a word is not chopped up into partial "tokens".
      1. altairprime · · focus · HN ↗
        Perhaps relevant: Chinese language is not more efficient than English in vibe coding: A preliminary study on token cost and problem-solving rate <a href="https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2604.14210v1" rel="nofollow">https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2604.14210v1
        1. treebeard901 · · focus · HN ↗
          Yeah that is one draw back they mention in the paper. So much of the full test stack is trained on and dependent on English.

          A few months ago I went down this same rabbit hole trying to determine if using Chinese has any inherent advantages due to the way more context is encoded into Chinese characters versus what you find in English.

          The total token size may not take into account any benefits gained by Chinese characters potentially encoding more context.

          I have been meaning to see if any Chinese native tokenizers have been created to account for this feature of the language.

          Unfortunately, many Chinese and Japanese characters can have many different meanings depending on how they are used. As a result they do not break down as well to be tokenized as a language like English.

          Maybe some kind of hybrid intermediate language would be most optimal.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.