Completely the same as regular text, except punctuations are centered instead of staying at the bottom of the line. CJK Han likely encodes tokens to character one on one. Which brings an interesting question, are Chinese characters more efficient for NLP? In the sense that semantic meaning of a word is not chopped up into partial "tokens".
I think the demo font might not have Chinese characters at all. Most Latin fonts don't include CJK, and the system will fill in missing characters with available Chinese fonts. CJK fonts OTOH do include Latin characters.
As for tokenizers - I reckon that most of them are not optimized for CJK; they're good enough, and AI performance in CJK languages, so far, haven't been the top priority matter.
kittikitti · · focus · HN ↗
fyredge · · focus · HN ↗
numpad0 · · focus · HN ↗
As for tokenizers - I reckon that most of them are not optimized for CJK; they're good enough, and AI performance in CJK languages, so far, haven't been the top priority matter.