Why do character and byte counts differ?
Unicode symbols can require several bytes in UTF-8 while appearing as one visible character.
Measure visible characters, non-whitespace characters, code points, and UTF-8 bytes.
Adjust the values below to get a clear estimate.
Grapheme segmentation estimates user-perceived characters, code points count Unicode scalar sequence elements, and TextEncoder reports UTF-8 byte length.
Visible characters = grapheme clusters; bytes = UTF-8 encoded lengthInputs: text: Hi ๐๐ฝ
Illustrative result: 5 graphemes ยท 9 code points ยท 11 UTF-8 bytes
The skin-tone emoji is one visible grapheme but contains multiple code points and several encoded bytes.
Unicode symbols can require several bytes in UTF-8 while appearing as one visible character.
The main count includes spaces and line breaks; a second count removes whitespace before measuring.