How Smartcat counts words in Chinese, Japanese, Thai and other non-spaced languages
Overview
Most languages separate words with spaces, so Smartcat counts them directly. Chinese, Cantonese, Japanese, Thai, Burmese, Tibetan, Khmer, Lao, and Vietnamese don't work this way — they either run without spaces, or use spaces to separate syllables, phrases, or sentences rather than words. For these languages, Smartcat estimates the word count from the number of characters or syllables in the text, using a ratio specific to each language or writing system.
Understanding Smartcat's counting method allows you to better understand your project statistics, quotes, and Smartwords charges.
How it works
| Language group | How Smartcat Counts Words |
|---|---|
| Languages with spaces between words (English, Russian, German, Korean, and others) | Words separated by spaces |
| Chinese (Simplified, Traditional), Cantonese, Japanese | Estimate based on character ratios |
| Thai, Burmese, Tibetan | Estimate based on character ratios |
| Vietnamese, Khmer, Lao | Estimate based on syllable ratios |
How Smartcat calculates the word count
Smartcat counts each segment separately and adds the segment totals together for your document total:
-
Smartcat counts the segment's main script in its own unit — characters, base letters, or syllables, depending on the language — multiplies that count by the language's ratio, and rounds to the nearest whole word
-
Any word outside the main script, such as a Latin word, a number, or a product name, counts as one word on its own
-
A segment that contains any text in its main script always has a minimum count of one word; it never counts as zero words.
-
Each URL or email link counts as one word, and its characters are left out of the rest of the calculation
Ratios by language
| Language | Unit | Ratio (Units per “word”) |
|---|---|---|
| Chinese (Simplified and Traditional), Cantonese | Han character | 0.70 |
| Japanese | Hiragana or katakana character | 0.30 |
| Japanese | Kanji (including 々) | 0.60 |
| Thai | Base letter (combining vowels and tone marks are not counted) | 0.25 |
| Lao | Base letter | 0.27 |
| Burmese | Base letter | 0.34 |
| Khmer | Base letter | 0.40 |
| Tibetan | Syllable (separated by the tsheg ་) | 0.60 |
| Vietnamese | Syllable | 0.70 |
Japanese has multiple ratios because it mixes scripts — each character is weighted by the script it belongs to.
Examples
| Language | Text | Estimated words | How |
|---|---|---|---|
| Chinese | 两个词 | 2 | 3 ideographs × 0.70 = 2.10 → 2 |
| Chinese | 我有3个iPhone | 4 | 3 ideographs → 2, plus “3” and “iPhone” |
| Japanese | 東京タワー | 2 | 2 kanji × 0.60 + 3 katakana × 0.30 = 2.10 → 2 |
| Japanese | Hello、世界! | 2 | 世界: 2 kanji × 0.60 = 1.20 → 1, plus “Hello” |
| Thai | สวัสดีครับ | 2 | 7 base letters × 0.25 = 1.75 → 2 |
| Thai | สวัสดี ABC 123 | 3 | 4 base letters → 1, plus “ABC” and “123” |
| Lao | ສະບາຍດີ | 2 | 6 base letters × 0.27 = 1.62 → 2 |
| Khmer | សួស្ដី | 1 | 3 base letters × 0.40 = 1.20 → 1 |
| Burmese | မြန်မာမြန်မာမြန်မာ | 3 | 9 base letters × 0.34 = 3.06 → 3 |
| Tibetan | བཀྲ་ཤིས་བདེ་ལེགས། | 2 | 4 syllables × 0.60 = 2.40 → 2 |
| Vietnamese | Tôi yêu tiếng Việt | 3 | 4 syllables × 0.70 = 2.80 → 3 |
| Chinese | 看 https://example.com | 2 | 1 ideograph → 1, plus the URL |
Project statistics and word count
For projects with a Chinese, Cantonese, or Japanese source, your project statistics and Excel export break your word count into three columns: words in alphabetic languages, Asian characters, and your combined total calculated with your document's method. The combined total can be lower than the number of Asian characters in the text.
Want to learn more?
See how Smartcat can transform your localization workflow.
Still need help?
Our support team responds within one business day.