How Smartcat counts words in Chinese, Japanese, Thai and other non-spaced languages

Overview

Most languages separate words with spaces, so Smartcat counts them directly. Chinese, Cantonese, Japanese, Thai, Burmese, Tibetan, Khmer, Lao, and Vietnamese don't work this way — they either run without spaces, or use spaces to separate syllables, phrases, or sentences rather than words. For these languages, Smartcat estimates the word count from the number of characters or syllables in the text, using a ratio specific to each language or writing system.

Understanding Smartcat's counting method allows you to better understand your project statistics, quotes, and Smartwords charges.


How it works

Language group How Smartcat Counts Words
Languages with spaces between words (English, Russian, German, Korean, and others) Words separated by spaces
Chinese (Simplified, Traditional), Cantonese, Japanese Estimate based on character ratios
Thai, Burmese, Tibetan Estimate based on character ratios
Vietnamese, Khmer, Lao Estimate based on syllable ratios

How Smartcat calculates the word count

Smartcat counts each segment separately and adds the segment totals together for your document total:

  1. Smartcat counts the segment's main script in its own unit — characters, base letters, or syllables, depending on the language — multiplies that count by the language's ratio, and rounds to the nearest whole word

  2. Any word outside the main script, such as a Latin word, a number, or a product name, counts as one word on its own

  3. A segment that contains any text in its main script always has a minimum count of one word; it never counts as zero words.

  4. Each URL or email link counts as one word, and its characters are left out of the rest of the calculation

Ratios by language

Language Unit Ratio (Units per “word”)
Chinese (Simplified and Traditional), Cantonese Han character 0.70
Japanese Hiragana or katakana character 0.30
Japanese Kanji (including 々) 0.60
Thai Base letter (combining vowels and tone marks are not counted) 0.25
Lao Base letter 0.27
Burmese Base letter 0.34
Khmer Base letter 0.40
Tibetan Syllable (separated by the tsheg ་) 0.60
Vietnamese Syllable 0.70

Japanese has multiple ratios because it mixes scripts — each character is weighted by the script it belongs to.

Examples

Language Text Estimated words How
Chinese 两个词 2 3 ideographs × 0.70 = 2.10 → 2
Chinese 我有3个iPhone 4 3 ideographs → 2, plus “3” and “iPhone”
Japanese 東京タワー 2 2 kanji × 0.60 + 3 katakana × 0.30 = 2.10 → 2
Japanese Hello、世界! 2 世界: 2 kanji × 0.60 = 1.20 → 1, plus “Hello”
Thai สวัสดีครับ 2 7 base letters × 0.25 = 1.75 → 2
Thai สวัสดี ABC 123 3 4 base letters → 1, plus “ABC” and “123”
Lao ສະບາຍດີ 2 6 base letters × 0.27 = 1.62 → 2
Khmer សួស្ដី 1 3 base letters × 0.40 = 1.20 → 1
Burmese မြန်မာမြန်မာမြန်မာ 3 9 base letters × 0.34 = 3.06 → 3
Tibetan བཀྲ་ཤིས་བདེ་ལེགས། 2 4 syllables × 0.60 = 2.40 → 2
Vietnamese Tôi yêu tiếng Việt 3 4 syllables × 0.70 = 2.80 → 3
Chinese 看 https://example.com 2 1 ideograph → 1, plus the URL

Project statistics and word count

For projects with a Chinese, Cantonese, or Japanese source, your project statistics and Excel export break your word count into three columns: words in alphabetic languages, Asian characters, and your combined total calculated with your document's method. The combined total can be lower than the number of Asian characters in the text.

Want to learn more?

See how Smartcat can transform your localization workflow.

Book a demo

Still need help?

Our support team responds within one business day.

Open a support case