Language Access in Healthcare
EvidenceE-0361Initial AI draft

Word sense (40%) and word order (22%) were the most common English-to-Chinese machine translation errors

2026-06-054 out · 0 in

Source

Turner (2015). Machine Translation of Public Health Materials From English to Chinese: A Feasibility Study. JMIR Public Health Surveill.

Description #

When 60 English public-health documents were machine-translated into Traditional Chinese with Google Translate and every error annotated against a linguistic categorization scheme, the errors were dominated by word-sense errors (translating a word's meaning incorrectly), which made up 40% of all annotated errors, followed by word-order errors (22%) and missing words (16%) (Table 2). Word sense and word order — the two most frequent categories — are the error types the authors note require the most cognitive effort to correct.

"For example, word sense errors (errors where the word meaning was translated incorrectly) constituted 40% of all annotated errors. The next most common error types involved word order (22%) and missing words (16%)." (Turner, 2015, p. 5)

Methods Context #

What? #

The observable: the frequency of each machine-translation error type, expressed as a percentage of all errors annotated across the document set.

"the right-hand column shows the corresponding frequency of the error type, computed as the percentage of all errors annotated in the total set of 60 documents." (Turner, 2015, p. 5)

How? #

The raw Google Translate output for each document was annotated by a native Chinese speaker with linguistics training using a predefined error-categorization scheme, and aggregate error statistics were computed.

"We developed a categorization scheme for MT errors, and all MTs were annotated based on this scheme by a native Chinese speaker with formal training in linguistics. Subsequently, aggregate error statistics were computed to gain insights into the most frequent error categories: word sense, word order, missing word, superfluous word, orthography/punctuation, particle error, untranslated word, pragmatic error, and other grammar error." (Turner, 2015, p. 4)

Who? #

The corpus was 60 health-promotion documents collected in both English and Traditional Chinese from US public-health agency websites (e.g., CDC, several state and county health departments), machine-translated via A-0016.

"We collected 60 health promotion documents available in English and Chinese (Traditional) from public health websites in the United States." (Turner, 2015, p. 3)

Other Notes #

Table 2 lists the full distribution: word sense 40%, word order 22%, missing word 16%, superfluous word 14%, other grammar error 3%, orthography/punctuation 3%, particle error 1%, untranslated word 0.03%, pragmatic error 0.01%.

Caveats #

  • Only a single machine-translation engine (Google Translate) was evaluated All machine translations, and hence the error analysis, were produced by a single engine, Google Translate. The measured error profile (word-sense and word-order errors dominating) could in principle be specific to that engine rather than characteristic of English-to-Chinese machine translation generally. The authors argue the limitation is modest because most statistical machine-translation systems of the era shared underlying models and would likely show similar error types and frequencies, but this remains an assumption rather than a tested comparison.