Machine translation was least accurate for South and Southeast Asian languages (Bengali Hindi Punjabi Vietnamese)
Source
Das (2019). Dangers of Machine Translation: The Need for Professionally Translated Anticipatory Guidance Resources for Limited English Proficiency Caregivers. Clinical Pediatrics.
Description #
Google Translate's accuracy varied sharply by target language: Bengali (1.39), Hindi (1.38), Punjabi (1.28), and Vietnamese (1.83) received the lowest mean accuracy scores, all in the "minimally useful" range, whereas Western European languages scored highest (Table 1). The authors note this gradient is consistent with prior work finding Google Translate most accurate for Western European languages and least accurate for Asian or African languages.
"Bengali, Hindi, Punjabi, and Vietnamese received the lowest accuracy scores, indicating only "minimally useful" translations." (Das, 2019, p. 248)
"The results of this study are consistent with the finding that, for health-related translations from English, Google Translate is most accurate for Western European languages and least accurate for Asian or African languages." (Das, 2019, p. 248)
Methods Context #
What? #ⓘ
The observable: the mean per-language translation accuracy score (rubric 1–5) for each of the 20 non-English languages, compared across language families.
"Mean translation accuracy scores and Wilcoxon test results for each language are noted in Table 1." (Das, 2019, p. 248)
How? #ⓘ
The nine AAP safety statements were machine-translated by A-0016ArtifactA-0016Initial AI draftGoogle TranslateA free, general-purpose machine translation app increasingly used as an ad-hoc communication tool in healthcare settings to bridge language barriers with LEP patients, including via voice-to-voice translation. "One such… into each language, back-translated to English by native-proficient professionals, scored on the ATA-adapted 5-point rubric, and averaged per language (Table 1).
"Nine bulleted statements obtained from the safety section of the AAP's English-language, 12-month well-child anticipatory guidance handout were translated into Spanish and 19 other foreign languages using Google Translate." (Das, 2019, p. 247)
Who? #ⓘ
The equivalence class: the top 20 non-English languages spoken in the United States, here contrasting the lowest-scoring South/Southeast Asian languages (Bengali, Hindi, Punjabi, Vietnamese) against higher-scoring Western European languages.
"This study assesses the accuracy of a popular free machine translation service, in translating AAP anticipatory guidance safety guidelines for the top 20 foreign languages spoken in the United States." (Das, 2019, p. 247)
Other Notes #
The lowest-scoring languages here (Bengali, Hindi, Punjabi, Vietnamese) are lower-resource for machine translation than Western European languages; the per-language means for these were roughly 1.3–1.8 out of 5 (Table 1), i.e., translations with frequent meaning-changing errors.
Caveats #
- Untrained small back-translator samples confound machine-translation accuracy scores Accuracy was measured indirectly, by scoring English back-translations rather than the machine translations themselves, and the back-translators — though native speakers — were mostly physicians or nurses not formally trained in translation, drawn from small samples that varied in size across languages. As a result, each language's score reflects the back-translators' abilities and errors as well as the quality of the machine translation, and the small, inconsistent samples limit precision and cross-language comparability.
- Back-translators familiarity with the guidance statements may overstate machine-translation accuracy All back-translators were employed at a pediatric hospital and were therefore likely already familiar with the anticipatory guidance safety statements being translated. That prior knowledge could have let them reconstruct the intended meaning from garbled machine output more readily than an actual limited-English-proficiency parent could, so the recorded accuracy scores may overstate how well the machine translations would communicate to the LEP caregivers the resources are meant for — biasing the study toward underestimating the real-world harm of machine translation.