Language Access in Healthcare
EvidenceE-0334Initial AI draft

Google Translate met the professional translation standard for only Spanish among 20 non-English languages

2026-06-054 out · 0 in

Source

Das (2019). Dangers of Machine Translation: The Need for Professionally Translated Anticipatory Guidance Resources for Limited English Proficiency Caregivers. Clinical Pediatrics.

Description #

When the professionally back-translated, AAP-published Spanish anticipatory guidance was used as a professional-quality benchmark (mean accuracy 5.00 on the 5-point rubric), Google Translate matched that professional standard for only one of the 20 non-English languages tested — Spanish, at 4.95 (Table 1). Spanish was also the only language whose accuracy scores did not differ significantly from the professional benchmark (Wilcoxon P = .080); every other language scored significantly lower (P ≤ .011 for all others, Table 1), with the next-best language, Portuguese, reaching only the "strong" range at 4.33.

"While the back-translations of AAP-published Spanish anticipatory guidance statements received a mean translation accuracy score of 5.00, Google Translate was unable to meet this professional standard for all but one language, Spanish (4.95)." (Das, 2019, p. 248)

"Portuguese fell into the "strong" translation range with a mean score of 4.33." (Das, 2019, p. 248)

Methods Context #

What? #

The observable: the accuracy of each machine translation, scored on a 5-point rubric (5 = "standard", 1 = "minimal") adapted from the American Translators Association.

"Back-translations were then evaluated on their accuracy using a 5-point rubric adapted from the American Translators Association." (Das, 2019, p. 247)

How? #

Nine AAP 12-month well-child safety statements were machine-translated with A-0016 into Spanish and 19 other languages; native-proficient health professionals back-translated them to English, and each language's scores were compared by Wilcoxon rank sum test to the back-translated AAP-published Spanish translation used as the professional control.

"Nine bulleted statements obtained from the safety section of the AAP's English-language, 12-month well-child anticipatory guidance handout were translated into Spanish and 19 other foreign languages using Google Translate." (Das, 2019, p. 247)

"To serve as a control, Spanish-speaking health care professionals back-translated the Spanish-language safety guidelines provided by the AAP." (Das, 2019, p. 247)

Who? #

The equivalence class: the AAP 12-month well-child anticipatory guidance safety statements, machine-translated into the top 20 non-English languages spoken in the United States and scored via native-proficient health-professional back-translators.

"Health care professionals with native proficiency in the target languages back-translated the machine-generated statements to English." (Das, 2019, p. 247)

Other Notes #

This is a translation-accuracy (quality) measurement, not a patient-comprehension outcome. Spanish is a special case in that the AAP already publishes a professional Spanish translation of this guidance; the finding highlights that for the 19 languages without such a professional resource, unsupervised machine translation did not reach professional quality.

Caveats #

  • Untrained small back-translator samples confound machine-translation accuracy scores Accuracy was measured indirectly, by scoring English back-translations rather than the machine translations themselves, and the back-translators — though native speakers — were mostly physicians or nurses not formally trained in translation, drawn from small samples that varied in size across languages. As a result, each language's score reflects the back-translators' abilities and errors as well as the quality of the machine translation, and the small, inconsistent samples limit precision and cross-language comparability.
  • Back-translators familiarity with the guidance statements may overstate machine-translation accuracy All back-translators were employed at a pediatric hospital and were therefore likely already familiar with the anticipatory guidance safety statements being translated. That prior knowledge could have let them reconstruct the intended meaning from garbled machine output more readily than an actual limited-English-proficiency parent could, so the recorded accuracy scores may overstate how well the machine translations would communicate to the LEP caregivers the resources are meant for — biasing the study toward underestimating the real-world harm of machine translation.