Language Access in Healthcare
EvidenceE-0335Initial AI draft

Nearly half of non-Spanish machine-translated safety statements were deficient or minimally useful

2026-06-054 out · 0 in

Source

Das (2019). Dangers of Machine Translation: The Need for Professionally Translated Anticipatory Guidance Resources for Limited English Proficiency Caregivers. Clinical Pediatrics.

Description #

Across the 690 individual statements machine-translated into the 19 non-Spanish languages, nearly half scored at the two lowest rubric levels: 28% (n = 193) were "minimally useful" (score 1, indicating frequent or serious meaning-changing errors) and 20.4% (n = 141) were "deficient" (score 2), so 48.4% fell at or below "deficient". Only 23.3% (n = 160) reached the top "standard" score, with a further 13.9% (n = 96) "strong" and 14.5% (n = 100) "acceptable" (Table 1).

"Of the 690 statements back-translated from non-Spanish languages, 28% (n = 193) were "minimally useful" (score of 1), 20.4% (n = 141) were "deficient" (score of 2), 14.5% (n = 100) were "acceptable" (score of 3), 13.9% (n = 96) were "strong" (score of 4), and 23.3% (n = 160) were "standard" (score of 5)." (Das, 2019, p. 248)

Methods Context #

What? #

The observable: the distribution of individual machine-translated safety statements across the five rubric accuracy levels, where a score of 1 ("minimal") denotes frequent/serious errors that changed the meaning with disruptive or inappropriate wording.

"and "minimal" (1, the lowest score) if the translation had frequent or serious errors that changed the meaning and had disruptive/inappropriate wording." (Das, 2019, p. 247)

How? #

Each of the nine AAP safety statements, machine-translated by A-0016 into 19 non-Spanish languages and back-translated to English by native-proficient professionals, was independently scored 1–5 on the ATA-adapted rubric and the scores tallied into accuracy categories.

"Back-translations were then evaluated on their accuracy using a 5-point rubric adapted from the American Translators Association." (Das, 2019, p. 247)

Who? #

The equivalence class: the pooled 690 back-translated statement instances (nine safety statements × the 19 non-Spanish languages among the top 20 spoken in the United States).

"Nine bulleted statements obtained from the safety section of the AAP's English-language, 12-month well-child anticipatory guidance handout were translated into Spanish and 19 other foreign languages using Google Translate." (Das, 2019, p. 247)

Other Notes #

This statement-level distribution complements the per-language means in Table 1: even languages with mid-range means contained many individual statements with meaning-changing errors, underscoring that machine output quality is uneven within a language, not only across languages.

Caveats #

  • Untrained small back-translator samples confound machine-translation accuracy scores Accuracy was measured indirectly, by scoring English back-translations rather than the machine translations themselves, and the back-translators — though native speakers — were mostly physicians or nurses not formally trained in translation, drawn from small samples that varied in size across languages. As a result, each language's score reflects the back-translators' abilities and errors as well as the quality of the machine translation, and the small, inconsistent samples limit precision and cross-language comparability.
  • Back-translators familiarity with the guidance statements may overstate machine-translation accuracy All back-translators were employed at a pediatric hospital and were therefore likely already familiar with the anticipatory guidance safety statements being translated. That prior knowledge could have let them reconstruct the intended meaning from garbled machine output more readily than an actual limited-English-proficiency parent could, so the recorded accuracy scores may overstate how well the machine translations would communicate to the LEP caregivers the resources are meant for — biasing the study toward underestimating the real-world harm of machine translation.