Language Access in Healthcare
EvidenceE-0338Initial AI draft

Google Translate and professional translation did not differ in adequacy (information preservation) of a Spanish warfarin brochure

2026-06-055 out · 0 in

Source

Khanna (2011). Performance of an online translation tool when applied to patient educational material. Journal of Hospital Medicine.

Description #

Google Translate (GT) and professional Spanish translations of the AHRQ warfarin brochure did not differ significantly on adequacy — the share of the original information preserved — with GT scoring 4.5 versus 4.8 for the professional translation (P = 0.19) (Table 1). Machine translation preserved information content comparably to professional translation.

"Comparisons for adequacy and meaning were not statistically significantly different." (Khanna, 2011, p. 523)

Methods Context #

What? #

The observable: adequacy, a Likert rating of how much of the source sentence's information was conveyed by the translation.

"adequacy being an assessment of information preservation ranging from 5 ("100% of information conveyed from the original") to 1 ("0% of information conveyed from the original")." (Khanna, 2011, p. 520)

How? #

Blinded scoring of A-0016 sentences against the professional translation, with evaluators having access to both the translated sentence and the original English sentence, compared via clustered linear regression.

"for adequacy, meaning, and severity, they had access to both the translated sentence and the original English sentence." (Khanna, 2011, p. 521)

Who? #

Forty-five sentences from the professionally prepared AHRQ warfarin-use instruction manual, each scored by three bilingual native-Spanish-speaking research assistants.

"A total of 45 sentences were evaluated by the bilingual research assistants." (Khanna, 2011, p. 522)

Other Notes #

This is a null (no-difference) result and supports an equivalence claim on information preservation, not a superiority claim. Note a source inconsistency: Table 1 reports adequacy as GT 4.5 vs professional 4.8, whereas the abstract reports the adequacy point estimates as 4.2 vs 4.5 ("similar adequacy (4.2 vs 4.5, P = 0.19)"); the adequacy and meaning point estimates appear transposed between the abstract and Table 1, though both sources agree the comparison was non-significant (P = 0.19).

Caveats #

  • Translation quality was judged by nonclinician evaluators, not patients, so comparable scores may not mean comparable patient comprehension Translation quality was rated by three college-educated, bilingual, nonclinician research assistants comparing sentences against the English source — not by the limited-English-proficiency patients who would actually read the material. Patients may have lower or more variable medical understanding and educational attainment than the evaluators, so the finding that machine and professional translations scored comparably on adequacy and meaning does not establish that patients would understand them equally well. Comparable expert-rated quality is an upstream proxy for, not a measurement of, patient comprehension.
  • The evaluated document was a professionally prepared brochure, not the spontaneous patient-specific text machine translation would actually be used for The material tested was a carefully, professionally prepared AHRQ patient-education brochure written at a 6th-grade reading level — precisely the kind of standardized document for which a professional translation would normally already exist. The authors note that online translation tools would most likely be deployed not for such material but for spontaneous, less-critical, patient-specific instructions, which may have different grammar, vocabulary, and error profiles. The comparable-quality finding therefore may not transfer to machine translation's intended real-world use case.