Language Access in Healthcare
EvidenceE-0340Initial AI draft

Google Translate sentences contained more errors of any severity than professional translation (39% vs 22%)

2026-06-054 out · 0 in

Source

Khanna (2011). Performance of an online translation tool when applied to patient educational material. Journal of Hospital Medicine.

Description #

Google Translate (GT) Spanish sentences from the AHRQ warfarin brochure were more likely to contain an error of any severity than the professionally translated sentences (39% vs 22%, P = 0.05) (Table 1). GT was more error-prone overall, though the difference reached only borderline significance.

"GT-translated sentences contained more errors of any severity as compared to the professional translations (39% vs 22%, P = 0.05), but a similar number of serious, clinically impactful errors (severity scores of 3, 2, or 1; 4% vs 2%, P = 0.61)." (Khanna, 2011, p. 523)

Methods Context #

What? #

The observable: the proportion of translated sentences flagged as containing an error of any kind, derived from the severity domain (any severity rating other than "Sentence basically accurate").

"Evaluators also assessed severity, a new measure of potential harm if a given sentence was assessed as having errors of any kind, ranging from 5 ("Error, no effect on patient care") to 1 ("Error, dangerous to patient") with an additional option of N/A ("Sentence basically accurate")." (Khanna, 2011, p. 520)

How? #

Blinded severity scoring of A-0016 and professional sentences; "any error" was defined as any sentence not assigned to the "N/A, Sentence basically accurate" category.

"assigned to the "N/A, Sentence basically accurate" category (ie, all sentences with a score between 5 and 1)" (Khanna, 2011, p. 523)

Who? #

Forty-five sentences from the professionally prepared AHRQ warfarin-use instruction manual, each scored by three bilingual native-Spanish-speaking research assistants.

"A total of 45 sentences were evaluated by the bilingual research assistants." (Khanna, 2011, p. 522)

Other Notes #

The difference in overall error frequency was borderline (P = 0.05). This "any error" rate is distinct from the error-severity finding: the two groups did not differ in the frequency of serious, clinically impactful errors (see the paired severe-error EVD).