Google Translate sentences contained more errors of any severity than professional translation (39% vs 22%)
Source
Khanna (2011). Performance of an online translation tool when applied to patient educational material. Journal of Hospital Medicine.
Description #
Google Translate (GT) Spanish sentences from the AHRQ warfarin brochure were more likely to contain an error of any severity than the professionally translated sentences (39% vs 22%, P = 0.05) (Table 1). GT was more error-prone overall, though the difference reached only borderline significance.
"GT-translated sentences contained more errors of any severity as compared to the professional translations (39% vs 22%, P = 0.05), but a similar number of serious, clinically impactful errors (severity scores of 3, 2, or 1; 4% vs 2%, P = 0.61)." (Khanna, 2011, p. 523)
Methods Context #
What? #ⓘ
The observable: the proportion of translated sentences flagged as containing an error of any kind, derived from the severity domain (any severity rating other than "Sentence basically accurate").
"Evaluators also assessed severity, a new measure of potential harm if a given sentence was assessed as having errors of any kind, ranging from 5 ("Error, no effect on patient care") to 1 ("Error, dangerous to patient") with an additional option of N/A ("Sentence basically accurate")." (Khanna, 2011, p. 520)
How? #ⓘ
Blinded severity scoring of A-0016ArtifactA-0016Initial AI draftGoogle TranslateA free, general-purpose machine translation app increasingly used as an ad-hoc communication tool in healthcare settings to bridge language barriers with LEP patients, including via voice-to-voice translation. "One such… and professional sentences; "any error" was defined as any sentence not assigned to the "N/A, Sentence basically accurate" category.
"assigned to the "N/A, Sentence basically accurate" category (ie, all sentences with a score between 5 and 1)" (Khanna, 2011, p. 523)
Who? #ⓘ
Forty-five sentences from the professionally prepared AHRQ warfarin-use instruction manual, each scored by three bilingual native-Spanish-speaking research assistants.
"A total of 45 sentences were evaluated by the bilingual research assistants." (Khanna, 2011, p. 522)
Other Notes #
The difference in overall error frequency was borderline (P = 0.05). This "any error" rate is distinct from the error-severity finding: the two groups did not differ in the frequency of serious, clinically impactful errors (see the paired severe-error EVD).