Google Translate scored lower on fluency than professional translation of a Spanish warfarin brochure (3.4 vs 4.7)
Source
Khanna (2011). Performance of an online translation tool when applied to patient educational material. Journal of Hospital Medicine.
Description #
On blinded evaluation by three bilingual research assistants, Google Translate (GT) Spanish sentences from an AHRQ warfarin-education brochure scored significantly lower on fluency (grammar and readability) than the professionally translated sentences (mean 3.4 vs 4.7 on a 5-point scale, P < 0.0001) (Table 1). GT was thus inferior to professional translation in grammatical quality.
"Sentences translated by GT received worse scores on fluency as compared to the professional translations (3.4 vs 4.7, P < 0.0001)." (Khanna, 2011, p. 523)
Methods Context #
What? #ⓘ
The observable: fluency, a Likert-scale rating of the grammatical correctness and readability of each translated Spanish sentence.
"with fluency being an assessment of grammar and readability ranging from 5 ("Perfect fluency; like reading a newspaper") to 1 ("No fluency; no appreciable grammar, not understandable")" (Khanna, 2011, p. 520)
How? #ⓘ
Blinded sentence-level comparison of A-0016ArtifactA-0016Initial AI draftGoogle TranslateA free, general-purpose machine translation app increasingly used as an ad-hoc communication tool in healthcare settings to bridge language barriers with LEP patients, including via voice-to-voice translation. "One such… against an independently produced professional Spanish translation, analyzed with clustered linear regression to account for each sentence being scored by all three evaluators.
"We compared the scores assigned to GT-translated sentences for each of the five manually scored domains as compared to the scores of the professionally translated sentences, as well as the impact of word count and sentence complexity on the scores achieved specifically by the GT-translated sentences, using clustered linear regression to account for the fact that each of the 45 sentences were scored by each of the three evaluators." (Khanna, 2011, p. 522)
Who? #ⓘ
Forty-five sentences drawn from a professionally prepared AHRQ warfarin-use instruction manual (6th-grade reading level), each scored by three nonclinician, bilingual, native-Spanish-speaking research assistants.
"We recruited three nonclinician, bilingual, native-Spanish-speaking research assistants as evaluators." (Khanna, 2011, p. 520)
Other Notes #
Fluency was the domain on which GT differed most clearly and significantly from professional translation, and it was the only one of the four content/quality domains with interrater reliability judged acceptable (intraclass correlation 0.70).