Evaluators had no overall preference between Google Translate and professional Spanish translation
Source
Khanna (2011). Performance of an online translation tool when applied to patient educational material. Journal of Hospital Medicine.
Description #
When bilingual evaluators directly compared paired Google Translate (GT) and professional Spanish sentences, they expressed no overall preference for the professional translation: the mean preference score was 3.2 (95% CI 2.7 to 3.7; P = 0.36) on a scale where 3 indicated no preference (Table 1).
"Evaluators had no overall preference for the professional translation (3.2, 95% confidence interval = 2.7 to 3.7, with 3 indicating no preference; P = 0.36) (Table 1)." (Khanna, 2011, p. 523)
Methods Context #
What? #ⓘ
The observable: a blinded head-to-head preference rating between the two translations of the same sentence, on a 5-point scale centered on "no preference."
"Finally, evaluators rated a blinded preference (also a new measure) for either of two translated sentences, ranging from "Strongly prefer translation #1" to "Strongly prefer translation #2."" (Khanna, 2011, p. 520)
How? #ⓘ
For a subset of sentences, evaluators saw the GT and professional translation of the same sentence (with the original English available as reference), blinded to which was machine-generated, and indicated a preference; scores were standardized so that 5 = strong preference for the professional translation and 1 = strong preference for GT.
"We subsequently converted this to preference for the professional translation, ranging from 5 ("Strongly prefer the professional translation") to 1 ("Strongly prefer the GT translation") in order to standardize the responses (Figures 1 and 2)." (Khanna, 2011, p. 520)
Who? #ⓘ
Nineteen total sentence pairs scored for preference (10 in initial and 9 in post-consolidation evaluation), drawn from the AHRQ warfarin brochure and rated by three bilingual native-Spanish-speaking research assistants.
"we pooled all 45 sentences (as well as the 19 total sentence pairs scored for preference) for the final analysis." (Khanna, 2011, p. 522)
Other Notes #
Preference had the poorest interrater reliability of all domains (intraclass correlation 0.37), so this null overall-preference result should be read alongside that measurement caveat. The overall null masks a complexity-dependent pattern captured in the paired complexity-mediation EVD.
Caveats #
- Interrater reliability was only moderate, with meaning and preference domains scoring particularly poorly The manual scoring system had only moderate interrater reliability, and the meaning and preference domains in particular scored poorly (intraclass correlations of 0.42 for meaning and 0.37 for preference, versus 0.70 for fluency). Because evaluators frequently disagreed on meaning and on which translation they preferred, the null findings on meaning preservation and on overall preference — and the complexity-dependent preference pattern — rest on comparatively noisy measurements and warrant more thorough assessment before routine use.