iTranslate matched human Spanish translators on the two simpler sentences but scored lower on the most difficult sentence
Source
Chen (2017). Machine or Human? Evaluating the Quality of a Language Translation Mobile App for Diabetes Education Material. JMIR Diabetes.
Description #
On blinded evaluation of English-to-Spanish translations of three diabetes-education sentences, the iTranslate app matched the professional human translator on the two relatively simple sentences (S2 and S3) but scored lower on the most difficult sentence (S1). For S1, iTranslate received slightly lower scores than the human on Adequacy and Meaning (3.33 vs 4) and larger gaps on Fluency (2 vs 4) and Severity (3.67 vs 4.67), whereas the two simpler sentences reached almost perfect fluency (Table 3; Fig. 2).
"Within the Fluency domain for iTranslate, the two relatively simple sentences (S2 and S3) had almost perfect fluency; however, the most difficult sentence (S1) had marginal fluency with several grammatical errors (Fluency=2)." (Chen, 2017, p. 6)
"For the most difficult sentence (S1), there was a slight difference between iTranslate and the Spanish human in the Adequacy and Meaning domains, where iTranslate received slightly lower scores (3.33 vs 4). We also noticed some gaps for S1 in the Fluency and Severity domains, where iTranslate received lower scores (2 vs 4 and 3.67 vs 4.67)." (Chen, 2017, p. 7)
Methods Context #
What? #ⓘ
The observable: translation quality of each sentence, scored by raters on four 5-point criteria — Fluency, Adequacy, Meaning, and Severity.
"instructed the raters to evaluate the translated sentences based on four criteria—Fluency, Adequacy, Meaning, and Severity—on a 5-point scale (1 indicates the lowest quality and 5 indicates the highest quality)." (Chen, 2017, p. 4)
How? #ⓘ
Certified medical translators scored the A-0015ArtifactA-0015Initial AI draftiTranslate voice-enabled mobile language translation appTo help patients and health care providers communicate across a language barrier by instantly converting spoken or written input into another language, so that individuals with limited English proficiency (LEP) can acces… version and the professional-human version of each sentence, blinded to which was which by labeling the audio files "version 1" (iTranslate) and "version 2" (human).
"To minimize rater bias and blind the evaluation process, the audio files were marked as version 1 (sentences translated by iTranslate) and version 2 (sentences translated by a human)." (Chen, 2017, p. 4)
Who? #ⓘ
Three American Translators Association (ATA)-certified Spanish medical translators evaluated the Spanish translations; the material was three spoken questions drawn from a publicly available diabetes patient-education pamphlet.
"We recruited six certified medical translators (three Spanish and three Chinese) to conduct blinded evaluations of the following versions: (1) sentences interpreted by iTranslate, and (2) sentences interpreted by the professional human translators." (Chen, 2017, p. 1)
"We chose a publicly available diabetes patient education pamphlet as a heuristic example for this pilot study." (Chen, 2017, p. 3)
Other Notes #
Sentence difficulty was indexed by Flesch-Kincaid grade level: S2 = 0.0, S3 = 1.0, S1 = 7.1; the quality gap appeared only on the hardest sentence (S1). This is a translation-quality (accuracy) rating by expert translators, not a comprehension outcome measured in LEP patients.
Caveats #
- Pilot evaluation of only three short sentences in two languages judged by translators not LEP patients This was a pilot study whose evidence rests on only three short spoken sentences from a single diabetes patient-education pamphlet, translated into just two languages (Spanish and Chinese). Because the sentence sample is so small and language-limited, the authors caution the findings should not be generalized to other languages or to longer, real-world clinical communication. In addition, translation quality was rated by certified translators rather than measured as comprehension or outcomes in actual LEP patients, so the app's usefulness in real patient encounters remains untested.