Language Access in Healthcare
EvidenceE-0333Initial AI draft

Bilingual raters preferred postedited MT and HT equivalently (37 vs 36 votes)

2026-06-053 out · 0 in

Source

Turner (2014). A comparison of human and machine translation of health promotion materials for public health practice: time, costs, and quality. Journal of Public Health Management and Practice.

Description #

When bilingual public health professionals blindly compared each postedited machine-translated (MT) document against its human-translated (HT) counterpart and voted for the version they preferred (or "equivalent"), preferences split almost evenly. Combining Phase 1 (25 matched pairs) and Phase 2 (3 documents), the tally was 37 votes for MT versus 36 votes for HT — a 1-vote edge for MT — indicating the two were equivalently preferred (Table 3).

"Reviewer preferences were combined from both Phase 1 and Phase 2, which resulted in 36 "votes" for HT and 37 "votes" for MT." (Turner, 2014, p. 527)

"A quality comparison by bilingual public health professionals showed that MT and HT were equivalently preferred." (Turner, 2014, p. 523)

Methods Context #

What? #

The observable: the rater's preference between the postedited MT and HT versions of the same document (or a judgment that they were equivalent), tallied as votes.

"Quality raters were instructed to review each translated document and indicate whether they preferred one translation to the other or felt they were equivalent." (Turner, 2014, p. 525)

How? #

A blinded pairwise comparison: raters saw the machine- and human-translated document pairs alongside the English original and cast preferences that were tallied voting-style (a preference = 1 vote; a judgment of equivalence = 1 vote for each).

"They were blindly presented with the machine/human-translated document pairs along with their English originals." (Turner, 2014, p. 525)

Who? #

Bilingual, native Spanish-speaking public health professionals served as quality raters; the quality analysis was conducted only on the Spanish-language documents.

"The quality analysis was performed only with Spanish documents." (Turner, 2014, p. 527)

Other Notes #

Quality here is a subjective, relative preference judgment by translation-experienced professionals, not an assessment against a formal linguistic standard (contrast the CIoL-scored accuracy in Hibbs 2026), and not a judgment by the lay target audience — a limitation captured as a caveat on this EVD. The near-tie (37 vs 36) is read by the authors as "equivalent"; it therefore opposes the cross-paper claim that human translation is preferred over MT-plus-postediting on quality (which the English-Chinese findings in Turner 2015 support). The Spanish–English pairing here is less linguistically divergent than the Chinese pairing, a plausible reason the equivalence held.

Caveats #

  • Quality was judged by public health translators not the target lay audiences The quality equivalence was established by public health professionals who were themselves involved in translation activities, not by members of the lay, limited-English-proficient audiences for whom the materials are intended. Professional translators may weigh linguistic fidelity differently from how a lay reader experiences comprehensibility, so equivalent expert preference does not guarantee equivalent real-world understanding by the target readers.