Language Access in Healthcare
EvidenceE-0343Initial AI draft

Sentence complexity predicted preference for professional over Google Translate translation (3.6 vs 2.6 for complex vs simple)

2026-06-054 out · 0 in

Source

Khanna (2011). Performance of an online translation tool when applied to patient educational material. Journal of Hospital Medicine.

Description #

Sentence complexity, but not length, was significantly associated with preference: evaluators preferred the professional translation over Google Translate (GT) for complex English sentences (>1 clause) while being ambivalent about simple ones (mean preference 3.6 vs 2.6, P = 0.03) (Table 1). For simple single-clause sentences the preference for professional translation disappeared.

"Complexity, however, was significantly associated with preference: evaluators' preferred the professional translation for complex English sentences while being more ambivalent about simple English sentences (3.6 vs 2.6, P = 0.03)." (Khanna, 2011, p. 523)

"Evaluators preferred the professionally translated sentences for complex sentences, but when the English source sentence was simple—containing a single clause—this preference disappeared." (Khanna, 2011, p. 524)

Methods Context #

What? #

The observable: the association between source-sentence complexity and evaluators' standardized preference for the professional over the GT translation.

"We found that sentence length was not associated with scores for fluency, adequacy, meaning, severity, or preference (P > 0.30 in each case)." (Khanna, 2011, p. 523)

How? #

Sentences were classified as simple vs complex by clause count; preference scores of A-0016 versus professional translations were modeled against complexity using clustered linear regression.

"Sentences were classified as "simple" if they contained one or fewer clauses and "complex" if they contained more than one clause." (Khanna, 2011, p. 522)

Who? #

The sentence pairs scored for preference, drawn from the AHRQ warfarin brochure, where 33% of English source sentences were simple and 67% complex, rated by three bilingual native-Spanish-speaking research assistants.

"Thirty-three percent of the English source sentences were "simple" and 67% were "complex."" (Khanna, 2011, p. 522)

Other Notes #

This moderation result gives the boundary condition for GT's comparable performance: GT looked equivalent to professional translation mainly on simple sentences, while the professional translation was preferred for complex, multi-clause sentences (which is also where GT's single "dangerous" error occurred).

Caveats #

  • Interrater reliability was only moderate, with meaning and preference domains scoring particularly poorly The manual scoring system had only moderate interrater reliability, and the meaning and preference domains in particular scored poorly (intraclass correlations of 0.42 for meaning and 0.37 for preference, versus 0.70 for fluency). Because evaluators frequently disagreed on meaning and on which translation they preferred, the null findings on meaning preservation and on overall preference — and the complexity-dependent preference pattern — rest on comparatively noisy measurements and warrant more thorough assessment before routine use.