Posteditors rated raw English-to-Chinese machine translation adequacy 3.32 and fluency 3.0 out of 5
Source
Turner (2015). Machine Translation of Public Health Materials From English to Chinese: A Feasibility Study. JMIR Public Health Surveill.
Description #
Posteditors rated the raw Google Translate output on 1–5 adequacy and fluency scales. Average adequacy was 3.32 (SD 0.90) — indicating much, but not all, of the source meaning was preserved — and average fluency was 3.0 (SD 0.84), corresponding to a grammar quality of non-native Chinese (Table 3). Ratings varied greatly by individual posteditor rather than by document type or length; notably, the more experienced posteditors rated adequacy and fluency lower than their less experienced counterparts.
"On average, the posteditors rated the adequacy of the translations at 3.32 (SD 0.90), suggesting that much of the original meaning of the source text was preserved in the MT. Average fluency rating was 3.0 (SD 0.84), which corresponds to a grammar quality level of non-native Chinese." (Turner, 2015, p. 6)
"Interestingly, the posteditors who had more experience with translation and health rated the adequacy and fluency lower than did their less experienced counterparts (Table 3)." (Turner, 2015, p. 6)
Methods Context #
What? #ⓘ
The observable: adequacy (how much of the English source meaning the machine translation retained) and fluency (grammatical quality of the Chinese), each scored 1–5.
"An adequacy of 1 indicated that none of the original meaning of the English source text was retained in the MT, while an adequacy of 5 indicated that all of the meaning was retained. A fluency rating of 1 indicated that the MT was incomprehensible, while a rating of 5 indicated flawless Chinese." (Turner, 2015, p. 4)
How? #ⓘ
After completing postediting of a document, each posteditor filled out a questionnaire rating the adequacy and fluency of the translation on the 1–5 scales; means and SDs were computed per document.
"After completing postediting, participants were asked to fill out a questionnaire to rate the adequacy and fluency of each MT+PE on a scale of 1-5. These rating scales are common in human evaluations of machine translation quality [10]." (Turner, 2015, p. 4)
Who? #ⓘ
The six native-Chinese-speaking posteditors (varying translation and health experience) who corrected the 25 machine-translated public-health documents (Table 1, Table 3).
"From the memberships of local Chinese cultural organizations, 6 Chinese translators were recruited for postediting and screened for language ability and health experience." (Turner, 2015, p. 4)
Other Notes #
Adequacy/fluency here are self-ratings of the raw MT by the same people who postedited it, so they index the starting quality of the machine output rather than the final MT+PE product. The inverse relationship with posteditor experience suggests more expert raters applied a stricter quality bar.