Language Access in Healthcare
EvidenceE-0223Initial AI draft

Machine and professional booklet translation did not differ in understanding of treatment intent (multivariate OR 0.99) among LEP adults

2026-06-053 out · 0 in

Source

Sp (2026). Translation approaches to support systemic anti-cancer therapy consent for individuals with limited English proficiency.. Supportive care in cancer : official journal of the Multinational Association of Supportive Care in Cancer.

Description #

Among LEP adults randomised to read either an unsupervised machine (Google Translate) or a professional (human) Bengali translation of the same SACT information booklet, translation type was not associated with the primary outcome of correctly understanding treatment intent. The effect was null on both univariate (OR = 1.15, 95% CI 0.43–3.07, p = 0.77) and covariate-adjusted multivariate models (OR = 0.99, 95% CI 0.33–2.99, p = 0.99), and no secondary comprehension outcome differed either (Supplementary Table 2). This null occurred despite a large difference in objective translation quality between the two arms (see companion translation-error EVD).

"Randomisation to professional or machine translations was not associated with the primary outcome (univariate OR=1.15, 95% CI 0.43–3.07, p=0.77; multivariate OR=0.99, 95% CI=0.33–2.99, p=0.99)" (Hibbs, 2026, p. 4)

"nor any secondary outcomes (Total Comprehension Score, confidence in explaining myeloma to a family member, or perceived clarity of language; multivariate p=0.46, p=0.13, p=0.74, respectively)." (Hibbs, 2026, p. 4)

Methods Context #

What? #

The observable: whether a participant correctly understood treatment intent — that the myeloma SACT regimen is non-curative — after reading the allocated translated booklet.

"We chose a primary outcome of the proportion of participants understanding treatment intent: specifically, that the myeloma SACT regimen would not cure myeloma but instead aims to increase lifespan and quality of life." (Hibbs, 2026, p. 3)

How? #

Participants were randomised 1:1 to a professional (human) or unsupervised machine translation of the booklet and completed the Booklet Comprehension Tool immediately after reading; associations were estimated with univariate and covariate-adjusted logistic regression.

"Participants were randomised to provision of a professional (human) translation or an unsupervised machine translation. Immediately after reading their allocated booklet, participants were asked to complete the Booklet Comprehension Tool." (Hibbs, 2026, p. 3)

"The association of randomisation with outcomes was assessed using both univariate and multivariate logistic regression models adjusted for self-reported age, gender, highest level of education, self-rated prior knowledge of myeloma and English language proficiency assessed over multiple domains." (Hibbs, 2026, p. 4)

Who? #

Healthy Bengali- or Sylheti-speaking adults with LEP recruited from community groups in East and North London; 121 of 123 randomised booklet participants completed at least part of the comprehension tool and were analysed (61 machine, 60 professional).

"123 individuals were recruited and randomised for the booklet component, of whom 121 completed at least part of their comprehension tool and were included in the analysis." (Hibbs, 2026, p. 4)

Other Notes #

This is a null result and is wired as contradicting the claim that professional translation improves comprehension over machine translation. The confidence intervals are wide (study was exploratory and not formally powered), so the null is imprecise rather than a demonstration of equivalence.

Caveats #

  • Bengali is low-resource, so machine-translation results may not generalize to higher-resource languages Machine-translation quality is language-specific. Bengali is a low-resource language with a limited linguistic corpus for training machine-translation tools, so it is likely to be translated worse than higher-resource languages (e.g. French, Spanish, German) for which such tools have far more training data. The poor machine-translation performance and high critical-error count observed here may therefore not generalize to other language groups, and conclusions about machine translation should not be extrapolated across languages.
  • Exploratory study not powered to its objectives, yielding wide confidence intervals This was an exploratory study with no pre-existing data on expected SACT-comprehension rates in this population, so no formal sample-size calculation was possible and the sample was set by available time, funding and community-partner capacity. Consequently the study was not definitively powered to its objectives and the effect-size estimates carry wide confidence intervals — the odds ratios and regression coefficients are imprecise and should be treated as hypothesis-generating rather than definitive.