Language Access in Healthcare
EvidenceE-0326Initial AI draft

96.7% of LEP PACU patients used machine translation successfully at least once with no need for human interpreters

2026-06-054 out · 0 in

Source

Kapoor (2022). Use of Neural Machine Translation Software for Patients With Limited English Proficiency to Assess Postoperative Pain and Nausea. JAMA.

Description #

Over the PACU stay, 96.7% of the 30 Spanish-speaking LEP patients (29 of 30) were able to use the Google Translate conversation mode successfully to answer all pain and nausea questions at least once, and no patient required the hospital's standard institutional human translation services (Table 2). The authors summarize this in the Discussion as more than 90% of patients being able to communicate their symptoms via the application.

"Most patients (83.3%) could communicate via the application on their first assessment attempt, and most (96.7%) were able to use the application successfully to answer all questions at least once during their PACU stay, with no patients needing standard institutional translation services." (Kapoor, 2022, p. 4)

"We observed that more than 90.0% of patients were able to communicate their pain and nausea using the translation application." (Kapoor, 2022, p. 4)

Methods Context #

What? #

The observable: whether a patient could use the application successfully to answer all questions at least once during the PACU stay, and whether standard institutional human translation services were needed.

"most (96.7%) were able to use the application successfully to answer all questions at least once during their PACU stay, with no patients needing standard institutional translation services." (Kapoor, 2022, p. 4)

How? #

Single-arm feasibility cohort using an iPad running the A-0016 to play preformatted Spanish pain/nausea questions in the PACU; patients answered yes/no and gave numeric ratings.

"Patients responded with yes or no and gave numbers for pain and nausea ratings." (Kapoor, 2022, p. 4)

Who? #

30 postoperative Spanish-speaking, Hispanic patients in the PACU at a single US cancer center (MD Anderson), enrolled July to October 2021.

"The study took place from July 6 to October 16, 2021, including 1-day follow-ups." (Kapoor, 2022, p. 4)

Other Notes #

This at-least-once metric (96.7%) is the figure the authors foreground as clearing the "more than 90%" bar, whereas the stricter every-time metric (76.7%, companion EVD) fell below the prespecified 90% feasibility threshold. "No patients needing standard institutional translation services" is the paper's evidence that the tool substituted for, rather than merely supplemented, human interpreters in this cohort.

Caveats #

  • Machine-translation assessment tested in only one language (Spanish) All assessments were conducted in a single language: every enrolled patient spoke only Spanish and self-identified as Hispanic. Neural machine translation quality varies substantially by language pair, and Spanish is one of the highest-resource, best-supported languages for engines such as Google Translate. Feasibility and satisfaction observed here therefore may not transfer to lower-resource languages where translation accuracy is poorer, so the findings should not be read as evidence that the tool works equally well across the 70 languages the application nominally supports.
  • Every-time success fell below the prespecified 90% feasibility threshold The study prespecified feasibility as at least 90% of patients being able to answer all 5 questions asked, but the rate at which the application actually assessed both pain and nausea every time was 76.7% — below that threshold. The headline "more than 90%" success rests on the more lenient at-least-once metric (96.7%), which counts a patient as a success even if the tool failed on some attempts. Readers should note that by the study's own stricter, prespecified criterion the tool did not clear the 90% bar, so the "feasible" conclusion depends on which metric is adopted.
  • Machine-translation feasibility shown only in a 30-patient single-center cohort The feasibility, usability, and satisfaction estimates all come from a single-arm cohort of only 30 patients at one US cancer center, with no comparison group. The authors themselves flag the small sample size, which yields wide confidence intervals (they estimated the 95% CI for a 90% feasibility rate spanned 73.5% to 97.9%) and limits how precisely any of the reported proportions can be interpreted or generalized to other settings.