Language Access in Healthcare
EvidenceE-0132Initial AI draft

Chatbot-enrolled LEP patients had a non-significant reduction in 90-day ED visits vs controls (0.9% vs 8.0%, P=.085)

2026-06-054 out · 0 in

Source

Joshua P Rainey (2023). A Multilingual Chatbot Can Effectively Engage Arthroplasty Patients Who Have Limited English Proficiency. Journal of Arthroplasty.

Description #

LEP TJA patients enrolled in the multilingual SMS chatbot had a lower rate of emergency department (ED) visits within 90 days than the non-enrolled historical LEP control (0.9% versus 8.0%), but the difference did not reach statistical significance (P = .085) (Fig. 4). The authors describe this as a "near significant reduction"; the point estimate favors the chatbot but the result is, by the conventional threshold, null.

"The LEP patients who were enrolled in the chatbot had fewer readmissions (0% versus 8.3%, P = .013) and a near significant reduction in ED visits (0.9% versus 8.0%, P = .085) compared to the historical control LEP patients who did not enroll in the chatbot, Figure 4." (Rainey, 2023, p. S79)

"Outcomes within 90 days of Surgery in terms of Emergency Department (ED) visits and Readmissions in LEP patients before and after the chatbot service." (Rainey, 2023, p. S81, Fig. 4)

Methods Context #

What? #

The observable: emergency department visits within 90 days of surgery, ascertained from scheduled postoperative telephone check-ins and clinic follow-up visits documented in the electronic health record.

"Independent t-tests, Fisher's exact tests, and chi-squared tests were performed to measure the effect that conversational engagement had on ED visits within 90 days, hospital readmissions within 90 days, and reoperations." (Rainey, 2023, p. S79)

How? #

Retrospective cohort comparison of LEP patients enrolled in the A-0004 against a historical control of non-enrolled LEP patients, with rates compared by Fisher's exact / chi-squared tests.

"a retrospective review was performed on all patients who underwent total hip or total knee arthroplasty who were also enrolled in a SMS chatbot from 2020 to 2022 at a single institution." (Rainey, 2023, p. S79)

Who? #

47 LEP TJA patients enrolled in the chatbot versus a historical control of 68 LEP TJA patients (surgery 2018-2019) not enrolled, at a single academic center.

"A historical control of 68 patients with LEP who did not enroll in the chatbot and underwent a TJA between 2018 and 2019 were identified for comparison purposes." (Rainey, 2023, p. S79)

Other Notes #

This result did not reach the conventional significance threshold (P = .085) and is therefore wired as contradicting the claim that the chatbot reduces ED visits, per the corpus convention that a non-significant result opposes the claim it bears on. The description preserves the authors' framing ("near significant reduction") and the direction of the point estimate.

Caveats #

  • Historical non-enrolled control differing on Medicaid and era, with near-zero event counts, confounds the chatbot-outcome comparison [Inferred:] The chatbot-versus-control outcome comparisons are confounded by the comparison design rather than by chatbot exposure alone. The control was a historical cohort operated on in 2018-2019, whereas the chatbot cohort was operated on in 2020-2022 — so the intervention is entangled with secular/era effects (including COVID-19-era changes in postoperative care and ED/readmission thresholds). The two groups were also not balanced: chatbot-enrolled patients were significantly more likely to have Medicaid (36.2% vs 16.2%, P = .0141). Finally, the favorable outcome rates rest on near-zero event counts (0/47 readmissions, 0/47 ED visits reported as 0.9%, 0/47 reoperations in the chatbot arm), making the rate estimates and P values fragile and imprecise. Together these features warrant treating the readmission benefit as association, not established causal effect.
  • Retrospective single-center chatbot study with only 90-day follow-up The study is a retrospective review at a single academic center, so its patient population may not represent other US regions, and it captured only short-term (90-day / at-least-3-month) outcomes — leaving long-term effects of chatbot engagement unassessed. The authors state that multicenter and prospective investigation, plus study of outcomes beyond 3 months, are needed to validate the findings. These features limit generalizability and preclude causal or durable-effect inference from the outcome comparisons.