Language Access in Healthcare
EvidenceE-0255Initial AI draft

Stratified by interview order only EQ-VAS with healthcare-interpreter-first differed significantly between methods (p 0.04)

2026-06-053 out · 0 in

Source

Xue (2019). Interpreter proxy versus healthcare interpreter for administration of patient surveys following arthroplasty: a pilot study. BMC Med Res Methodol.

Description #

When mean differences were stratified by interview order, the differences between the two methods remained non-significant for every measure except the EQ-VAS: when the first call was made with a healthcare interpreter, the healthcare-interpreter score was slightly higher, giving a significant order effect (mean difference −2.51, Wilcoxon p = 0.04; Table 4). For the EQ-VAS with proxy-first order the difference was non-significant (mean difference 3.73, p = 0.11). This is the single significant result in the study and is a marginal, order-dependent EQ-VAS effect rather than a general method difference.

"When results were stratified according the interview order, the differences between mean scores remained insignificant except for the EQ-VAS mean difference where if the first call was made with a healthcare interpreter this resulted in a slightly higher score (Table 4)." (Xue, 2019, pp. 4–5)

"EQ-VAS Proxy first 3.73 -25.20 to 32.60 0.11 / HC Interpreter first -2.51 -15.70 to 10.60 0.04" (Xue, 2019, p. 8, Table 4)

Methods Context #

What? #

The observable: the mean difference between methods for each measure, stratified by which interpreter type conducted the first interview, and its Wilcoxon significance.

"Table 4: Mean differences stratified by interview order" (Xue, 2019, p. 8)

How? #

The crossover data were split by randomised interview order (proxy first vs healthcare interpreter first) and each stratum's mean difference was tested with a Wilcoxon paired rank-sum test. See A-0009.

"In addition, a Wilcoxon paired ranked sum test was performed on each measure to assess the statistical significance of the differences obtained between the two methods of interview administration." (Xue, 2019, p. 3)

Who? #

85 LEP arthroplasty patients randomised to interpreter-proxy-first (n = 44 analysed) or healthcare-interpreter-first (n = 41 analysed) order.

"89 patients provided consent and were given a random allocation of call order (n = 46, interpreter proxy first; n = 43, health care interpreter first)." (Xue, 2019, p. 3)

Other Notes #

The authors read this marginal EQ-VAS order effect as reflecting normal week-to-week variation in self-rated health rather than a systematic method bias; they conclude patients were not continuing to improve between interviews.

Caveats #

  • The two arms differed in call structure (2-way proxy vs 3-way interpreter conference) confounding interpreter type with modality [Inferred: not flagged by the authors as a limitation.] The two administration methods differed not only in who interpreted but in the call structure itself: proxy interviews were 2-way calls (research officer + proxy relaying to/from the patient), whereas certified-interpreter interviews were 3-way conference calls connecting officer, interpreter, and patient. The observed agreement therefore conflates the interpreter-type effect with a call-modality effect, so agreement cannot be cleanly attributed to interpreter type alone.
  • The one-month window around the 6-month follow-up may have let true health change confound test-retest reliability Because the crossover interviews were spaced up to ~2 weeks apart and the first could fall anywhere within one month either side of the 6-month post-operative date, some of the disagreement attributed to interpreter method may actually be true week-to-week change in the patient's health. The authors acknowledge this window may have confounded the test-retest reliability, so the method-agreement estimates partly conflate interpreter effects with genuine health variation over time.