Language Access in Healthcare
EvidenceE-0246Initial AI draft

Interpreter proxies and healthcare interpreters showed substantial-to-almost-perfect agreement across most arthroplasty PROMs

2026-06-054 out · 0 in

Source

Xue (2019). Interpreter proxy versus healthcare interpreter for administration of patient surveys following arthroplasty: a pilot study. BMC Med Res Methodol.

Description #

In a randomised crossover of 85 LEP arthroplasty patients each surveyed by both a family/carer interpreter proxy and a certified healthcare interpreter, inter-rater agreement between the two administration methods was at least substantial (agreement score > 0.60) for every patient-reported outcome measure except the EQ-5D anxiety/depression domain (0.57) (Table 3). Reporting of the exact agreement range is internally inconsistent: the abstract states "kappa = 0.69–0.87 and ICCs above 0.74," whereas the Results text states the kappa scores ranged 0.66–1 and the ICCs 0.66–0.87 (several ICCs, e.g. personal care 0.66 and usual activities 0.68, therefore fall below the abstract's "above 0.74").

"There was substantial to excellent inter-rater agreement (kappa = 0.69–0.87 and ICCs above 0.74) for all but one measure." (Xue, 2019, p. 1)

"Agreement between methods of interview was at least substantial (agreement score > 0.60) for all outcomes except for the anxiety/depression section of EQ-5D which scored 0.57 as shown in Table 3. The remainder of the kappa scores ranged from 0.66 to 1, ICCs ranged from 0.66 to 0.87 and CCCs from 0.66 to 0.87." (Xue, 2019, pp. 3–4)

Methods Context #

What? #

The observable: inter-rater agreement between the two interview-administration methods (interpreter proxy vs certified healthcare interpreter) on each patient-reported outcome measure, quantified by Cohen's kappa, ICC, and Lin's concordance correlation coefficient.

"The outcomes assessed were the levels of agreement between the two methods of language interpreting as determined by Cohen's kappa coefficients, Intraclass correlation coefficient (ICC) and Concordance correlation coefficient (CCC) statistics where appropriate." (Xue, 2019, p. 3)

How? #

Randomised crossover design: each participant was surveyed once by an interpreter proxy (2-way call) and once by a certified healthcare interpreter (3-way conference call) in randomised order within ~2 weeks; agreement between the paired scores was then computed. See A-0009 and A-0010.

"We used a randomised crossover study design to compare survey outcomes between two groups: surveys conducted using interpreter proxies (family members and carers), and those conducted using certified healthcare interpreters." (Xue, 2019, p. 2)

Who? #

85 LEP patients (of 89 consented) due for their routine 6-month post-arthroplasty telephone follow-up in the ACORN clinical quality registry, drawn from a multi-hospital elective hip/knee arthroplasty population in South Western Sydney.

"Eighty-five patients successfully received both methods of follow-up calls and were included in the data analysis." (Xue, 2019, p. 3)

Other Notes #

Agreement metrics were interpreted per Landis and Koch: 0.61–0.80 substantial, >0.80 almost perfect. This overall EVD is the headline for the paper; the per-measure findings are captured in separate EVDs.

Caveats #

  • The two arms differed in call structure (2-way proxy vs 3-way interpreter conference) confounding interpreter type with modality [Inferred: not flagged by the authors as a limitation.] The two administration methods differed not only in who interpreted but in the call structure itself: proxy interviews were 2-way calls (research officer + proxy relaying to/from the patient), whereas certified-interpreter interviews were 3-way conference calls connecting officer, interpreter, and patient. The observed agreement therefore conflates the interpreter-type effect with a call-modality effect, so agreement cannot be cleanly attributed to interpreter type alone.
  • English-version surveys (not language-specific translations) plus variable proxy linguistic skill may have affected accuracy The validated surveys were administered in their English versions and translated on the fly by the interpreter or proxy, rather than using language-specific validated translations. Because of this, variation in the linguistic skill of the (untrained) proxy interpreters may have affected accuracy — an uncontrolled source of measurement error that could inflate or deflate the observed agreement.
  • Per-language effects on proxy-interpreter agreement could not be assessed (insufficient n, many languages) The study could not assess whether the level of proxy-versus-interpreter agreement varied by the patient's specific language, because the sample (85 patients across at least 13 language groups) was too small and too linguistically fragmented for per-language analysis. Agreement may plausibly differ across languages (and the cultural framing of survey items), so the pooled agreement estimates could mask language-specific reliability problems.
  • Findings limited to arthroplasty follow-up and may not generalize to socially sensitive survey topics The reliability findings come only from routine, low-stigma arthroplasty outcome questions. The authors caution that they may not generalise to other settings, and in particular to socially sensitive topics, where using a family or carer as the interpreter proxy could compromise disclosure and thereby reduce the accuracy of proxy interpreting. The favourable agreement should therefore not be assumed to hold for sensitive clinical content.