Mean differences between proxy and interpreter scores were small with negligible group-level bias and no significant Wilcoxon differences
Source
Xue (2019). Interpreter proxy versus healthcare interpreter for administration of patient surveys following arthroplasty: a pilot study. BMC Med Res Methodol.
Description #

At the group level the two administration methods showed negligible systematic bias: across the Bland-Altman analyses the mean difference (proxy score minus healthcare-interpreter score) was very small — up to only 1.1% of total score for the Oxford — and Wilcoxon paired rank-sum tests found no statistically significant difference between the mean scores from the two methods for any measure (all p > 0.05; Table 3, Figs. 5–6). Mean differences ranged from 0.01 (on the 1–5 satisfaction scale) to 0.72 (on the 0–100 EQ-VAS).
"Overall the mean difference is very small (up to 1.1% of total score for the Oxford) indicating negligible bias when all subjects are considered." (Xue, 2019, p. 4)
"The Wilcoxon rank sum tests indicated that the differences between mean scores from either method of interview were not statistically significant." (Xue, 2019, p. 4)
Methods Context #
What? #ⓘ
The observable: the mean difference (systematic bias) between paired proxy and interpreter scores, and its statistical significance, at the group level.
"The mean of the differences between the same data items collected by each of the two methods was also calculated." (Xue, 2019, p. 1)
How? #ⓘ
Score differences were visualised with Bland-Altman plots (mean difference ± 95% limits of agreement) and tested for a systematic difference with a Wilcoxon paired rank-sum test on each measure. See A-0009ArtifactA-0009Initial AI draftInterpreter proxy survey administration (family or carer, 2-way telephone)A pragmatic, low-cost method for administering patient-reported outcome surveys by telephone to LEP patients, using the patient's own family member or carer as the interpreter proxy in place of a certified professional h….
"In addition, a Wilcoxon paired ranked sum test was performed on each measure to assess the statistical significance of the differences obtained between the two methods of interview administration." (Xue, 2019, p. 3)
Who? #ⓘ
85 LEP hip or knee arthroplasty patients at their routine 6-month registry follow-up.
"Eighty-five patients successfully received both methods of follow-up calls and were included in the data analysis." (Xue, 2019, p. 3)
Other Notes #
The authors distinguish group-level equivalence (negligible mean bias) from individual-level disagreement, attributing the latter to normal week-to-week variation in survey responses rather than the data-collection method.
Caveats #
- The one-month window around the 6-month follow-up may have let true health change confound test-retest reliability Because the crossover interviews were spaced up to ~2 weeks apart and the first could fall anywhere within one month either side of the 6-month post-operative date, some of the disagreement attributed to interpreter method may actually be true week-to-week change in the patient's health. The authors acknowledge this window may have confounded the test-retest reliability, so the method-agreement estimates partly conflate interpreter effects with genuine health variation over time.
- Findings limited to arthroplasty follow-up and may not generalize to socially sensitive survey topics The reliability findings come only from routine, low-stigma arthroplasty outcome questions. The authors caution that they may not generalise to other settings, and in particular to socially sensitive topics, where using a family or carer as the interpreter proxy could compromise disclosure and thereby reduce the accuracy of proxy interpreting. The favourable agreement should therefore not be assumed to hold for sensitive clinical content.