Language Access in Healthcare
ClaimC-0240Initial AI draft

In voice-based translation tools the speech-recognition stage, not the translation stage, is the primary source of failure

2026-06-052 out · 10 in

Narrative synthesis #

In voice-based translation tools the speech-recognition stage, not the translation stage, is the primary failure point: two qualitative studies independently name background noise, dialect, and accent as what breaks (Panayiotou, Hwang), and two evaluations quantify the same axis — S-MINDS had lower word error rates than commercial systems across quiet, noisy, and disfluent conditions, and VERAA's mapping accuracy fell for Spanish and for the on-device pipeline. The practical content is a design trade-off: users prefer fixed-phrase apps because they avoid the ASR stage, but fixed-phrase apps cannot carry the patient's reply back, leaving communication one-directional — reliability against bidirectionality.

Soller's favourable result is for the vendor's own system (S-MINDS) against commercial comparators, so treat that direction as suggestive rather than settled. Merge candidates (proposal only — do not touch): C-0212, C-0219, C-0211, C-0228.