In voice-based translation tools the speech-recognition stage, not the translation stage, is the primary source of failure
Narrative synthesis #
In voice-based translation tools the speech-recognition stage, not the translation stage, is the primary failure point: two qualitative studies independently name background noise, dialect, and accent as what breaks (Panayiotou, Hwang), and two evaluations quantify the same axis — S-MINDS had lower word error rates than commercial systems across quiet, noisy, and disfluent conditions, and VERAA's mapping accuracy fell for Spanish and for the on-device pipeline. The practical content is a design trade-off: users prefer fixed-phrase apps because they avoid the ASR stage, but fixed-phrase apps cannot carry the patient's reply back, leaving communication one-directional — reliability against bidirectionality.
Soller's favourable result is for the vendor's own system (S-MINDS) against commercial comparators, so treat that direction as suggestive rather than settled. Merge candidates (proposal only — do not touch): C-0212ClaimC-0212Initial AI draftFree real-time machine translation apps are unreliable for clinical communication being slow and inaccurate with accents dialects and noiseFree real-time machine translation apps (e.g. Google Translate) are unreliable for clinical communication: their free-text voice translation is slow and difficult to use and its accuracy is degraded by patients' accents,…, C-0219ClaimC-0219Initial AI draftFixed-phrase translation apps are preferred over real-time voice-to-voice translation apps for healthcare communicationFor healthcare communication, users prefer fixed-phrase (preset, curated) translation apps over open, real-time voice-to-voice translation apps, because the preset apps are easier to use and avoid the accuracy failures o…, C-0211ClaimC-0211Initial AI draftFixed-phrase translation apps cannot convey patients' spoken responses leaving communication one-directionalFixed-phrase (phrasebook) translation apps translate in only one direction and cannot render a patient's spoken response, leaving communication one-directional and perpetuating a communication barrier. The limitation is…, C-0228ClaimC-0228Initial AI draftA concept-based domain-tuned speech translation system is more robust to noise and disfluency than general-purpose commercial systemsA concept-based, domain-tuned speech translation system is more robust to background noise, speech disfluency, and rapid speech — in both speech-recognition word error rate and translation accuracy — than general-purpose….