Language Access in Healthcare
EvidenceE-0387Initial AI draft

S-MINDS had lower speech-recognition word error rates than three commercial systems across quiet noisy and disfluent conditions

2026-06-053 out · 0 in

Source

Soller (2012). Performance of a new speech translation device in translating verbal recommendations of medication action plans for patients with diabetes. Journal of Diabetes Science and Technology.

Description #

In a controlled laboratory comparison of speech-recognition word error rate (WER) on common glycemic-control recommendations, the Fluential S-MINDS system had markedly lower WER than the three commercial systems across all three sound settings (Table 10). Fluential's WER was 0.74% in the clean setting, 6.40% in the noisy setting, and 13.54% in the disfluent setting, versus 10–12% (clean), 15–25% (noisy), and 27–30% (disfluent) for Dragon, Jibbigo, and Google. Differences from Fluential in the noisy setting were significant (p < .0001).

"The Fluential S-MINDS system outperformed the other three commercial systems for speech recognition in both the quiet–fluent and the noisy–fluent conditions as well as in the disfluency test (Table 10)." (Soller, 2012, p. 933)

Methods Context #

What? #

The observable: automatic speech recognition word error rate — the fraction of spoken words rendered incorrectly as on-screen text — computed per Jurafsky and Martin.

"Word error rate was calculated per Jurafsky and Martin." (Soller, 2012, p. 930)

"Automatic speech recognition word error rate is defined as researcher's speech appearing as text on the iPhone screen." (Soller, 2012, p. 933)

How? #

Four automatic speech recognition and three machine translation systems (Table 1) were tested simultaneously in fixed position under quiet–fluent, noisy–fluent, and quiet–disfluent conditions, with differences from Fluential tested by unpaired t-test.

"Test conditions using the same male native English speaker were (1) quiet environment, fluent speech (quiet–fluent); (2) noisy environment, fluent speech (noisy–fluent); (3) quiet environment, disfluent speech (quiet–disfluent)." (Soller, 2012, p. 930)

Who? #

The systems compared were Fluential S-MINDS versus Dragon Dictation, Jibbigo, and Google Translate iPhone apps, evaluated on common practitioner glycemic-control recommendations rather than on live patients.

"Four automatic speech recognition and three machine translation systems were compared independently and with Fluential's concept-based translation processing performance integrated with other systems' automatic speech recognition component (see Table 1 for system configurations)." (Soller, 2012, p. 929)

Other Notes #

This is a laboratory benchmark using one male native-English speaker reading utterances, not a patient-facing outcome (see qualifying caveat). Companion Table 11 reports the parallel translation-accuracy comparison.

Caveats #

  • The commercial-system comparison used one native-English speaker in a simulated laboratory not live bilingual clinical exchanges The head-to-head comparison of S-MINDS against Dragon, Jibbigo, and Google was a laboratory benchmark: a single male native-English speaker read pre-selected utterances into devices held in fixed position, and output was scored by a human observer. It did not involve real LEP patients, patient-side Spanish speech, accents/dialects, or live bidirectional counseling. The measured superiority therefore reflects controlled recognition/translation of one speaker's English rather than performance in the messier bilingual clinical exchanges the device is intended for.