Language Access in Healthcare
EvidenceE-0388Initial AI draft

S-MINDS scored higher translation accuracy than commercial speech translation systems across sound environments

2026-06-054 out · 0 in

Source

Soller (2012). Performance of a new speech translation device in translating verbal recommendations of medication action plans for patients with diabetes. Journal of Diabetes Science and Technology.

Description #

On a human-scored translation-quality scale (1 = good … 5 = no translation), the Fluential S-MINDS system scored highest (closest to 1.0) across all sound environments and was statistically significantly better than the stand-alone Google and Jibbigo systems (Table 11). Fluential's mean quality was 1.03 (quiet), 1.00 (noisy), and 1.24 (disfluent), versus 2.06–3.60 for Jibbigo and 2.21–3.70 for Google (all p ≤ .0001 vs Fluential). Layering Fluential's concept translation onto the other systems' speech recognition improved their scores relative to their stand-alone performance.

"Fluential was statistically significantly better than Google and Jibbigo (note Dragon is not a speech-translation system) and, when combined with the other commercial systems, improved them compared with their stand-alone scores. Fluential alone scored highest across all sound environments." (Soller, 2012, p. 934)

Methods Context #

What? #

The observable: translation quality of each conversational segment, scored on a five-point ordinal scale where a lower score is better.

"Nominal scores were converted to an ordinal quality score based on a five-point scale with good = 1, fair = 2, poor = 3, mistranslation = 4, or not translated = 5." (Soller, 2012, p. 930)

How? #

A human observer scored each system's translations using the method of Laws and colleagues, comparing stand-alone systems and hybrid configurations that added Fluential's concept translation to the other systems' automatic speech recognition (Table 1).

"The translations of each system were evaluated by a human observer using the scoring method of Laws and colleagues." (Soller, 2012, p. 933)

Who? #

The A-0020 compared against the Google Translate and Jibbigo commercial speech-translation iPhone apps (Dragon being speech-recognition only), on 102 quiet, 34 noisy, and 35 disfluent utterances read by one male native-English speaker.

"Four automatic speech recognition and three machine translation systems were compared independently and with Fluential's concept-based translation processing performance integrated with other systems' automatic speech recognition component (see Table 1 for system configurations)." (Soller, 2012, p. 929)

Other Notes #

Quality was judged by a University-based human observer scoring English-side output of simulated exchanges; the comparison does not involve real LEP patients or bidirectional live counseling (see qualifying caveat).

Caveats #

  • The commercial-system comparison used one native-English speaker in a simulated laboratory not live bilingual clinical exchanges The head-to-head comparison of S-MINDS against Dragon, Jibbigo, and Google was a laboratory benchmark: a single male native-English speaker read pre-selected utterances into devices held in fixed position, and output was scored by a human observer. It did not involve real LEP patients, patient-side Spanish speech, accents/dialects, or live bidirectional counseling. The measured superiority therefore reflects controlled recognition/translation of one speaker's English rather than performance in the messier bilingual clinical exchanges the device is intended for.