Sogang University researchers win acceptance for two dysarthria studies at INTERSPEECH 2026
Translated from Korean and summarized by DistantNews. Read the original for the full story.
At a glance
- Two papers by Sogang Universityโs NLP and ISDS research team were accepted for oral presentation at INTERSPEECH 2026 in Sydney.
- One study uses cross-lingual retrieval to improve dysarthria severity assessment when clinical speech data are limited.
- A second pipeline uses speech-recognition tools and large language models to detect inappropriate pauses, improving classification performance when pause-related features are added.
A Sogang University research team will present two studies on AI-assisted dysarthria assessment at INTERSPEECH 2026, a leading international conference on speech and language processing.
One paper proposes a cross-lingual retrieval-augmented classification method for assessing dysarthria severity. The approach searches speech from another language with similar severity characteristics and uses those cases alongside the speech being analyzed, addressing the shortage of clinical voice data.
Experiments using Korean post-stroke dysarthria data and Italian dysarthria data from patients with amyotrophic lateral sclerosis recorded balanced accuracy of 87.3% for Korean data and 86.7% for Italian data. The results improved on methods that used only data from each language, by 8.4 and 20.0 percentage points respectively.
We expect that these methods for effectively using clinical data across different languages and automatically analyzing inappropriate pauses in patientsโ speech can expand into AI technologies that support the assessment and rehabilitation of speech and language disorders.
The second paper presents a pipeline for automatically identifying inappropriate pauses in dysarthric speech. It combines Whisper-based speech recognition with forced alignment and voice activity detection to locate pauses and measure their duration. A large language model then examines the surrounding speech context and provides an assessment of whether each pause is appropriate, along with its reasoning.
Adding the extracted pause features to dysarthria classification improved macro-accuracy by 8.4 percentage points and macro-F1 by 7.2 points. The team said the work could support more objective and explainable voice-based clinical assessment and rehabilitation tools. Both papers received government-backed support through Koreaโs Institute of Information and Communications Technology Planning and Evaluation and are scheduled for oral presentation in Sydney from September 27 to October 1, 2026.
We will continue researching reliable and explainable speech AI technologies that can be used in real clinical settings.
Originally published by Hankyoreh in Korean. Translated, summarized, and contextualized automatically by DistantNews, with a note on how the source frames the story. Not individually reviewed before publishing. How this works.