Wake-up word spotting (WWS) for dysarthric speech remains challenging because it results in highly variable articulation, unstable timing, and irregular prosodic rhythm. Moreover, existing systems remain largely speaker-dependent, lacking adequate generalization to speaker-independent settings. Motivated by neurophysiological findings that intentional utterances induce preparatory motor activity that facilitates regular rhythmic patterns, we have developed a Rhythm-Aware Wake-up Word Spotting (RAWS) framework that explicitly leverages these cues for dysarthric speech. RAWS comprises three components: (1) a Temporal Prosody Structure Encoder (TPSE) that models speaking-rate and pause-durations via feature extraction, positional encoding, and transformer-based temporal processing; (2) a Large Language Model–Guided Auxiliary Learning (LLM-GAL) mechanism that provides perceptual rhythm-naturalness scores as auxiliary supervision; and (3) a Progressive Adapter-based Domain Alignment (PADA) strategy that enables effective non-dysarthric-to-dysarthric speech transfer while reducing cross-speaker variability. To the best of our knowledge, this study represents the first investigation of speaker-independent dysarthric WWS, with experimental validations on the Mandarin dysarthric speech corpus (MDSC) and its extended version (MDSC v2) demonstrating that RAWS outperforms strong competitive systems, achieving state-of-the-art results.
Gao, M., Chen, H., Du, J., Siniscalchi, S.M. (2026). Rhythm-Aware Modeling for Speaker-Independent Dysarthric Wake-Up Word Spotting. IEEE TRANSACTIONS ON AUDIO, SPEECH, AND LANGUAGE PROCESSING, 34, 3371-3383 [10.1109/TASLPRO.2026.3705688].
Rhythm-Aware Modeling for Speaker-Independent Dysarthric Wake-Up Word Spotting
Gao M.
Primo
Formal Analysis
;Siniscalchi S. M.Ultimo
Supervision
2026-01-01
Abstract
Wake-up word spotting (WWS) for dysarthric speech remains challenging because it results in highly variable articulation, unstable timing, and irregular prosodic rhythm. Moreover, existing systems remain largely speaker-dependent, lacking adequate generalization to speaker-independent settings. Motivated by neurophysiological findings that intentional utterances induce preparatory motor activity that facilitates regular rhythmic patterns, we have developed a Rhythm-Aware Wake-up Word Spotting (RAWS) framework that explicitly leverages these cues for dysarthric speech. RAWS comprises three components: (1) a Temporal Prosody Structure Encoder (TPSE) that models speaking-rate and pause-durations via feature extraction, positional encoding, and transformer-based temporal processing; (2) a Large Language Model–Guided Auxiliary Learning (LLM-GAL) mechanism that provides perceptual rhythm-naturalness scores as auxiliary supervision; and (3) a Progressive Adapter-based Domain Alignment (PADA) strategy that enables effective non-dysarthric-to-dysarthric speech transfer while reducing cross-speaker variability. To the best of our knowledge, this study represents the first investigation of speaker-independent dysarthric WWS, with experimental validations on the Mandarin dysarthric speech corpus (MDSC) and its extended version (MDSC v2) demonstrating that RAWS outperforms strong competitive systems, achieving state-of-the-art results.| File | Dimensione | Formato | |
|---|---|---|---|
|
Rhythm-Aware_Modeling_for_Speaker-Independent_Dysarthric_Wake-Up_Word_Spotting.pdf
Solo gestori archvio
Descrizione: Main Paper
Tipologia:
Pre-print
Dimensione
4.56 MB
Formato
Adobe PDF
|
4.56 MB | Adobe PDF | Visualizza/Apri Richiedi una copia |
|
Rhythm-Aware_Modeling_for_Speaker-Independent_Dysarthric_Wake-Up_Word_Spotting.pdf
Solo gestori archvio
Tipologia:
Versione Editoriale
Dimensione
2.13 MB
Formato
Adobe PDF
|
2.13 MB | Adobe PDF | Visualizza/Apri Richiedi una copia |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


