Archivio istituzionale della ricerca dell'Università degli Studi di Palermo

We propose a multi-dimensional structured state space (S4) approach to speech enhancement. To better capture the spectral dependencies across the frequency axis, we focus on modifying the multi-dimensional S4 layer with whitening transformation to build new small-footprint models that also achieve good performance. We explore several S4-based deep architectures in time (T) and time-frequency (TF) domains. The 2-D S4 layer can be considered a particular convolutional layer with an infinite receptive field although it utilizes fewer parameters than a conventional convolutional layer. Evaluated on the VoiceBank-DEMAND data set, when compared with the conventional U-net model based on convolutional layers, the proposed TF-domain S4-based model is 78.6% smaller in size, yet it still achieves competitive results with a PESQ score of 3.15 with data augmentation. By increasing the model size, we can even reach a PESQ score of 3.18.

Ku P.-J., Yang C.-H.H., Siniscalchi S.M., Lee C.-H. (2023). A Multi-dimensional Deep Structured State Space Approach to Speech Enhancement Using Small-footprint Models. In Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH 2023 (pp. 2453-2457). International Speech Communication Association [10.21437/Interspeech.2023-1084].

A Multi-dimensional Deep Structured State Space Approach to Speech Enhancement Using Small-footprint Models

Yang C. -H. H.;Siniscalchi S. M.^Supervision;Lee C. -H.

2023-01-01

Abstract

We propose a multi-dimensional structured state space (S4) approach to speech enhancement. To better capture the spectral dependencies across the frequency axis, we focus on modifying the multi-dimensional S4 layer with whitening transformation to build new small-footprint models that also achieve good performance. We explore several S4-based deep architectures in time (T) and time-frequency (TF) domains. The 2-D S4 layer can be considered a particular convolutional layer with an infinite receptive field although it utilizes fewer parameters than a conventional convolutional layer. Evaluated on the VoiceBank-DEMAND data set, when compared with the conventional U-net model based on convolutional layers, the proposed TF-domain S4-based model is 78.6% smaller in size, yet it still achieves competitive results with a PESQ score of 3.15 with data augmentation. By increasing the model size, we can even reach a PESQ score of 3.18.

Scheda breve

Scheda completa

Scheda completa (DC)

	Data
	
				2023
			
	DOI del contributo 
DATO PREVISTO SU LOGINMIUR
	
				https://dx.doi.org/10.21437/Interspeech.2023-1084
			
	URL dell'editore (Open access ove possibile)
	
				https://www.isca-archive.org/interspeech_2023/ku23_interspeech.html
			
	Citazione
	
				Ku P.-J.,  Yang C.-H.H.,  Siniscalchi S.M.,  Lee C.-H. (2023). A Multi-dimensional Deep Structured State Space Approach to Speech Enhancement Using Small-footprint Models. In Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH 2023 (pp. 2453-2457). International Speech Communication Association [10.21437/Interspeech.2023-1084].
			
	Appare nelle tipologie:
	
				2.07 Contributo in atti di convegno pubblicato in volume

File in questo prodotto:

File	Dimensione	Formato
ku23_interspeech.pdf Solo gestori archvio Descrizione: Il testo pieno dell’articolo è disponibile al seguente link: https://www.isca-archive.org/interspeech_2023/ku23_interspeech.html Tipologia: Versione Editoriale Dimensione 352.32 kB Formato Adobe PDF Visualizza/Apri Richiedi una copia	352.32 kB	Adobe PDF	Visualizza/Apri Richiedi una copia

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10447/637526

Citazioni

ND

5

3

social impact