Archivio istituzionale della ricerca dell'Università degli Studi di Palermo

We propose three techniques to improve mispronunciation detection of Mandarin tones of second language (L2) learners using tone-based extended recognition network (ERN). First, we extend our model from deep neural network (DNN) to bidirectionallon-short-term memory (BLSTM) in order to model tone-level co-articulation influenced by a broader temporal context (e.g., two or three consecutive Mandarin syllables). Second, we relax the hard labels to characterize the situations when a single tone class label is not enough because L2 learners' pronunciations are often between two canonical tone categories. Therefore, soft targets (a probabilistic transcription) are proposed for acoustic model training in place of conventional hard targets (one-hot targets). Third, we average tone scores produced by BLSTM models trained with hard and soft targets to seek the complementarity from modeling at the tone-target levels. Compared to our previous system based on the DNN-trained ERNs, the BLSTM-trained system with soft targets reduces the equal error rate (ERR) from 5.77% to 4.86%, and system combination decreases EER further to 4.34%, achieving a 24.78% relative error reduction.

Wei Li, Nancy F. Chen, Sabato Marco Siniscalchi, Chin-Hui Lee (2018). Improving Mandarin Tone Mispronunciation Detection for Non-Native Learners with Soft-Target Tone Labels and BLSTM-Based Deep Models. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 6249-6253). IEEE [10.1109/ICASSP.2018.8461629].

Improving Mandarin Tone Mispronunciation Detection for Non-Native Learners with Soft-Target Tone Labels and BLSTM-Based Deep Models

Wei Li;Nancy F. Chen;Sabato Marco Siniscalchi;Chin-Hui Lee

2018-01-01

Abstract

We propose three techniques to improve mispronunciation detection of Mandarin tones of second language (L2) learners using tone-based extended recognition network (ERN). First, we extend our model from deep neural network (DNN) to bidirectionallon-short-term memory (BLSTM) in order to model tone-level co-articulation influenced by a broader temporal context (e.g., two or three consecutive Mandarin syllables). Second, we relax the hard labels to characterize the situations when a single tone class label is not enough because L2 learners' pronunciations are often between two canonical tone categories. Therefore, soft targets (a probabilistic transcription) are proposed for acoustic model training in place of conventional hard targets (one-hot targets). Third, we average tone scores produced by BLSTM models trained with hard and soft targets to seek the complementarity from modeling at the tone-target levels. Compared to our previous system based on the DNN-trained ERNs, the BLSTM-trained system with soft targets reduces the equal error rate (ERR) from 5.77% to 4.86%, and system combination decreases EER further to 4.34%, achieving a 24.78% relative error reduction.

Scheda breve

Scheda completa

Scheda completa (DC)

	Data
	
				2018
			
	ISBN della monografia 
DATO PREVISTO SU LOGINMIUR
	
				978-1-5386-4658-8
			
	DOI del contributo 
DATO PREVISTO SU LOGINMIUR
	
				https://dx.doi.org/10.1109/ICASSP.2018.8461629
			
	Citazione
	
				Wei Li,  Nancy F. Chen,  Sabato Marco Siniscalchi,  Chin-Hui Lee (2018). Improving Mandarin Tone Mispronunciation Detection for Non-Native Learners with Soft-Target Tone Labels and BLSTM-Based Deep Models. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 6249-6253). IEEE [10.1109/ICASSP.2018.8461629].
			
	Appare nelle tipologie:
	
				2.07 Contributo in atti di convegno pubblicato in volume

File in questo prodotto:

File	Dimensione	Formato
0006249.pdf Solo gestori archvio Tipologia: Versione Editoriale Dimensione 618.96 kB Formato Adobe PDF Visualizza/Apri Richiedi una copia	618.96 kB	Adobe PDF	Visualizza/Apri Richiedi una copia

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10447/649573

Citazioni

ND

8

7

social impact