In this paper, we propose a domain adaptation framework to address the device mismatch issue in acoustic scene classification leveraging upon neural label embedding (NLE) and relational teacher student learning (RTSL). Taking into account the structural relationships between acoustic scene classes, our proposed framework captures such relationships which are intrinsically device-independent. In the training stage, transferable knowledge is condensed in NLE from the source domain. Next in the adaptation stage, a novel RTSL strategy is adopted to learn adapted target models without using paired source-target data often required in conventional teacher student learning. The proposed framework is evaluated on the DCASE 2018 Task1b data set. Experimental results based on AlexNet-L deep classification models confirm the effectiveness of our proposed approach for mismatch situations. NLE-alone adaptation compares favourably with the conventional device adaptation and teacher student based adaptation techniques. NLE with RTSL further improves the classification accuracy

Hu, H.u., Siniscalchi, S.M., Wang, Y., Lee, C. (2020). Relational Teacher Student Learning with Neural Label Embedding for Device Adaptation in Acoustic Scene Classification. In 1st Annual Conference of the International Speech Communication Association (pp. 1196-1200) [10.21437/Interspeech.2020-2038].

Relational Teacher Student Learning with Neural Label Embedding for Device Adaptation in Acoustic Scene Classification

Siniscalchi, Sabato Marco
Supervision
;
2020-01-01

Abstract

In this paper, we propose a domain adaptation framework to address the device mismatch issue in acoustic scene classification leveraging upon neural label embedding (NLE) and relational teacher student learning (RTSL). Taking into account the structural relationships between acoustic scene classes, our proposed framework captures such relationships which are intrinsically device-independent. In the training stage, transferable knowledge is condensed in NLE from the source domain. Next in the adaptation stage, a novel RTSL strategy is adopted to learn adapted target models without using paired source-target data often required in conventional teacher student learning. The proposed framework is evaluated on the DCASE 2018 Task1b data set. Experimental results based on AlexNet-L deep classification models confirm the effectiveness of our proposed approach for mismatch situations. NLE-alone adaptation compares favourably with the conventional device adaptation and teacher student based adaptation techniques. NLE with RTSL further improves the classification accuracy
2020
Settore ING-INF/05 - Sistemi Di Elaborazione Delle Informazioni
Hu, H.u., Siniscalchi, S.M., Wang, Y., Lee, C. (2020). Relational Teacher Student Learning with Neural Label Embedding for Device Adaptation in Acoustic Scene Classification. In 1st Annual Conference of the International Speech Communication Association (pp. 1196-1200) [10.21437/Interspeech.2020-2038].
File in questo prodotto:
File Dimensione Formato  
2038.pdf

Solo gestori archvio

Descrizione: Il testo pieno dell’articolo è disponibile al seguente link: https://www.isca-archive.org/interspeech_2020/hu20e_interspeech.html
Tipologia: Versione Editoriale
Dimensione 366.64 kB
Formato Adobe PDF
366.64 kB Adobe PDF   Visualizza/Apri   Richiedi una copia

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10447/636621
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 5
  • ???jsp.display-item.citation.isi??? 3
social impact