TY - GEN
T1 - Deep recurrent mixture of experts for speech enhancement
AU - Chazan, Shlomo E.
AU - Goldberger, Jacob
AU - Gannot, Sharon
N1 - Publisher Copyright: © 2017 IEEE.
PY - 2017/12/7
Y1 - 2017/12/7
N2 - Deep neural networks (DNNs) have recently became a viable methodology for single microphone speech enhancement. The most common approach, is to feed the noisy speech features into a fully-connected DNN to either directly enhance the speech signal or to infer a mask which can be used for the speech enhancement. In this case, one network has to deal with the large variability of the speech signal. Most approaches also discard the speech continuity. In this paper, we propose a deep recurrent mixture of experts (DRMoE) architecture that addresses these two issues. In order to reduce the large speech variability, we split the network into a mixture of networks (denoted experts), each of which specializes in a specific and simpler task and a gating network. The time-continuity of the speech signal is taken into account by implementing the experts and the gating network as a recurrent neural network (RNN). Experimental study shows that the proposed algorithm produces higher objective measurements scores compared to both a single RNN and a deep mixture of experts (DMoE) architectures.
AB - Deep neural networks (DNNs) have recently became a viable methodology for single microphone speech enhancement. The most common approach, is to feed the noisy speech features into a fully-connected DNN to either directly enhance the speech signal or to infer a mask which can be used for the speech enhancement. In this case, one network has to deal with the large variability of the speech signal. Most approaches also discard the speech continuity. In this paper, we propose a deep recurrent mixture of experts (DRMoE) architecture that addresses these two issues. In order to reduce the large speech variability, we split the network into a mixture of networks (denoted experts), each of which specializes in a specific and simpler task and a gating network. The time-continuity of the speech signal is taken into account by implementing the experts and the gating network as a recurrent neural network (RNN). Experimental study shows that the proposed algorithm produces higher objective measurements scores compared to both a single RNN and a deep mixture of experts (DMoE) architectures.
KW - long short-Term memory
KW - recurrent neural network
KW - speech presence probability
UR - https://www.scopus.com/pages/publications/85042366929
U2 - 10.1109/waspaa.2017.8170055
DO - 10.1109/waspaa.2017.8170055
M3 - Conference contribution
T3 - IEEE Workshop on Applications of Signal Processing to Audio and Acoustics
SP - 359
EP - 363
BT - 2017 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, WASPAA 2017
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2017 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, WASPAA 2017
Y2 - 15 October 2017 through 18 October 2017
ER -