arXiv:1901.00295 Abstract | arXiv Analytics

arXiv:1901.00295 [cs.SD]Abstract References Reviews Resources

End-to-End Model for Speech Enhancement by Consistent Spectrogram Masking

Xingjian Du, Mengyao Zhu, Xuan Shi, Xinpeng Zhang, Wen Zhang, Jingdong Chen

Published 2019-01-02Version 1

Recently, phase processing is attracting increasinginterest in speech enhancement community. Some researchersintegrate phase estimations module into speech enhancementmodels by using complex-valued short-time Fourier transform(STFT) spectrogram based training targets, e.g. Complex RatioMask (cRM) [1]. However, masking on spectrogram would violentits consistency constraints. In this work, we prove that theinconsistent problem enlarges the solution space of the speechenhancement model and causes unintended artifacts. ConsistencySpectrogram Masking (CSM) is proposed to estimate the complexspectrogram of a signal with the consistency constraint in asimple but not trivial way. The experiments comparing ourCSM based end-to-end model with other methods are conductedto confirm that the CSM accelerate the model training andhave significant improvements in speech quality. From ourexperimental results, we assured that our method could enha

Categories: cs.SD, cs.AI, cs.MM, eess.AS

Keywords: end-to-end model, consistent spectrogram masking, speech enhancement, researchersintegrate phase estimations module, model training andhave significant improvements

Related articles: Most relevant | Search more

arXiv:2405.06573 [cs.SD] (Published 2024-05-10)

An Investigation of Incorporating Mamba for Speech Enhancement

Rong Chao, Wen-Huang Cheng, Moreno La Quatra, Sabato Marco Siniscalchi, Chao-Han Huck Yang, Szu-Wei Fu, Yu Tsao

arXiv:1902.03926 [cs.SD] (Published 2019-02-08)

Speech enhancement with variational autoencoders and alpha-stable distributions

Simon Leglaive, Umut Simsekli, Antoine Liutkus, Laurent Girin, Radu Horaud

arXiv:2305.08292 [cs.SD] (Published 2023-05-15)

ForkNet: Simultaneous Time and Time-Frequency Domain Modeling for Speech Enhancement

Feng Dang, Qi Hu, Pengyuan Zhang, Yonghong Yan