Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones
Di cosa parla
Si cerca di stimare una descrizione, chiamata Relative Transfer Matrix, che rappresenta come il suono di più sorgenti arriva a più microfoni, per aiutare a migliorare la voce in ambienti rumorosi. Finora questa stima sfruttava statistiche delle registrazioni multi-microfono; qui gli autori provano approcci basati su reti neurali supervisionate, proponendo tre architetture diverse — alcune analizzano il segnale nel tempo o nella frequenza e una mantiene memoria temporale. I test mostrano che questi modelli stimano la matrice in modo più accurato rispetto al metodo statistico e che, applicati al miglioramento della voce, ottengono risultati paragonabili al riferimento.
Cosa permette di osservare
Permette di esplorare se le reti neurali possono sostituire i metodi basati sulle statistiche delle registrazioni per ricostruire come il suono si trasferisce tra sorgenti e microfoni, e se ciò migliora la pulizia della voce in situazioni rumorose.
Dalla fonte
The Relative Transfer Matrix (ReTM), recently introduced as a generalization of the relative transfer function for multiple receivers and sources, shows promising performance when applied to speech enhancement in noisy environments. Estimating the ReTM of sound sources by exploiting the covariance matrices of multichannel recordings is highly beneficial for practical applications and, to date, remains the only proposed approach. This paper investigates deep learning-based ReTM estimation. We propose three novel supervised learning frameworks using time and short-time frequency transform domain convolutional networks, and a Long Short-Term Memory-based recurrent neural network. Experimental results demonstrate that the proposed models achieve more accurate estimation of the ReTM using five objective metrics compared to the covariance-based method. We also show the effectiveness of the pr…