Welcome to Journal of University of Chinese Academy of Sciences,Today is
Research Articles

Multi-channel voice activity detection in low signal-to-noise ratio environment

  • XIAO Si ,
  • GONG Jie ,
  • LI Baoqing
Expand
  • 1. School of Microelectronics, University of Chinese Academy of Sciences, Beijing 100049, China;
    2. Key Laboratory of Microsystem Technology, Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences, Shanghai 201800, China

Received date: 2021-12-16

  Revised date: 2022-02-08

  Online published: 2022-02-08

Abstract

Traditional voice activity detection algorithm only uses the time-frequency information, hence the detection accuracy will reduce rapidly in the low signal-to-noise environment, especially when the noise is non-stationary. Multi-channel speech signal has rich spatial information, which helps to improve the accuracy of detection as a supplement to time-frequency information. In this paper, on the basis of multi-channel spatial feature research, we propose a new multi-channel voice activity detection algorithm, by leveraging the maximum eigenvalue of the multi-channel covariance matrix (covariance matrix maximum eigenvalue, CMME) of the received array signals. First, we extract the CMME of the array signal as the feature of detection frame by frame, to track the speech signal. Then the double threshold method is adopted to determine whether the current frame is a speech frame. The results show that, compared with Mel energy ratio and the improved energy zero-entropy algorithm, the proposed algorithm has higher detection accuracy in VCTK and laboratory corpus, and thus is more robust in the low signal-to-noise ratio and non-stationary noise environment.

Cite this article

XIAO Si , GONG Jie , LI Baoqing . Multi-channel voice activity detection in low signal-to-noise ratio environment[J]. Journal of University of Chinese Academy of Sciences, 2023 , 40(5) : 687 -693 . DOI: 10.7523/j.ucas.2022.011

References

[1] Wang H K, Ye Z F, Chen J D. A speech enhancement system for automotive speech recognition with a hybrid voice activity detection method[C]//2018 16th International Workshop on Acoustic Signal Enhancement (IWAENC). September 17-20, 2018, Tokyo, Japan. IEEE, 2018:1-9. DOI:10.1109/IWAENC.2018.8521410.
[2] Bisio I, Garibotto C, Grattarola A, et al. Smart and robust speaker recognition for context-aware in-vehicle applications[J]. IEEE Transactions on Vehicular Technology, 2018, 67(9):8808-8821. DOI:10.1109/TVT.2018.2849577.
[3] G T Y, Vinay H C, Nayana T R, et al. Speech enhancement and encoding using SS-VAD and LPC[C]//2019 4th International Conference on Electrical, Electronics, Communication, Computer Technologies and Optimization Techniques (ICEECCOT). December 13-14, 2019, Mysuru, India. IEEE, 2019:151-157. DOI:10.1109/ICEECCOT46775.2019.9114541.
[4] Chen S H, Wu H T, Chen C H, et al. Robust voice activity detection algorithm based on the perceptual wavelet packet transform[C]//2005 International Symposium on Intelligent Signal Processing and Communication Systems. December 13-16, 2005, Hong Kong, China. IEEE, 2005:45-48. DOI:10.1109/ISPACS.2005.1595342.
[5] Haigh J A, Mason J S. Robust voice activity detection using cepstral features[C]//Proceedings of TENCON '93. IEEE Region 10 International Conference on Computers, Communications and Automation. October 19-21, 1993, Beijing, China. IEEE, 1993:321-324. DOI:10.1109/TENCON.1993.327987.
[6] 陈振锋,吴蔚澜,刘加,等. 基于Mel倒谱特征顺序统计滤波的语音端点检测算法[J].中国科学院大学学报,2014, 31(4):524-529. DOI:10.7523/j.issn.2095-6134.2014.04.012.
[7] Sohn J, Kim N S, Sung W. A statistical model-based voice activity detection[J]. IEEE Signal Processing Letters, 1999, 6(1):1-3. DOI:10.1109/97.736233.
[8] Davis A, Nordholm S, Togneri R. Statistical voice activity detection using low-variance spectrum estimation and an adaptive threshold[J]. IEEE Transactions on Audio, Speech, and Language Processing, 2006, 14(2):412-424. DOI:10.1109/TSA.2005.855842.
[9] Ghosh P K, Tsiartas A, Narayanan S. Robust voice activity detection using long-term signal variability[J]. IEEE Transactions on Audio, Speech, and Language Processing, 2011, 19(3):600-613. DOI:10.1109/TASL.2010.2052803.
[10] Ma Y N, Nishihara A. Efficient voice activity detection algorithm using long-term spectral flatness measure[J]. EURASIP Journal on Audio, Speech, and Music Processing, 2013, 2013:87. DOI:10.1186/1687-4722-2013-21.
[11] 张君昌, 张丹, 崔力. 一种鲁棒自适应阈值的语音端点检测方法[J]. 西安电子科技大学学报, 2015, 42(5):115-119. DOI:10.3969/j.issn.1001-2400.2015.05.020.
[12] 张涛, 刘阳, 任相赢. 基于长时信号功率谱变化的语音端点检测[J].计算机科学与探索, 2019, 13(9):1534-1542. DOI:10.3778/j.issn.1673-9418.1809029.
[13] Hoffman M W, Li Z, Khataniar D. GSC-based spatial voice activity detection for enhanced speech coding in the presence of competing speech[J]. IEEE Transactions on Speech and Audio Processing, 2001, 9(2):175-178. DOI:10.1109/89.902284.
[14] Huang S H, Park J, Chang J H. Dual-microphone voice activity detection based on using optimally weighted maximum a posteriori probabilities[C]//2016 IEEE International Conference on Acoustics, Speech and Signal Processing. March 20-25, 2016, Shanghai, China. IEEE, 2016:5360-5364. DOI:10.1109/ICASSP.2016.7472701.
[15] 赵益波, 蒋祎, 吴礼福, 等. 基于麦克风阵列自适应非线性滤波的语音信号端点检测方法[J]. 科技通报, 2017, 33(4):199-203. DOI:10.13774/j.cnki.kjtb.2017.04.045.
[16] Schwartz O, David A, Shahen-Tov O, et al. Multi-microphone voice activity and single-talk detectors based on steered-response power output entropy[C]//2018 IEEE International Conference on the Science of Electrical Engineering in Israel. December 12-14, 2018, Eilat, Israel. IEEE, 2018:1-4. DOI:10.1109/ICSEE.2018.8646089.
[17] 黄镇坤, 章小兵, 朱俞清. 低信噪比环境下改进的新能零熵语音端点检测[J]. 微电子学与计算机, 2020, 37(6):19-23, 29. DOI:10.19304/j.cnki.issn1000-7180.2020.06.004.
[18] 柏顺, 颜夕宏, 张生平, 等. 基于梅尔频率倒谱系数与短时能量的低信噪比语音端点检测[J]. 南京师大学报(自然科学版), 2021, 44(2):117-120. DOI:10.3969/j.issn.1001-4616.2021.02.016.
[19] Hegde R, Muralishankar R. Voice activity detection using novel teager energy based band spectral entropy[C]//2019 International Conference on Communication and Electronics Systems (ICCES). July 17-19, 2019, Coimbatore, India. IEEE, 2019:1272-1278. DOI:10.1109/ICCES45898.2019.9002565.
Outlines

/