欢迎访问中国科学院大学学报,今天是
信息与电子科学

低数据资源条件下基于Bottleneck特征与SGMM模型的语音识别系统

  • 吴蔚澜 ,
  • 蔡猛 ,
  • 田垚 ,
  • 杨晓昊 ,
  • 陈振锋 ,
  • 刘加 ,
  • 夏善红
展开
  • 1. 中国科学院大学, 北京 100190;
    2. 中国科学院电子学研究所 传感技术国家重点实验室, 北京 100190;
    3. 清华大学电子工程系 清华信息科学与技术国家实验室, 北京 100084

收稿日期: 2014-02-27

  修回日期: 2014-03-07

  网络出版日期: 2015-01-15

基金资助

国家自然科学基金(61005019,61273268,61370034,90920302)和北京市自然科学基金(KZ201110005005)资助

Bottleneck features and subspace Gaussian mixture models for low-resource speech recognition

  • WU Weilan ,
  • CAI Meng ,
  • TIAN Yao ,
  • YANG Xiaohao ,
  • CHEN Zhenfeng ,
  • LIU Jia ,
  • XIA Shanhong
Expand
  • 1. University of Chinese Academy of Sciences, Beijing 100190, China;
    2. State Key Laboratory of Transducer Technology, Institute of Electronics, Chinese Academy of Sciences, Beijing 100190, China;
    3. Tsinghua National Laboratory for Information Science and Technology, Department of Electronic Engineering, Tsinghua University, Beijing 100084, China

Received date: 2014-02-27

  Revised date: 2014-03-07

  Online published: 2015-01-15

摘要

语音识别系统需要大量有标注训练数据,在低数据资源条件下的识别性能往往不理想.针对数据匮乏问题,本文先研究子空间高斯混合声学模型通过参数共享减少待估计的参数规模,并使用基于最大互信息准则的区分型训练技术提高识别精度;而后在特征层面应用基于深度神经网络的Bottleneck特征来达到特征提取和降维的目的;最后将上述研究成果结合并构建了低资源条件下的语音识别系统.在国际标准的OpenKWS 2013数据库上的实验结果表明,本文的技术能够有效改善低资源条件下的系统识别性能,相比基线系统有12%左右的词错误率降低.

本文引用格式

吴蔚澜 , 蔡猛 , 田垚 , 杨晓昊 , 陈振锋 , 刘加 , 夏善红 . 低数据资源条件下基于Bottleneck特征与SGMM模型的语音识别系统[J]. 中国科学院大学学报, 2015 , 32(1) : 97 -102 . DOI: 10.7523/j.issn.2095-6134.2015.01.016

Abstract

State-of-the-art speech recognition systems often depend on a lot of training data, but perform poorly when limited data is available. In this paper, we study speech recognition systems under low-resource condition. The subspace Gaussian mixture (SGMM) model is first applied to reduce the number of parameters. The model is further enhanced by discriminative training based on maximum mutual information criterion. The bottleneck features based on deep neural networks are then studied to make robust feature extraction. The SGMM model and the bottleneck features are finally combined to produce a novel speech recognition system under low-resource condition. On the standard OpenKWS 2013 evaluation corpus, experimental results show the combination of the two technologies brings substantial relative improvement of about 12% over the baseline system.

参考文献

[1] Cui X, Xue J, Dognin P L, et al. Acoustic modeling with bootstrap and restructuring for low-resourced languages[C]//Interspeech. 2010: 2 974-2 977.

[2] Vu N T, Schlippe T, Kraus F, et al. Rapid bootstrapping of five eastern european languages using the rapid language adaptation toolkit[C]//Interspeech. 2010: 865-868.

[3] Rabiner L R. A Tutorial on hidden markov models and selected applications in speech recognition[J].Proceedings of IEEE, 1989, 77(2):257-286.

[4] Davis S, Mermelstein P. Comparison of parametric representations formonosyllable word recognition in continuously spoken sentences[J]. IEEE Transactions on Acoustics, Speech, and Signal Processing, 1980, 28(4):357-366.

[5] Povey D, Burget L, Agarwal M, et al. Subspace Gaussian mixture models for speech recognition[C]//Acoustics Speech and Signal Processing (ICASSP), 2010 IEEE International Conference on. IEEE, 2010: 4 330-4 333.

[6] Povey D, Burget L, Agarwal M, et al. The subspace Gaussian mixture model: a structured model for speech recognition[J]. Computer Speech and Language, 2011, 25(2):404-439.

[7] Dahl G, Yu D, Deng L, et al. Context-dependent pre-trained deep neural networks for large vocabulary speech recognition [J]. IEEE Trans on Audio, Speech and Language Processing, 2012, 20(1): 30-42.

[8] Seide F, Li G, Yu D. Conversational speech transcription using context-dependent deep neural networks//Interspeech. 2011: 437-440.

[9] Normandin Y. Hidden Markov models, maximum mutual information estimation, and the speech recognition problem . Canada: McGill University, 1991.

[10] He X D, Deng L, Chou W. Discriminative learning in sequential pattern recognition[J]. IEEE Signal Processing Magazine, 2008, 14(1):14-36.

[11] Yu D, Seltzer M L. Improved bottleneck features using pretrained deep neural networks[C]//INTERSPEECH. 2011: 237-240.

[12] IARPA. OpenKWS13 keyword search evaluation . (2013-01-25) . http://www.nist.gov/itl/iad/mig/upload/OpenKWS13.

[13] 单煜翔. 高效大词汇量连续语音识别解码算法研究与工程化实现[D]. 北京: 清华大学, 2012.

[14] Povey D, Ghoshal A, Boulianne G, et al. The Kaldi speech recognition toolkit[C]//Proc ASRU. 2011: 1-4.

[15] 钱彦旻. 低数据资源条件下的语音识别技术新方法研究[D]. 北京: 清华大学, 2013.

[16] Hinton G, Deng L, Yu D, et al. Deep neural networks for acoustic modeling in speech recognition: the shared views of four research groups[J]. IEEE, Signal Processing Magazine, 2012, 29(6): 82-97.

[17] Hinton G E, Osindero S, Teh Y W. A fast learning algorithm for deep belief nets[J]. Neural computation, 2006, 18(7): 1 527-1 554.

[18] Hinton G E, Salakhutdinov R R. Reducing the dimensionality of data with neural networks[J]. Science, 2006, 313(5786): 504-507.

[19] Fontaine V, Ris C, Boite J M. Nonlinear discriminant analysis for improved speech recognition[C]//Eurospeech. 1997.

[20] Grézl F, Karafiát M, Kontár S, et al. Probabilistic and bottle-neck features for LVCSR of meetings[C]//Proc ICASSP. 2007(4):757-761.

文章导航

/