欢迎访问中国科学院大学学报,今天是
论文

一种快速中文分词词典机制

  • 吴晶晶 ,
  • 荆继武 ,
  • 聂晓峰 ,
  • 王平建
展开
  • 1. 中国科学技术大学电子工程与信息科学系,合肥 230027;
    2. 中国科学院研究生院信息安全国家重点实验室, 北京100049

收稿日期: 2008-10-16

  修回日期: 2009-04-21

  网络出版日期: 2009-09-15

基金资助

国家高技术研究发展计划(863)(2006AA01Z454)、国家信息安全242计划(2005B23)和国家自然科学基金(60573015)资助 

Fast dictionary mechanism for Chinese word segmentation

  • WU Jing-Jing ,
  • JING Ji-Wu ,
  • NIE Xiao-Feng ,
  • Wang Ping-Jian
Expand
  • 1. Department of Electronic Engineering and Information Science, University of Science and Technology of China, Hefei 230027, China;
    2. State Key Laboratory of Information Security, Graduate University of the Chinese Academy of Sciences, Beijing 100049, China

Received date: 2008-10-16

  Revised date: 2009-04-21

  Online published: 2009-09-15

摘要

通过研究目前中文分词领域各类分词机制,注意到中文快速分词机制的关键在于对单双字词的识别,在这一思想下,提出了一种快速中文分词机制:双字词-长词哈希机制,通过提高单双字词的查询效率来实现对中文分词机制的改进.实验证明,该机制提高了中文文本分词的效率.

本文引用格式

吴晶晶 , 荆继武 , 聂晓峰 , 王平建 . 一种快速中文分词词典机制[J]. 中国科学院大学学报, 2009 , 26(5) : 703 -711 . DOI: 10.7523/j.issn.2095-6134.2009.5.017

Abstract

With the development of global networking through Internet, the amount of articles in Chinese or other native languages is increasing rapidly. As the lack of explicit separator, word segmentation is a precondition for the processing of these character-based languages and thus it affects the whole system in performance. In this paper, we propose a new solution for Chinese word segmentation problem based on Lexicon named double-character-and-long-word-hash-indexing (DCLWHI).Compared with traditional lexicon mechanism, DCLWHI improves the speed and efficiency of word segmentation without extra memory spending and gains the same accuracy.

参考文献


[1] Huang C N, Zhao H. Chinese word segmentation: A decade review
[J]. Journal of Chinese Information Processing,2007,21(3):8-19(in Chinese). 黄昌宁,赵 海. 中文分词十年回顾
[J].中文信息学报,2007,21(3):8-19.

[2] Asahara M,Goh C L, Wang X J,et al.Combining segmenter and chunker for Chinese word segmentation //Proceedings of Second SIGHAN Workshop on Chinese Language Processing.2003:144-147.

[3] Li M, Gao J F, Huang C N,et al.Unsupervised training for overlapping ambiguity resolution in Chinese word segmentation //Proceedings of Second SIGHAN Workshop on Chinese Language Processing.2003:1-7.

[4] Li J B, Zhou Q, Chen Z S. A study on fast algorithm for Chinese dictionary lookup
[J]. Journal of Chinese Information Processing,2006,20(5):31-39(in Chinese). 李江波,周 强,陈祖舜.汉语词典快速查询算法研究
[J].中文信息学报,2006,20(5):31-39.

[5] Sun M S, Zuo Z P, Huang C N. An experimental study on dictionary mechanism for Chinese word segmentation
[J]. Journal of Chinese Information Processing, 2000,14(1):1-6(in Chinese). 孙茂松,左正平,黄昌宁.汉语自动分词词典机制的实验研究
[J].中文信息学报,2000,14(1):1-6.

[6] Yang W F, Chen G Y, Li X. PATRICIA-tree based dictionary mechanism for Chinese word segmentation
[J]. Journal of Chinese Information Processing, 2001,15(3):44-49(in Chinese). 杨文峰,陈光英,李 星.基于PATRICIA tree自动分词词典机制
[J].中文信息学报,2001,15(3):44-49.

[7] Wang S L, Zhang H P, Wang B. Research of optimization on double-array trie and its application
[J]. Journal of Chinese Information Processing,2006,20(5):24-30(in Chinese). 王思力,张华平,王 斌.双数组Trie树算法优化及其应用研究
[J].中文信息学报,2006,20(5):24-30.

[8] Li Q H, Chen Y J, Sun J G. A new dictionary mechanism for Chinese word segmentation
[J]. Journal of Chinese Information Processing, 2003,17(4):13-18(in Chinese). 李庆虎,陈玉健,孙家广.中文分词词典新机制—双字哈希机制
[J].中文信息学报,2003,17(4):13-18.

[9] Choi A, Cheng C H, Ko Y L. Word extraction from Chinese documents by occurrence counts //International Conference on Computer Processing of Chinese and Oriental Languages.Toronto, Canada,1988:488-491.

[10] Honglan Jin, Kam-Fai.A Chinese dictionary construction algorithm for information retrieval //Proceedings of the ACM Transactions on Asian Language Information Processing(TALIP).2002(4):281-296.

[11] Information Technology Lab, NIST . . http://www.nist.gov/speech/tests/tdt/.

[12] TDT3 Multilanguage Text Corpus, Version 2.0; LDC Catalog Number LDC2001T58, isbn: 158563-193-0.

文章导航

/