收稿日期: 2008-10-16
修回日期: 2009-04-21
网络出版日期: 2009-09-15
基金资助
国家高技术研究发展计划(863)(2006AA01Z454)、国家信息安全242计划(2005B23)和国家自然科学基金(60573015)资助
Fast dictionary mechanism for Chinese word segmentation
Received date: 2008-10-16
Revised date: 2009-04-21
Online published: 2009-09-15
通过研究目前中文分词领域各类分词机制,注意到中文快速分词机制的关键在于对单双字词的识别,在这一思想下,提出了一种快速中文分词机制:双字词-长词哈希机制,通过提高单双字词的查询效率来实现对中文分词机制的改进.实验证明,该机制提高了中文文本分词的效率.
关键词: 文本实时处理; 中文分词; 词典法分词; 双字词-长词哈希机制
吴晶晶 , 荆继武 , 聂晓峰 , 王平建 . 一种快速中文分词词典机制[J]. 中国科学院大学学报, 2009 , 26(5) : 703 -711 . DOI: 10.7523/j.issn.2095-6134.2009.5.017
With the development of global networking through Internet, the amount of articles in Chinese or other native languages is increasing rapidly. As the lack of explicit separator, word segmentation is a precondition for the processing of these character-based languages and thus it affects the whole system in performance. In this paper, we propose a new solution for Chinese word segmentation problem based on Lexicon named double-character-and-long-word-hash-indexing (DCLWHI).Compared with traditional lexicon mechanism, DCLWHI improves the speed and efficiency of word segmentation without extra memory spending and gains the same accuracy.
[1] Huang C N, Zhao H. Chinese word segmentation: A decade review
[J]. Journal of Chinese Information Processing,2007,21(3):8-19(in Chinese). 黄昌宁,赵 海. 中文分词十年回顾
[J].中文信息学报,2007,21(3):8-19.
[2] Asahara M,Goh C L, Wang X J,et al.Combining segmenter and chunker for Chinese word segmentation //Proceedings of Second SIGHAN Workshop on Chinese Language Processing.2003:144-147.
[3] Li M, Gao J F, Huang C N,et al.Unsupervised training for overlapping ambiguity resolution in Chinese word segmentation //Proceedings of Second SIGHAN Workshop on Chinese Language Processing.2003:1-7.
[4] Li J B, Zhou Q, Chen Z S. A study on fast algorithm for Chinese dictionary lookup
[J]. Journal of Chinese Information Processing,2006,20(5):31-39(in Chinese). 李江波,周 强,陈祖舜.汉语词典快速查询算法研究
[J].中文信息学报,2006,20(5):31-39.
[5] Sun M S, Zuo Z P, Huang C N. An experimental study on dictionary mechanism for Chinese word segmentation
[J]. Journal of Chinese Information Processing, 2000,14(1):1-6(in Chinese). 孙茂松,左正平,黄昌宁.汉语自动分词词典机制的实验研究
[J].中文信息学报,2000,14(1):1-6.
[6] Yang W F, Chen G Y, Li X. PATRICIA-tree based dictionary mechanism for Chinese word segmentation
[J]. Journal of Chinese Information Processing, 2001,15(3):44-49(in Chinese). 杨文峰,陈光英,李 星.基于PATRICIA tree自动分词词典机制
[J].中文信息学报,2001,15(3):44-49.
[7] Wang S L, Zhang H P, Wang B. Research of optimization on double-array trie and its application
[J]. Journal of Chinese Information Processing,2006,20(5):24-30(in Chinese). 王思力,张华平,王 斌.双数组Trie树算法优化及其应用研究
[J].中文信息学报,2006,20(5):24-30.
[8] Li Q H, Chen Y J, Sun J G. A new dictionary mechanism for Chinese word segmentation
[J]. Journal of Chinese Information Processing, 2003,17(4):13-18(in Chinese). 李庆虎,陈玉健,孙家广.中文分词词典新机制—双字哈希机制
[J].中文信息学报,2003,17(4):13-18.
[9] Choi A, Cheng C H, Ko Y L. Word extraction from Chinese documents by occurrence counts //International Conference on Computer Processing of Chinese and Oriental Languages.Toronto, Canada,1988:488-491.
[10] Honglan Jin, Kam-Fai.A Chinese dictionary construction algorithm for information retrieval //Proceedings of the ACM Transactions on Asian Language Information Processing(TALIP).2002(4):281-296.
[11] Information Technology Lab, NIST . . http://www.nist.gov/speech/tests/tdt/.
[12] TDT3 Multilanguage Text Corpus, Version 2.0; LDC Catalog Number LDC2001T58, isbn: 158563-193-0.
/
| 〈 |
|
〉 |