With the rapid developments of Internet technologies and popularization of Internet among daily activities, we are surrounded by all kinds of information every moment. Hence, to mine valuable information from massive data has always been a hotspot of research at home and abroad. In this environment, relationship extraction is an important subtask of information extraction, which purpose is to identify the relationship between entities from the text, so as to mine the structured information in the text, that is, fact triplet. In the text, entity overlapping and relationship overlapping are very common phenomena, but the existing joint extraction model cannot effectively solve such problems, so the paper proposes a new joint extraction model, which regards the relationship extraction task as consisting of entity recognition and relationship recognition of two subtasks. The two subtasks are identified using sequence labeling method and multi-classification method, respectively. In the joint extraction process, in order to fully mine the semantic information of the text, the part of speech (POS) and syntactic dependency (Deprel) features were added to the input layer of the model. Attention mechanism is also introduced in the model, which can eliminate the problem of long-distance dependence as sentence length increases. Finally, the paper conducts relationship extraction experiments on the NYT dataset and the WebNLG dataset. The experimental results show that the model proposed in the paper can effectively solve the problem of overlapping relationships and obtain the best extraction effect.
ZHAO Minjun
,
ZHAO Yawei
,
ZHAO Yajie
,
LUO Gang
. A new joint model for extracting overlapping relations based on deep learning[J]. Journal of University of Chinese Academy of Sciences, 2022
, 39(2)
: 240
-251
.
DOI: 10.7523/j.ucas.2020.0026
[1] Zelenko D, Aone C, Richardella A. Kernel methods for relation extraction[C]//Proceedings of the ACL-02 conference on Empirical methods in natural language processing-EMNLP’02. Not Known. Morristown, NJ, USA: Association for Computational Linguistics, 2002: 71-78. DOI:10.3115/1118693.1118703.
[2] Chan Y S, Roth D. Exploiting syntactico-semantic structures for relation extraction[C]//Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume 1. Portland, Oregon, 2011: 551-560.
[3] Li Q, Ji H. Incremental joint extraction of entity mentions and relations[C]//Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Baltimore, Maryland. Stroudsburg, PA,USA: Association for Computational Linguistics, 2014: 402-412. DOI:10.3115/v1/p14-1038.
[4] Yu X, Lam W. Jointly identifying entities and extracting relations in encyclopedia text via a graphical model approach[C]//Proceedings of the 23rd International Conference on Computational Linguistics: Posters. Association for Computational Linguistics. August 23-27,2010, Beijing, China. 2010: 1399-1407.
[5] Miwa M, Bansal M. End-to-end relation extraction using LSTMs on sequences and tree structures[C]// Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Berlin, Germany. Stroudsburg, PA, USA: Association for Computational Linguistics, 2016: 1105-1116. DOI:10.18653/v1/p16-1105.
[6] Dai D, Xiao X Y, Lyu Y J, et al. Joint extraction of entities and overlapping relations using position-attentive sequence labeling[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2019, 33:6300-6308. DOI:10.1609/aaai.v33i01.33016300.
[7] Zeng X R, Zeng D J, He S Z, et al. Extracting relational facts by an end-to-end neural model with copy mechanism[C]//Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Melbourne, Australia. Stroudsburg, PA, USA: Association for Computational Linguistics, 2018: 506-514. DOI:10.18653/v1/p18-1047.
[8] Collobert R, Weston J, Bottou L, et al. Natural language processing (almost) from scratch[J]. IEICE Transactions on Fundamentals of Electronics, Communications and Computer Sciences, 2011, abs/1103.0398: 2493-2537.
[9] Ma X Z, Hovy E. End-to-end sequence labeling via Bi-directional LSTM-CNNs-CRF[EB/OL]. 2016: arXiv: 1603.01354[cs.LG].(2016-03-04) [2020-05-20] https://arxiv.org/abs/1603.01354.
[10] Auer S, Bizer C, Kobilarov G, et al. DBpedia: a nucleus for a web of open data[C]//The Semantic Web. Springer, Berlin, Heidelberg, 2007: 722-735. DOI:10.1007/978-3-540-76298-0_52.
[11] Bollacker K, Evans C, Paritosh P, et al. Freebase: a collaboratively created graph database for structuring human knowledge[C]//SIGMOD ′08: Proceedings of the 2008 ACM SIGMOD international conference on Management of data. 2008: 1247-1250. DOI:10.1145/1376616.1376746.
[12] Nadeau D, Sekine S. A survey of named entity recognition and classification[J]. Lingvisticae Investigationes, 2007, 30(1): 3-26. DOI:10.1075/li.30.1.03nad.
[13] Rink B, Harabagiu S. Utd: classifying semantic relations by combining lexical and semantic resources[C]//Proceedings of the 5th International Workshop on Semantic Evaluation. Association for Computational Linguistics, 2010: 256-259.
[14] Luo G, Huang X J, Lin C Y, et al. Joint entity recognition and disambiguation[C]//Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Lisbon, Portugal. Stroudsburg, PA, USA: Association for Computational Linguistics, 2015: 879-888. DOI:10.18653/v1/d15-1104.
[15] Lample G, Ballesteros M, Subramanian S, et al. Neural architectures for named entity recognition[C]//Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. San Diego, California. Stroudsburg, PA, USA: Association for Computational Linguistics, 2016: 260-270. DOI:10.18653/v1/n16-1030.
[16] Huang Z H, Xu W, Yu K. Bidirectional LSTM-CRF models for sequence tagging[J]. arXiv:1508.01991,2015.
[17] Zheng S C, Wang F, Bao H Y, et al. Joint extraction of entities and relations based on a novel tagging scheme[C]//Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Vancouver, Canada. Stroudsburg, PA, USA: Association for Computational Linguistics, 2017: 1227-1236. DOI:10.18653/v1/p17-1113.
[18] Fu T J, Li P H, Ma W Y. GraphRel: modeling text as relational graphs for joint entity and relation extraction[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence, Italy. Stroudsburg, PA, USA: Association for Computational Linguistics, 2019: 1409-1418. DOI:10.18653/v1/p19-1136.
[19] Mikolov T, Chen K, Corrado G, et al. Efficient estimation of word representations in vector space[J]. arXiv:1301.3781,2013.
[20] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[C]//Advances in neural information processing systems. 2017: 5998-6008.
[21] Riedel S, Yao L M, McCallum A. Modeling relations and their mentions without labeled text[C]//Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer, Berlin, Heidelberg, 2010: 148-163. DOI:10.1007/978-3-642-15939-8_10.
[22] Gardent C, Shimorina A, Narayan S, et al. Creating training corpora for NLG micro-planners[C]//Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).Vancouver, Canada. Stroudsburg, PA, USA: Association for Computational Linguistics, 2017: 179-188. DOI:10.18653/v1/p17-1017.
[23] Hochreiter S, Schmidhuber J. Long short-term memory[J]. Neural Computation, 1997, 9(8):1735-1780. DOI:10.1162/neco.1997.9.8.1735.
[24] Kingma D P, Ba J. Adam: a method for stochastic optimization[J]. arXiv:1412.6980,2014.