In this paper, a new method for recognizing caption texts in videos is proposed. Due to varying font sizes, colors, styles, and resolutions and complex backgrounds in videos, it is still a challenging problem to recognize overlaid texts in videos. Most existing overlaid text recognition methods are based on the combination of text binarization and traditional OCR engine. However, the process of text binarization may incur noises and text stroke information loss. Additionally, techniques of traditional OCRs are mainly focused on high-resolution scans of printed documents, which have the characteristics of single color background, little noise, and more complete stroke information. Hence, traditional OCR engines might not be robust enough to recognize the binarization results of overlaid text images. In order to solve this problem, we directly extract Gabor features from overlaid text images without binarization for training the two-level character recognizer. The final experimental results demonstrate that the proposed method makes a great progress in overlaid Chinese text recognition with multiple fonts.
TIAN Jie
,
WANG Weiqiang
,
SUN Yi
. Recognition of overlaid Chinese characters in videos without binarization[J]. Journal of University of Chinese Academy of Sciences, 2018
, 35(3)
: 402
-408
.
DOI: 10.7523/j.issn.2095-6134.2018.03.015
[1] Yang H, Meinel C. Content based lecture video retrieval using speech and video text information[J]. IEEE Transactions on Learning Technologies, 2014, 7(2):142-154.
[2] Jung K, Kim K I, Jain A K. Text information extraction in images and video:a survey[J]. Pattern Recognition, 2004, 37(5):977-997.
[3] Wang Z, Yang L, Wu X, et al. A survey on video caption extraction technology[C]//Multimedia Information Networking and Security (MINES), 2012 Fourth International Conference on. IEEE, 2012:713-716.
[4] Yusufu T, Wang Y, Fang X. A video text detection and tracking system[C]//Multimedia (ISM), 2013 IEEE International Symposium on. IEEE, 2013:522-529.
[5] Lienhart R W, Stuber F. Automatic text recognition in digital videos[C]//Electronic Imaging:Science & Technology. International Society for Optics and Photonics, 1996:180-188.
[6] Niblack W. An introduction to digital image processing[M]. Birkeroed:Strandberg Publishing Company, 1985.
[7] Sauvola J, Pietikäinen M. Adaptive document image binarization[J]. Pattern Recognition, 2000, 33(2):225-236.
[8] Wolf C, Jolion J M, Chassaing F. Text localization, enhancement and binarization in multimedia documents[C]//Pattern Recognition, 2002. Proceedings. 16th International Conference on. IEEE, 2002, 2:1037-1040.
[9] Huang X, Ma H, Zhang H. A new video text extraction approach[C]//Multimedia and Expo, 2009. ICME 2009. IEEE International Conference on. IEEE, 2009:650-653.
[10] Kita K, Wakahara T. Binarization of color characters in scene images using k-means clustering and support vector machines[C]//Pattern Recognition (ICPR), 201020th International Conference on. IEEE, 2010:3183-3186.
[11] Peng X, Setlur S, Govindaraju V, et al. Markov random field based binarization for hand-held devices captured document images[C]//Proceedings of the Seventh Indian Conference on Computer Vision, Graphics and Image Processing. ACM, 2010:71-76.
[12] Wang Y, Shi C, Xiao B, et al. MRF based text binarization in complex images using stroke feature[C]//Document Analysis and Recognition (ICDAR), 201513th International Conference on. IEEE, 2015:821-825.
[13] Mancas-Thillou C, Gosselin B. Character segmentation-by-recognition using log-Gabor filters[C]//Pattern Recognition, 2006. ICPR 2006.18th International Conference on. IEEE, 2006, 2:901-904.
[14] Huang X. A novel video text extraction approach based on Log-Gabor filters[C]//Image and Signal Processing (CISP), 20114th International Congress on. IEEE, 2011, 1:474-478.
[15] Mishra A, Alahari K, Jawahar C V. Unsupervised refinement of color and stroke features for text binarization[J]. International Journal on Document Analysis and Recognition, 2017,20(2):1-17.
[16] Hao Q, Feng Z D, Ge Y. A study on the use of Gabor features for Chinese OCR[C]//Intelligent Multimedia, Video and Speech Processing, 2001. Proceedings of 2001 International Symposium on. IEEE, 2001:389-392.
[17] Wang X, Ding X, Liu C. Gabor filters-based feature extraction for character recognition[J]. Pattern recognition, 2005, 38(3):369-379.
[18] Jin X B, Liu C L, Hou X. Regularized margin-based conditional log-likelihood loss for prototype learning[J]. Pattern Recognition, 2010, 43(7):2428-2438.
[19] Smith R. An overview of the Tesseract OCR engine[C]//Document Analysis and Recognition, 2007. ICDAR 2007. Ninth International Conference on. IEEE, 2007, 2:629-633.
[20] Otsu N. A threshold selection method from gray-level histograms[J]. Automatica, 1975, 11(285-296):23-27.
[21] Zhang Z, Wang W, Lu K. Video text extraction using the fusion of color gradient and log-Gabor filter[C]//Pattern Recognition (ICPR), 201422nd International Conference on. IEEE, 2014:2938-2943.