Preprocessing of low-quality handwritten documents using Markov random fields

IEEE Trans Pattern Anal Mach Intell. 2009 Jul;31(7):1184-94. doi: 10.1109/TPAMI.2008.126.

Abstract

This paper presents a statistical approach to the preprocessing of degraded handwritten forms including the steps of binarization and form line removal. The degraded image is modeled by a Markov Random Field (MRF) where the hidden-layer prior probability is learned from a training set of high-quality binarized images and the observation probability density is learned on-the-fly from the gray-level histogram of the input image. We have modified the MRF model to drop the preprinted ruling lines from the image. We use the patch-based topology of the MRF and Belief Propagation (BP) for efficiency in processing. To further improve the processing speed, we prune unlikely solutions from the search space while solving the MRF. Experimental results show higher accuracy on two data sets of degraded handwritten images than previously used methods.

Publication types

  • Research Support, U.S. Gov't, Non-P.H.S.

MeSH terms

  • Algorithms*
  • Artificial Intelligence*
  • Electronic Data Processing / methods*
  • Handwriting*
  • Image Enhancement / methods
  • Image Interpretation, Computer-Assisted / methods*
  • Information Storage and Retrieval / methods*
  • Markov Chains
  • Models, Statistical
  • Pattern Recognition, Automated / methods*
  • Reading
  • Reproducibility of Results
  • Sensitivity and Specificity
  • Subtraction Technique