Content Recognition and Context Modeling for Document Analysis and Retrieval

dc.contributor.advisorChellappa, Ramaen_US
dc.contributor.advisorDoermann, David Sen_US
dc.contributor.authorZhu, Guangyuen_US
dc.contributor.departmentElectrical Engineeringen_US
dc.contributor.publisherDigital Repository at the University of Marylanden_US
dc.contributor.publisherUniversity of Maryland (College Park, Md.)en_US
dc.date.accessioned2010-07-02T05:30:45Z
dc.date.available2010-07-02T05:30:45Z
dc.date.issued2009en_US
dc.description.abstractThe nature and scope of available documents are changing significantly in many areas of document analysis and retrieval as complex, heterogeneous collections become accessible to virtually everyone via the web. The increasing level of diversity presents a great challenge for document image content categorization, indexing, and retrieval. Meanwhile, the processing of documents with unconstrained layouts and complex formatting often requires effective leveraging of broad contextual knowledge. In this dissertation, we first present a novel approach for document image content categorization, using a lexicon of shape features. Each lexical word corresponds to a scale and rotation invariant local shape feature that is generic enough to be detected repeatably and is segmentation free. A concise, structurally indexed shape lexicon is learned by clustering and partitioning feature types through graph cuts. Our idea finds successful application in several challenging tasks, including content recognition of diverse web images and language identification on documents composed of mixed machine printed text and handwriting. Second, we address two fundamental problems in signature-based document image retrieval. Facing continually increasing volumes of documents, detecting and recognizing unique, evidentiary visual entities (\eg, signatures and logos) provides a practical and reliable supplement to the OCR recognition of printed text. We propose a novel multi-scale framework to detect and segment signatures jointly from document images, based on the structural saliency under a signature production model. We formulate the problem of signature retrieval in the unconstrained setting of geometry-invariant deformable shape matching and demonstrate state-of-the-art performance in signature matching and verification. Third, we present a model-based approach for extracting relevant named entities from unstructured documents. In a wide range of applications that require structured information from diverse, unstructured document images, processing OCR text does not give satisfactory results due to the absence of linguistic context. Our approach enables learning of inference rules collectively based on contextual information from both page layout and text features. Finally, we demonstrate the importance of mining general web user behavior data for improving document ranking and other web search experience. The context of web user activities reveals their preferences and intents, and we emphasize the analysis of individual user sessions for creating aggregate models. We introduce a novel algorithm for estimating web page and web site importance, and discuss its theoretical foundation based on an intentional surfer model. We demonstrate that our approach significantly improves large-scale document retrieval performance.en_US
dc.identifier.urihttp://hdl.handle.net/1903/10214
dc.subject.pqcontrolledComputer Scienceen_US
dc.subject.pqcontrolledInformation Technologyen_US
dc.subject.pquncontrolledClickRanken_US
dc.subject.pquncontrolledContent recognitionen_US
dc.subject.pquncontrolledContext modelingen_US
dc.subject.pquncontrolledDocument analysis and retrievalen_US
dc.subject.pquncontrolledLanguage identificationen_US
dc.subject.pquncontrolledSignature detection and matchingen_US
dc.titleContent Recognition and Context Modeling for Document Analysis and Retrievalen_US
dc.typeDissertationen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Zhu_umd_0117E_10987.pdf
Size:
30.54 MB
Format:
Adobe Portable Document Format