As We May Perceive: Finding the Boundaries of Compound Documents on the Web
- Pavel Dmitriev(Cornell University)
This paper considers the problem of identifying on the Web compound documents (cDocs) - groups of web pages that in aggregate constitute semantically coherent information entities. Examples of cDocs are a news article consisting of several html pages, or a set of pages describing specifications, price, and reviews of a digital camera. Being able to identify cDocs would be useful in many applications including web and intranet search, user navigation, automated collection generation, and information extraction. In the past, several heuristic approaches have been proposed to identify cDocs . However, heuristics fail to capture the variety of types, styles and goals of information on the web, and do not account for the fact that the definition of a cDoc often depends on the context. This paper presents an experimental evaluation of three machine learning-based algorithms for cDoc discovery. These algorithms are responsive to the varying structure of cDocs and adaptive to their application-specific nature. Based on our previous work , this paper proposes a different scenario for discovering cDocs, and compares in this new setting the local machine learned clustering algorithm from  to a global purely graph based approach  and a Conditional Markov Network approach previously applied to noun coreference task . The results show that the approach of  outperforms the other algorithms, suggesting that global relational characteristics of web sites are too noisy for cDoc identification purposes.
Inquiries can be sent to: