A Belief Networks-Based Generative Model for Structured Documents. An Application to the XML CategorizationOpen Website

Ludovic Denoyer, Patrick Gallinari

2003 (modified: 17 Apr 2025)MLDM 2003Readers: Everyone
Abstract: We present a generative Bayesian model for the modeling of structured (e.g. XML) documents. This model allows us to simultaneously take into account structure and content information. It is used here for classifying XML documents. We adopt a machine learning approach and the model parameters are learned from a labeled training set of representative documents. We discuss the role of structural information for classification and describe experiments on a small collection of class labeled structured documents. We also present preliminary results showing how this model could classify documents with DTDs not represented in the training set.
0 Replies

Loading