Unsupervised Feature Generation using Knowledge Repositories for Effective Text Categorization

Published: 01 Jan 2010, Last Modified: 18 Apr 2024ECAI 2010EveryoneRevisionsBibTeXCC BY-SA 4.0
Abstract: We propose an unsupervised feature generation algorithm using the repositories of human knowledge for effective text categorization. Conventional bag of words (BOW) depends on the presence / absence of keywords to classify the documents. To understand the actual context behind these keywords, we use knowledge concepts / hyperlinks from external knowledge sources through content and structure mining on Wikipedia. Then, the features of knowledge concepts are clustered to generate knowledge cluster vectors with which the input text documents are mapped into a high dimensional feature space and the classification is performed. The simulation results show that the proposed approach identifies associated features in the text collection and yields an improved classification accuracy.