Comparison of text features election policies and using an adaptive framework


Taşci Ş., GÜNGÖR T.

Expert Systems with Applications, cilt.40, sa.12, ss.4871-4886, 2013 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 40 Sayı: 12
  • Basım Tarihi: 2013
  • Doi Numarası: 10.1016/j.eswa.2013.02.019
  • Dergi Adı: Expert Systems with Applications
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus
  • Sayfa Sayıları: ss.4871-4886
  • Anahtar Kelimeler: Adaptive keyword selection, Document categorization, Feature selection, Local and global policies, Support vector machines
  • Boğaziçi Üniversitesi Adresli: Evet

Özet

Text categorization is the task of automatically assigning unlabeled text documents to some predefined category labels by means of an induction algorithm. Since the data in text categorization are high-dimensional, often feature selection is used for reducing the dimensionality. In this paper, we make an evaluation and comparison of the feature selection policies used in text categorization by employing some of the popular feature selection metrics. For the experiments, we use datasets which vary in size, complexity, and skewness. We use support vector machine as the classifier and tf-idf weighting for weighting the terms. In addition to the evaluation of the policies, we propose new feature selection metrics which show high success rates especially with low number of keywords. These metrics are two-sided local metrics and are based on the difference of the distributions of a term in the documents belonging to a class and in the documents not belonging to that class. Moreover, we propose a keyword selection framework called adaptive keyword selection. It is based on selecting different number of terms for each class and it shows significant improvement on skewed datasets that have a limited number of training instances for some of the classes. © 2013 Elsevier Ltd. All rights reserved.