Two-stage feature selection for text classification


Özgür L., GÜNGÖR T.

30th International Symposium on Computer and Information Sciences, ISCIS 2015, London, İngiltere, 21 - 24 Eylül 2015, cilt.363, ss.329-337, (Tam Metin Bildiri)

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Cilt numarası: 363
  • Doi Numarası: 10.1007/978-3-319-22635-4_30
  • Basıldığı Şehir: London
  • Basıldığı Ülke: İngiltere
  • Sayfa Sayıları: ss.329-337
  • Boğaziçi Üniversitesi Adresli: Evet

Özet

In this paper, we focus on feature coverage policies used for feature selection in the text classification domain. Two alternative policies are discussed and compared: corpus-based and class-based selection of features. We make a detailed analysis of pruning and keyword selection by varying the parameters of the policies and obtain the optimal usage patterns. In addition, by combining the optimal forms of these methods, we propose a novel two-stage feature selection approach. The experiments on three independent datasets showed that the proposed method results in a statistically significant increase over the traditional methods in the success rates of the classifier.