Text categorization with class-based and corpus-based keyword selection
20th International Symposium on Computer and Information Sciences, ISCIS 2005, İstanbul, Türkiye, 26 - 28 Ekim 2005, cilt.3733 LNCS, ss.606-615, (Tam Metin Bildiri)
- Yayın Türü: Bildiri / Tam Metin Bildiri
- Cilt numarası: 3733 LNCS
- Doi Numarası: 10.1007/11569596_63
- Basıldığı Şehir: İstanbul
- Basıldığı Ülke: Türkiye
- Sayfa Sayıları: ss.606-615
- Anahtar Kelimeler: Keyword selection, Reuters-21578, SVM, Text categorization
- Boğaziçi Üniversitesi Adresli: Evet
Özet
In this paper, we examine the use of keywords in text categorization with SVM. In contrast to the usual belief, we reveal that using keywords instead of all words yields better performance both in terms of accuracy and time. Unlike the previous studies that focus on keyword selection metrics, we compare the two approaches for keyword selection. In corpus-based approach, a single set of keywords is selected for all classes. In class-based approach, a distinct set of keywords is selected for each class. We perform the experiments with the standard Reuters-21578 dataset, with both boolean and tf-idf weighting. Our results show that although tf-idf weighting performs better, boolean weighting can be used where time and space resources are limited. Corpus-based approach with 2000 keywords performs the best. However, for small number of keywords, class-based approach outperforms the corpus-based approach with the same number of keywords. © Springer-Verlag Berlin Heidelberg 2005.