Classification of skewed and homogenous document corpora with class-based and corpus-based keywords
29th Annual German Conference on Artificial Intelligence (KI 2006), Bremen, Almanya, 14 - 17 Haziran 2006, cilt.4314 LNAI, ss.91-101, (Tam Metin Bildiri)
- Yayın Türü: Bildiri / Tam Metin Bildiri
- Cilt numarası: 4314 LNAI
- Doi Numarası: 10.1007/978-3-540-69912-5_8
- Basıldığı Şehir: Bremen
- Basıldığı Ülke: Almanya
- Sayfa Sayıları: ss.91-101
- Boğaziçi Üniversitesi Adresli: Evet
Özet
In this paper, we examine the performance of the two policies for keyword selection over standard document corpora of varying properdes. While in corpus-based policy a single set of keywords is selected for all classes globally, in class-based policy a distinct set of keywords is selected for each class locally. We use SVM as the learning method and perform experiments with boolean and tf-idf weighting. In contrast to the common belief, we show that using keywords instead of all words generally yields better performance and tf-idf weighting does not always outperform boolean weighting. Our results reveal that corpus-based approach performs better for large number of keywords while class-based approach performs better for small number of keywords. In skewed datasets, class-based keyword selection performs consistently better than corpus-based approach in terms of macro-averaged F-measure. In homogenous datasets, performances of class-based and corpus-based approaches are similar except for small number of keywords. © Springer-Verlag Berlin Heidelberg 2007.