Optimization of dependency and pruning usage in text classification


Özgür L., GÜNGÖR T.

Pattern Analysis and Applications, cilt.15, sa.1, ss.45-58, 2012 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 15 Sayı: 1
  • Basım Tarihi: 2012
  • Doi Numarası: 10.1007/s10044-010-0195-5
  • Dergi Adı: Pattern Analysis and Applications
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus
  • Sayfa Sayıları: ss.45-58
  • Anahtar Kelimeler: Lexical dependency, Pruning analysis, Stanford parser, Text classification
  • Boğaziçi Üniversitesi Adresli: Evet

Özet

In this study, a comprehensive analysis of the lexical dependency and pruning concepts for the text classification problem is presented. Dependencies are included in the feature vector as an extension to the standard bag-of-words approach. The pruning process filters features with low frequencies so that fewer but more informative features remain in the solution vector. The pruning levels for words, dependencies, and dependency combinations for different datasets are analyzed in detail. The main motivation in this work is to make use of dependencies and pruning efficiently in text classification and to achieve more successful results using much smaller feature vector sizes. Three different datasets were used in the experiments and statistically significant improvements for most of the proposed approaches were obtained. © 2010 Springer-Verlag London Limited.