Turkish labeled text corpus Türkçe etiketli metin derlemi
2014 22nd Signal Processing and Communications Applications Conference, SIU 2014, Trabzon, Türkiye, 23 - 25 Nisan 2014, ss.1395-1398, (Tam Metin Bildiri)
- Yayın Türü: Bildiri / Tam Metin Bildiri
- Doi Numarası: 10.1109/siu.2014.6830499
- Basıldığı Şehir: Trabzon
- Basıldığı Ülke: Türkiye
- Sayfa Sayıları: ss.1395-1398
- Anahtar Kelimeler: Classification, Corpus, Inverse Document Frequency, Latent Dirichlet Allocation, Natural Language Processing, NLP, Paper, Term Frequcney, TF-IDF, Turkish
- Boğaziçi Üniversitesi Adresli: Evet
Özet
A labeled text corpus made up of Turkish papers' titles, abstracts and keywords is collected. The corpus includes 35 number of different disciplines, and 200 documents per subject. This study presents the text corpus' collection and content. The classification performance of Term Frequcney - Inverse Document Frequency (TF-IDF) and topic probabilities of Latent Dirichlet Allocation (LDA) features are compared for the text corpus. The text corpus is shared as open source so that it could be used for natural language processing applications with academic purposes. © 2014 IEEE.