Resources for Turkish morphological processing
Language Resources and Evaluation, cilt.45, sa.2, ss.249-261, 2011 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 45 Sayı: 2
- Basım Tarihi: 2011
- Doi Numarası: 10.1007/s10579-010-9128-6
- Dergi Adı: Language Resources and Evaluation
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus
- Sayfa Sayıları: ss.249-261
- Anahtar Kelimeler: Turkish language resources, Morphological parser, Morphological disambiguation, Web corpus
- Boğaziçi Üniversitesi Adresli: Evet
Özet
We present a set of language resources and tools-a morphological parser, a morphological disambiguator, and a text corpus-for exploiting Turkish morphology in natural language processing applications. The morphological parser is a state-of-the-art finite-state transducer-based implementation of Turkish morphology. The disambiguator is based on the averaged perceptron algorithm and has the best accuracy reported for Turkish in the literature. The text corpus has been compiled from the web and contains about 500 million tokens. This is the largest Turkish web corpus published. © 2010 Springer Science+Business Media B.V.