Morphological disambiguation of turkish text with perceptron algorithm
8th International Conference on Computational Linguistics and Intelligent Text Processing, CICLing 2007, Mexico City, Meksika, 18 - 24 Şubat 2007, cilt.4394 LNCS, ss.107-118, (Tam Metin Bildiri)
- Yayın Türü: Bildiri / Tam Metin Bildiri
- Cilt numarası: 4394 LNCS
- Doi Numarası: 10.1007/978-3-540-70939-8_10
- Basıldığı Şehir: Mexico City
- Basıldığı Ülke: Meksika
- Sayfa Sayıları: ss.107-118
- Boğaziçi Üniversitesi Adresli: Evet
Özet
This paper describes the application of the perceptron algorithm to the morphological disambiguation of Turkish text. Turkish has a productive derivational morphology. Due to the ambiguity caused by complex morphology, a word may have multiple morphological parses, each with a different stem or sequence of morphemes. The methodology employed is based on ranking with perceptron algorithm which has been successful in some NLP tasks in English. We use a baseline statistical trigram-based model of a previous work to enumerate an n-best list of candidate morphological parse sequences for each sentence. We then apply the perceptron algorithm to rerank the n-best list using a set of 23 features. The perceptron trained to do morphological disambiguation improves the accuracy of the baseline model from 93.61% to 96.80%. When we train the perceptron as a POS tagger, the accuracy is 98.27%. Turkish morphological disambiguation and POS tagging results that we obtained is the best reported so far. © Springer-Verlag Berlin Heidelberg 2007.