Morphological disambiguation of turkish text with perceptron algorithm


Sak H., GÜNGÖR T., SARAÇLAR M.

8th International Conference on Computational Linguistics and Intelligent Text Processing, CICLing 2007, Mexico City, Meksika, 18 - 24 Şubat 2007, cilt.4394 LNCS, ss.107-118, (Tam Metin Bildiri)

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Cilt numarası: 4394 LNCS
  • Doi Numarası: 10.1007/978-3-540-70939-8_10
  • Basıldığı Şehir: Mexico City
  • Basıldığı Ülke: Meksika
  • Sayfa Sayıları: ss.107-118
  • Boğaziçi Üniversitesi Adresli: Evet

Özet

This paper describes the application of the perceptron algorithm to the morphological disambiguation of Turkish text. Turkish has a productive derivational morphology. Due to the ambiguity caused by complex morphology, a word may have multiple morphological parses, each with a different stem or sequence of morphemes. The methodology employed is based on ranking with perceptron algorithm which has been successful in some NLP tasks in English. We use a baseline statistical trigram-based model of a previous work to enumerate an n-best list of candidate morphological parse sequences for each sentence. We then apply the perceptron algorithm to rerank the n-best list using a set of 23 features. The perceptron trained to do morphological disambiguation improves the accuracy of the baseline model from 93.61% to 96.80%. When we train the perceptron as a POS tagger, the accuracy is 98.27%. Turkish morphological disambiguation and POS tagging results that we obtained is the best reported so far. © Springer-Verlag Berlin Heidelberg 2007.