Turkish verbal multiword expressions corpus Türkçe çok sözcüklü fiil ifadeleri derlemi


Berk G., Erden B., GÜNGÖR T.

26th IEEE Signal Processing and Communications Applications Conference, SIU 2018, İzmir, Türkiye, 2 - 05 Mayıs 2018, ss.1-4, (Tam Metin Bildiri)

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Doi Numarası: 10.1109/siu.2018.8404583
  • Basıldığı Şehir: İzmir
  • Basıldığı Ülke: Türkiye
  • Sayfa Sayıları: ss.1-4
  • Anahtar Kelimeler: Corpus, Corpus Annotation, Multiword Expression, Natural Language Processing, Turkish
  • Boğaziçi Üniversitesi Adresli: Evet

Özet

In this study, a Turkish corpus with labeled verbal multiword expressions was built. The verbal multiword expressions in the corpus were annotated according to their subcategories. The Turkish train and test corpora that was published in PARSEME Shared Task 1.0 were updated as train and development corpora based on PARSEME Annotation Guidelines. Additionally, a new Turkish test corpus was created by following the guidelines. The corpus consists of newspaper articles on politics, world, life, art and columns. The corpus will be released in PARSEME Shared Task 1.1. The corpus will be an important source to be used in many Turkish natural languages processing applications such as syntactic parsing, machine translation and n-gram language modeling.