A detailed analysis and improvement of feature-based named entity recognition for turkish
21st International Conference on Speech and Computer, SPECOM 2019, İstanbul, Türkiye, 20 - 25 Ağustos 2019, cilt.11658 LNAI, ss.9-19, (Tam Metin Bildiri)
- Yayın Türü: Bildiri / Tam Metin Bildiri
- Cilt numarası: 11658 LNAI
- Doi Numarası: 10.1007/978-3-030-26061-3_2
- Basıldığı Şehir: İstanbul
- Basıldığı Ülke: Türkiye
- Sayfa Sayıları: ss.9-19
- Anahtar Kelimeler: Conditional Random Fields, Dependency Parsing, Named Entity Recognition, Turkish
- Boğaziçi Üniversitesi Adresli: Evet
Özet
Named Entity Recognition (NER) is an important task in Natural Language Processing (NLP) with a wide range of applications. Recently, word embedding based systems that does not rely on hand-crafted features dominate the task as in the case of many other sequence labeling tasks in NLP. However, we are also observing the emergence of hybrid models that make use of hand crafted features through data augmentation to improve performance of such NLP systems. Such hybrid systems are especially important for less resourced languages such as Turkish as deep learning models require a large dataset to achieve good performance. In this paper, we first give a detailed analysis of the effect of various syntactic, semantic and orthographic features on NER for Turkish. We also improve the performance of the best feature based models for Turkish using additional features. We believe that our results will guide the research in this area and help making use of the key features for data augmentation.