Sentiment analysis in Turkish: Supervised, semi-supervised, and unsupervised techniques


Aydln C. R., GÜNGÖR T.

Natural Language Engineering, cilt.27, sa.4, ss.455-483, 2021 (SCI-Expanded, AHCI, SSCI, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 27 Sayı: 4
  • Basım Tarihi: 2021
  • Doi Numarası: 10.1017/s1351324920000200
  • Dergi Adı: Natural Language Engineering
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Arts and Humanities Citation Index (AHCI), Social Sciences Citation Index (SSCI), Scopus, Applied Science & Technology Source, Compendex, Computer & Applied Sciences, INSPEC, Linguistics & Language Behavior Abstracts, Psycinfo, DIALNET
  • Sayfa Sayıları: ss.455-483
  • Anahtar Kelimeler: Machine learning, Morphological analysis, Opinion mining, Sentiment analysis, Text classification
  • Boğaziçi Üniversitesi Adresli: Evet

Özet

Although many studies on sentiment analysis have been carried out for widely spoken languages, this topic is still immature for Turkish. Most of the works in this language focus on supervised models, which necessitate comprehensive annotated corpora. There are a few unsupervised methods, and they utilize sentiment lexicons either built by translating from English lexicons or created based on corpora. This results in improper word polarities as the language and domain characteristics are ignored. In this paper, we develop unsupervised (domain-independent) and semi-supervised (domain-specific) methods for Turkish, which are based on a set of antonym word pairs as seeds. We make a comprehensive analysis of supervised methods under several feature weighting schemes. We then form ensemble of supervised classifiers and also combine the unsupervised and supervised methods. Since Turkish is an agglutinative language, we perform morphological analysis and use different word forms. The methods developed were tested on two datasets having different styles in Turkish and also on datasets in English to show the portability of the approaches across languages. We observed that the combination of the unsupervised and supervised approaches outperforms the other methods, and we obtained a significant improvement over the state-of-the-art results for both Turkish and English.