Analysis of Subword Tokenization Approaches for Turkish Language Türkçe Dili için Kelime Bölme Yaklaşimlarinin Incelenmesi


Erkaya E., GÜNGÖR T.

31st IEEE Conference on Signal Processing and Communications Applications, SIU 2023, İstanbul, Türkiye, 5 - 08 Temmuz 2023, (Tam Metin Bildiri)

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Doi Numarası: 10.1109/siu59756.2023.10223973
  • Basıldığı Şehir: İstanbul
  • Basıldığı Ülke: Türkiye
  • Anahtar Kelimeler: morphology, natural language processing, subword tokenizers, Turkish
  • Boğaziçi Üniversitesi Adresli: Evet

Özet

Various tokenization approaches have been proposed in natural language processing research. These approaches have further evolved from character- and word-level representations to subword-level representations. However, the impact of tokenizations on model performance has not been thoroughly discussed, especially for morphologically rich languages. In this paper, we comprehensively analyze subword tokenizers for Turkish which is a highly inflected and morphologically rich language and we propose a morphologically-based approach. Also, we examine how the tokenizer parameters like vocabulary and corpus sizes change the characteristics of tokenizers.