Developing a concept extraction system for Turkish
2011 International Conference on Artificial Intelligence, ICAI 2011, Las Vegas, NV, Amerika Birleşik Devletleri, 18 - 21 Temmuz 2011, cilt.2, ss.855-860, (Tam Metin Bildiri)
- Yayın Türü: Bildiri / Tam Metin Bildiri
- Cilt numarası: 2
- Basıldığı Şehir: Las Vegas, NV
- Basıldığı Ülke: Amerika Birleşik Devletleri
- Sayfa Sayıları: ss.855-860
- Anahtar Kelimeler: Concept extraction, Natural language processing
- Boğaziçi Üniversitesi Adresli: Evet
Özet
In recent years, due to the vast amount of available electronic media and data, the necessity of analyzing electronic documents automatically was increased. In order to assess if a document contains valuable information or not, concepts, key phrases or main idea of the document have to be known. There are some studies on extracting key phrases or main ideas of documents for Turkish. However, to the best of our knowledge, there is no concept extraction system for Turkish although such systems exist for well-known languages. In this paper, a concept extraction system is proposed for Turkish. By applying some statistical and Natural Language Processing methods, documents are identified by concepts. As a result, the system generates concepts with 51% success, but it generates more concepts than it should be. Since concepts are abstract entities, in other words they do not have to be written in the texts as they appear, assigning concepts is a very difficult issue. Moreover, if we take into account the complexity of the Turkish language this result can be seen as quite satisfactory.