Conversational open-domain question answering for resource-constrained languages


Budur E., GÜNGÖR T.

Turkish Journal of Electrical Engineering and Computer Sciences, cilt.33, sa.2, ss.203-223, 2025 (SCI-Expanded, Scopus, TRDizin)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 33 Sayı: 2
  • Basım Tarihi: 2025
  • Doi Numarası: 10.55730/1300-0632.4122
  • Dergi Adı: Turkish Journal of Electrical Engineering and Computer Sciences
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Compendex, TR DİZİN (ULAKBİM)
  • Sayfa Sayıları: ss.203-223
  • Anahtar Kelimeler: Open-domain question answering, conversational open-domain question answering, low-resource languages, machine translation
  • Boğaziçi Üniversitesi Adresli: Evet

Özet

The growing interest in Conversational AI has led to the development of Conversational OpenQA systems as a crucial step for meeting users’ information needs in real world scenarios. Conversational OpenQA systems enhance standard OpenQA performance by leveraging conversation history of the users. However, building effective Conversational OpenQA systems requires large-scale Conversational OpenQA datasets, often limited to the English language, hindering progress in low-resource languages. We present a robust Conversational OpenQA system enhanced by conversational context, designed for languages with limited resources and exemplified in our case study for Turkish. To address data limitations in a cost-effective way, we repurpose existing datasets like SQuAD-TR and XQuAD-TR, treating them as if they were constructed within a conversational context. Our findings indicate that incorporating conversation signals in the retriever models results in up to an absolute increase of 18.82% in Success@1 for retrievers. This improvement extends to the reader models enhanced by the conversational context, narrowing the gap in EM/F1 scores up to 4.12%/4.43%, respectively, compared to Standard QA readers.