A Unified Turkic Idiom Understanding Benchmark: Idiom Detection and Semantic Retrieval Across Five Turkic Languages


Creative Commons License

Aslantaş G., GÜNGÖR T.

2nd Workshop on Natural Language Processing for Turkic Languages, SIGTURK 2026, Rabat, Fas, 29 Mart 2026, ss.38-51, (Tam Metin Bildiri)

Özet

Idiomatic expressions are culturally grounded, semantically opaque, and challenging for multilingual natural language processing systems. Despite the large speaker population of Turkic languages, resources for monolingual and cross-lingual idiom understanding remain scarce. We introduce the first unified benchmark for idiom understanding across Turkish, Azerbaijani, Turkmen, Gagauz, and Uzbek, featuring token-level idiom span annotations. We evaluate seven models for idiom identification and nine embedding models for semantic retrieval under multiple fine-tuning schemes. Our benchmark enables systematic analysis of how idiomatic meanings are shared, transformed, or diverge across Turkic languages.