A Multilingual Evaluation Dataset for Monolingual Word Sense Alignment Sina Ahmadi*, John P. McCrae*, Sanni Nimb1, Fahad Khan3, Monica Monachini3, Bolette S. Pedersen8, Thierry Declerck2,12, Tanja Wissik2, Andrea Bellandi3, Irene Pisani4, Thomas Troelsgard˚ 1, Sussi Olsen8, Simon Krek5, Veronika Lipp6, Tamas´ Varadi´ 6,Laszl´ o´ Simon6, Andras´ Gyorffy˝ 6, Carole Tiberius9, Tanneke Schoonheim9, Yifat Ben Moshe10, Maya Rudich10, Raya Abu Ahmad10, Dorielle Lonke10, Kira Kovalenko11, Margit Langemets13, Jelena Kallas13, Oksana Dereza7, Theodorus Fransen7, David Cillessen7, David Lindemann14, Mikel Alonso14, Ana Salgado15 Jose´ Luis Sancho16, Rafael-J. Urena-Ruiz˜ 16, Jordi Porta Zamorano16, Kiril Simov17, Petya Osenova17, Zara Kancheva17, Ivaylo Radev17, Ranka Stankovic´18, Andrej Perdih19, Dejan Gabrovsekˇ 19 *Insight Centre for Data Analytics, National University of Ireland, Galway fsina.ahmadi,
[email protected] (other affiliations in Appendix A) Abstract Aligning senses across resources and languages is a challenging task with beneficial applications in the field of natural language process- ing and electronic lexicography. In this paper, we describe our efforts in manually aligning monolingual dictionaries. The alignment is carried out at sense-level for various resources in 15 languages. Moreover, senses are annotated with possible semantic relationships such as broadness, narrowness, relatedness, and equivalence. In comparison to previous datasets for this task, this dataset covers a wide range of languages and resources and focuses on the more challenging task of linking general-purpose language. We believe that our data will pave the way for further advances in alignment and evaluation of word senses by creating new solutions, particularly those notoriously requiring data such as neural networks. Our resources are publicly available at https://github.com/elexis-eu/MWSA.