Proceedings of the 1St Workshop on Sense, Concept and Entity Representations and Their Applications, Pages 1–11, Valencia, Spain, April 4 2017
Total Page:16
File Type:pdf, Size:1020Kb
Load more
Recommended publications
-
The Main Features of the E-Glava Online Valency Dictionary
The Main Features of the e-Glava Online Valency Dictionary Matea Birtić, Ivana Brač, Siniša Runjaić Institute of Croatian Language and Linguistics, Ulica Republike Austrije 16, HR-10000 Zagreb, Croatia E.mail: [email protected], [email protected], [email protected] Abstract E-Glava is an online valency dictionary of Croatian verbs. The theoretical approach to valency follows the German tradition, particularly that of the VALBU dictionary, with some minor changes and adjustments. The main principle of our valency approach is to link valency patterns to specific verb meanings. The verb list is compiled semi-automatically on the basis of the Croatian Frequency Dictionary and Croatian language textbooks. Currently, e-Glava contains descriptions of 57 psychological verbs with 187 meanings and 375 valency patterns. The lexicographic articles are written in Tschwanelex. A Document Type Definition editing module has been used, and the description of verbs follows a three-level linguistic schema prepared for lexicographers. Verbs are distributed throughout 34 semantic classes, and examples are extracted manually from Croatian corpora. Fully processed data for each semantic class will be publicly available in the form of a browsable HTML dictionary. The paper also presents a comparison between e-Glava and other cognate resources, as well as a summary of its main advantages, disadvantages, and potential applied uses. Keywords: Croatian language; valency dictionary; e-dictionary; syntax 1. Introduction Sentence structure and the syntactic behaviour of verbs were perhaps the most intriguing and interesting topics for early grammatical descriptions and, later, linguistic descriptions of language. Valency properties are relevant to both theoretical and applied linguistic considerations. -
FERSIWN GYMRAEG ISOD the National Corpus of Contemporary
FERSIWN GYMRAEG ISOD The National Corpus of Contemporary Welsh Project Report, October 2020 Authors: Dawn Knight1, Steve Morris2, Tess Fitzpatrick2, Paul Rayson3, Irena Spasić and Enlli Môn Thomas4. 1. Introduction 1.1. Purpose of this report This report provides an overview of the CorCenCC project and the online corpus resource that was developed as a result of work on the project. The report lays out the theoretical underpinnings of the research, demonstrating how the project has built on and extended this theory. We also raise and discuss some of the key operational questions that arose during the course of the project, outlining the ways in which they were answered, the impact of these decisions on the resource that has been produced and the longer-term contribution they will make to practices in corpus-building. Finally, we discuss some of the applications and the utility of the work, outlining the impact that CorCenCC is set to have on a range of different individuals and user groups. 1.2. Licence The CorCenCC corpus and associated software tools are licensed under Creative Commons CC-BY-SA v4 and thus are freely available for use by professional communities and individuals with an interest in language. Bespoke applications and instructions are provided for each tool (for links to all tools, refer to section 10 of this report). When reporting information derived by using the CorCenCC corpus data and/or tools, CorCenCC should be appropriately acknowledged (see 1.3). § To access the corpus visit: www.corcencc.org/explore § To access the GitHub site: https://github.com/CorCenCC o GitHub is a cloud-based service that enables developers to store, share and manage their code and datasets. -
The Workshop Programme
The Workshop Programme Monday, May 26 14:30 -15:00 Opening Nancy Ide and Adam Meyers 15:00 -15:30 SIGANN Shared Corpus Working Group Report Adam Meyers 15:30 -16:00 Discussion: SIGANN Shared Corpus Task 16:00 -16:30 Coffee break 16:30 -17:00 Towards Best Practices for Linguistic Annotation Nancy Ide, Sameer Pradhan, Keith Suderman 17:00 -18:00 Discussion: Annotation Best Practices 18:00 -19:00 Open Discussion and SIGANN Planning Tuesday, May 27 09:00 -09:10 Opening Nancy Ide and Adam Meyers 09:10 -09:50 From structure to interpretation: A double-layered annotation for event factuality Roser Saurí and James Pustejovsky 09:50 -10:30 An Extensible Compositional Semantics for Temporal Annotation Harry Bunt, Chwhynny Overbeeke 10:30 -11:00 Coffee break 11:00 -11:40 Using Treebank, Dictionaries and GLARF to Improve NomBank Annotation Adam Meyers 11:40 -12:20 A Dictionary-based Model for Morpho-Syntactic Annotation Cvetana Krstev, Svetla Koeva, Du!ko Vitas 12:20 -12:40 Multiple Purpose Annotation using SLAT - Segment and Link-based Annotation Tool (DEMO) Masaki Noguchi, Kenta Miyoshi, Takenobu Tokunaga, Ryu Iida, Mamoru Komachi, Kentaro Inui 12:40 -14:30 Lunch 14:30 -15:10 Using inheritance and coreness sets to improve a verb lexicon harvested from FrameNet Mark McConville and Myroslava O. Dzikovska 15:10 -15:50 An Entailment-based Approach to Semantic Role Annotation Voula Gotsoulia 16:00 -16:30 Coffee break 16:30 -16:50 A French Corpus Annotated for Multiword Expressions with Adverbial Function Eric Laporte, Takuya Nakamura, Stavroula Voyatzi -
The National Corpus of Contemporary Welsh 1. Introduction
The National Corpus of Contemporary Welsh Project Report, October 2020 1. Introduction 1.1. Purpose of this report This report provides an overview of the CorCenCC project and the online corpus resource that was developed as a result of work on the project. The report lays out the theoretical underpinnings of the research, demonstrating how the project has built on and extended this theory. We also raise and discuss some of the key operational questions that arose during the course of the project, outlining the ways in which they were answered, the impact of these decisions on the resource that has been produced and the longer-term contribution they will make to practices in corpus-building. Finally, we discuss some of the applications and the utility of the work, outlining the impact that CorCenCC is set to have on a range of different individuals and user groups. 1.2. Licence The CorCenCC corpus and associated software tools are licensed under Creative Commons CC-BY-SA v4 and thus are freely available for use by professional communities and individuals with an interest in language. Bespoke applications and instructions are provided for each tool (for links to all tools, refer to section 10 of this report). When reporting information derived by using the CorCenCC corpus data and/or tools, CorCenCC should be appropriately acknowledged (see 1.3). § To access the corpus visit: www.corcencc.org/explore § To access the GitHub site: https://github.com/CorCenCC o GitHub is a cloud-based service that enables developers to store, share and manage their code and datasets. -
Overabundance in Croatian Dual-Class Verbs FLUMINENSIA, God
Tomislava Bošnjak Botica, Gordana Hržica, Overabundance in Croatian dual-class verbs FLUMINENSIA, god. 28 (2016), br. 1 Tomislava Bošnjak Botica, Gordana Hržica OVERABUNDANCE IN CROATIAN DUAL-CLASS VERBS dr. sc. Tomislava Bošnjak Botica, Institut za hrvatski jezik i jezikoslovlje, [email protected], Zagreb dr. sc. Gordana Hržica, Edukacijsko-rehabilitacijski fakultet, [email protected], Zagreb izvorni znanstveni članak UDK 811.163.42’367.625 rukopis primljen: 5. 4. 2016.; prihvaćen za tisak: 21. 6. 2016. Croatian verbal inflection morphology is typically described using verb class distinctions. The number of classes differs among approaches, but the basic criterion for class division is the presence or absence and the type of suppletion in verb stems. Generally, one verb belongs to one inflectional class or paradigm only. However, some verbs belong to two classes, i.e. they have two parallel sets of stems. In such dual-class verbs, one infinitive form is realizable in two present forms in all cells within a class, i.e. there is an overabundance (Thorton 2011). Inevitably, one of the stem forming paradigms is a class with categorial suppletion. The present stem of a categorial suppletion class has a greater phonological distance from the infinitive stem than the present stem of the other class. Using a different terminology one class can be described as more transparent, while the other is less transparent (more opaque) in forming the present stem. This study attempts to present overabundance in dual-class verbs and to determine whether competition in such forms can be explained by their tendency to conform to one default class or by other factors, specifically, by the phonological distance between the two paradigms of dual-class verbs. -
Book of Abstracts
Book of Abstracts - Journée d’études AFLiCo JET Corpora and Representativeness jeudi 3 et vendredi 4 mai 2018 Bâtiment Max Weber (W) Keynote speakers Dawn Knight et Thomas Egan Plus d’infos : bit.ly/2FPGxvb 1 Table of contents DAY 1 – May 3rd 2018 Thomas Egan, Some perils and pitfalls of non-representativeness 1 Daniel Henke, De quoi sont représentatifs les corpus de textes traduits au juste ? Une étude de corpus comparable-parallèle 1 Ilmari Ivaska, Silvia Bernardini, Adriano Ferraresi, The comparability paradox in multilingual and multi-varietal corpus research: Coping with the unavoidable 2 Antonina Bondarenko, Verbless Sentences: Advantages and Challenges of a Parallel Corpus-based Approach 3 Adeline Terry, The representativeness of the metaphors of death, disease, and sex in a TV show corpus 5 Julien Perrez, Pauline Heyvaert, Min Reuchamps, On the representativeness of political corpora in linguistic research 6 Joshua M. Griffiths, Supplementing Maximum Entropy Phonology with Corpus Data 8 Emmanuelle Guérin, Olivier Baude, Représenter la variation – Revisiter les catégories et les variétés dans le corpus ESLO 9 Caroline Rossi, Camille Biros, Aurélien Talbot, La variation terminologique en langue de spécialité : pour une analyse à plusieurs niveaux 10 Day 2 – May 4rd 2018 Dawn Knight, Representativeness in CorCenCC: corpus design in minoritised languages 11 Frederick Newmeyer, Conversational corpora: When 'big is beautiful' 12 Graham Ranger, How to get "along": in defence of an enunciative and corpus- based approach 13 Thi Thu Trang -
Overabundance in Croatian Dual-Class Verbs FLUMINENSIA, God
Tomislava Bošnjak Botica, Gordana Hržica, Overabundance in Croatian dual-class verbs FLUMINENSIA, god. 28 (2016), br. 1, str. 83-106 83 Tomislava Bošnjak Botica, Gordana Hržica OVERABUNDANCE IN CROATIAN DUAL-CLASS VERBS dr. sc. Tomislava Bošnjak Botica, Institut za hrvatski jezik i jezikoslovlje, [email protected], Zagreb dr. sc. Gordana Hržica, Edukacijsko-rehabilitacijski fakultet, [email protected], Zagreb izvorni znanstveni članak UDK 811.163.42’367.625 rukopis primljen: 5. 4. 2016.; prihvaćen za tisak: 21. 6. 2016. Croatian verbal inflection morphology is typically described using verb class distinctions. The number of classes differs among approaches, but the basic criterion for class division is the presence or absence and the type of suppletion in verb stems. Generally, one verb belongs to one inflectional class or paradigm only. However, some verbs belong to two classes, i.e. they have two parallel sets of stems. In such dual-class verbs, one infinitive form is realizable in two present forms in all cells within a class, i.e. there is an overabundance (Thorton 2011). Inevitably, one of the stem forming paradigms is a class with categorial suppletion. The present stem of a categorial suppletion class has a greater phonological distance from the infinitive stem than the present stem of the other class. Using a different terminology one class can be described as more transparent, while the other is less transparent (more opaque) in forming the present stem. This study attempts to present overabundance in dual-class verbs and to determine whether competition in such forms can be explained by their tendency to conform to one default class or by other factors, specifically, by the phonological distance between the two paradigms of dual-class verbs. -
The Spoken British National Corpus 2014
The Spoken British National Corpus 2014 Design, compilation and analysis ROBBIE LOVE ESRC Centre for Corpus Approaches to Social Science Department of Linguistics and English Language Lancaster University A thesis submitted to Lancaster University for the degree of Doctor of Philosophy in Linguistics September 2017 Contents Contents ............................................................................................................. i Abstract ............................................................................................................. v Acknowledgements ......................................................................................... vi List of tables .................................................................................................. viii List of figures .................................................................................................... x 1 Introduction ............................................................................................... 1 1.1 Overview .................................................................................................................................... 1 1.2 Research aims & structure ....................................................................................................... 3 2 Literature review ........................................................................................ 6 2.1 Introduction .............................................................................................................................. -
Language, Communities & Mobility
The 9th Inter-Varietal Applied Corpus Studies (IVACS) Conference Language, Communities & Mobility University of Malta June 13 – 15, 2018 IVACS 2018 Wednesday 13th June 2018 The conference is to be held at the University of Malta Valletta Campus, St. Paul’s Street, Valletta. 12:00 – 12:45 Conference Registration 12:45 – 13:00 Conferencing opening welcome and address University of Malta, Dean of the Faculty of Arts: Prof. Dominic Fenech Room: Auditorium 13:00 – 14:00 Plenary: Dr Robbie Love, University of Leeds Overcoming challenges in corpus linguistics: Reflections on the Spoken BNC2014 Room: Auditorium Rooms Auditorium (Level 2) Meeting Room 1 (Level 0) Meeting Room 2 (Level 0) Meeting Room 3 (Level 0) 14:00 – 14:30 Adverb use in spoken ‘Yeah, no, everyone seems to On discourse markers in Lexical Bundles in the interaction: insights and be saying that’: ‘New’ Lithuanian argumentative Description of Drug-Drug implications for the EFL pragmatic markers as newspaper discourse: a Interactions: A Corpus-Driven classroom represented in fictionalized corpus-based study Study - Pascual Pérez-Paredes & Irish English - Anna Ruskan - Lukasz Grabowski Geraldine Mark - Ana Mª Terrazas-Calero & Carolina Amador-Moreno 1 14:30 – 15:00 “Um so yeah we’re <laughs> So you have to follow The Study of Ideological Bias Using learner corpus data to you all know why we’re different methods and through Corpus Linguistics in inform the development of a here”: hesitation in student clearly there are so many Syrian Conflict News from diagnostic language tool and -
Book of Abstracts
xx Table of Contents Session O1 - Machine Translation & Evaluation . 1 Session O2 - Semantics & Lexicon (1) . 2 Session O3 - Corpus Annotation & Tagging . 3 Session O4 - Dialogue . 4 Session P1 - Anaphora, Coreference . 5 Session: Session P2 - Collaborative Resource Construction & Crowdsourcing . 7 Session P3 - Information Extraction, Information Retrieval, Text Analytics (1) . 9 Session P4 - Infrastructural Issues/Large Projects (1) . 11 Session P5 - Knowledge Discovery/Representation . 13 Session P6 - Opinion Mining / Sentiment Analysis (1) . 14 Session P7 - Social Media Processing (1) . 16 Session O5 - Language Resource Policies & Management . 17 Session O6 - Emotion & Sentiment (1) . 18 Session O7 - Knowledge Discovery & Evaluation (1) . 20 Session O8 - Corpus Creation, Use & Evaluation (1) . 21 Session P8 - Character Recognition and Annotation . 22 Session P9 - Conversational Systems/Dialogue/Chatbots/Human-Robot Interaction (1) . 23 Session P10 - Digital Humanities . 25 Session P11 - Lexicon (1) . 26 Session P12 - Machine Translation, SpeechToSpeech Translation (1) . 28 Session P13 - Semantics (1) . 30 Session P14 - Word Sense Disambiguation . 33 Session O9 - Bio-medical Corpora . 34 Session O10 - MultiWord Expressions . 35 Session O11 - Time & Space . 36 Session O12 - Computer Assisted Language Learning . 37 Session P15 - Annotation Methods and Tools . 38 Session P16 - Corpus Creation, Annotation, Use (1) . 41 Session P17 - Emotion Recognition/Generation . 43 Session P18 - Ethics and Legal Issues . 45 Session P19 - LR Infrastructures and Architectures . 46 xxi Session I-O1: Industry Track - Industrial systems . 48 Session O13 - Paraphrase & Semantics . 49 Session O14 - Emotion & Sentiment (2) . 50 Session O15 - Semantics & Lexicon (2) . 51 Session O16 - Bilingual Speech Corpora & Code-Switching . 52 Session P20 - Bibliometrics, Scientometrics, Infometrics . 54 Session P21 - Discourse Annotation, Representation and Processing (1) . 55 Session P22 - Evaluation Methodologies . -
Pregled Razvoja Hrvatske E-Leksikografije
Pregled razvoja hrvatske e-leksikografije Kristina Štrkalj Despot i Ana Ostroški Anić U radu se daje kratak pregled razvoja i trenutačnoga stanja hrvatske e-leksikografije od prvih korpusno utemeljenih rječnika do najrecentnijega, izvorno digitalnoga rječnika Mrežnika. Detaljnije se opisuju leksikografski projekti kojima je Maja Bratanić dala važan prinos, a koji su se pokazali prijelomnim točkama u tome razvoju: čestotni rječnik Milana Moguša, Maje Bratanić i Marka Tadića, višejezični leksikografski projekt Johna Sinclaira te projekt STRUNA u Institutu za hrvatski jezik i jezikoslovlje. 1. Uvod Danas je vrlo raširen i gotovo potpuno prihvaćen stav o (e-)leksikografiji kao autonomnoj disciplini u odnosu na lingvistiku, i to disciplini koja ima razvijenu vlastitu teoriju i praksu. Za takav pristup leksikografiji osobito je zaslužna utjecajna danska funkcionalna leksikografska škola, tj. aarhuška škola (s Bergenholtzom kao najistaknutijim predstavnikom), koja je takav stav najjasnije artikulirala ističući ne samo autonomnost te discipline u odnosu na lingvistiku nego i njezinu pripadnost informacijskim znanostima (v. npr. Fuertes-Olivera i Bergenholtz 2011, Tarp 2012, Štrkalj Despot i Möhrs 2015). Takav je pristup izrastao na tradicionalnim i utjecajnim leksikografskim školama poput ruske (npr. Scerba 1940, Sorokoletov 1978) i njemačke (npr. Duda i dr. 1986, Wiegand 1999) te na pogledima istaknutih leksikografa poput Gouwsa (2011) ili Zguste (1992). No ovako artikuliran stav o autonomnosti leksikografije izazvao je i snažan otpor nekih vrlo uglednih leksikografa (najistaknutiji su Atkins i Rundell 2008) te metaleksikografa poput Béjointa 2010, koji drže da se ne može govoriti o teoriji leksikografije, nego samo o praksi izrade rječnika. Tako npr. Béjoint (2010) ističe kako znanost može imati teoriju, ali praktični rad ne može jer prirodni fenomeni trebaju svoje teorije, a zasigurno ne može biti teorije o stvaranju artefakata. -
Hrvatsko Društvo Za Primijenjenu Lingvistiku
Hrvatsko društvo za primijenjenu lingvistiku Croatian Applied Linguistics Society Kroatische Gesellschaft für Angewandte Linguistik Association croate de linguistique appliquée Associazione croata di linguistica applicata JEZIK I UM XXXII. međunarodni znanstveni skup 3. – 5. svibnja 2018. Rijeka, Hrvatska LANGUAGE AND MIND 32nd International Conference 3rd – 5th May 2018 Rijeka, Croatia Knjiga sažetaka Book of Abstracts XXXII. međunarodni znanstveni skup 32nd International Conference 3. – 5. svibnja 2018. / 3rd – 5th May 2018 Rijeka, Hrvatska / Rijeka, Croatia JEZIK I UM LANGUAGE AND MIND Knjiga sažetaka Book of Abstracts Uredile / Editors Diana Stolac / Magdalena Nigoević HDPL /CALS 2018. IZDAVAČ Srednja Europa Hrvatsko društvo za primijenjenu lingvistiku ISBN 978-953-7963-77-4 Jezik i um Language and mind Mjesto održavanja / Conference venue Filozofski fakultet u Rijeci, Sveučilište u Rijeci Faculty of Humanities and Social Sciences, University of Rijeka Organizacijski odbor / Organizing Committee Mihaela Matešić (Rijeka), Anita Memišević (Rijeka), Anastazija Vlastelić (Rijeka), Sanda Lucija Udier (Zagreb), Diana Stolac (Rijeka), Magdalena Nigoević (Split), Borana Morić-Mohorovičić (Rijeka), Benedikt Perak (Rijeka), Jasna Bićanić (Rijeka) Programski odbor / Programme Committee Lada Badurina (Rijeka), Tatjana Balazic-Bulc (Ljubljana), Branimir Belaj (Osijek), Silvija Batoš (Dubrovnik), Goranka Blagus Bartolec (Zagreb), Tomislava Bošnjak Botica (Zagreb), Maja Brala-Vukanović (Rijeka), Mario Brdar (Osijek), Kristina Cergol Kovačević (Zagreb),