Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell

Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell Djamé Seddah1 Farah Essaidi1 Amal Fethi1 Matthieu Futeral1 Benjamin Muller1;2 Pedro Javier Ortiz Suárez1;2 Benoît Sagot1 Abhishek Srivastava1 1Inria, Paris, France 2Sorbonne Université, Paris, France [email protected] Abstract sourced, namely Haitian Creole and Algerian dialectal Arabic in these cases. No readily avail- We introduce the first treebank for a romanized able parsing and machine translations systems are user-generated content variety of Algerian, a available for such languages. Taking as an ex- North-African Arabic dialect known for its fre- quent usage of code-switching. Made of 1500 ample the Arabic dialects spoken in North-Africa, sentences, fully annotated in morpho-syntax mostly from Morocco to Tunisia, sometimes called and Universal Dependency syntax, with full Maghribi, sometimes Darija, these idioms notori- translation at both the word and the sentence ously contain various degrees of code-switching levels, this treebank is made freely available. with languages of former colonial powers such as It is supplemented with 50k unlabeled sen- French, Spanish, and, to a much lesser extent, Ital- tences collected from Common Crawl and web- ian, depending on the area of usage (Habash, 2010; crawled data using intensive data-mining tech- Cotterell et al., 2014; Saadane and Habash, 2015). niques. Preliminary experiments demonstrate its usefulness for POS tagging and dependency They share Modern Standard Arabic (MSA) as parsing. We believe that what we present in their matrix language (Myers-Scotton, 1993), and this paper is useful beyond the low-resource of course present a rich morphology. In conjunc- language community. This is the first time that tion with the resource scarcity issue, the code- enough unlabeled and annotated data is pro- switching variability displayed by these languages vided for an emerging user-generated content challenges most standard NLP pipelines, if not all. dialectal language with rich morphology and What makes these dialects especially interesting code switching, making it an challenging test- bed for most recent NLP approaches. is their widespread use in user-generated content found on social media platforms, where they are 1 Introduction generally written using a romanized version of the Arabic script, called Arabizi, which is neither stan- Until the rise of fully unsupervised techniques that dardized nor formalized. The absence of standard- would free our field from its addiction to anno- ization for this script adds another layer of varia- tated data, the question of building useful data tion in addition to well-known user generated con- sets for under-resourced languages at a reasonable tent idiosyncrasies, making the processing of this cost is still crucial. Whether the lack of labeled kind of text an even more challenging task. data originates from being a minority language In this work, we present a new data set of about status, its almost oral-only nature or simply its 1500 sentences randomly sampled from the ro- programmed political disappearance, geopolitical manized Algerian dialectal Arabic corpus of Cot- events are a factor highlighting a language defi- terell et al. (2014) and from a small corpus of ciency in terms of natural language processing re- lyrics coming from Algerian dialectal Arabic Hip- sources that can have an important societal impact. Hop and Raï music genre that had the advan- Events such as the Haïti crisis in 2010 (Munro, tage of having already available translations and 2010) and the current Algerian revolts (Nossiter, of being representative of Algerian vernacular ur- 1 2019) are massively reflected on social media, yet ban youth language. We manually annotated this often in languages or dialects that are poorly re- data set with morpho-syntactic information (parts- 1https://www.nytimes.com/2019/03/01/world/ of-speech and morphological features), together africa/algeria-protests-bouteflika.html with glosses and code-switching labels at the word 1139 Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1139–1150 July 5 - 10, 2020. c 2020 Association for Computational Linguistics level, as well as sentence-level translations. Fur- Gloss Attested forms Lang thermore, we added an additional manual annota- why wa3lach w3alh 3alach 3lache NArabizi tion layer following the Universal Dependencies all ekl kal kolach koulli kol NArabizi annotation scheme (Nivre et al., 2018), making many beaucoup boucoup bcp French of this corpus, to the best of our knowledge, the Table 1: Examples of lexical variation in NArabizi first user-generated content treebank in romanized dialectal Arabic. This treebank contains 36% of French tokens, making it a valuable resource to As stated above, this dialect is mostly spoken measure and study the impact of code-switching on and has even been dubbed with disdain as a Creole NLP tools. We supplement this annotated corpus language by the higher levels of the Algerian po- 3 with about 50k unlabeled sentences extracted from litical hierarchy. Still, its usage is ubiquitous in both Common Crawl and additional web crawled the society and, by extension, in social media user- data, making of this data set an important mile- generated content. Interestingly, the lack of Arabic stone in North-African dialectal Arabic NLP. This support in input devices led to the rise of a roman- corpus is made freely available under a Creative ized written form of this dialect, which makes use Commons license.2 of alphanumeric letters as additional graphemes to represent phonemes that the Latin script does not 2 The Language naturally cover. Not limited to North-African dialectal Arabic, this non-standard “transliteration” As stated by Habash (2010), Arabic languages are concurrently emerged all over the Arabic-speaking often classified into three categories : (i) Classical world, and is often called Arabizi. Whether or not Arabic, as found in the Qur’an and related canon- written in Arabizi, the inter-dialectal divergences ical texts, (ii) Modern Standard Arabic, the offi- between all Arabic dialects remain. cial language of the vast majority of Arabic speak- The following list highlights some of the main ing countries and (iii) Dialectal Arabic, whose in- properties of Arabizi compared to MSA written in stances exhibit so much variations that they are the Arabic script. not mutually understandable across geographically • Unlike in MSA written in the Arabic script, distant regions. As space is missing for an exhaus- where short vowels are marked using optional tive description of Arabic language variations, we diacritics, all vowels are explicitly written; refer the reader to Habash (2010), Samih (2017) • Digits are used to cope with Arabic phonemes and especially to Saadane and Habash (2015) for that have no counterpart in the Latin script; a thorough account of Algerian dialectal Arabic, for instance, the digit “3” is often used to de- which is the focus of this work. In short, the key note the ayin consonant, because it is graphi- properties of North-African dialectal Arabic are: cally similar to its rendition in Arabic script; • It is a Semitic language, non codified, mostly • No norms exist, resulting in a high degree of spoken; variability between people writing in Arabizi. • It has a rich-inflexion system, which quali- From now on, we will call NArabizi the Algerian fies this dialect as a morphologically-rich lan- dialect of Arabic when written in Arabizi, thereby guage (Tsarfaty et al., 2010), even though simultaneously referring to the language variety Saadane and Habash (2015) write that many and to the script itself. Table 1 presents several properties present in Classical Arabic are ab- examples of lexical variation within NArabizi. In- sent from this dialect (e.g. it has simplified terestingly, this variability also affects the code- nominal and verbal case systems); switched vocabulary, which is mostly French in • It displays a high degree of variability at the case of NArabizi. A typical example of NAra- all levels: spelling and transliteration conven- bizi that also exhibits code-switching with non- tions, phonology, morphology, lexicon; standard French spelling can be seen in Example 1. • It exhibits a high degree of code-switching; due to historical reasons and cultural influ- (1) ence of French in the media circles, the Alge- Source: salem 3alikoum inchalah le pondium rian dialect, as well as Tunisian and Morocco, et les midailes d’or is known for its heavy use of French words. 3https://www.lesoirdalgerie.com/articles/ 2http://almanach-treebanks.fr/NArabizi 2010/02/17/article.php?sid=95823&cid=2 1140 Norm.: Assalamu alaykum inshallah le contracted into one token).5 podium et les médailles d’or Trans.: Peace be on you God willing [we will Morphological Analysis This layer consists of get] the podium and the gold medals two sets of part-of-speech tags, one following the Universal POS tagset (Petrov et al., 2011) 3 Corpus and the other the FTB-cc tagset extended to deal with user-generated content (Seddah et al., As other North-African Arabic dialects, NArabizi 2012). In cases of word contractions, we fol- is a resource-poor language, with, to the best of our lowed their guidelines and used multiple POS knowledge, only one available corpus developed as in cetait (`itwas')/PRON+VERB/CLS+V. In by Cotterell et al. (2014) for language identifica- addition, we added several morphological fea- tion purposes. tures following the Universal Dependency annotation scheme (Nivre et al., 2018), namely gender, 3.1 Data Collection number, tense and verbal mood. Note that instead Cotterell et al. (2014)’s corpus was collected in of adding lemmas, we included French glosses 2012 from an Algerian newspaper’s web forums for two reasons: firstly for practical reasons, as and covers a wide range of topics (from discussion they helped manual corrections done by non-native about football events to politics).

Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell

How Do BERT Embeddings Organize Linguistic Knowledge?

Treebanks, Linguistic Theories and Applications Introduction to Treebanks

Senserelate::Allwords - a Broad Coverage Word Sense Tagger That Maximizes Semantic Relatedness

Deep Linguistic Analysis for the Accurate Identification of Predicate

The Procedure of Lexico-Semantic Annotation of Składnica Treebank

Unified Language Model Pre-Training for Natural

Building a Treebank for French

Converting an HPSG-Based Treebank Into Its Parallel Dependency-Based Treebank

Corpus Based Evaluation of Stemmers

Lecture 5: Part-Of-Speech Tagging

Merging Propbank, Nombank, Timebank, Penn Discourse Treebank and Coreference James Pustejovsky, Adam Meyers, Martha Palmer, Massimo Poesio

Fasttext-Based Intent Detection for Inflected Languages †