Eslema. Towards a Corpus for Asturian
Total Page:16
File Type:pdf, Size:1020Kb
Eslema. Towards a Corpus for Asturian Xulio Viejoz, Roser Saur´ı∗, Angel´ Neiray zDepartamento de Filolog´ıa Espanola˜ yComputer Science Department Universidad de Oviedo fjviejo, [email protected] ∗Computer Science Department Brandeis University [email protected] Abstract We present Eslema, the first project devoted to building a corpus for Asturian, which is carried out at Oviedo University. Eslema receives minor funding from the Spanish government, which is fundamental for basic issues such as equipment acquisition. However, it is insufficient for hiring researchers for a reasonable period of time. The scarcity of funding prompted us to look for much needed resources in entities with no institutional relation to the project, such as publishing companies and radio stations. In addition, we have started collaborations with external research groups. We are for example initiating a project devoted to developing a wiki-based platform, to be used by the community of Asturian speakers, for loading and annotating texts in Eslema. That will benefit both our project, allowing to enlarge the corpus at a minimum cost, and the Asturian community, causing a stronger presence of Asturian in information technologies and, as a consequence, boosting the confidence of speakers in their language, which will hopefully contribute to slow down the serious process of substitution it is currently undergoing. 1. Introduction 1998. The most reliable estimates of the status and vitality We present Eslema, the first project devoted to building a of Asturian nowadays calculate the community of speak- corpus for Asturian, which is carried out by the Research ers corresponds to approximately a third of the population. Group on Asturian Philology (Seminariu de Filolox´ıa That is, there are about 300,000 speakers of Asturian in a Asturiana) within the Spanish Philology Department at community of 1,000,000 people, including about 50,000 Oviedo University (http://www.uniovi.es/eslema). Its goal speakers of Galego-Asturian concentrated in the western- is compiling a corpus representative of this Romance lan- most part of Asturias. These figures bode ill for the future guage, hence containing documents of a varied typology, of the language since Asturian competence is notably re- including both written and oral discourse, and texts from duced among young people, in such a way that the genera- the different historical periods of the Asturian language, tional transmission of the language is seriously threatened. starting from the medieval ages. In fact, there has been a 20% loss of native speakers during The name of the project, Eslema, is an Asturian word not the last decade (Llera Ramo, 2002). documented in any dictionary but in a collection of oral As- The legal status of Asturian is very different from that of turian folk tales (Suarez´ Lopez,´ 2000). Eslema refers to an other languages in Spain, or other European minority lan- impression drawn in a stone indicating the existence of a guages such as Romansh, Irish, and Frisian. The Asturias hidden treasure. Aside from its obvious poetic force, this Autonomy Statute (1981) recognizes the Asturian linguistic word also expresses the great value that a complete col- specificity in its fourth article, and establishes that “Bable lection of documents, adequately structured and organized, will enjoy protection. Its use, diffusion in the mass me- has for research in linguistic-related disciplines. dia, and teaching will be promoted, respecting in every case The paper is organized as follows: The next section offers a the local variants and the voluntary nature of its learning.” small introduction to Asturian and its current situation, and However, it does not bestow oficial status upon Asturian, section 3 presents the project in some detail. Then, section or in other words, it does not recognize the civil rights of 4 comments on the resources (both material an human) that its speakers as such. Increasing social demand for linguis- make Eslema possible. And finally, section 5 touches upon tic normalization triggered an intense political debate about the social contribution of the project. the inclusion of the official status of the language in the ref- ormation of the Autonomy Statute in 1998 (blocked by the 2. The Asturian language mainstream parties). It also led local politicians to enact Asturian (or, equivalently, Asturian-Leonese, Asturleonese, the so-called “Law for the protection of Bable/Asturian”, or Bable) is the autochthonous language of most of the ter- which explicitly includes some linguistic rights (e.g., street ritory in the Principality of Asturias, the provinces of Leon signs in Asturian, publication of administrative documen- and Zamora, in Spain, and the district of Miranda do Douro tation and official advertising brochures in Asturian, etc.) (Portugal). Nowadays it is most actively used in the As- and was unanimously approved by all political parties. So turian territory proper and in the north of Portugal, where far, however, this law’s normative intentions have not been the local version, Mirandese, achieved co-oficial status in significantly developed. One of the main linguistic demands since the 1970s has Asturian is generally seen through a Castilian perspective been the possibility of teaching Asturian at schools. So so- and considered a mere dialect, an approach that underval- cially deep-rooted is that demand that, as a matter of fact, it ues, or simply ignores, the long and solid written tradition is specifically contained in the fourth article of the Auton- that exists in that language. Sometimes its historical texts omy Statute, as mentioned above. The first school-teaching have been presented (with scarce scientific rigor) as either initiatives took place in 1984. Nowadays, the teaching of samples of ”Old Spanish” or ”dialectal literature” epiphe- Asturian has attained a remarkable expansion in primary nomena. As a matter of fact, most historical corpora for schools (86% of the public schools, with 15,227 students in Spanish include medieval Asturian-Leonese texts, adhering 2002), and a much more modest one in secondary school to no linguistic criteria other than a unitarizing discourse (20% of the institutions, with 2,171 students in that same of clear ideological and nationalistic objective, often times year). Nonetheless, Asturian is conceived as a marginal fostered by the Spanish philological tradition itself. In that subject and not as part of the core school curriculum, and it sense, documenting in a coherent way the linguistic tradi- is often taught at inconvenient or unpopular times. In fact, tion of Asturian and its corresponding historic evolution is its optional status implies that students frequently have to a priority issue. choose between studying Asturian or other subjects essen- There is yet an additional reason for compiling a historical tial for their background, such as a foreign second language corpus. Current Asturian, mainly relegated to oral linguis- or computer programming. Viejo Fernandez´ (2004) offers tic exchanges, is affected by the strong pressure of Span- a detailed view of Asturian socio-linguistic situation. ish, with which it tends to blend in an ongoing process Currently, Asturias is undergoing a new reformation of its of dissolution. Amidst this adverse situation, Asturian is statute, during which the issue of the Asturian legal sta- currently undergoing the codification process that healthier tus will probably be debated again. In spite of some slight languages (e.g., French, Spanish, English) experienced at advancements concerning the normalization of Asturian in an earlier time, mostly during the Renaissance. And for the the past few years, its officilization still appears to be a re- success of that process, the existence of historical text has mote possibility given the opposing stance of the two main proven to be crucial. As it happens, oftentimes establishing political parties. This situation is indeed of absolute rel- the autochthonous vocabulary is approached from the mere evance to Eslema. Not having a satisfactory political and intuition of intellectuals and writers, leading to unnatural legal framework seriously restricts the possibilities of an or unpopular solutions which end up being rejected by the adequate linguistic normalization and, consequently, of any speakers. Therefore, having historical text dating previous potential research on Asturian aiming at effective and prac- to the Castilianization suffered in the past half century will tical outcomes. On the other hand, this precarious situation substantially help codify and stabilize the language. What makes it even more necessary developing new strategies for is more, medieval and classical Asturian texts provide con- ensuring the use of Asturian in everyday situations, as well solidated models of specific domains and their correspond- as for boosting its prestige as a a modern language, capable ing terminology, such as administrative and juridic regis- of coping with the technology-based communication needs ters, which are fundamental when normalizing a language of present-day society. used mainly in oral contexts. 3. Eslema. A General Corpus for Asturian 3.2. Subcorpora 3.1. Project goal 3.2.1. Corpus of Medieval Asturian-Leonese Eslema was conceived as a framework to develop several It contains several hundreds of texts in medieval Asturian subcorpora –hence the title General Corpus for the Asturian and Leonese (two dialectal variations of the same lan- Language (Corpus xeneral de la llingua asturiana), which guage), all of them of legal typology. These documents is part of the project’s official name. Our initial goal is are not available online yet, given that they are still pend- the construction of three of these subcorpora: the Corpus ing annotation. Currently, we are working on the tag set of Medieval Asturian-Leonese (including documents from design. the 13th-15th centuries), the Corpus of Classical Literature Aside from specific issues related to medieval legal lan- (between 1639 and 1950), and the Corpus of Present-day guage (e.g., what is an adequate textual typology, etc.), we Asturian. need to address two additional problems.