View metadata, citation and similar papers at core.ac.uk brought to you by CORE provided by Repositório Comum Annotating COMPARA, a Grammar-aware Parallel Corpus Diana Santos, Susana Inácio Linguateca, www.linguateca.pt, Oslo node at SINTEF ICT SINTEF ICT, Pb 124 Blindern, N-0314 Oslo, Norway
[email protected],
[email protected] Abstract In this paper we describe the annotation of COMPARA, currently the largest post-edited parallel corpora which includes Portuguese. We describe the motivation, the results so far, and the way the corpus is being annotated. We also provide the first grounded results about syntactical ambiguity in Portuguese. Finally, we discuss some interesting problems in this connection. syntactic function (not revised). The annotation was 1. Introduction automatically performed by the PALAVRAS parser COMPARA (www.linguateca.pt/COMPARA/) is a large developed by Eckhard Bick, followed by a thorough parallel corpus based on a collection of intellectual revision which has resulted in parser Portuguese-English and English-Portuguese literary improvements and new versions being applied to the same source texts and translations, which has been developed material and to new additions to COMPARA. Work on and post-edited ever since 1999 (Frankenberg-Garcia & annotating the English side of COMPARA, using the Santos, 2003). COMPARA has been designed with a view CLAWS PoS tagger (Rayson & Garside, 1998) is just to be an aid in language learning, translation training, starting. contrastive and monolingual linguistic research and Let us give the readers a flavour of what COMPARA language engineering. offers as new search functionalities: it is possible to look This paper has two aims: first, to present for the first time for the part of speech distribution of forms known to be the syntactic annotation of COMPARA and its intellectual ambiguous between grammatical categories, as well as revision (or post-edition), after its automatic annotation select concordances of only one grammatical with PALAVRAS (Bick, 2000) and a post-processing interpretation.