2 the Corpus As Object: Design and Purpose

2 The corpus as object: Design and purpose As corpora have become larger and more diverse, and as they are more frequently used to make de®nitive statements about language, issues of how they are designed have become more important. Four aspects of corpus design are discussed in this chapter: size, content, representativeness and permanence. The chapter also summarises some types of corpus investigation, each of which treats the corpus as a different kind of object. Issues in corpus design Size As computer technology has advanced since the 1960s, so it has become feasible to store and access corpora of ever-increasing size. Whereas the LOB Corpus and Brown Corpus seemed as big as anyone would ever want, at the time, nowadays 1 million words is fairly small in terms of corpora. The British National Corpus is 100 million words; the Bank of English is currently about 400 million. CANCODE is 5 million words. The feasible size of a corpus is not limited so much by the capacity of a computer to store it, as by the speed and ef®ciency of the access software. If, for example, counting the number of past and present tense forms of the verb BE in a given corpus takes longer than a few minutes, the researcher may prefer to use a smaller corpus whose results might be considered to be just as reliable but on which the software would work much more speedily. Sheer quantity of information can be overwhelming for the observer. In the Bank of English corpus, for example, a search for the word point gives almost 143,000 hits. Few researchers can obtain useful information simply by looking at so many concordance lines. One solution is to use a smaller corpus, which will give less data. Another is to keep the large corpus, but to use software which will make a selection of data from the whole, either by selecting at random a proportion of the total concordance lines, or by identifying and allowing selection of the most signi®cant collocates, or other 25 26 Corpora in applied linguistics signi®cant features. In this way, the data from a very large corpus can be reduced to a manageable scale whilst retaining the advantages of coverage of the large corpus. The question of corpus size can be a contentious one. There are many arguments to support the view that a small corpus can be valuable under certain circumstances. Carter and McCarthy 41995: 143), for example, argue that for the purposes of studying grammar in spoken language a relatively small corpus is suf®cient 4because grammar words tend to be very frequent). Some kinds of corpus investigation that depend on the manual annotation of the corpus are also, necessarily, restricted in terms of the size of the corpus that can feasibly be annotated. Similarly, many commercially available concordance software packages restrict the size of the corpus they can be used with 4see Biber et al 1998: 285±286 for details). Someone aiming to build a balanced corpus may restrict the amount of data from one source in order to match data from another source 4for example, the amount of written data may be kept smaller than it need be because it is dif®cult to obtain large amounts of spoken data). A more extreme view is that for some purposes there is an optimum corpus size that should not be exceeded even if practical considerations would allow it, that is, that a relatively small corpus is not just suf®cient but also necessary. However, this argument usually refers to the dif®culties of processing large amounts of information about very frequent words, as discussed above 4e.g. Carter and McCarthy 1995: 143). The opposing view would be that it is preferable to select from a large amount of data than to restrict the amount of data available 4Sinclair 1992). Arguments about optimum corpus size tend to be academic for most people. Most corpus users simply make use of as much data as is available, without worrying too much about what is not available. As well as the very large, general corpora designed to assist in writing dictionaries and other reference books, there are thousands of smaller corpora around the world, some comprising only a few thousand words and designed for a particular piece of research. Content It is a truism that a corpus is neither good nor bad in itself, but suited or not suited to a particular purpose. Decisions about what should go into a corpus are based on what the corpus is going to be used for, but also about what is available. A researcher who wishes to study The corpus as object: Design and purpose 27 academic articles in History, for example, may plan to build a corpus consisting simply of published articles. Even with such a straight- forward design plan, however, a number of issues will have to be resolved, such as: . Will the articles come from one journal, or from a range of journals? How will the journal4s) be selected? . Will the articles chosen be restricted to those apparently written by native speakers of the target language, or not? . Will articles other than research articles, such as book reviews, be included? . If the articles are to be stored electronically, and copyright clearance is required, which publishers are willing to give permis- sion for their articles to be used? . If the articles are required in electronic form, which publishers are willing to make electronic versions of their journals available? To the ®rst three of these questions, there are no `right' answers. The best choice in each case will depend on the precise purpose of the research. The last two questions refer to purely pragmatic issues, but these may in the end affect the design of the corpus more profoundly than the ®rst three. Publishers' policies are not the only issues affecting availability for corpora. Parallel corpora, for example, are composed of texts that are translations of each other. They can make use only of texts that have been translated between the two languages involved. Other issues of selection come second to the question of what translations exist. Using unpublished material in a corpus does not obviate the issue of availability. For example, a researcher wishing to study student writing may ask students to submit electronic versions of their essays to assist in this work. If this submission is voluntary, the researcher will collect only the essays whose authors are willing for their essays to be used. For some purposes, corpus design is scarcely an issue. For example, if a corpus is being used to encourage learners to investigate language data for themselves, the precise contents of that corpus may be relatively unimportant. Any collection of newspapers written in the target form of the language, for example, will be adequate for some investigation purposes, even if the newspapers do not represent all aspects of language use 4see, for example, Dodd 1997). Where the corpus is being taken as representative of a language or a language variety, the notion of design and balance is very important. This will be considered separately below. 28 Corpora in applied linguistics Balance and representativeness Corpora are very often intended to be representative of a particular kind of language. If the object of study is academic prose, or casual conversation, or the language of newspapers, or American English, for example, an attempt must be made to build a corpus that is representative of the whole. 4See Kennedy 1991: 52; Crowdy 1993; Burnard ed. 1995 for discussions of representativeness in the British National Corpus.) Usually, this involves breaking the whole down into component parts and aiming to include equal amounts of data from each of the parts. For example, under `the language of newspapers', it may be decided to include a range of newspaper types 4broadsheet and tabloid, for instance) and a range of article types 4hard news, human interest, editorials, letters, sports, business, advertisements and so on). A balanced corpus might be said to consist of equal numbers of words in each category: broadsheet hard news; tabloid hard news; broadsheet human interest; tabloid human interest and so on. This notion of balance is, however, open to question. It might be argued that broadsheet newspapers contain more words, on average, than tabloids, and that therefore the corpus should contain more from the broadsheets than from the tabloids. Alternatively, it might be argued that tabloids are read by more people than broadsheets and that therefore they should comprise a larger part of the corpus. If the broadsheets contain more hard news and editorial than the tabloids, and the tabloids contain more human interest and sports news, should the proportions in the corpus be adjusted accordingly? In the case of newspapers, the solution may be to include all issues of a selection of publications from a given week, month or year. This will allow the proportions to determine themselves. It does mean, however, that more data will probably be gathered from one type of publication than another. Similar problems with relation to the collection of academic prose or casual conversation, for example, are not open to so easy a solution. The problem is that `being representative' inevitably involves knowing what the character of the `whole' is. Where the proportions of that character are unknowable, attempts to be representative tend to rest on little more than guesswork. The problem of representation and balance becomes more dif®cult when the corpus is supposed to represent a regional variety of English 4Singapore English, or British English, or Australian English), with all its complexity of internal variation.

2 the Corpus As Object: Design and Purpose

Corpus Studies in Applied Linguistics

Download (237Kb)

1. Introduction

Distributed Memory Bound Word Counting for Large Corpora

Unit 3: Available Corpora and Software

Lexicography Between NLP and Linguistics: Aspects of Theory and Practice

Types of Corpora and Some Famous (English) Examples

Intégration De Verbnet Dans Un Réalisateur Profond

The Spoken British National Corpus 2014

The BT Correspondence Corpus 1853- 1982: the Development and Analysis of Archival Language Resources

The Corpus of Contemporary American English As the First Reliable Monitor Corpus of English

An Analysis of Factive Cognitive Verb Complementation Patterns Used by Elt Students