Managing Order Relations in Mlbibtex∗
Total Page:16
File Type:pdf, Size:1020Kb
Managing order relations in MlBibTEX∗ Jean-Michel Hufflen LIFC (EA CNRS 4157) University of Franche-Comté 16, route de Gray 25030 BESANÇON CEDEX France hufflen (at) lifc dot univ-fcomte dot fr http://lifc.univ-fcomte.fr/~hufflen Abstract Lexicographical order relations used within dictionaries are language-dependent. First, we describe the problems inherent in automatic generation of multilin- gual bibliographies. Second, we explain how these problems are handled within MlBibTEX. To add or update an order relation for a particular natural language, we have to program in Scheme, but we show that MlBibTEX’s environment eases this task as far as possible. Keywords Lexicographical order relations, dictionaries, bibliographies, colla- tion algorithm, Unicode, MlBibTEX, Scheme. Streszczenie Porządek leksykograficzny w słownikach jest zależny od języka. Najpierw omó- wimy problemy powstające przy automatycznym generowaniu bibliografii wielo- języcznych. Następnie wyjaśnimy, jak są one traktowane w MlBibTEX-u. Do- danie lub zaktualizowanie zasad sortowania dla konkretnego języka naturalnego umożliwia program napisany w języku Scheme. Pokażemy, jak bardzo otoczenie MlBibTEX-owe ułatwia to zadanie. Słowa kluczowe Zasady sortowania leksykograficznego, słowniki, bibliografie, algorytmy sortowania leksykograficznego, Unikod, MlBibTEX, Scheme. 0 Introduction these items w.r.t. the alphabetical order of authors’ Looking for a word in a dictionary or for a name names. But the bst language of bibliography styles in a phone book is a common task. We get used [14] only provides a SORT function [13, Table 13.7] to the lexicographic order over a long time. More suitable for the English language, the commands for precisely, we get used to our own lexicographic or- accents and other diacritical signs being ignored by der, because it belongs to our cultural background. this sort operation. It depends on languages. This problem is particu- The purpose of this article is to show how this larly acute when we deal with managing multilin- problem is solved in MlBibTEX’s first public release. In practice, this version deals only with European gual bibliographies, as in our program MlBibTEX— languages using the Latin alphabet. The MlBibT X for ‘MultiLingual BibTEX’. Let us recall that this E program is written using the Scheme programming program aims to be a ‘better’ BibTEX [15], the bibli- language [10]. A standardised library providing sup- ography processor usually associated with the LATEX word processor [12]. When it builds a ‘References’ port for Unicode [22] has been designed [18, §§ 1.1 & 1.2], but we cannot say that the present version section for a LATEX document, BibTEX uses a bib- liography style to determine the layout. Some bib- of Scheme is Unicode-compliant, even if some imple- 1 liography styles are unsorted, that is, the order of mentations include partial support. So some parts bibliographical items within the bibliography is the of our present implementation of order relations are order of first citations of these items throughout the temporary, but we think that this implementation document. However, most of BibTEX’s styles sort 1 At the time of finishing the revised version of this article, the proposal for Scheme’s next standard has just been ratified ∗ Title in Polish: Zarządzanie zasadami sortowania lek- and is now the ‘official’ sixth version of this language [19, 18]. sykograficznego w MlBIBTEX-u. See http://www.r6rs.org for more details. TUGboat, Volume 29, No. 1 — XVII European TEX Conference, 2007 101 Jean-Michel Hufflen • The Czech alphabet is: a < b < c < č < d < ... < h < ch < i < ... < r < ř < s < š < t < . z < ž. • In Danish, accented letters are grouped at the end of the alphabet: a < ... < z < æ < ø < a. • The Estonian language does not use the same order for unaccented letters as the usual Latin order; in addition, accented letters are either inserted into the alphabet or alphabeticised like the corresponding unaccented letter: a < ... < s ∼ š < z ∼ ž < t < ... < w < õ < ä < ö < ü < x < y. • Here are the accented letters in the French language: à ∼ â, ç, è ∼ é ∼ ê ∼ ë, î ∼ ï, ô, ù ∼ û ∼ ü, ÿ. When two words differ by an acccent, the unaccented letter takes precedence, but w.r.t. a right-to-left order:a cote < côte < coté < côté. Two ligatures are used: ‘æ’ (resp. ‘œ’), alphabeticised like ‘ae’ (resp. ‘oe’). • There are three accented letters in German — ‘ä’, ‘ö’, ‘ü’ — and three lexicographic orders: – DINb-1: a ∼ ä, o ∼ ö, u ∼ ü; – DIN-2: ae ∼ ä, oe ∼ ö, ue ∼ ü; – Austrian: a < ä < ... < o < ö < ... < u < ü < v < ... < z. • The Hungarian alphabet is: a ∼ á < b < c < cs < d < dz < dzs < e ∼ é < f < g < gy < h < i ∼ í < j < k < l < ly < m < n < ny < o ∼ ó < ö ∼ ő < p < ... < s < sz < t < ty < u ∼ ú < ü ∼ ű < v < ... < z < zs Some double digraphs must be restored before sorting: ccs ! cs+cs, ddz ! dz+dz, ggy ! gy+gy, lly ! ly+ly, nny ! ny+ny, ssz ! sz+sz, tty ! ty+ty The same for the double trigraph: ddzs ! dzs+dzs. • The Polish alphabet is: a < ą < b < c < ć < d < e < ę < ... < l < ł < m < n < ń < o < ó < p < ... < s < ś < t < ... < z < ż • The Romanian alphabet is: a < ă < â < b < ... < i < î < j < . s < ş < t < ţ < u < ... < z. • The Slovak alphabet is: a < á < ä < b < c < č < d < ď < dz < dž < e < é < f < g < h < ch < i < í < j < k < l < ĺ < ľ < m < n < ň < o < ó < ô < p < q < r < ŕ < s < š < t < ť < u < ú < ... < y < ý < z < ž • The Spanish alphabet was a < b < c < ch < d < ... < l < ll < m < n < ñ < o < ... < z until 1994. Now the digraphs ‘ch’ and ‘ll’ are no longer viewed as single letters in modern dictionaries, and the words using ‘ñ’ are interleaved with words using ‘n’. • In Swedish, accented letters are grouped at the end of the alphabet: a < ... < z < a < ä < ö. a Using a left-to-right order for this step is common mistake even for French people. But the accurate order is right-to-left, as specified in [7]. b Deutsche Institut für Normung (German Institute of normalisation). Figure 1: Some order relations used in European languages. could be easily updated for future versions dealing programmer should be able to specify a new order with Unicode. relation. We give more technical details in an an- In the first section, we show how diverse lex- nex, for users that would like to experiment more icographic orders used throughout European coun- themselves. In particular, we explain how to deal tries are. This allows readers to estimate this diver- with languages using the Latin 2 encoding, despite sity and to realise the complexity of this task. We our implementation being based on Latin 1. also explain why this problem is made more diffi- cult when we consider it within the framework of 1 European languages and bibliographies. Then we show how order relations lexicographic orders operate in MlBibTEX and how they are built. We Figure 1 gives an idea of the diversity of order re- also give some details about the common and differ- lations used throughout some European countries. ◦ ent points between xındy [13, § 11.3] and MlBibTEX, In this figure, ‘a < b’ denotes that the words be- both being programs using multilingual order rela- ginning with a are ‘less than’ the words beginning tions. Reading this article does not require advanced with b, whereas ‘a ∼ b’ expresses that the letters knowledge of Scheme;2 in fact, we think that a non- a and b are interleaved, except that a takes prece- dence over b if two words differ only by these two let- 2 Readers can refer to [20] for an introductory book about ters. Roughly speaking, there are two families of lan- Scheme. 102 TUGboat, Volume 29, No. 1 — XVII European TEX Conference, 2007 Managing order relations in MlBibTEX guages in the realm of associated lexicographic or- can be decomposed into these ‘simpler’ characters: ders. Accented letters are sometimes fully viewed as latin small letter o, U+006F ‘real’ letters, distinct from unaccented ones: exam- combining circumflex accent, U+0302 ples are given by some Slavonic languages. In other The sort algorithm requires several passes. To de- languages, accented letters are sorted as if there scribe it roughly, an information about weight, given were no accent. The precedence of a unaccented by sort keys, is associated with each string. Then letter over an accented one is determined in various this information is re-arranged according to sort lev- ways: it follows a left-to-right order in Irish, Ital- els, w.r.t. letters, w.r.t. accents, etc. Finally, a binary ian, and Portuguese, a right-to-left order in French. comparison between bytes is done, level by level, un- The Estonian language ‘mixes’ the two approaches: til the two strings can be distinguished. This algo- some accented letters — ‘õ’, ‘ä’ — are alphabeticised, rithm can be refined for a particular language, by some — ‘š’, ‘ž’ — are interleaved. Last, some letter using a specialised sort key table, possibly including groups may be viewed as a single letter and alpha- sort keys for accented letters and digraphs viewed as beticised as another letter. For example, the Hun- single letters. This modus operandi would be diffi- garian words beginning with ‘cs’ are alphabeticised cult to put into action within MlBibT X. First, we separately from the words beginning with ‘c’. In E do not have complete support for Unicode:4 for ex- fact, the ‘c-’ entry in a Hungarian dictionary con- ample, we cannot directly deal with characters such tains words beginning with ‘c’ and not with ‘cs’. as the ‘combining circumflex accent’, not included The ‘c-’ entry is followed by the ‘cs-’ entry, before in the Latin-1 encoding.