Semantics of body parts in African WordNet: a case of Northern Sotho

Mampaka Lydia Mojapelo University of Department of African Languages

[email protected]

Abstract involved. The premise in building African Wordnets is to model it on the Princeton This paper presents a linguistic account of the structure while staying true to the African lexical semantics of body parts in African context. WordNet, with special reference to Northern Sotho. It focuses on external human body Among the challenges that were encountered in parts synsets in Northern Sotho. The paper the process of building African Wordnets was seeks to support the effectiveness of African that some of the synsets extracted from WordNet as a resource for services such as in Princeton for development of African WordNet the healthcare and medical field in South Africa. It transpired from this exploration that did not make immediate sense for African there is either a one-to-one correspondence or languages and the African context, for a number some form of misalignment of lexicalisation of reasons. For example, among them are with regard to the sample of examined synsets for concepts that are geographically synsets. The paper concludes by making distant from the South African context, such as suggestions on how African WordNet can animal and plant species. This situation would deal with such semantic misalignments in result in non-lexicalised concepts. Some non- order to improve its efficiency as a resource lexicalised concepts were left blank and for for the targeted purpose. some it was decided that available linguistic resources would be used for coinage and 1 Introduction borrowing. The envisaged convenience of African WordNet became clearer to the writer African WordNet is a project that aims to build (a linguist, project translator or lexicographer) a lexical database for all indigenous official through other synsets of a more general nature languages of South Africa, which will be linked that were easy to work with. One of the to one another. It is modelled on Princeton 1 semantic domains that was considered generally WordNet through the expand approach applicable to any context was Anatomy, (Vossen, 1998). The approach was informed by BodyPart. It was assumed that this kind of experiences shared by earlier Wordnets such as domain would have relatively fewer gaps BalkaNet, MultiWordNet, and other languages compared to domains that are geographically or in the EuroWordNet, to name but a few. The culturally more restricted. BodyPart also ranks expand approach takes synonym sets (synsets) ninth among the 50 most frequently suggested from Princeton WordNet, with their relations, upper merged ontologies (SUMOs) in Princeton and convert them into the target language. The WordNet (PWN), as at 2014-03-11. approach already lends the development of African Wordnets to the use of more than one The downside of BodyPart was that the synsets language, that is, English and the target extracted from Northern Sotho showed that language concerned. African WordNet is further none of the synsets done so far were aimed at internally multilingual with five out of nine the human anatomy. The SUMO_BodyPart official African languages of South Africa that consisted of words that were unrelated to are currently part of the project. Northern Sotho 2 humans, such as ‘scale’ (as in fish-scale), (Sesotho sa Leboa) is one of the languages ‘shell’, ‘paw’, ‘feather’ and ‘wool’. Other examples to illustrate unrelatedness to humans 1 is that the senses of the word seatla ‘hand’ were 2 http://wordnet.princeton.edu Cf. Guthrie’s zone S30 limited to Domain_Transport, SUMO_Device and Domain_Factotum, SUMO_Constant other concepts through semantic relations such Quantity, and denotation to parts of the human as hyponymy and meronymy, in the case of body did not feature. Similarly, the senses of nouns. African WordNet will further link the leoto ‘leg’ were limited to Domain_Factotum, languages spoken in the country to each other. SUMO_Shape-Attribute and Domain_Zoology, SUMO_Mammal, which is a different synset from Domain_Anatomy, SUMO_BodyPart. 2 About the body parts lexicon in This paper was premised on the understanding Northern Sotho that, comparatively speaking, non-human body parts and other domains mentioned here may Since the available body-parts synsets in the not demonstrate the immediate and direct Northern Sotho Wordnet were deemed not societal impact of African WordNet to the immediately useful for human healthcare and extent that may be achieved with human body medicine purposes, the writer considered parts. exploring external human body parts, which will later be followed by internal ones to complete the South Africa is a multilingual and multicultural healthcare and medical intent. A list was drawn, country. According to the latest South African verified and augmented against Northern Sotho statistics (Statistics South Africa, 2012) on the Language Board (1988) as well as Ziervogel & use of home languages only 9,6% of the general Mokgokong (1975) and a paper in progress on population speak English as their home verbs expressing physical pain. The list had language (L1), while the majority speak the Northern Sotho and English equivalents. Already other ten official languages and their dialects as when giving equivalents outside WordNet it L1. The remainder (>90%) speak English either emerged that there may be misalignment in the as a second, third or fourth language or not at form of general-specific lexicalisation of senses. all. Among this vast majority are healthcare For example, Northern Sotho uses the same word workers, medical students and practitioners, as for ‘finger’ and ‘toe’. Unless the difference is well as individuals and communities who readily apparent from the context a descriptive should receive healthcare and medical services. phrase is used for ease denotation. The question Another issue is that studies incidental to most is: How big is the misalignment and how are we academic qualifications in South Africa are going to solve the problem linguistically? The presented through the medium of English, sample used here is used as an index of which inevitably means that most students learn misalignments, as well as possible solutions, for through a foreign medium. For some English the rest of the development of the Northern Sotho schooling starts before they have duly mastered Wordnet. The next step was to match the body their L1. This apparent disadvantage is balanced parts on the list with English synsets. by the foundation laid in English, which will give the student a significant headstart in his or her academic career, still with insufficient knowledge of his or her L1. L1 English 3 Lexical entries in Northern Sotho speakers on the other hand are not motivated to Wordnet learn other languages until they have completed In keeping with Princeton the lexical entries in their studies and happen to find themselves in African WordNet are guided by information such an occupational environment where they have to as part of speech (POS), domain, SUMO, adjust to a different language medium. It may definition, usage and the English ID. This paper therefore be useful to provide a multilingual focuses on the Northern Sotho nouns under the platform for accessing domain lexicons on a Domain_Anatomy, SUMO_BodyPart. According level that is more than just a dictionary. to the definition and usage provided in English as Terminology lists and glossaries are being well as the ID, only body parts that are developed for various purposes in South Africa, specifically human were picked out. Fellbaum including healthcare and medicine, but none of (1998) contends that although the majority of these is an African language Wordnet. African lexicalised concepts are shared among languages, WordNet will not only provide definitions and not every language will have words denoting contextual usages of words, but will be based on concepts that are lexicalised in other languages. synsets. Synsets are sets of lexicalisations of a Therefore it is expected that some concepts may particular concept, and WordNet links them to be lexicalised in English and not in Northern mechanism may prove to be more practical than Sotho, and vice versa. It is deemed necessary for another. For the purpose of this paper it is this semantic domain to have as many lexicalised assumed that Northern Sotho may be differently concepts as possible, given the envisaged use in resourced, given the object to explore how the the healthcare sector. The paper will also look project can try to solve extant misalignment into these semantic relations and ensure that the challenges without losing the Princeton structure Northern Sotho synsets are presented in a manner while remaining true to the African context, a that is not misconstrued. manoeuvre requiring a certain amount of fineness. Lexicalisation is defined as realisation of meaning in a single word or morpheme where words are already present in a language, as well as the addition of new words as new concepts 4 Queries and results enter the languages in due course. The addition of new words involves strategies of word formation To begin, the items on the list were queried from such as compounding, derivation and borrowing. the English dictionary in DEBVisDic (WordNet Another issue to lexicalisation is some level of editor and browser). Only sense 1 of acceptability among the speakers of a language, SUMO_BodyPart under Domain_Anatomy was which will lead to general acceptability. The selected. The definitions, usages and synset IDs body-parts synsets in Northern Sotho reflect were used to obtain correct matches. General different types of lexicalisation, including personal knowledge of Northern Sotho, as a addition of new words by the strategies mother tongue speaker, was complemented and mentioned above. There are also cases of non- verified against the Northern Sotho-English lexicalisation which have yet to be resolved. bilingual and Northern Sotho-English- trilingual dictionaries. The results gained from Although the expand approach has proved to be the queries confirmed some degree of most expedient for new wordnets, lexicalisation misalignment between Northern Sotho and challenges are inevitable for most of them. For English. Clearly no comment is required on the example, in building the Konkani WordNet from one-to-one matches. The examples used here Hindi WordNet (Walawalikar et. al 2010), which represent one-to-many and many-to-one is a closely related language, some challenges mappings as well as lexicalisation gaps. were experienced. The challenges also involved the English source and they include linking A sample of words representing 88 Northern errors, missing entries, definitions, concept Sotho concepts, with English equivalents, was misalignment and lexicalisation. The issue of used. The list is not exhaustive, but it is a fair culture-specificity is also reported as one of the representation of external human body parts. causes of misalignment. In dealing with Also, not all possible connections have been alignment in the Hebrew WordNet, which was indicated in the illustrations. While the initial also built on the expand approach; Ordan and focus was on external body parts, parts of the oral Winter (2007) distinguish between contingent cavity were included as they are too close to the and systematic instances of non-equivalence. external facial body-parts and not as concealed as The two cases attest to the fact that lexicons of other internal body-parts. The English different languages mirror misalignments of both equivalents of the Northern Sotho words on the cultural and internal language structural nature. list were browsed and their IDs noted in order that their definitions and usages establish correct Vincze and Almási (2014) also treat matches. lexicalisation challenges encountered in dealing with the Hungarian WordNet. The intention of Queried senses in English (anatomy, human body this paper is not to reinvent the wheel but to learn part) were not found for the following words: from others’ experiences in the realisation that languages may be dissimilarly resourced, materially and structurally. Northern Sotho is a head Bantu language of the Niger-Congo language family, which is agglutinating with productive big hair morphology. Therefore one lexicalisation type or hair on arms and legs protruding forehead Therefore the issue of misalignment is not only a matter of lexical items, but of concepts as well. eye ridge The following diagrams provide reference for the cheek current discussion. For every Northern Sotho lexical item, an English translation equivalent is tongue provided. For combined connections, refer to adam’s apple appendix 1.

below the buttock (where the thigh starts)

back of hand

back

back of knee

foot

heel

When queried, the relevant senses of the words above could not be matched with the IDs found in DEBVisDic. A peculiar gap in English on human body parts relates to ‘head’, ‘cheek’, ‘tongue’, ‘adam’s apple’, ‘back’, ‘foot’ and ‘heel’. It is assumed that the rest of the words may be more physiologically or culturally relevant in Northern Sotho than in English. While it is still peculiar to some extent that ‘back’ was not found because physiologically, especially in Diagram 1: Arm connections the healthcare and medical context, the concept should have the same denotative significance in both languages, the gap was understood in the context of possible cultural dissimilarities. Mokokotlo ‘back’, as in the ‘back part of the human torso’, is one of the most recognisable lexical items in Northern Sotho due to what the concept represents. It is the part of the body that a baby or toddler is carried and strapped on for guaranteed safety and protection. In this context the back is culturally associated with care, nurturing, raising, acceptance and protection. The concept (and therefore the word) is culturally significant. With regard to setšhitšhi ‘big hair’ (not the same as ‘long hair’, which would be natural in the English lexicon) the gap in English is understood to be due to physiological difference.

Halliday et. al. (2004) explicate at length problems of cross-language mapping even for concepts that seem simple such as kinship terms. The examples of siblings and cousins between Diagram 2: Leg connections English and Australian Pitjantjatjara resonate with Northern Sotho and other . contexts (Fellbaum 1998). Meronymy is explained by Croft and Cruse (2004) as a sense relation between meanings rather than between individual entities, that is when the meaning of one word is part of the meaning of another. The word for ‘hand’ in Northern Sotho is seatla. It expresses the same concept expressed by [POS: n ID: ENG 20-05246212-n BCS: 3], which is sense 1 of the Domain_ Anatomy, SUMO_Bodypart and defined in English as “the (prehensile) extremity of the superior limb”. Letsogo is Northern Sotho for ‘arm’ [POS: n ID: ENG 20- 05245410-n BCS: 3], Arm: 1, defined in English as “a human limb; technically part of the superior

limb between the shoulder and the elbow but Diagram 3: Torso connections commonly used to refer to the whole superior limb”. In Northern Sotho letsogo refers to the whole superior limb, which includes the hand. According to the definition provided above the common usage of the English ‘arm’ is the same as the Northern Sotho letsogo, but the technical usage is not. In Northern Sotho the word letsogo is also used to refer to seatla ‘hand’, but the whole limb is never called seatla. That is, while seatla ‘hand’ is a meronym of letsogo ‘arm’, the two are also synonymous. Similarly leoto ‘leg’ [ENG20-05242579-n] is used for both ‘leg’ and ‘foot’ while a separate specific word for ‘foot’ is lenao. These examples illustrate lexicalisation that reflects the occurrence of meronymy between lexical items that are also synonymous.

Another scenario relates to the case of monwana for both ‘toe’ [ENG20-05258265-n] and ‘finger’ [ENG20-05247839-n], and ntši for ‘eyebrow’ [ENG20-05007503-n] and ‘eyelash’ [ENG20-

05008887-n]. In this case Northern Sotho uses Diagram 4: Head connections one word to express separate concepts, or concepts that are viewed as separate in English. These two examples illustrate that the words 4.1 One-to-many and many-to-one monwana and ntši are used in Northern Sotho as hypernyms. Descriptive phrases ‘of the foot’ and Two types of misalignment will be used for ‘of the hand’ are used as hyponyms of monwana illustration here. There are cases of Northern in cases where distinction is deemed necessary. A Sotho lexicalisation of human body parts that similar descriptive strategy is not used for ntši; it mingle synonymy and meronymy, not in a may also be cumbersome as both ‘eyebrow’ and confusing way though. In the context of WordNet ‘eyelash’ belong to the eye. words are synonymous if they express the same concept and can be interchanged in some 4.2 Possible non-lexicalisation in English the form of descriptive phrases to distinguish ‘finger’ and ‘toe’. The descriptions wa lenao and Another concept that is lexicalised in Northern wa leoto ‘of the foot’; wa seatla and wa letsogo Sotho but could not be found from querying the ‘of the hand’ are consistent with language usage English in DEBVisDic is nyaraga (Mokgokong and are not expected to pose any problems. The and Ziervogel 1975), also pronounced nyarago. same solution cannot work in the case of ntši The English trees relating to ‘leg’ and ‘buttock’ since eyebrow and eyelash are both ‘of the eye’. were examined as the concept is understood to be Northern Board (1988) uses either a body part below the buttock or the compounding as a strategy to distinguish the two. uppermost back part of the leg. Its absence in the While they are both ntši the source coined two trees pointed to possible non-lexicalisation. ntšikgolo as additional lexicalisation for ‘eyebrow’. The second component of the The following section proposes possible compound -kgolo (-golo) ‘big’ suggests that an linguistic means of catering for the misalignment eyebrow is dominant. The source was produced issues mentioned above in African WordNet. by a standardising body (Northern Sotho Language Board) which was obviously cognisant 5 Handling misalignments of the gaps in terms of lexicalisation. They It is necessary to provide linguistic solutions to probably considered either the overaching the misalignment challenges mentioned above. position of the eyebrow in relation to the Vincze and Almási (2014) suggest a number of eyelashes or the perceived amount of hair in both, strategies for the Hungarian lexicalisation issues, to come up with a suggestion that an eyebrow is namely to shorten the tree, flatten the tree, the main ntši. Another example of compounding restructure the tree and lexicalize the concepts. from the same source is khurumelakhuru for They are also of the opinion that the merge ‘kneecap’. -khurumela is a verb stem which approach would have alleviated some of the means to close or to cover. Khuru is ‘knee’. challenges. For Konkani Walawalikar et. al Therefore conceptualisation points to something (2010) suggest, among others, that the target that covers, closes off or protects the knee. language synsets for which there were gaps in the Lexicalisation strategies such as these provide source language could be used to fill the gaps, promising resources for African WordNet. What thereby strengthening the HWN. Ordan and remains is whether or not such lexical items will Winter (2007) detail strategies for building filter down to everyday usage. Hebrew synsets, which include linking Hebrew The last issue relates to the apparent English non- word senses to related PWN sysnsets from lexicalisation of concepts that are lexicalised in Hebrew to English and from English to Hebrew. Northern Sotho, and vice versa. Nyaraga ‘below Lexical gaps from both sides are acknowledged the buttocks’ is part of the Northern Sotho and used to preserve and link semantic lexicon whose lexicalisation could not be information. ascertained in English. The English equivalent is This paper takes a linguistic view to addressing provided in Northern Sotho dictionaries as a the challenges mentioned above, which relate to phrase. The English lexicalisation of the Northern lexicalisation of the concepts. The first group of Sotho ntahle ‘back of hand’ could also not be Northern Sotho words which could not be ascertained. Over and above being a body part, matched from English seem to be a matter of part of a hand, ntahle has an added connotation misses which can be addressed if probed further. relating to slapping (backhand slap). That is, The next situation concerns seatla ‘hand’ and slapping someone with the inner part of a hand lenao ‘foot’ which are meronyms of letsogo and the outer part of a hand would be reflected by ‘arm’ and leoto ‘leg’, respectively, and proved to the use of different lexical items. Such words be synonymous as well. Therefore lexical items need to be added as they represent concepts that seatla and letsogo will be in the same synset are intertwined with the idiom of the language. while they are meronymically related as well. An expected scenario of the expand approach The same applied to lenao ‘foot’ and leoto ‘leg’. where English is the source language would The next issue concerns monwana ‘finger’ and obviously reveal Northern Sotho non- ‘toe’ and ntši ‘eyelash’ and ‘eyebrow’. In the lexicalisation of concepts that are lexicalised in language synonyms for monwana are provided in English. With regard to the domain under discussion descriptive phrases are common, for were different types of lexical misalignment example ‘nose’ is nko and ‘nostril’ is lešoba la presented, but also lexicalisation mechanisms nko, literally ‘hole of nose’. ‘Pubis’ is lerapo la that are used in the language. While the pele la noka, literally ‘bone of front of waist’. mentioned mechanisms may be grammatically Another lexicalisation mechanism that is sound and fill lexicalisation gaps, the words also productive in Bantu languages, which was need to receive general acceptability to the point nevertheless not observed in the current sample, of being in reasonably high frequency use rather is derivation. Affixes are used productively to than merely existing. form words from different word categories. Direct borrowing is also not evident in the It is envisaged that the proposed strategies will current sample, but it is commonly used in the fill the gaps, and that inclusion of internal body lexicalisation of technological concepts and parts and functions, as well as verbs of specific disease names. From this sample an expressing physical pain will produce trees that example of indirect borrowing is evident in mirror the language. It remains to be seen how coinage that resembles the English formations far the translators in the project will go in such as khurumelakhuru above and moropana wa utilising the lexicalisation strategies mentioned in tsebe literally ‘small drum of ear’ for ‘eardrum’. this paper. To assist with acceptability and Lexicalisation mechanisms that were employed standardisation the synsets will also be shared for this sample hint at linguistic routes to follow with selected practitioners in the target field for in dealing with further development of human comments. body parts.

Acknowledgement 6 Challenges Ms Marissa Griesel, for support with the While the linguistic side of the project may prove illustrations exciting, there are challenges of an IT nature. The challenges include changes in the IT infrastructure at the hosting institutions, as well as problems with the DEBVisDic editor. Such References challenges hamper the development of the Christiane Fellbaum (Ed). 1998. Wordnet: an wordnets, as they result in interrupted access to electronic lexical database. Cambridge, Mass: the server and inconsistent functionality of the The MIT Press. editor. This becomes a challenge if one wants to browse and edit existing synsets, or add new Dirk Ziervogel and Pothinus C. Mokgokong. 1975. synstes. Nonetheless, manual and semi-automatic Pukuntšu ye kgolo ya Sesotho sa Leboa/ data gathering methods are used so that when a Comprehensive Northern Sotho dictionary/ permanent IT solution is reached there is enough Groot Noord-Sotho woordeboek. Pretoria: J. L. linguistic data to fast-track the development of Van Schaik. the wordnets. M.A.K Halliday, Wolfgang Teubert, Colin Yallop and Anna Čermáková. 2004. Lexicology and Corpus Linguistics: an introduction. London: Continuum. 7 Conclusion Noam Ordan and Shuly Wintner. 2007. Hebrew The paper presented actual and possible scenarios WordNet: a test case of aligning lexical that may pose challenges when developing the databases across languages. International Northern Sotho Wordnet on Domain_Anatomy, Journal of Translation 19(1):39-58. SUMO_BodyPart. Human body parts are targeted in this paper due to their connection to Northern Sotho Language Board. 1988. Sesotho sa human health care and medicine. Many speakers Leboa Mareo le Mongwalo No. 4/ Northern Sotho Terminology and Orthography No. 4/ whose L1 is not Northern Sotho may benefit Noord-Sotho Terminologie en Spelreëls No. 4. from the database as it will be linking Northern Pretoria: Government Printer. Sotho not only to English but to other South African indigenous languages as well. Not only Piek Vossen (ed.) 1998. EuroWordNet:A Multilingual Database with Lexical Semantic Networks. Kluwer Academic Publishers, Dordrecht.

Shantaram Walawalikar, Shilpa Desai, Ramdas Karmali, Sushant Naik, Damodar Ghanekar, Chandralekha D'Souza and Jyoti Pawar. 2010. Experiences in Building the Konkani WordNet Using the Expansion Approach. In Proceedings of the Fifth Global WordNet Conference, January 2010. Mumbai, India.

Statistics South Africa http://www.stassa.gov.za accessed on 19 August 2015

Verinika Vincze and Attila Almási (2014) In Proceedings of the Seventh Global WordNet Conference, January 2014. University of Tartu, Estonia.

William Croft and D. Alan Cruse 2004. Cognitive linguistics. Cambridge: University Press.

Appendix 1: Combined connections