Information Retrieval and Knowledge Organization: a Perspective from the Philosophy of Science
Total Page:16
File Type:pdf, Size:1020Kb
information Article Information Retrieval and Knowledge Organization: A Perspective from the Philosophy of Science Birger Hjørland Department of Communication, University of Copenhagen, 8 Karen Blixens Plads, DK-2300 Copenhagen S, Denmark; [email protected] Abstract: Information retrieval (IR) is about making systems for finding documents or information. Knowledge organization (KO) is the field concerned with indexing, classification, and representing documents for IR, browsing, and related processes, whether performed by humans or computers. The field of IR is today dominated by search engines like Google. An important difference between KO and IR as research fields is that KO attempts to reflect knowledge as depicted by contemporary scholarship, in contrast to IR, which is based on, for example, “match” techniques, popularity measures or personalization principles. The classification of documents in KO mostly aims at reflecting the classification of knowledge in the sciences. Books about birds, for example, mostly reflect (or aim at reflecting) how birds are classified in ornithology. KO therefore requires access to the adequate subject knowledge; however, this is often characterized by disagreements. At the deepest layer, such disagreements are based on philosophical issues best characterized as “paradigms”. No IR technology and no system of knowledge organization can ever be neutral in relation to paradigmatic conflicts, and therefore such philosophical problems represent the basis for the study of IR and KO. Keywords: information retrieval; knowledge organization; philosophy of science; classification; knowledge organization systems; ontologies; Kuhnian paradigm theory; pragmatism Citation: Hjørland, B. Information Retrieval and Knowledge Organization: A Perspective from the Philosophy of Science. Information 1. Introduction 2021, 12, 135. https://doi.org/ Information retrieval (IR) and knowledge organization (KO) are two research fields 10.3390/info12030135 that, on the one hand, are separate fields of study, but on the other hand, have the same aim: to facilitate the findability of documents, knowledge, and information. Anderson Academic Editor: Claudio Gnoli and Pérez-Carballo’s [1] handbook of KO has the title Information Retrieval Design, which indicates the close connection between KO and IR, where IR is about search processes, Received: 23 February 2021 while KO is about designing optimal structures for IR (in addition to other purposes KO Accepted: 8 March 2021 Published: 20 March 2021 may serve). Today, the field of IR has mainly migrated from information science to computer Publisher’s Note: MDPI stays neutral science and has developed systems used in search engines such as Google, which are with regard to jurisdictional claims in extremely successful. There is, however, a surprisingly simple objection to the underlying published maps and institutional affil- principles in mainstream IR research: they are not based on scientific or scholarly norms on iations. which documents or knowledge claims have the best scientific or scholarly foundation and correspond to our best theories and findings. It seems obvious that users need to retrieve what is regarded as true knowledge (or to be more precise, what is considered our best substantiated knowledge claims). The main approaches in IR are discussed in Section3. In this respect, the field of KO has often had the goal, at the least implicitly, to Copyright: © 2021 by the author. Licensee MDPI, Basel, Switzerland. classify and represent documents and knowledge in accordance with updated scientific This article is an open access article knowledge (e.g., to classify books about birds in accordance with the classification of birds distributed under the terms and in ornithology). Despite this aim, the theory of KO has lacked an adequate toolset of conditions of the Creative Commons conceptualization to cope with it (and often, approaches have dominated that do not deal Attribution (CC BY) license (https:// adequately with the problems. This is discussed in Section2). creativecommons.org/licenses/by/ The purpose of this article is to provide arguments and conceptualization for creating 4.0/). methods and principles to develop systems and processes in IR and KO that aim at Information 2021, 12, 135. https://doi.org/10.3390/info12030135 https://www.mdpi.com/journal/information Information 2021, 12, 135 2 of 25 providing users with our best substantiated knowledge. To do so, this article argues, we need to base IR on KO, and we need to base KO on insights derived from the philosophy of science. 2. The Field of Knowledge Organization (KO) KO is sometimes termed “information organization” and [2] (p. 391) found that “information organization” is now by far the most used name for courses in educational programs in North American ALA-accredited programs. It is also sometimes termed “knowledge representation”, as in the subtitle of the official journal of ISKO, Knowledge Organization: “International Journal devoted to Concept Theory, Classification, Indexing and Knowledge Representation”. KO is about knowledge organizing processes such as indexing, tagging, classifying, describing, and organizing documents and information, and about knowledge organization systems (KOS) such as classification systems, thesauri, and ontologies. KO has a practical aim. Although it is also a theoretical field, the aim of both practical activities and empirical and theoretical research is to support practical activities. The practical aims are the sole justification for KO. It has been claimed, for example, that KOS such as thesauri and controlled vocabularies are obsolete today, cf., [3]. KO must show that its processes and systems are necessary. If, for example, search engines such as Google can fulfill all practical needs without relying on KO, there is no need for this field. KO must always consider itself in relation to all possible alternative approaches, whether they have been developed inside or outside the field itself. Specifically, we must demonstrate, why, for example, Google is not enough (cf., [4]). This article contains arguments about the necessity of KO. The KOS developed for organizing documents and information are primarily about organizing concepts. The conceptual structures used in KOS are to a large degree derived from specific domains of knowledge. Documents about birds, for example, are often organized according to how ornithologists organize the birds themselves. This principle was recognized by Henry Bliss [5], who argued (p. 37): “To make the [bibliographic] classification conform to the scientific and educational organization of knowledge is to make it the more practical”, and this is probably the main reason the field was and still is named KO. The roots of KO lie in the following: (1) the practical classification and indexing of books and other kinds of documents in libraries and bibliographical databases; (2) philosophical principles, including Aristotle’s logic and Francis Bacon’s classification of knowledge, among many others; (The reader should be warned, however, that there are many misunderstandings about philosophical issues, including Aristotle’s role in classification. What is commonly attributed to Aristotle is a myth (see, e.g., [6], Chapter 2: “The Aristotelian Framework”). (3) scientific and scholarly contributions, including, for example, the contributions of Aristotle, Carl Linnaeus, Charles Darwin and many other scientists to the classification of living organisms and all other things in the world; (4) developments in information technology (IT), such as databases, communication networks and social media. The connections between these four fields are important for the development of KO as a field of study and practice but have often been neglected. This article argues that subject knowledge and its foundation in the philosophy of science is the most important perspective for KO, but this perspective has so far not been very influential in KO. Alternatively, influential perspectives have been: • To consider KO an intuitive process that does not need deeper justification. The original classification of journals in the citation indexes published by the Institute for Scientific Information, for example, was just intuitive (cf., [7], p. 602). In practice, many KO activities have been driven by this view: you simply start constructing a classification and stop when it seems to suit the purpose. A criticism of this perspective Information 2021, 12, 135 3 of 25 may state that any classification is always serving some interests at the expense of other interests, and if unanalyzed, it cannot be optimized for the purpose it is going to serve (cf., [8]). Of course, this view challenges the whole idea of KO as a needed field of study. • To claim that since only the individual user knows their own “information need”, they are the only person qualified to set principles of what should be found in IR and how documents should be indexed or classified. This is the view underlying user-oriented and cognitive approaches in information science and KO and is discussed a little in the present article and more detailed in [9]. • To say that the best way, both economically and qualitatively, is to base KO on user tagging or related “social technologies”, implying that a deeper theoretical under- standing of KO is unnecessary. User tagging is by some both seen as “democratic” and economically preferrable, while there are also critical voices (for a discussion, see [10]). • To claim that KO is basically a technological problem,