Arxiv:2106.11410V2 [Cs.CL] 15 Jul 2021 Inferiority of Racialized People’S Language Practices Underrepresented in STEM and Academia

Total Page:16

File Type:pdf, Size:1020Kb

Load more

A Survey of Race, Racism, and Anti-Racism in NLP Anjalie Field Su Lin Blodgett Carnegie Mellon University Microsoft Research [email protected] [email protected] Zeerak Waseem Yulia Tsvetkov University of Sheffield University of Washington [email protected] [email protected] Abstract While researchers and activists have increasingly Despite inextricable ties between race and lan- drawn attention to racism in computer science and guage, little work has considered race in NLP academia, frequently-cited examples of racial bias research and development. In this work, we in AI are often drawn from disciplines other than survey 79 papers from the ACL anthology that NLP, such as computer vision (facial recognition) mention race. These papers reveal various (Buolamwini and Gebru, 2018) or machine learn- types of race-related bias in all stages of NLP ing (recidivism risk prediction) (Angwin et al., model development, highlighting the need for 2016). Even the presence of racial biases in search proactive consideration of how NLP systems engines like Google (Sweeney, 2013; Noble, 2018) can uphold racial hierarchies. However, per- sistent gaps in research on race and NLP re- has prompted little investigation in the ACL com- main: race has been siloed as a niche topic munity. Work on NLP and race remains sparse, and remains ignored in many NLP tasks; most particularly in contrast to concerns about gender work operationalizes race as a fixed single- bias, which have led to surveys, workshops, and dimensional variable with a ground-truth label, shared tasks (Sun et al., 2019; Webster et al., 2019). which risks reinforcing differences produced In this work, we conduct a comprehensive sur- by historical racism; and the voices of histor- vey of how NLP literature and research practices ically marginalized people are nearly absent in NLP literature. By identifying where and how engage with race. We first examine 79 papers from NLP literature has and has not considered race, the ACL Anthology that mention the words ‘race’, especially in comparison to related fields, our ‘racial’, or ‘racism’ and highlight examples of how work calls for inclusion and racial justice in racial biases manifest at all stages of NLP model NLP research practices. pipelines (§3). We then describe some of the limi- 1 Introduction tations of current work (§4), specifically showing that NLP research has only examined race in a nar- Race and language are tied in complicated ways. row range of tasks with limited or no social context. Raciolinguistics scholars have studied how they are Finally, in §5, we revisit the NLP pipeline with a fo- mutually constructed: historically, colonial pow- cus on how people generate data, build models, and ers construct linguistic and racial hierarchies to are affected by deployed systems, and we highlight justify violence, and currently, beliefs about the current failures to engage with people traditionally arXiv:2106.11410v2 [cs.CL] 15 Jul 2021 inferiority of racialized people’s language practices underrepresented in STEM and academia. continue to justify social and economic exclusion While little work has examined the role of race 1 (Rosa and Flores, 2017). Furthermore, language in NLP specifically, prior work has discussed race is the primary means through which stereotypes in related fields, including human-computer in- and prejudices are communicated and perpetuated teraction (HCI) (Ogbonnaya-Ogburu et al., 2020; (Hamilton and Trolier, 1986; Bar-Tal et al., 2013). Rankin and Thomas, 2019; Schlesinger et al., However, questions of race and racial bias 2017), fairness in machine learning (Hanna et al., have been minimally explored in NLP literature. 2020), and linguistics (Charity Hudley et al., 2020; 1We use racialization to refer the process of “ascribing and Motha, 2020). We draw comparisons and guid- prescribing a racial category or classification to an individual ance from this work and show its relevance to NLP or group of people . based on racial attributes including but not limited to cultural and social history, physical features, research. Our work differs from NLP-focused re- and skin color” (Charity Hudley, 2017). lated work on gender bias (Sun et al., 2019), ‘bias’ generally (Blodgett et al., 2020), and the adverse U.S. However, as race and racism are global con- impacts of language models (Bender et al., 2021) structs, some aspects of our analysis are applicable in its explicit focus on race and racism. to other contexts. We suggest that future studies In surveying research in NLP and related fields, on racialization in NLP ground their analysis in the we ultimately find that NLP systems and research appropriate geo-cultural context, which may result practices produce differences along racialized lines. in findings or analyses that differ from our work. Our work calls for NLP researchers to consider the social hierarchies upheld and exacerbated by 3 Survey of NLP literature on race NLP research and to shift the field toward “greater 3.1 ACL Anthology papers about race inclusion and racial justice” (Charity Hudley et al., 2020). In this section, we introduce our primary survey data—papers from the ACL Anthology3—and we 2 What is race? describe some of their major findings to empha- size that NLP systems encode racial biases. We It has been widely accepted by social scientists that searched the anthology for papers containing the race is a social construct, meaning it “was brought terms ‘racial’, ‘racism’, or ‘race’, discarding ones into existence or shaped by historical events, social that only mentioned race in the references section forces, political power, and/or colonial conquest” or in data examples and adding related papers cited rather than reflecting biological or ‘natural’ differ- by the initial set if they were also in the ACL An- ences (Hanna et al., 2020). More recent work has thology. In using keyword searches, we focus on criticized the “social construction” theory as circu- papers that explicitly mention race and consider lar and rooted in academic discourse, and instead papers that use euphemistic terms to not have sub- referred to race as “colonial constituted practices”, stantial engagement on this topic. As our focus including “an inherited western, modern-colonial is on NLP and the ACL community, we do not in- practice of violence, assemblage, superordination, clude NLP-related papers published in other venues exploitation and segregation” (Saucier et al., 2016). in the reported metrics (e.g. Table1), but we do The term race is also multi-dimensional and draw from them throughout our analysis. can refer to a variety of different perspectives, in- Our initial search identified 165 papers. How- cluding racial identity (how you self-identify), ob- ever, reviewing all of them revealed that many do served race (the race others perceive you to be), not deeply engage on the topic. For example, 37 and reflected race (the race you believe others per- papers mention ‘racism’ as a form of abusive lan- ceive you to be) (Roth, 2016; Hanna et al., 2020; guage or use ‘racist’ as an offensive/hate speech Ogbonnaya-Ogburu et al., 2020). Racial catego- label without further engagement. 30 papers only rizations often differ across dimensions and depend mention race as future work, related work, or mo- on the defined categorization schema. For exam- tivation, e.g. in a survey about gender bias, “Non- ple, the United States census considers Hispanic binary genders as well as racial biases have largely an ethnicity, not a race, but surveys suggest that been ignored in NLP” (Sun et al., 2019). After 2/3 of people who identify as Hispanic consider discarding these types of papers, our final analysis 2 it a part of their racial background. Similarly, set consists of 79 papers.4 the census does not consider ‘Jewish’ a race, but Table1 provides an overview of the 79 papers, some NLP work considers anti-Semitism a form manually coded for each paper’s primary NLP task of racism (Hasanuzzaman et al., 2017). Race de- and its focal goal or contribution. We determined pends on historical and social context—there are task/application labels through an iterative process: no ‘ground truth’ labels or categories (Roth, 2016). listing the main focus of each paper and then col- As the work we survey primarily focuses on the lapsing similar categories. In cases where papers United States, our analysis similarly focuses on the 3The ACL Anthology includes papers from all official 2https://www:census:gov/mso/ ACL venues and some non-ACL events listed in AppendixA, www/training/pdf/race-ethnicity- as of December 2020 it included 6; 200 papers onepager:pdf/, https://www:census:gov/ 4We do not discard all papers about abusive language, only topics/population/race/about:html, ones that exclusively use racism/racist as a classification label. https://www:pewresearch:org/fact-tank/ We retain papers with further engagement, e.g. discussions 2015/06/15/is-being-hispanic-a-matter- of how to define racism or identification of racial bias in hate of-race-ethnicity-or-both/ speech classifiers. Detect Bias Debias Total Collect Corpus Analyze Corpus Develop Model Survey/Position Abusive Language 6 4 2 5 2 2 21 Social Science/Social Media 2 10 6 1 - 1 20 Text Representations (LMs, embeddings) - 2 - 9 2 - 13 Text Generation (dialogue, image captions, story gen. ) - - 1 5 1 1 8 Sector-specific NLP applications (edu., law, health) 1 2 - - 1 3 7 Ethics/Task-independent Bias 1 - 1 1 1 2 6 Core NLP Applications (parsing, NLI, IE) 1 - 1 1 1 - 4 Total 11 18 11 22 8 9 79 Table 1: 79 papers on race or racism from the ACL anthology, categorized by NLP application and focal task. could rightfully be included in multiple categories, vealing under-representation in training data, some- we assign them to the best-matching one based on times tangentially to primary research questions: stated contributions and the percentage of the paper Rudinger et al.(2017) suggest that gender bias may devoted to each possible category.
Recommended publications
  • Ai Governance in 2019 a Year in Review

    Ai Governance in 2019 a Year in Review

    The report editor can be reached at [email protected] We welcome any comments on this report and any communication related to AI AI GOVERNANCE IN 2019 governance. A YEAR IN REVIEW OBSERVATIONS OF 50 GLOBAL EXPERTS SHANGHAI INSTITUTE FOR SCIENCE OF SCIENCE April, 2020 Shanghai Institute for Science of Science ALL LIVING THINGS ARE NOURISHED WITHOUT INJURING ONE ANOTHER, AND ALL ROADS RUN PARALLEL WITHOUT INTERFERING WITH ONE ANOTHER. CHUNG YUNG, SECTION OF THE LI CHI TABLE OF CONTENTS FOREWORD China Initiative: Applying Long-Cycle, Multi-Disciplinary Social Experimental on Exploring 21 the Social Impact of Artificial Intelligence By SHI Qian By SU Jun Going Beyond AI Ethics Guidelines 23 INTRODUCTION 01 By Thilo Hagendorff By LI Hui and Brian Tse 25 PART 1 TECHNICAL PERSPECTIVES FROM WORLD-CLASS SCIENTISTS 07 By Petra Ahrweiler 27 The Importance of Talent in the Information Age 07 By John Hopcroft The Impact of Journalism 29 From the Standard Model of AI to Provably Beneficial Systems 09 By Colin Allen By Stuart Russell and Caroline Jeanmaire Future of Work in Singapore: Staying on Task 31 The Importance of Federated Learning 11 By Poon King Wang By YANG Qiang Developing AI at the Service of Humanity 33 Towards A Formal Process of Ethical AI 13 By Ferran Jarabo Carbonell By Pascale Fung Enhance Global Cooperation in AI Governance on the Basis of Further Cultural Consensus 35 From AI Governance to AI Safety 15 By WANG Xiaohong By Roman Yampolskiy Three Modes of AI Governance 37 By YANG Qingfeng PART 2 INTERDISCIPLINARY ANALYSES FROM PROFESSIONAL RESEARCHERS 17 PART 3 RESPONSIBLE LEADERSHIP FROM THE INDUSTRY 39 The Rapid Growth in the Field of AI Governance 17 By Allan Dafoe & Markus Anderljung Companies Need to Take More Responsibilities in Advancing AI Governance 39 Towards Effective Value Alignment in AI: From "Should" to "How" 19 By YIN Qi By Gillian K.
  • Our Digital Future in a Divided World 22Nd APRU ANNUAL JUNE 24-26, 2018 PRESIDENTS’ MEETING National TAIWAN UNIVERSIT Y

    Our Digital Future in a Divided World 22Nd APRU ANNUAL JUNE 24-26, 2018 PRESIDENTS’ MEETING National TAIWAN UNIVERSIT Y

    Our Digital Future in a Divided World 22nd APRU ANNUAL JUNE 24-26, 2018 PRESIDENTS’ MEETING NATIONAL TAIWAN UNIVERSIT Y MEETING REPORT APRU Members Australia Korea Australian National University KAIST The University of Melbourne Korea University The University of Sydney POSTECH UNSW Sydney Seoul National University Yonsei University Canada The University of British Columbia Malaysia University of Malaya Chile University of Chile Mexico Tecnológico de Monterrey China and Hong Kong SAR Fudan University New Zealand Nanjing University The University of Auckland Peking University The Chinese University of Hong Kong Philippines The Hong Kong University of Science and Technology University of the Philippines The University of Hong Kong Tsinghua University Russia University of Chinese Academy of Sciences Far Eastern Federal University University of Science and Technology of China Zhejiang University Singapore Nanyang Technological University Taiwan National University of Singapore National Taiwan University National Tsing Hua University Thailand Chulalongkorn University Indonesia University of Indonesia United States of America California Institute of Technology Japan Stanford University Keio University University of California, Berkeley Nagoya University University of California, Davis Osaka University University of California, Irvine Tohoku University University of California, Los Angeles The University of Tokyo University of California, San Diego Waseda University University of California, Santa Barbara University of Hawaiʻi at Mãnoa University
  • The Black Box, Unlocked PREDICTABILITY and UNDERSTANDABILITY in MILITARY AI

    The Black Box, Unlocked PREDICTABILITY and UNDERSTANDABILITY in MILITARY AI

    The Black Box, Unlocked PREDICTABILITY AND UNDERSTANDABILITY IN MILITARY AI Arthur Holland Michel UNIDIR Acknowledgements Support from UNIDIR’s core funders provides the foundation for all the Institute’s activities. This study was produced by the Security and Technology Programme, which is funded by the Governments of Germany, the Netherlands, Norway and Switzerland, and by Microsoft. The author wishes to thank the study’s external reviewers, Dr. Pascale Fung, Lindsey R. Sheppard, Dr. Martin Hagström, and Maaike Verbruggen, as well as the numerous subject matter experts who provided valuable input over the course of the research process. Design and layout by Eric M. Schulz. Note The designations employed and the presentation of the material in this publication do not imply the expression of any opinion whatsoever on the part of the Secretariat of the United Nations concerning the legal status of any country, territory, city or area, or of its authorities, or concerning the delimitation of its frontiers or boundaries. The views expressed in the publication are the sole responsibility of the individual authors. They do not necessary reflect the views or opinions of the United Nations, UNIDIR, its staff members or sponsors. Citation Holland Michel, Arthur. 2020. ‘The Black Box, Unlocked: Predictability and Understand- ability in Military AI.’ Geneva, Switzerland: United Nations Institute for Disarmament Research. doi: 10.37559/SecTec/20/AI1 About UNIDIR The United Nations Institute for Disarmament Research (UNIDIR) is a voluntarily fund- ed, autonomous institute within the United Nations. One of the few policy institutes worldwide focusing on disarmament, UNIDIR generates knowledge and promotes dialogue and action on disarmament and security.
  • 1 Mona T. Diab

    1 Mona T. Diab

    Mona T. Diab, PhD Associate Professor Department of Computer Science School of Engineering and Applied Science George Washington University [email protected] http://www.seas.gwu.edu/~mtdiab Office: +1(202) 994.8109 RESEARCH FOCUS & INTERESTS Computational lexical semantics, multilingual processing, computational sociolinguistics, computational pragmatics, social media analytics, health analytics, low resource language processing, resource building, applied machine learning techniques, text analytics, information extraction, sentiment and emotion analysis, Arabic computational linguistics. PROFESSIONAL EXPERIENCE 01.2013-present Associate Professor, Department of Computer Science, The George Washington University, Washington DC, USA 01.2013-present Director, GW NLP Lab (CARE4Lang), The George Washington University, Washington DC, USA (~20 active members) 06.2005-present Co-Director, Computational Approaches for Arabic Dialect Modeling (CADIM) Group, Columbia University, The George Washington University, NYU-Abu Dhabi (~10 active members) 09.2009-12.2012 Research Scientist (Principal Investigator), Center for Computational Learning Systems (CCLS), Columbia University, New York NY, USA 09.2009-12.2012 Adjunct Associate Professor, Department of Computer Science, Columbia University, New York NY, USA 09.2007-08.2009 Adjunct Assistant Professor, Department of Computer Science, Columbia University, New York NY, USA 02.2005-08.2009 Associate Research Scientist (Principal Investigator), Center for Computational Learning Systems (CCLS), Columbia University,
  • The Miracle Continues School of Engineering Celebrates HKUST’S Th Anniversary

    The Miracle Continues School of Engineering Celebrates HKUST’S Th Anniversary

    Snapshot Engineering Newsletter No.20 Summer 2011 The Miracle Continues School of Engineering Celebrates HKUST’s th Anniversary Calendar of Events September 24-25, 2011 December 2011 HKUST Information Day HKUST Engineering Day HKUST Campus HKUST Campus October 14-23, 2011 December 8-11, 2011 HKUST “Bring Technology to Community” Exhibition The 3rd International Symposium on Plasticity and Impact Hong Kong Science Museum HKUST Clear Water Bay Campus (Dec 8-10) and HKUST Fok Ying Tung Graduate School, Nansha Campus (Dec 10-11) The above events are subject to change without prior notice. Don't be the Missing Link ... Editors: Diana Liu, Dorothy Yip Alumni relationships are invaluable assets to the School and Contributing Editor: Sally Course alumni. To foster the growth of our alumni network, please Address: School of Engineering keep us informed of your recent news and send us your The Hong Kong University of Science and Technology updated contact information via email to [email protected]. Clear Water Bay, Kowloon, Hong Kong Stay connected and keep in touch! Phone: (852) 2358-5917 Fax: (852) 2358-1458 Email: [email protected] Website: www.seng.ust.hk Facebook: www.facebook.com/SENG.HKUST PTC-G14143 24 Dean's Message New Appointments Research 2011 marks another milestone in the history of HKUST as the technology play in making the world a better place, from Concurrent Prof Ben Letaief Honored as University reaches its 20th year, and the School of Engineering improved healthcare to more efficient buildings and bridges to ■ Prof Mordecai Golin (SENG) is fully playing its part in commemorating the new consumer electronic breakthroughs.
  • Kathleen Mckeown

    Kathleen Mckeown

    Kathleen McKeown Department of Computer Science Columbia University New York, NY. 10027 U.S.A. Phone: 212-939-7118 Fax: 212-666-0140 email: [email protected] website: http://www.cs.columbia.edu/ kathy Current position Henry and Gertrude Rothschild Professor of Computer Science, Columbia University, New York Research Interests Computational Linguistics/Natural-Language Processing: Text Summarization; Language Genera- tion; Social Media Analysis; Open-ended question answering; Sentiment Analysis Education 1982 Ph.D., Computer and Information Science, University of Pennsylvania 1979 M.S., Computer and Information Science, University of Pennsylvania 1976 A.B., Comparative Literature, Brown University Appointments held 2012-2018 Founding Director, Columbia Data Science Institute 2011-2012 Vice Dean for Research, School of Engineering and Applied Science 2005-present Henry and Gertrude Rothschild Professor of Computer Science, Columbia University 1997-present Professor, Columbia University 7/2003-12/2003 Acting Chair, Department of Computer Science, Columbia University 1997-2002 Chair, Department of Computer Science, Columbia University 1987-1997 Associate Professor, Columbia University 1982-1987 Assistant Professor, Columbia University Honors & awards 2016 Keynote Speaker, International Semantic Web Conference, Kobe, Japan 20162014 Grace Hopper Distinguished Lecture, University of Pennsylvania Philadelphia, Pa., Nov. 2016. Distin- guished Lecturer, University of Edinburgh, Launch of the Centre for Doctoral Training in Data Science, Edinburgh,
  • Big Data Institute Educational Technology Programs Hong Kong and Nation’S Transformation Big Data Research Center

    Big Data Institute Educational Technology Programs Hong Kong and Nation’S Transformation Big Data Research Center

    ©BDI Jan 2021 Big Data Institute Educational Technology Programs Hong Kong and nation’s Transformation Big Data Research Center Funding Support Industry Applications Scientific Research Talent Training Industrialize scientific Mission Recruit achievements and talents in create long-term and academia impact to the world Vision Navigate Aspire innovation research and breakthrough development Benefit the Greater Bay Area (HK / Macau / Guangdong), Greater China Region and the world 50+ faculty interdisciplinary Schools and Departments for from 120+ students 16 collaborative R&D projects and 4 joint labs Over HK$ 80 million & US$ 1 million industrial sponsorship and donation with one of the largest government-funded ITF Smart City Project BDI WeChat Overseas Joint Lab in Asia ACHIEVEMENT st 1 Line/Naver Overseas Joint Lab in the World HIGHLIGHTS Trained 400+ BDT Master Students Partnership with leading companies and more… BDI Team At present, there are more than 50 faculty members and over 120 students involved BDI’s 16 research projects from • School of Engineering • School of Science • School of Business & Management • Department of Computer Science and Engineering • Department of Industrial Engineering and Decision Analytics • Department of Electronic and Computer Engineering, • Department of Chemical and Biological Engineering • Department of Information Systems, Business Statistics and Operations Management • Department of Mathematics • Division of Life Science • Division of Social Science Prof Lei Chen Director of BDI Prof Nancy Ip Vice-President