On Faithfulness and Factuality in Abstractive Summarization

On Faithfulness and Factuality in Abstractive Summarization Joshua Maynez∗ Shashi Narayan∗ Bernd Bohnet Ryan McDonald Google Research fjoshuahm,shashinarayan,bohnetbd,[email protected] Abstract understanding how maximum likelihood training and approximate beam-search decoding in these It is well known that the standard likelihood models lead to less human-like text in open-ended training and approximate decoding objectives text generation such as language modeling and in neural text generation models lead to less human-like responses for open-ended tasks story generation (Holtzman et al., 2020; Welleck such as language modeling and story gener- et al., 2020; See et al., 2019). In this paper we ation. In this paper we have analyzed limi- investigate how these models are prone to gener- tations of these models for abstractive docu- ate hallucinated text in conditional text generation, ment summarization and found that these mod- specifically, extreme abstractive document summa- els are highly prone to hallucinate content that rization (Narayan et al., 2018a). is unfaithful to the input document. We con- Document summarization — the task of produc- ducted a large scale human evaluation of sev- eral neural abstractive summarization systems ing a shorter version of a document while preserv- to better understand the types of hallucinations ing its information content (Mani, 2001; Nenkova they produce. Our human annotators found and McKeown, 2011) — requires models to gener- substantial amounts of hallucinated content in ate text that is not only human-like but also faith- all model generated summaries. However, our ful and/or factual given the document. The exam- analysis does show that pretrained models are ple in Figure1 illustrates that the faithfulness and better summarizers not only in terms of raw factuality are yet to be conquered by conditional metrics, i.e., ROUGE, but also in generating text generators. The article describes an event faithful and factual summaries as evaluated by humans. Furthermore, we show that textual en- of “Conservative MP Zac Smith winning the pri- tailment measures better correlate with faith- mary for 2016 London mayoral election”, but sum- fulness than standard metrics, potentially lead- maries often forge entities (e.g., “Nigel Goldsmith” ing the way to automatic evaluation metrics as or “Zac Goldwin”) or information (e.g., “UKIP well as training and decoding criteria.1 leader Nigel Goldsmith”, “Nigel Goldsmith winning the mayoral election”, “Sadiq Khan being the 1 Introduction former London mayor” or “Zac Goldwin being the Current state of the art conditional text generation Labour’s candidate”) that are not supported by the models accomplish a high level of fluency and co- document or are factually wrong. Interestingly, all arXiv:2005.00661v1 [cs.CL] 2 May 2020 herence, mostly thanks to advances in sequence- summaries are topical and fluent, and perform well to-sequence architectures with attention and copy in terms of ROUGE scores (Lin and Hovy, 2003). (Sutskever et al., 2014; Bahdanau et al., 2015; Gu We conducted a large-scale human evaluation et al., 2016), fully attention-based Transformer ar- of hallucinated content in systems that use Re- chitectures (Vaswani et al., 2017; Dai et al., 2019) current Neural Network (RNN) (See et al., 2017), and more recently pretrained language modeling Convolutional Neural Network (CNN) (Narayan for natural language understanding (Devlin et al., et al., 2018a), and Transformers (Radford et al., 2019; Radford et al., 2018; Yang et al., 2019; Liu 2019; Rothe et al., 2020), as well as human et al., 2019). There has been a growing interest in written summaries for the recently introduced eXtreme SUMmarization task (XSUM, Narayan ∗ The first two authors contributed equally. et al., 2018a). We seek to answer the following 1Our human annotated summaries for faithfulness and factuality will be released at https://github.com/google-research- questions: (i) How frequently do abstractive sum- datasets/xsum hallucination annotations. marizers hallucinate content?; (ii) Do models hal- GOLD Zac Goldsmith will contest the 2016 London mayoral election for the Conservatives, it has been announced. DOCUMENT: The Richmond Park and North Kingston MP said he was ”honoured” after winning 70% of the 9,227 votes cast using an online primary system. He beat London Assembly Member Andrew Boff, MEP Syed Kamall and London’s deputy mayor for crime and policing Stephen Greenhalgh. Mr Goldsmith’s main rival is likely to be Labour’s Sadiq Khan. (2 sentences with 59 words are abbreviated here.) Mr Goldsmith, who was the favourite for the Tory nomination, balloted his constituents earlier this year to seek permission to stand. At the very point of his entry into the race for London mayor, Zac Goldsmith’s decision revealed two big characteristics. (5 sentences with 108 words are abbreviated here.) Mr Goldsmith - who first entered Parliament in 2010 - told the BBC’s Daily Politics that he hoped his environmental record would appeal to Green and Lib Dem voters and he also hoped to ”reach out” to UKIP supporters frustrated with politics as usual and the UK’s relationship with the EU. Zac Goldsmith Born in 1975, educated at Eton and the Cambridge Centre for Sixth-form Studies (5 sentences with 76 words are abbreviated here.) Mr Goldsmith, who has confirmed he would stand down from Parliament if he became mayor, triggering a by-election, said he wanted to build on current mayor Boris Johnson’s achievements. (3 sentences with 117 words are abbreviated here.) Both Mr Khan and Mr Goldsmith oppose a new runway at Heathrow airport, a fact described by the British Chambers of Commerce as ”depressing”. (1 sentences with 31 words is abbreviated here.) Current mayor Boris Johnson will step down next year after two terms in office. He is also currently the MP for Uxbridge and South Ruislip, having been returned to Parliament in May. Some Conservatives have called for an inquiry into the mayoral election process after only 9,227 people voted - compared with a 87,884 turnout for the Labour contest. (4 sentences with 121 words are abbreviated here.) PTGEN UKIP leader Nigel Goldsmith has been elected as the new mayor of London to elect [45.7, 6.1, 28.6] a new Conservative MP. TCONVS2S Former London mayoral candidate Zac Goldsmith has been chosen to stand in the [50.0, 26.7, 37.5] London mayoral election. TRANS2S Former London mayor Sadiq Khan has been chosen as the candidate to be the next [35.3, 12.5, 23.5] mayor of London. GPT-TUNED Conservative MP Zac Goldwin’s bid to become Labour’s candidate in the 2016 [42.4, 25.8, 36.4] London mayoral election. BERTS2S Zac Goldsmith has been chosen to contest the London mayoral election. [66.7, 40.0, 51.9] Figure 1: Hallucinations in extreme document summarization: the abbreviated article, its gold summary and the abstractive model generated summaries (PTGEN, See et al. 2017; TCONVS2S, Narayan et al. 2018a; and, GPT- TUNED,TRANS2S and BERTS2S, Rothe et al. 2020) for a news article from the extreme summarization dataset (Narayan et al., 2018a). The dataset and the abstractive models are described in Section3 and4. We also present the [ROUGE-1, ROUGE-2, ROUGE-L]F1 scores relative to the reference gold summary. Words in red correspond to hallucinated information whilst words in blue correspond to faithful information. lucinate by manipulating the information present argue that large-scale pretrained models are merely in the input document (intrinsic hallucinations) or better at learning data-specific regularities (Niven by adding information not directly inferable from and Kao, 2019), at least on in-domain summa- the input document (extrinsic hallucinations)?; (iii) rization the gains in automatic metrics are real- How much hallucinated content is factual, even ized in observable differences by humans. Fourth, when unfaithful?; and (iv) Are there automatic ROUGE (Lin and Hovy, 2003) and BERTScore means of measuring these hallucinations? (Zhang et al., 2020) correlates less with faithfulness/factuality than metrics derived from automatic Our main conclusions are as follows: First, semantic inference systems, specifically the degree intrinsic and extrinsic hallucinations happen fre- to which a summary is entailed by the source docu- quently – in more than 70% of single-sentence sum- ment. This presents an opportunity for improved maries. Second, the majority of hallucinations are automatic evaluation measures as well as model extrinsic, which potentially could be valid abstrac- training and decoding objectives. We show prelim- tions that use background knowledge. However, inary experiments in this direction. our study found that over 90% of extrinsic hallucinations were erroneous. Thus, hallucinations hap- 2 Hallucinations in Summarization pen in most summaries and the majority of these are neither faithful nor factual. Third, models ini- Open-ended generation — the task of generating tialized with pretrained parameters perform best text that forms a natural continuation from the input both on automatic metrics and human judgments of text — requires the model to hallucinate text; hence faithfulness/factuality. Furthermore, they have the the focus has been to ensure that the model learns highest percentage of extrinsic hallucinations that to generate text that is more human-like (i.e., less are factual. This suggests that while some studies repetitive or dull with more content-related words) (Holtzman et al., 2020; Welleck et al., 2020; See extrinsic hallucinations; these terms are not intro- et al., 2019). In contrast, tasks such as document duced in the document. A model with a poorly- summarization (Nenkova and McKeown, 2011; See informed decoder and that is agnostic to the di- et al., 2017; Paulus et al., 2018) and data-to-text vergence issue between the source and target texts generation (Lebret et al., 2016; Wiseman et al., (Wiseman et al., 2017; Dhingra et al., 2019), will 2017) which are not open-ended, require models to function more as an open-ended language model be factual and/or faithful to the source text.

On Faithfulness and Factuality in Abstractive Summarization

Details

Download

Copyright

We respect the copyrights and intellectual property rights of all users. All uploaded documents are either original works of the uploader or authorized works of the rightful owners.

Support