Statistical Challenges in Assessing and Fostering the Reproducibility of Scientific Results: Summary of a Workshop

Total Page:16

File Type:pdf, Size:1020Kb

Statistical Challenges in Assessing and Fostering the Reproducibility of Scientific Results: Summary of a Workshop THE NATIONAL ACADEMIES PRESS This PDF is available at http://www.nap.edu/21915 SHARE Statistical Challenges in Assessing and Fostering the Reproducibility of Scientific Results: Summary of a Workshop DETAILS 132 pages | 7 x 10 | PAPERBACK ISBN 978-0-309-39202-0 | DOI: 10.17226/21915 AUTHORS BUY THIS BOOK Michelle Schwalbe, Rapporteur; Committee on Applied and Theoretical Statistics; Board on Mathematical Sciences and Their Applications; Division on Engineering and Physical Sciences; FIND RELATED TITLES National Academies of Sciences, Engineering, and Medicine Visit the National Academies Press at NAP.edu and login or register to get: – Access to free PDF downloads of thousands of scientific reports – 10% off the price of print titles – Email or social media notifications of new titles related to your interests – Special offers and discounts Distribution, posting, or copying of this PDF is strictly prohibited without written permission of the National Academies Press. (Request Permission) Unless otherwise indicated, all materials in this PDF are copyrighted by the National Academy of Sciences. Copyright © National Academy of Sciences. All rights reserved. Statistical Challenges in Assessing and Fostering the Reproducibility of Scientific Results: Summary of a Workshop Statistical Challenges in Assessing and Fostering the Reproducibility of Scientifi c Results Summary of a Workshop Michelle Schwalbe, Rapporteur Committee on Applied and Theoretical Statistics Board on Mathematical Sciences and Their Applications Division on Engineering and Physical Sciences Copyright © National Academy of Sciences. All rights reserved. Statistical Challenges in Assessing and Fostering the Reproducibility of Scientific Results: Summary of a Workshop THE NATIONAL ACADEMIES PRESS 500 Fifth Street, NW Washington, DC 20001 This workshop was supported by Grant No. DMS-1351163 between the National Academies of Sci- ences and the National Science Foundation. Any opinions, findings, or conclusions expressed in this publication do not necessarily reflect the views of any organization or agency that provided support for the project. International Standard Book Number-13: 978-0-309-39202-0 International Standard Book Number-10: 0-309-39202-0 Digital Object Identifier: 10.17226/21915 Cover: The photo “Peas in a Pod” by Andrew_Writer is used with Creative Commons Attribution 2.0 Generic license (https://www.flickr.com/photos/dragontomato/3698617520/in/gallery-106928713@ N04-72157638971870524/). This report is available in limited quantities from: Board on Mathematical Sciences and Their Applications National Academies of Sciences, Engineering, and Medicine 500 Fifth Street NW Washington, DC 20001 [email protected] http://www.nas.edu/bmsa Additional copies of this workshop summary are available for sale from the National Academies Press, 500 Fifth Street, NW, Keck 360, Washington, DC 20001; (800) 624-6242 or (202) 334-3313; http://www.nap.edu/. Copyright 2016 by the National Academy of Sciences. All rights reserved. Printed in the United States of America Suggested citation: National Academies of Sciences, Engineering, and Medicine. 2016. Statistical Challenges in Assessing and Fostering the Reproducibility of Scientific Results: Summary of a Workshop. Washington, DC: The National Academies Press. doi: 10.17226/21915. Copyright © National Academy of Sciences. All rights reserved. Statistical Challenges in Assessing and Fostering the Reproducibility of Scientific Results: Summary of a Workshop The National Academy of Sciences was established in 1863 by an Act of Congress, signed by President Lincoln, as a private, nongovernmental institution to advise the nation on issues related to science and technology. Members are elected by their peers for outstand- ing contributions to research. Dr. Ralph J. Cicerone is president. The National Academy of Engineering was established in 1964 under the charter of the National Academy of Sciences to bring the practices of engineering to advising the nation. Members are elected by their peers for extraordinary contributions to engineer- ing. Dr. C. D. Mote, Jr., is president. The National Academy of Medicine (formerly the Institute of Medicine) was estab lished in 1970 under the charter of the National Academy of Sciences to advise the nation on medical and health issues. Members are elected by their peers for distinguished contribu- tions to medicine and health. Dr. Victor J. Dzau is president. The three Academies work together as the National Academies of Sciences, Engineering, and Medicine to provide independent, objective analysis and advice to the nation and conduct other activities to solve complex problems and inform public policy decisions. The Academies also encourage education and research, recognize outstanding contribu- tions to knowledge, and increase public understanding in matters of science, engineering, and medicine. Learn more about the National Academies of Sciences, Engineering, and Medicine at www.national-academies.org. Copyright © National Academy of Sciences. All rights reserved. Statistical Challenges in Assessing and Fostering the Reproducibility of Scientific Results: Summary of a Workshop Copyright © National Academy of Sciences. All rights reserved. Statistical Challenges in Assessing and Fostering the Reproducibility of Scientific Results: Summary of a Workshop PLANNING COMMITTEE ON STATISTICAL CHALLENGES IN ASSESSING AND FOSTERING THE REPRODUCIBILITY OF SCIENTIFIC RESULTS CONSTANTINE GATSONIS, Brown University, Co-Chair GIOVANNI PARMIGIANI, Dana-Farber Cancer Institute, Co-Chair STEPHEN FIENBERG, Carnegie Mellon University STEVEN N. GOODMAN, Stanford University School of Medicine JOHN H. HOLMES, University of Pennsylvania Perelman School of Medicine ALAN F. KARR, RTI International JELENA KOVAČEVIĆ, Carnegie Mellon University XIHONG LIN, Harvard University ROGER PENG, Johns Hopkins Bloomberg School of Public Health VICTORIA STODDEN, University of Illinois, Urbana-Champaign Staff MICHELLE K. SCHWALBE, Senior Program Officer RODNEY N. HOWARD, Administrative Assistant LINDA CASOLA, Senior Program Assistant and Staff Editor v Copyright © National Academy of Sciences. All rights reserved. Statistical Challenges in Assessing and Fostering the Reproducibility of Scientific Results: Summary of a Workshop COMMITTEE ON APPLIED AND THEORETICAL STATISTICS CONSTANTINE GATSONIS, Brown University, Chair DEEPAK AGARWAL, LinkedIn KATHERINE BENNETT ENSOR, Rice University MONTSERRAT (MONTSE) FUENTES, North Carolina State University ALFRED O. HERO III, University of Michigan DAVID M. HIGDON, Virginia Bioinformatics Institute ROBERT E. KASS, Carnegie Mellon University JOHN LAFFERTY, University of Chicago XIHONG LIN, Harvard University JOSÉ M. F. MOURA, Carnegie Mellon University SHARON-LISE T. NORMAND, Harvard University GIOVANNI PARMIGIANI, Dana-Farber Cancer Institute ADRIAN RAFTERY, University of Washington LANCE WALLER, Emory University EUGENE WONG, University of California, Berkeley Staff MICHELLE K. SCHWALBE, Director RODNEY N. HOWARD, Administrative Assistant LINDA CASOLA, Senior Program Assistant vi Copyright © National Academy of Sciences. All rights reserved. Statistical Challenges in Assessing and Fostering the Reproducibility of Scientific Results: Summary of a Workshop BOARD ON MATHEMATICAL SCIENCES AND THEIR APPLICATIONS DONALD SAARI, University of California, Irvine, Chair DOUGLAS N. ARNOLD, University of Minnesota JOHN B. BELL, E. O. Lawrence Berkeley National Laboratory VICKI M. BIER, University of Wisconsin, Madison JOHN R. BIRGE, University of Chicago L. ANTHONY COX, JR., Cox Associates, Inc. MARK L. GREEN, University of California, Los Angeles BRYNA KRA, Northwestern University JOSEPH A. LANGSAM, Morgan Stanley (retired) ANDREW W. LO, Massachusetts Institute of Technology DAVID MAIER, Portland State University WILLIAM A. MASSEY, Princeton University JUAN C. MEZA, University of California, Merced CLAUDIA NEUHAUSER, University of Minnesota FRED S. ROBERTS, Rutgers University GUILLERMO R. SAPIRO, Duke University CARL P. SIMON, University of Michigan KATEPALLI SREENIVASAN, New York University ELIZABETH A. THOMPSON, University of Washington Staff SCOTT T. WEIDMAN, Director NEAL GLASSMAN, Senior Program Officer MICHELLE K. SCHWALBE, Senior Program Officer RODNEY N. HOWARD, Administrative Assistant BETH DOLAN, Financial Associate vii Copyright © National Academy of Sciences. All rights reserved. Statistical Challenges in Assessing and Fostering the Reproducibility of Scientific Results: Summary of a Workshop Copyright © National Academy of Sciences. All rights reserved. Statistical Challenges in Assessing and Fostering the Reproducibility of Scientific Results: Summary of a Workshop Acknowledgment of Reviewers This workshop summary has been reviewed in draft form by individuals chosen for their diverse perspectives and technical expertise, in accordance with procedures approved by the Report Review Committee. The purpose of this independent review is to provide candid and critical comments that will assist the institution in making its published workshop summary as sound as possible and to ensure that the workshop summary meets institutional standards for objectivity, evidence, and responsiveness to the project’s charge. The review comments and draft manuscript remain confidential to protect the integrity of the study process. We wish to thank the following individuals for their review of this workshop summary: C. Glenn Begley, TetraLogic Pharmaceuticals, Joel Greenhouse, Carnegie Mellon University, Joelle
Recommended publications
  • Cebm-Levels-Of-Evidence-Background-Document-2-1.Pdf
    Background document Explanation of the 2011 Oxford Centre for Evidence-Based Medicine (OCEBM) Levels of Evidence Introduction The OCEBM Levels of Evidence was designed so that in addition to traditional critical appraisal, it can be used as a heuristic that clinicians and patients can use to answer clinical questions quickly and without resorting to pre-appraised sources. Heuristics are essentially rules of thumb that helps us make a decision in real environments, and are often as accurate as a more complicated decision process. A distinguishing feature is that the Levels cover the entire range of clinical questions, in the order (from top row to bottom row) that the clinician requires. While most ranking schemes consider strength of evidence for therapeutic effects and harms, the OCEBM system allows clinicians and patients to appraise evidence for prevalence, accuracy of diagnostic tests, prognosis, therapeutic effects, rare harms, common harms, and usefulness of (early) screening. Pre-appraised sources such as the Trip Database (1), (2), or REHAB+ (3) are useful because people who have the time and expertise to conduct systematic reviews of all the evidence design them. At the same time, clinicians, patients, and others may wish to keep some of the power over critical appraisal in their own hands. History Evidence ranking schemes have been used, and criticised, for decades (4-7), and each scheme is geared to answer different questions(8). Early evidence hierarchies(5, 6, 9) were introduced primarily to help clinicians and other researchers appraise the quality of evidence for therapeutic effects, while more recent attempts to assign levels to evidence have been designed to help systematic reviewers(8), or guideline developers(10).
    [Show full text]
  • The Hierarchy of Evidence Pyramid
    The Hierarchy of Evidence Pyramid Available from: https://www.researchgate.net/figure/Hierarchy-of-evidence-pyramid-The-pyramidal-shape-qualitatively-integrates-the- amount-of_fig1_311504831 [accessed 12 Dec, 2019] Available from: https://journals.lww.com/clinorthop/Fulltext/2003/08000/Hierarchy_of_Evidence__From_Case_Reports_to.4.aspx [accessed 14 March 2020] CLINICAL ORTHOPAEDICS AND RELATED RESEARCH Number 413, pp. 19–24 © 2003 Lippincott Williams & Wilkins, Inc. Hierarchy of Evidence: From Case Reports to Randomized Controlled Trials Brian Brighton, MD*; Mohit Bhandari, MD, MSc**; Downloaded from https://journals.lww.com/clinorthop by BhDMf5ePHKav1zEoum1tQfN4a+kJLhEZgbsIHo4XMi0hCywCX1AWnYQp/IlQrHD3oaxD/v Paul Tornetta, III, MD†; and David T. Felson, MD* In the hierarchy of research designs, the results This hierarchy has not been supported in two re- of randomized controlled trials are considered cent publications in the New England Journal of the highest level of evidence. Randomization is Medicine which identified nonsignificant differ- the only method for controlling for known and ences in results between randomized, controlled unknown prognostic factors between two com- trials, and observational studies. The current au- parison groups. Lack of randomization predis- thors provide an approach to organizing pub- poses a study to potentially important imbal- lished research on the basis of study design, a hi- ances in baseline characteristics between two erarchy of evidence, a set of principles and tools study groups. There is a hierarchy of evidence, that help clinicians distinguish ignorance of evi- with randomized controlled trials at the top, con- dence from real scientific uncertainty, distin- trolled observational studies in the middle, and guish evidence from unsubstantiated opinions, uncontrolled studies and opinion at the bottom.
    [Show full text]
  • Downloading of a Human Consciousness Into a Digital Computer Would Involve ‘A Certain Loss of Our Finer Feelings and Qualities’89
    When We Can Trust Computers (and When We Can’t) Peter V. Coveney1,2* Roger R. Highfield3 1Centre for Computational Science, University College London, Gordon Street, London WC1H 0AJ, UK, https://orcid.org/0000-0002-8787-7256 2Institute for Informatics, Science Park 904, University of Amsterdam, 1098 XH Amsterdam, Netherlands 3Science Museum, Exhibition Road, London SW7 2DD, UK, https://orcid.org/0000-0003-2507-5458 Keywords: validation; verification; uncertainty quantification; big data; machine learning; artificial intelligence Abstract With the relentless rise of computer power, there is a widespread expectation that computers can solve the most pressing problems of science, and even more besides. We explore the limits of computational modelling and conclude that, in the domains of science and engineering that are relatively simple and firmly grounded in theory, these methods are indeed powerful. Even so, the availability of code, data and documentation, along with a range of techniques for validation, verification and uncertainty quantification, are essential for building trust in computer generated findings. When it comes to complex systems in domains of science that are less firmly grounded in theory, notably biology and medicine, to say nothing of the social sciences and humanities, computers can create the illusion of objectivity, not least because the rise of big data and machine learning pose new challenges to reproducibility, while lacking true explanatory power. We also discuss important aspects of the natural world which cannot be solved by digital means. In the long-term, renewed emphasis on analogue methods will be necessary to temper the excessive faith currently placed in digital computation.
    [Show full text]
  • Caveat Emptor: the Risks of Using Big Data for Human Development
    1 Caveat emptor: the risks of using big data for human development Siddique Latif1,2, Adnan Qayyum1, Muhammad Usama1, Junaid Qadir1, Andrej Zwitter3, and Muhammad Shahzad4 1Information Technology University (ITU)-Punjab, Pakistan 2University of Southern Queensland, Australia 3University of Groningen, Netherlands 4National University of Sciences and Technology (NUST), Pakistan Abstract Big data revolution promises to be instrumental in facilitating sustainable development in many sectors of life such as education, health, agriculture, and in combating humanitarian crises and violent conflicts. However, lurking beneath the immense promises of big data are some significant risks such as (1) the potential use of big data for unethical ends; (2) its ability to mislead through reliance on unrepresentative and biased data; and (3) the various privacy and security challenges associated with data (including the danger of an adversary tampering with the data to harm people). These risks can have severe consequences and a better understanding of these risks is the first step towards mitigation of these risks. In this paper, we highlight the potential dangers associated with using big data, particularly for human development. Index Terms Human development, big data analytics, risks and challenges, artificial intelligence, and machine learning. I. INTRODUCTION Over the last decades, widespread adoption of digital applications has moved all aspects of human lives into the digital sphere. The commoditization of the data collection process due to increased digitization has resulted in the “data deluge” that continues to intensify with a number of Internet companies dealing with petabytes of data on a daily basis. The term “big data” has been coined to refer to our emerging ability to collect, process, and analyze the massive amount of data being generated from multiple sources in order to obtain previously inaccessible insights.
    [Show full text]
  • Causality in Medicine with Particular Reference to the Viral Causation of Cancers
    Causality in medicine with particular reference to the viral causation of cancers Brendan Owen Clarke A dissertation submitted in partial fulfillment of the requirements for the degree of Doctor of Philosophy of UCL Department of Science and Technology Studies UCL August 20, 2010 Revised version January 24, 2011 I, Brendan Owen Clarke, confirm that the work presented in this thesis is my own. Where information has been derived from other sources, I confirm that this has been indicated in the thesis. ................................................................ January 24, 2011 2 Acknowledgements This thesis would not have been written without the support and inspiration of Donald Gillies, who not only supervised me, but also introduced me to history and philosophy of science in the first instance. I have been very privileged to be the beneficiary of so much of his clear thinking over the last few years. Donald: thank you for everything, and I hope this thesis lives up to your expectations. I am also extremely grateful to Michela Massimi, who has acted as my second supervisor. She has provided throughout remarkably insightful feedback on this project, no matter how distant it is from her own research interests. I’d also like to thank her for all her help and guidance on teaching matters. I should also thank Vladimir Vonka for supplying many important pieces of both the cervical cancer history, as well as his own work on causality in medicine from the perspective of a real, working medical researcher. Phyllis McKay-Illari provided a critical piece of the story, as well as many interesting and stim- ulating discussions about the more philosophical aspects of this thesis.
    [Show full text]
  • A Meta-Review of Transparency and Reproducibility-Related Reporting Practices in Published Meta- Analyses on Clinical Psychological Interventions (2000-2020)
    A meta-review of transparency and reproducibility-related reporting practices in published meta- analyses on clinical psychological interventions (2000-2020). Rubén López-Nicolás1, José Antonio López-López1, María Rubio-Aparicio2 & Julio Sánchez- Meca1 1 Universidad de Murcia (Spain) 2 Universidad de Alicante (Spain) Author note: This paper has been published at Behavior Research Methods. The published version is available at: https://doi.org/10.3758/s13428-021-01644-z. All materials, data, and analysis script coded have been made publicly available on the Open Science Framework: https://osf.io/xg97b/. Correspondence should be addressed to Rubén López-Nicolás, Facultad de Psicología, Campus de Espinardo, Universidad de Murcia, edificio nº 31, 30100 Murcia, España. E-mail: [email protected] 1 Abstract Meta-analysis is a powerful and important tool to synthesize the literature about a research topic. Like other kinds of research, meta-analyses must be reproducible to be compliant with the principles of the scientific method. Furthermore, reproducible meta-analyses can be easily updated with new data and reanalysed applying new and more refined analysis techniques. We attempted to empirically assess the prevalence of transparency and reproducibility-related reporting practices in published meta-analyses from clinical psychology by examining a random sample of 100 meta-analyses. Our purpose was to identify the key points that could be improved with the aim to provide some recommendations to carry out reproducible meta-analyses. We conducted a meta-review of meta-analyses of psychological interventions published between 2000 and 2020. We searched PubMed, PsycInfo and Web of Science databases. A structured coding form to assess transparency indicators was created based on previous studies and existing meta-analysis guidelines.
    [Show full text]
  • Reproducibility: Promoting Scientific Rigor and Transparency
    Reproducibility: Promoting scientific rigor and transparency Roma Konecky, PhD Editorial Quality Advisor What does reproducibility mean? • Reproducibility is the ability to generate similar results each time an experiment is duplicated. • Data reproducibility enables us to validate experimental results. • Reproducibility is a key part of the scientific process; however, many scientific findings are not replicable. The Reproducibility Crisis • ~2010 as part of a growing awareness that many scientific studies are not replicable, the phrase “Reproducibility Crisis” was coined. • An initiative of the Center for Open Science conducted replications of 100 psychology experiments published in prominent journals. (Science, 349 (6251), 28 Aug 2015) - Out of 100 replication attempts, only 39 were successful. The Reproducibility Crisis • According to a poll of over 1,500 scientists, 70% had failed to reproduce at least one other scientist's experiment or their own. (Nature 533 (437), 26 May 2016) • Irreproducible research is a major concern because in valid claims: - slow scientific progress - waste time and resources - contribute to the public’s mistrust of science Factors contributing to Over 80% of respondents irreproducibility Nature | News Feature 25 May 2016 Underspecified methods Factors Data dredging/ Low statistical contributing to p-hacking power irreproducibility Technical Bias - omitting errors null results Weak experimental design Underspecified methods Factors Data dredging/ Low statistical contributing to p-hacking power irreproducibility Technical Bias - omitting errors null results Weak experimental design Underspecified methods • When experimental details are omitted, the procedure needed to reproduce a study isn’t clear. • Underspecified methods are like providing only part of a recipe. ? = Underspecified methods Underspecified methods Underspecified methods Underspecified methods • Like baking a loaf of bread, a “scientific recipe” should include all the details needed to reproduce the study.
    [Show full text]
  • Evidence Based Policymaking’ in a Decentred State Abstract
    Paul Cairney, Professor of Politics and Public Policy, University of Stirling, UK [email protected] Paper accepted for publication in Public Policy and Administration 21.8.19 Special Issue (Mark Bevir) The Decentred State The myth of ‘evidence based policymaking’ in a decentred state Abstract. I describe a policy theory story in which a decentred state results from choice and necessity. Governments often choose not to centralise policymaking but they would not succeed if they tried. Many policy scholars take this story for granted, but it is often ignored in other academic disciplines and wider political debate. Instead, commentators call for more centralisation to deliver more accountable, ‘rational’, and ‘evidence based’ policymaking. Such contradictory arguments, about the feasibility and value of government centralisation, raise an ever-present dilemma for governments to accept or challenge decentring. They also accentuate a modern dilemma about how to seek ‘evidence-based policymaking’ in a decentred state. I identify three ideal-type ways in which governments can address both dilemmas consistently. I then identify their ad hoc use by UK and Scottish governments. Although each government has a reputation for more or less centralist approaches, both face similar dilemmas and address them in similar ways. Their choices reflect their need to appear to be in control while dealing with the fact that they are not. Introduction: policy theory and the decentred state I use policy theory to tell a story about the decentred state. Decentred policymaking may result from choice, when a political system contains a division of powers or a central government chooses to engage in cooperation with many other bodies rather than assert central control.
    [Show full text]
  • Rejecting Statistical Significance Tests: Defanging the Arguments
    Rejecting Statistical Significance Tests: Defanging the Arguments D. G. Mayo Philosophy Department, Virginia Tech, 235 Major Williams Hall, Blacksburg VA 24060 Abstract I critically analyze three groups of arguments for rejecting statistical significance tests (don’t say ‘significance’, don’t use P-value thresholds), as espoused in the 2019 Editorial of The American Statistician (Wasserstein, Schirm and Lazar 2019). The strongest argument supposes that banning P-value thresholds would diminish P-hacking and data dredging. I argue that it is the opposite. In a world without thresholds, it would be harder to hold accountable those who fail to meet a predesignated threshold by dint of data dredging. Forgoing predesignated thresholds obstructs error control. If an account cannot say about any outcomes that they will not count as evidence for a claim—if all thresholds are abandoned—then there is no a test of that claim. Giving up on tests means forgoing statistical falsification. The second group of arguments constitutes a series of strawperson fallacies in which statistical significance tests are too readily identified with classic abuses of tests. The logical principle of charity is violated. The third group rests on implicit arguments. The first in this group presupposes, without argument, a different philosophy of statistics from the one underlying statistical significance tests; the second group—appeals to popularity and fear—only exacerbate the ‘perverse’ incentives underlying today’s replication crisis. Key Words: Fisher, Neyman and Pearson, replication crisis, statistical significance tests, strawperson fallacy, psychological appeals, 2016 ASA Statement on P-values 1. Introduction and Background Today’s crisis of replication gives a new urgency to critically appraising proposed statistical reforms intended to ameliorate the situation.
    [Show full text]
  • The Statistical Crisis in Science
    The Statistical Crisis in Science Data-dependent analysis— a "garden of forking paths"— explains why many statistically significant comparisons don't hold up. Andrew Gelman and Eric Loken here is a growing realization a short mathematics test when it is This multiple comparisons issue is that reported "statistically sig expressed in two different contexts, well known in statistics and has been nificant" claims in scientific involving either healthcare or the called "p-hacking" in an influential publications are routinely mis military. The question may be framed 2011 paper by the psychology re Ttaken. Researchers typically expressnonspecifically as an investigation of searchers Joseph Simmons, Leif Nel the confidence in their data in terms possible associations between party son, and Uri Simonsohn. Our main of p-value: the probability that a per affiliation and mathematical reasoning point in the present article is that it ceived result is actually the result of across contexts. The null hypothesis is is possible to have multiple potential random variation. The value of p (for that the political context is irrelevant comparisons (that is, a data analysis "probability") is a way of measuring to the task, and the alternative hypoth whose details are highly contingent the extent to which a data set provides esis is that context matters and the dif on data, invalidating published p-val- evidence against a so-called null hy ference in performance between the ues) without the researcher perform pothesis. By convention, a p-value be two parties would be different in the ing any conscious procedure of fishing low 0.05 is considered a meaningful military and healthcare contexts.
    [Show full text]
  • Replicability Problems in Science: It's Not the P-Values' Fault
    The replicability problems in Science: It’s not the p-values’ fault Yoav Benjamini Tel Aviv University NISS Webinar May 6, 2020 1. The Reproducibility and Replicability Crisis Replicability with significance “We may say that a phenomenon is experimentally demonstrable when we know how to conduct an experiment which will rarely fail to give us statistically significant results.” Fisher (1935) “The Design of Experiments”. Reproducibility/Replicability • Reproduce the study: from the original data, through analysis, to get same figures and conclusions • Replicability of results: replicate the entire study, from enlisting subjects through collecting data, and analyzing the results, in a similar but not necessarily identical way, yet get essentially the same results. (Biostatistics, Editorial 2010, Nature Editorial 2013, NSA 2019) “ reproducibilty is the ability to replicate the results…” in a paper on “reproducibility is not replicability” We can therefore assure reproducibility of a single study but only enhance its replicability Opinion shared by 2019 report of National Academies on R&R Outline 1. The misguided attack 2. Selective inference: The silent killer of replicability 3. The status of addressing evident selective inference 6 2. The misguided attack Psychological Science “… we have published a tutorial by Cumming (‘14), a leader in the new-statistics movement…” • 9. Do not trust any p value. • 10. Whenever possible, avoid using statistical significance or p- values; simply omit any mention of null hypothesis significance testing (NHST). • 14. …Routinely report 95% CIs… Editorial by Trafimow & Marks (2015) in Basic and Applied Social Psychology: From now on, BASP is banning the NHSTP… 7 Is it the p-values’ fault? ASA Board’s statement about p-values (Lazar & Wasserstein Am.
    [Show full text]
  • Just How Easy Is It to Cheat a Linear Regression? Philip Pham A
    Just How Easy is it to Cheat a Linear Regression? Philip Pham A THESIS in Mathematics Presented to the Faculties of the University of Pennsylvania in Partial Fulfillment of the Requirements for the Degree of Master of Arts Spring 2016 Robin Pemantle, Supervisor of Thesis David Harbater, Graduate Group Chairman Abstract As of late, the validity of much academic research has come into question. While many studies have been retracted for outright falsification of data, perhaps more common is inappropriate statistical methodology. In particular, this paper focuses on data dredging in the case of balanced design with two groups and classical linear regression. While it is well-known data dredging has pernicious e↵ects, few have attempted to quantify these e↵ects, and little is known about both the number of covariates needed to induce statistical significance and data dredging’s e↵ect on statistical power and e↵ect size. I have explored its e↵ect mathematically and through computer simulation. First, I prove that in the extreme case that the researcher can obtain any desired result by collecting nonsense data if there is no limit on how much data he or she collects. In practice, there are limits, so secondly, by computer simulation, I demonstrate that with a modest amount of e↵ort a researcher can find a small number of covariates to achieve statistical significance both when the treatment and response are independent as well as when they are weakly correlated. Moreover, I show that such practices lead not only to Type I errors but also result in an exaggerated e↵ect size.
    [Show full text]