A peer-reviewed version of this preprint was published in PeerJ on 7 July 2017. View the peer-reviewed version (peerj.com/articles/3544), which is the preferred citable publication unless you specifically need to cite this preprint. Amrhein V, Korner-Nievergelt F, Roth T. 2017. The earth is flat (p > 0.05): significance thresholds and the crisis of unreplicable research. PeerJ 5:e3544 https://doi.org/10.7717/peerj.3544 The earth is flat (p>0.05): Significance thresholds and the crisis of unreplicable research Valentin Amrhein1,2,3, Fränzi Korner-Nievergelt3,4, Tobias Roth1,2 1Zoological Institute, University of Basel, Basel, Switzerland 2Research Station Petite Camargue Alsacienne, Saint-Louis, France 3Swiss Ornithological Institute, Sempach, Switzerland 4Oikostat GmbH, Ettiswil, Switzerland Email address:
[email protected] Abstract The widespread use of 'statistical significance' as a license for making a claim of a scientific finding leads to considerable distortion of the scientific process (according to the American Statistical Association). We review why degrading p-values into 'significant' and 'nonsignificant' contributes to making studies irreproducible, or to making them seem irreproducible. A major problem is that we tend to take small p-values at face value, but mistrust results with larger p-values. In either case, p-values tell little about reliability of research, because they are hardly replicable even if an alternative hypothesis is true. Also significance (p≤0.05) is hardly replicable: at a good statistical power of 80%, two studies will be 'conflicting', meaning that one is significant and the other is not, in one third of the cases if there is a true effect.