Towards Reproducible Bioinformatics: The OpenBio-C Scientific Workflow Environment Alexandros Kanterakis1, Galateia Iatraki1, Konstantina Pityanou1,2, Lefteris Koumakis1,Nikos Kanakaris3, Nikos Karacapilidis3, George Potamias2 1Institute of Computer Science, Foundation for Research and Technology-Hellas (FORTH), Heraklion, Greece {kantale,giatraki,koumakis,potamias)@ics.forth.gr 2Dept. of Electrical and Computer Engineering, Hellenic Mediterranean University Heraklion, Crete, Greece
[email protected] 3Industrial Management and Information Systems Lab, MEAD, University of Patras, Patras, Greece
[email protected],
[email protected] [email protected] Abstract—The highly competitive environment of post- Reproducibility “now”. It is reported that published and genomics biomedical research is pushing towards the highly-ranked biomedical research results are very often not production of highly qualified results and their utilization in reproducible [3]. In a relevant investigation, about 92% clinical practice. Due to this pressure, an important factor of (11/12) of interviewed researchers could not reproduce scientific progress has been underestimated namely, the results even if the methods presented in the original papers reproducibility of the results. To this end, it is critical to design were exactly replicated [4]. In addition, there are estimates and implement computational platforms that enable seamless that preclinical research is dominated by more than 50% of and unhindered access to distributed bio-data, software and irreproducible results at a cost of about 28 billion dollars [5]. computational infrastructures. Furthermore, the need to We argue that this unnatural separation between scientific support collaboration and form synergistic activities between reports and analysis pipelines is one of the origins of the the engaged researchers is also raised, as the mean to achieve reproducibility crisis [6].