LLNL Computation Directorate Annual Report (2014)

Total Page:16

File Type:pdf, Size:1020Kb

LLNL Computation Directorate Annual Report (2014) PRODUCTION TEAM LLNL Associate Director for Computation Dona L. Crawford Deputy Associate Directors James Brase, Trish Damkroger, John Grosh, and Michel McCoy Scientific Editors John Westlund and Ming Jiang Art Director Amy Henke Production Editor Deanna Willis Writers Andrea Baron, Rose Hansen, Caryn Meissner, Linda Null, Michelle Rubin, and Deanna Willis Proofreader Rose Hansen Photographer Lee Baker LLNL-TR-668095 3D Designer Prepared by LLNL under Contract DE-AC52-07NA27344. Ryan Chen This document was prepared as an account of work sponsored by an agency of the United States government. Neither the United States government nor Lawrence Livermore National Security, LLC, nor any of their employees makes any warranty, expressed or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, Print Production manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States government or Lawrence Livermore National Security, LLC. The views and opinions of authors expressed herein do not necessarily state or reflect those of the Charlie Arteago, Jr., and Monarch Print Copy and Design Solutions United States government or Lawrence Livermore National Security, LLC, and shall not be used for advertising or product endorsement purposes. CONTENTS Message from the Associate Director . 2 An Award-Winning Organization . 4 CORAL Contract Awarded and Nonrecurring Engineering Begins . 6 Preparing Codes for a Technology Transition . 8 Flux: A Framework for Resource Management . 10 Improved Performance Data Visualization for Extreme-Scale Systems . 12 Machine Learning Strengthens Performance Predictions . 14 Planning HPC Resources for the Institution . 16 Enhancing Data-Intensive Computing at Livermore . 18 Interweaving Timelines to Save Time . 20 Remapping Algorithm Boosts BLAST Simulations . 22 Catching Bugs with the Automated Testing System . 24 NIF Deploys New Advanced Radiographic Capability . 26 Leveraging Data-Intensive Computing for Sleuthing Seismic Signals . 28 Managing Application Portability for Next-Generation Platforms . 30 Training Tomorrow’s Cybersecurity Specialists . 32 New Capabilities for Information Technology Service Management . 34 Appendices . 36 Publications . 50 Industrial Collaborators . 55 National Laboratory Collaborators . 62 MESSAGE FROM THE ASSOCIATE DIRECTOR or more than six decades, Lawrence This requires that Computation advance the state from proposal to reality. LLNL and Oak Ridge interface—with the 20-petaflop Sequoia Blue 2 FLivermore National Laboratory (LLNL) of the art in both high performance computing National Laboratory have partnered with the Gene/Q machine providing a virtual laboratory has pioneered the computational capabilities (HPC) and the supporting computer science. hardware vendors IBM, NVIDIA, and Mellanox for these explorations. required to advance science and address the to form several working groups responsible nation’s toughest and most pressing national Larger and more complex simulations require for co-designing the new architecture. Co- As these codes evolve, our teams also must security challenges. This effort, dating to the world-class platforms. To sustain the nation’s design is an essential part of developing next- ensure that they remain stable, fast, and Laboratory’s earliest years, has been punctuated nuclear weapons deterrent without the need generation architectures well suited to key DOE accurate. For more than a decade, Livermore by the acquisition of leading computer systems for full-scale underground nuclear tests, we are applications. By incorporating the expertise of has been developing the Automated Testing and their application to an ever-broadening preparing for another major advance in HPC that the vendors’ hardware architects, the system System, which runs approximately 4,000 tests spectrum of scientific and technological will make our physics and engineering simulation software developers, and the DOE laboratory nightly across Livermore Computing’s (LC’s) problems. In 2014, this tradition continued models more predictive. The Collaboration of Oak experts—including domain scientists, computer Linux and IBM Blue Gene systems. These tests as the Computation Directorate provided the Ridge, Argonne, and Livermore (CORAL) is making scientists, and applied mathematicians—Sierra generate diagnostic information used to ensure computing architectures, system software, this advance possible. As a result of CORAL, will be well equipped to handle the most that Livermore’s applications are ready for future productivity tools, algorithmic innovations, and three DOE systems—one for each laboratory— demanding computational problems. systems and challenges. application codes necessary to fulfill mission- are being acquired. LLNL’s system acquisition critical national security objectives for the contract for Sierra was signed in November In addition to the design and development of As large-scale systems continue to explode in U.S. Departments of Energy (DOE), Homeland 2014 with an expected delivery date late in CY17. the platform, other preparations are underway parallelism, simulation codes must find additional Security, and Defense, as well as other federal Sierra is a next-generation supercomputer and to ensure that Sierra can hit the computer ways to exploit the increasingly complex design. and state agencies. Computation’s next capability platform for room floor running. An important aspect of a This can be especially difficult for simulations that the National Nuclear Security Administration successful launch is having applications that are must replicate phenomena that evolve, change, As Livermore scientists continue to push the and the Advanced Simulation and Computing capable of utilizing such an immense system. We and propagate over time. Time is inherently boundaries of what is scientifically possible, they Program. At the moment, Sierra only exists as a are adapting our existing large applications using sequential, but thanks to novel work by require computational simulations that are more proposal within contract language. A significant incremental improvements such as fine-grained researchers at Livermore, Memorial University, precise, that cover longer periods of time, and that amount of research, development, design, and threading, use of accelerators, and scaling to and Belgium’s Katholieke Universiteit Leuven, visualize higher fidelity, more complex systems. testing must still take place to move Sierra millions of nodes using message processing existing large-scale codes are becoming capable COMPUTATION 2014 ANNUAL REPORT LLNL of parallelizing in time. By using a new multilevel some of our most powerful supercomputers that will push the performance of Hadoop even students, 49 undergraduates, 1 high school algorithm called XBraid, some applications can is an important investment the Laboratory further. For example, 800 gigabytes of flash student, and 5 faculty members—from now solve problems up to 10 times faster. As makes in our science and technology. One of memory has been installed on each of the Catalyst 89 universities and 8 countries. Specific specialties its name suggests, XBraid allows simulations to the ways Computation supports LLNL projects system’s 304 compute nodes. Memory and within the scholar program have been emerging “braid” together multiple timelines, eliminating and programs is through the annual Computing storage in the system allow Catalyst, a first-of-a- over the past several years—the Cyber Defenders the need to be solved sequentially. These braided Grand Challenge program. Now in its ninth year, kind supercomputer, to appear as a 300-terabyte and the Co-Design Summer School. Now in its solutions are solved much more coarsely, and the Grand Challenge program awarded more memory system. In this configuration, our fifth year, the Cyber Defenders matched each the results are fed back into the algorithm until than 15 million central processing unit-hours scientists reduced the TeraSort runtime, a of the 21 computer science and engineering they converge on a solution that matches the per week to projects that address compelling, big data benchmark, to a little more than 230 students with an LLNL mentor and assigned them expected results from traditional sequential large-scale problems, push the envelope of seconds—much shorter than traditional Hadoop. a real-world technical project wherein they apply algorithms, within defined tolerances. capability computing, and advance science. Also this year, Catalyst was made available to technologies, develop solutions to computer- This year marked the first time that institutional industry collaborators through Livermore’s High security-related problems of national interest, and Taking advantage of LC’s valuable resources in demand for cycles on the 5-petaflop Vulcan Performance Computing Innovation Center to explore new technologies that can be applied to an orderly manner is no small task. With a sizable system exceeded the available allocation, an further test big data technologies, architectures, computer security. The 2014 inaugural Co-Design number of complex applications, Livermore encouraging indication of the growing interest and applications.
Recommended publications
  • The Reliability Wall for Exascale Supercomputing Xuejun Yang, Member, IEEE, Zhiyuan Wang, Jingling Xue, Senior Member, IEEE, and Yun Zhou
    . IEEE TRANSACTIONS ON COMPUTERS, VOL. *, NO. *, * * 1 The Reliability Wall for Exascale Supercomputing Xuejun Yang, Member, IEEE, Zhiyuan Wang, Jingling Xue, Senior Member, IEEE, and Yun Zhou Abstract—Reliability is a key challenge to be understood to turn the vision of exascale supercomputing into reality. Inevitably, large-scale supercomputing systems, especially those at the peta/exascale levels, must tolerate failures, by incorporating fault- tolerance mechanisms to improve their reliability and availability. As the benefits of fault-tolerance mechanisms rarely come without associated time and/or capital costs, reliability will limit the scalability of parallel applications. This paper introduces for the first time the concept of “Reliability Wall” to highlight the significance of achieving scalable performance in peta/exascale supercomputing with fault tolerance. We quantify the effects of reliability on scalability, by proposing a reliability speedup, defining quantitatively the reliability wall, giving an existence theorem for the reliability wall, and categorizing a given system according to the time overhead incurred by fault tolerance. We also generalize these results into a general reliability speedup/wall framework by considering not only speedup but also costup. We analyze and extrapolate the existence of the reliability wall using two representative supercomputers, Intrepid and ASCI White, both employing checkpointing for fault tolerance, and have also studied the general reliability wall using Intrepid. These case studies provide insights on how to mitigate reliability-wall effects in system design and through hardware/software optimizations in peta/exascale supercomputing. Index Terms—fault tolerance, exascale, performance metric, reliability speedup, reliability wall, checkpointing. ! 1INTRODUCTION a revolution in computing at a greatly accelerated pace [2].
    [Show full text]
  • Biomolecular Simulation Data Management In
    BIOMOLECULAR SIMULATION DATA MANAGEMENT IN HETEROGENEOUS ENVIRONMENTS Julien Charles Victor Thibault A dissertation submitted to the faculty of The University of Utah in partial fulfillment of the requirements for the degree of Doctor of Philosophy Department of Biomedical Informatics The University of Utah December 2014 Copyright © Julien Charles Victor Thibault 2014 All Rights Reserved The University of Utah Graduate School STATEMENT OF DISSERTATION APPROVAL The dissertation of Julien Charles Victor Thibault has been approved by the following supervisory committee members: Julio Cesar Facelli , Chair 4/2/2014___ Date Approved Thomas E. Cheatham , Member 3/31/2014___ Date Approved Karen Eilbeck , Member 4/3/2014___ Date Approved Lewis J. Frey _ , Member 4/2/2014___ Date Approved Scott P. Narus , Member 4/4/2014___ Date Approved And by Wendy W. Chapman , Chair of the Department of Biomedical Informatics and by David B. Kieda, Dean of The Graduate School. ABSTRACT Over 40 years ago, the first computer simulation of a protein was reported: the atomic motions of a 58 amino acid protein were simulated for few picoseconds. With today’s supercomputers, simulations of large biomolecular systems with hundreds of thousands of atoms can reach biologically significant timescales. Through dynamics information biomolecular simulations can provide new insights into molecular structure and function to support the development of new drugs or therapies. While the recent advances in high-performance computing hardware and computational methods have enabled scientists to run longer simulations, they also created new challenges for data management. Investigators need to use local and national resources to run these simulations and store their output, which can reach terabytes of data on disk.
    [Show full text]
  • Measuring Power Consumption on IBM Blue Gene/P
    View metadata, citation and similar papers at core.ac.uk brought to you by CORE provided by Springer - Publisher Connector Comput Sci Res Dev DOI 10.1007/s00450-011-0192-y SPECIAL ISSUE PAPER Measuring power consumption on IBM Blue Gene/P Michael Hennecke · Wolfgang Frings · Willi Homberg · Anke Zitz · Michael Knobloch · Hans Böttiger © The Author(s) 2011. This article is published with open access at Springerlink.com Abstract Energy efficiency is a key design principle of the Top10 supercomputers on the November 2010 Top500 list IBM Blue Gene series of supercomputers, and Blue Gene [1] alone (which coincidentally are also the 10 systems with systems have consistently gained top GFlops/Watt rankings an Rpeak of at least one PFlops) are consuming a total power on the Green500 list. The Blue Gene hardware and man- of 33.4 MW [2]. These levels of power consumption are al- agement software provide built-in features to monitor power ready a concern for today’s Petascale supercomputers (with consumption at all levels of the machine’s power distribu- operational expenses becoming comparable to the capital tion network. This paper presents the Blue Gene/P power expenses for procuring the machine), and addressing the measurement infrastructure and discusses the operational energy challenge clearly is one of the key issues when ap- aspects of using this infrastructure on Petascale machines. proaching Exascale. We also describe the integration of Blue Gene power moni- While the Flops/Watt metric is useful, its emphasis on toring capabilities into system-level tools like LLview, and LINPACK performance and thus computational load ne- highlight some results of analyzing the production workload glects the fact that the energy costs of memory references at Research Center Jülich (FZJ).
    [Show full text]
  • Blue Gene®/Q Overview and Update
    Blue Gene®/Q Overview and Update November 2011 Agenda Hardware Architecture George Chiu Packaging & Cooling Paul Coteus Software Architecture Robert Wisniewski Applications & Configurations Jim Sexton IBM System Technology Group Blue Gene/Q 32 Node Board Industrial Design BQC DD2.0 4-rack system 5D torus 3 © 2011 IBM Corporation IBM System Technology Group Top 10 reasons that you need Blue Gene/Q 1. Ultra-scalability for breakthrough science – System can scale to 256 racks and beyond (>262,144 nodes) – Cluster: typically a few racks (512-1024 nodes) or less. 2. Highest capability machine in the world (20-100PF+ peak) 3. Superior reliability: Run an application across the whole machine, low maintenance 4. Highest power efficiency, smallest footprint, lowest TCO (Total Cost of Ownership) 5. Low latency, high bandwidth inter-processor communication system 6. Low latency, high bandwidth memory system 7. Open source and standards-based programming environment – Red Hat Linux distribution on service, front end, and I/O nodes – Lightweight Compute Node Kernel (CNK) on compute nodes ensures scaling with no OS jitter, enables reproducible runtime results – Automatic SIMD (Single-Instruction Multiple-Data) FPU exploitation enabled by IBM XL (Fortran, C, C++) compilers – PAMI (Parallel Active Messaging Interface) runtime layer. Runs across IBM HPC platforms 8. Software architecture extends application reach – Generalized communication runtime layer allows flexibility of programming model – Familiar Linux execution environment with support for most
    [Show full text]
  • TMA4280—Introduction to Supercomputing
    Supercomputing TMA4280—Introduction to Supercomputing NTNU, IMF January 12. 2018 1 Outline Context: Challenges in Computational Science and Engineering Examples: Simulation of turbulent flows and other applications Goal and means: parallel performance improvement Overview of supercomputing systems Conclusion 2 Computational Science and Engineering (CSE) What is the motivation for Supercomputing? Solve complex problems fast and accurately: — efforts in modelling and simulation push sciences and engineering applications forward, — computational requirements drive the development of new hardware and software. 3 Computational Science and Engineering (CSE) Development of computational methods for scientific research and innovation in engineering and technology. Covers the entire spectrum of natural sciences, mathematics, informatics: — Scientific model (Physics, Biology, Medical, . ) — Mathematical model — Numerical model — Implementation — Visualization, Post-processing — Validation ! Feedback: virtuous circle Allows for larger and more realistic problems to be simulated, new theories to be experimented numerically. 4 Outcome in Industrial Applications Figure: 2004: “The Falcon 7X becomes the first aircraft in industry history to be entirely developed in a virtual environment, from design to manufacturing to maintenance.” Dassault Systèmes 5 Evolution of computational power Figure: Moore’s Law: exponential increase of number of transistors per chip, 1-year rate (1965), 2-year rate (1975). WikiMedia, CC-BY-SA-3.0 6 Evolution of computational power
    [Show full text]
  • Replication, History, and Grafting in the Ori File System
    Replication, History, and Grafting in the Ori File System Ali Jose´ Mashtizadeh, Andrea Bittau, Yifeng Frank Huang, David Mazieres` Stanford University Abstract backup, versioning, access from any device, and multi- user sharing—over storage capacity. But is the cloud re- Ori is a file system that manages user data in a modern ally the best place to implement data management fea- setting where users have multiple devices and wish to tures, or could the file system itself directly implement access files everywhere, synchronize data, recover from them better? disk failure, access old versions, and share data. The In terms of technological changes, disk space has in- key to satisfying these needs is keeping and replicating creased dramatically and has outgrown the increase in file system history across devices, which is now prac- wide-area bandwidth. In 1990, a typical desktop machine tical as storage space has outpaced both wide-area net- had a 60 MB hard disk, whose entire contents could tran- work (WAN) bandwidth and the size of managed data. sit a 9,600 baud modem in under 14 hours [2]. Today, Replication provides access to files from multiple de- $120 can buy a 3 TB disk, which requires 278 days to vices. History provides synchronization and offline ac- transfer over a 1 Mbps broadband connection! Clearly, cess. Replication and history together subsume backup cloud-based storage solutions have a tough time keep- by providing snapshots and avoiding any single point of ing up with disk capacity. But capacity is also outpac- failure. In fact, Ori is fully peer-to-peer, offering oppor- ing the size of managed data—i.e., the documents on tunistic synchronization between user devices in close which users actually work (as opposed to large media proximity and ensuring that the file system is usable so files or virtual machine images that would not be stored long as a single replica remains.
    [Show full text]
  • Introduction to the Practical Aspects of Programming for IBM Blue Gene/P
    Introduction to the Practical Aspects of Programming for IBM Blue Gene/P Alexander Pozdneev Research Software Engineer, IBM June 22, 2015 — International Summer Supercomputing Academy IBM Science and Technology Center Established in 2006 • Research projects I HPC simulations I Security I Oil and gas applications • ESSL math library for IBM POWER • IBM z Systems mainframes I Firmware I Linux on z Systems I Software • IBM Rational software • Smarter Commerce for Retail • Smarter Cities 2 c 2015 IBM Corporation Outline 1 Blue Gene architecture in the HPC world 2 Blue Gene/P architecture 3 Conclusion 4 References 3 c 2015 IBM Corporation Outline 1 Blue Gene architecture in the HPC world Brief history of Blue Gene architecture Blue Gene in TOP500 Blue Gene in Graph500 Scientific applications Gordon Bell awards Blue Gene installation sites 2 Blue Gene/P architecture 3 Conclusion 4 References 4 c 2015 IBM Corporation Brief history of Blue Gene architecture • 1999 US$100M research initiative, novel massively parallel architecture, protein folding • 2003, Nov Blue Gene/L first prototype TOP500, #73 • 2004, Nov 16 Blue Gene/L racks at LLNL TOP500, #1 • 2007 Second generation Blue Gene/P • 2009 National Medal of Technology and Innovation • 2012 Third generation Blue Gene/Q • Features: I Low frequency, low power consumption I Fast interconnect balanced with CPU I Light-weight OS 5 c 2015 IBM Corporation Blue Gene in TOP500 Date # System Model Nodes Cores Jun’14 / Nov’14 3 Sequoia Q 96k 1.5M Jun’13 / Nov’13 3 Sequoia Q 96k 1.5M Nov’12 2 Sequoia
    [Show full text]
  • Sistemas De Archivos Distribuido Para Clúster HPC Utilizando Ceph
    Departamento de Telecomunicaciones y Electrónica Título: Sistemas de archivos distribuido para Clúster HPC utilizando Ceph Autor: Daniel Placencia Alvarez Tutor: Ing. Javier Antonio Ruiz Bosch , Junio, 2019 Este documento es Propiedad Patrimonial de la Universidad Central “Marta Abreu” de Las Villas, y se encuentra depositado en los fondos de la Biblioteca Universitaria “Chiqui Gómez Lubian” subordinada a la Dirección de Información Científico Técnica de la mencionada casa de altos estudios. Se autoriza su utilización bajo la licencia siguiente: Atribución- No Comercial- Compartir Igual Para cualquier información contacte con: Dirección de Información Científico Técnica. Universidad Central “Marta Abreu” de Las Villas. Carretera a Camajuaní. Km 5½. Santa Clara. Villa Clara. Cuba. CP. 54 830 Teléfonos.: +53 01 42281503-1419 Hago constar que el presente trabajo de diploma fue realizado en la Universidad Central “Marta Abreu” de Las Villas como parte de la culminación de estudios de la especialidad de Ingeniería en Telecomunicaciones y Electrónica, autorizando a que el mismo sea utilizado por la Institución, para los fines que estime conveniente, tanto de forma parcial como total y que además no podrá ser presentado en eventos, ni publicados sin autorización de la Universidad. Firma del Autor Los abajo firmantes certificamos que el presente trabajo ha sido realizado según acuerdo de la dirección de nuestro centro y el mismo cumple con los requisitos que debe tener un trabajo de esta envergadura referido a la temática señalada. Firma del Tutor Firma del Jefe de Departamento donde se defiende el trabajo Firma del Responsable de Información Científico-Técnica i PENSAMIENTO Muchos de los fracasos en la vida lo experimentan personas que no se dan cuenta de cuan cerca estuvieron del éxito cuando decidieron darse por vencidos.
    [Show full text]
  • The Blue Gene/Q Compute Chip
    The Blue Gene/Q Compute Chip Ruud Haring / IBM BlueGene Team © 2011 IBM Corporation Acknowledgements The IBM Blue Gene/Q development teams are located in – Yorktown Heights NY, – Rochester MN, – Hopewell Jct NY, – Burlington VT, – Austin TX, – Bromont QC, – Toronto ON, – San Jose CA, – Boeblingen (FRG), – Haifa (Israel), –Hursley (UK). Columbia University . University of Edinburgh . The Blue Gene/Q project has been supported and partially funded by Argonne National Laboratory and the Lawrence Livermore National Laboratory on behalf of the United States Department of Energy, under Lawrence Livermore National Laboratory Subcontract No. B554331 2 03 Sep 2012 © 2012 IBM Corporation Blue Gene/Q . Blue Gene/Q is the third generation of IBM’s Blue Gene family of supercomputers – Blue Gene/L was announced in 2004 – Blue Gene/P was announced in 2007 . Blue Gene/Q – was announced in 2011 – is currently the fastest supercomputer in the world • June 2012 Top500: #1,3,7,8, … 15 machines in top100 (+ 101-103) – is currently the most power efficient supercomputer architecture in the world • June 2012 Green500: top 20 places – also shines at data-intensive computing • June 2012 Graph500: top 2 places -- by a 7x margin – is relatively easy to program -- for a massively parallel computer – and its FLOPS are actually usable •this is not a GPGPU design … – incorporates innovations that enhance multi-core / multi-threaded computing • in-memory atomic operations •1st commercial processor to support hardware transactional memory / speculative execution •… – is just a cool chip (and system) design • pushing state-of-the-art ASIC design where it counts • innovative power and cooling 3 03 Sep 2012 © 2012 IBM Corporation Blue Gene system objectives .
    [Show full text]
  • R00456--FM Getting up to Speed
    GETTING UP TO SPEED THE FUTURE OF SUPERCOMPUTING Susan L. Graham, Marc Snir, and Cynthia A. Patterson, Editors Committee on the Future of Supercomputing Computer Science and Telecommunications Board Division on Engineering and Physical Sciences THE NATIONAL ACADEMIES PRESS Washington, D.C. www.nap.edu THE NATIONAL ACADEMIES PRESS 500 Fifth Street, N.W. Washington, DC 20001 NOTICE: The project that is the subject of this report was approved by the Gov- erning Board of the National Research Council, whose members are drawn from the councils of the National Academy of Sciences, the National Academy of Engi- neering, and the Institute of Medicine. The members of the committee responsible for the report were chosen for their special competences and with regard for ap- propriate balance. Support for this project was provided by the Department of Energy under Spon- sor Award No. DE-AT01-03NA00106. Any opinions, findings, conclusions, or recommendations expressed in this publication are those of the authors and do not necessarily reflect the views of the organizations that provided support for the project. International Standard Book Number 0-309-09502-6 (Book) International Standard Book Number 0-309-54679-6 (PDF) Library of Congress Catalog Card Number 2004118086 Cover designed by Jennifer Bishop. Cover images (clockwise from top right, front to back) 1. Exploding star. Scientific Discovery through Advanced Computing (SciDAC) Center for Supernova Research, U.S. Department of Energy, Office of Science. 2. Hurricane Frances, September 5, 2004, taken by GOES-12 satellite, 1 km visible imagery. U.S. National Oceanographic and Atmospheric Administration. 3. Large-eddy simulation of a Rayleigh-Taylor instability run on the Lawrence Livermore National Laboratory MCR Linux cluster in July 2003.
    [Show full text]
  • IBM Research Report Designing a Highly-Scalable Operating System
    RC24037 (W0608-087) August 28, 2006 Computer Science IBM Research Report Designing a Highly-Scalable Operating System: The Blue Gene/L Story José Moreira1, Michael Brutman2, José Castaños1, Thomas Engelsiepen3, Mark Giampapa1, Tom Gooding2, Roger Haskin3, Todd Inglett2, Derek Lieber1, Pat McCarthy2, Mike Mundy2, Jeff Parker2, Brian Wallenfelt4 1IBM Research Division Thomas J. Watson Research Center P.O. Box 218 Yorktown Heights, NY 10598 2IBM Systems and Technology Group Rochester, MN 55901 3IBM Research Division Almaden Research Center 650 Harry Road San Jose, CA 95120-6099 4GeoDigm Corporation Chanhassen, MN 55317 Research Division Almaden - Austin - Beijing - Haifa - India - T. J. Watson - Tokyo - Zurich LIMITED DISTRIBUTION NOTICE: This report has been submitted for publication outside of IBM and will probably be copyrighted if accepted for publication. It has been issued as a Research Report for early dissemination of its contents. In view of the transfer of copyright to the outside publisher, its distribution outside of IBM prior to publication should be limited to peer communications and specific requests. After outside publication, requests should be filled only by reprints or legally obtained copies of the article (e.g. , payment of royalties). Copies may be requested from IBM T. J. Watson Research Center , P. O. Box 218, Yorktown Heights, NY 10598 USA (email: [email protected]). Some reports are available on the internet at http://domino.watson.ibm.com/library/CyberDig.nsf/home . Designing a Highly-Scalable Operating System: The Blue Gene/L Story José Moreira*, Michael Brutman**, José Castaños*, Thomas Engelsiepen***, Mark Giampapa*, Tom Gooding**, Roger Haskin***, Todd Inglett**, Derek Lieber*, Pat McCarthy**, Mike Mundy**, Jeff Parker**, Brian Wallenfelt**** *{jmoreira,castanos,giampapa,lieber}@us.ibm.com ***{engelspn,roger}@almaden.ibm.com IBM Thomas J.
    [Show full text]
  • The Evolution of File Systems
    The Evolution of File Systems Thomas Rivera, Hitachi Data Systems Craig Harmer, April 2011 SNIA Legal Notice The material contained in this tutorial is copyrighted by the SNIA. Member companies and individuals may use this material in presentations and literature under the following conditions: Any slide or slides used must be reproduced without modification The SNIA must be acknowledged as source of any material used in the body of any document containing material from these presentations. This presentation is a project of the SNIA Education Committee. Neither the Author nor the Presenter is an attorney and nothing in this presentation is intended to be nor should be construed as legal advice or opinion. If you need legal advice or legal opinion please contact an attorney. The information presented herein represents the Author's personal opinion and current understanding of the issues involved. The Author, the Presenter, and the SNIA do not assume any responsibility or liability for damages arising out of any reliance on or use of this information. NO WARRANTIES, EXPRESS OR IMPLIED. USE AT YOUR OWN RISK. The Evolution of File Systems 2 © 2012 Storage Networking Industry Association. All Rights Reserved. 2 Abstract The File Systems Evolution Over time additional file systems appeared focusing on specialized requirements such as: data sharing, remote file access, distributed file access, parallel files access, HPC, archiving, security, etc. Due to the dramatic growth of unstructured data, files as the basic units for data containers are morphing into file objects, providing more semantics and feature- rich capabilities for content processing This presentation will: Categorize and explain the basic principles of currently available file system architectures (e.g.
    [Show full text]