ATI Stream Computing Update

Total Page:16

File Type:pdf, Size:1020Kb

ATI Stream Computing Update Stream Computing on ATI Radeon™ Embedded Graphics Processors Peter Mandl Sr. Product Mgr, Embedded Graphics January 2010 1 Stream Computing Harnessing the Computational Power of GPUs • Hundreds of processing elements in a Graphics Processor Unit (GPU) • Enormous computational capability for data parallel workloads • Enables significant performance/watt and performance/$ improvements • Works together with the CPU Software Applications Graphics Workloads Serial and Task Parallel Data Parallel Workloads Workloads 2 ATI Stream Technology is… Heterogeneous: Developers leverage AMD GPUs and CPUs for optimal application performance and user experience High performance: Massively parallel, programmable GPU architecture enables superior performance and power efficiency Industry Standards: OpenCL™ and DirectCompute 11 enable cross- platform development Gaming Digital Content Productivity Creation Sciences Government Engineering 3 The World’s Most Efficient GPU* 16 14 14.47 GFLOPS/W 12 GFLOPS/W GFLOPS/mm2 10 7.50 8 4.50 7.90 6 GFLOPS/mm2 2.01 2.21 4 4.56 1.07 2.24 2 0.42 1.06 0.92 0 Nov-05 Jan-06 Sep-07 Nov-07 Jun-08 Oct-09 ATI Radeon™ ATI Radeon™ ATI Radeon™ HD ATI Radeon™ HD ATI Radeon™ HD ATI Radeon™ HD X1800 XT X1900 XTX 2900 PRO 3870 4870 5870 *Based on comparison of consumer client single-GPU configurations as of 12/08/09. ATI Radeon™ HD 5870 provides 4 14.47 GFLOPS/W and 7.90 GFLOPS/mm2 vs. NVIDIA GTX 285 at 5.21 GFLOPS/W and 2.26 GFLOPS/mm2. Enhancing GPUs for Computation RV770 Architecture • 800 stream processing units • Fast double-precision processing – up to 240 GFLOPS • 115 GB/sec GDDR5 memory interface • Integer bit shift operations for all units – useful for crypto, compression, video processing • Data sharing between threads • More aggressive clock gating for improved performance per watt 5 5 Stream Software Development 6 ATI Stream Technology Development and Deployment Deployment Development ATI Radeon™ graphics on A PC for consumer applications ATI FirePro™ graphics for professional graphics Industry AMD FireStream™ Standards: compute accelerator OpenCL™ and for computation Microsoft DirectCompute AMD Embedded Radeon™ Graphics for long availability and small footprint 7 7 Moving Past Proprietary Solutions for Ease of Cross-Platform Programming Open and Custom Tools High Level Tools High Level Language Application Specific Compilers Libraries Industry Standard Interfaces OpenCL™ DirectX® OpenGL® AMD AMD Other GPUs CPUs CPUs/GPUs •Cross-platform development •Interoperability with OpenGL and DirectX •CPU/GPU backends enable balanced platform approach 8 ATI Stream SDK v2.0 (Production): OpenCL™ For Multicore x86 CPUs and GPUs The Power of Fusion: Developers leverage heterogeneous architecture to enable a superior user experience • First complete OpenCL™ development platform • Certified OpenCL™ 1.0 compliant by the Khronos Group1 • Write code that can scale well on multi-core CPUs and GPUs • AMD delivers on the promise of support for OpenCL™, with both high-performance CPU and GPU technologies • Available for download now – includes documentation, samples, and developer support Software Applications Graphics Workloads Serial and Task Parallel Data Parallel Workloads Workloads Product Page: http://developer.amd.com/stream 1 Conformance logs submitted for the ATI Radeon™ HD 5800 series GPUs, ATI Radeon™ HD 5700 series GPUs, ATI Radeon™ HD 4800 series GPUs, ATI FirePro™ V8700 series GPUs, AMD FireStream™ 9200 series GPUs, ATI Mobility 9 Radeon™ HD 4800 series GPUs and x86 CPUs with SSE3. ATI Stream SDK v2.0 (Production): Key Features • New: Support for OpenCL™ ICD (Installable Client Driver) • New: Support for atomic functions for 32-bit integers • New: Microsoft® Visual Studio® 2008-integrated ATI Stream Profiler performance analysis tool • Preview: Support for OpenCL™ / OpenGL® interoperability • Preview: Support for OpenCL™ / Microsoft® DirectX® 10 interoperability • Preview: Support for double-precision floating point basic arithmetic in OpenCL™ C kernels • Added support for ATI Radeon™ HD 5970 GPU1 • Supported OSes: Microsoft® XP / Vista® / 7 and Linux® openSUSE™ 11.0 / Ubuntu® 9.04 • Supported Compilers: Microsoft® Visual Studio® 2008 Pro / GCC 4.3 or later / ICC 11.x Product Page: http://developer.amd.com/stream 1 Conformance logs submitted for the ATI Radeon™ HD 5800 series GPUs, ATI Radeon™ HD 5700 series GPUs, ATI Radeon™ HD 4800 series GPUs, ATI FirePro™ V8700 series GPUs, AMD FireStream™ 9200 series GPUs, ATI Mobility 10 Radeon™ HD 4800 series GPUs and x86 CPUs with SSE3. OpenCL™: Game-Changing Development Enabling Broad Adoption of GP-GPU Capabilities . Industry standard API: Open, multiplatform development platform for heterogeneous architectures . The power of Fusion: Leverages CPUs and GPUs for balanced system approach . Broad industry support: Created by architects from AMD, Apple, IBM, Intel, Nvidia, Sony, etc. Fast track development: Ratified in December; system vendors working on implementations now . Momentum: Enormous interest from mainstream developers More stream-enabled applications across all markets 11 Stream Computing Applications 12 ATI Stream Technology Enabled Multimedia Applications MediaShow 5 PowerDirector 8 MediaShow Espresso PowerDirector 7 SimHD™ Plug-in for TotalMedia Theatre Roxio Creator™ 2010 Roxio Creator™ 2010 Pro 13 High Performance With ATI Stream Technology: Cryptography through Distributed.net . Distributed.net provides a distributed compute model allowing users to donate compute cycles to large compute-intensive projects . RC5-72 – Cryptography algorithm that searches for encryption keys RC5-72 1400 1200 1000 800 600 MKeys/sec 400 200 0 ATI Radeon ATI Radeon ATI Radeon Nvidia GEForce Nvidia GEForce Nvidia GEForce HD5870 HD4870 HD4850 GTX285 GTX280 GTX260 *Based on AMD internal testing using RC5-72 clients as of 9/04/09. Results shown in millions of keys evaluated per second [Mkeys/sec]. Configuration: AMD Phenom™ X4 9950 Black Edition processor, 8GB DDR2 RAM, Windows Vista® 32-bit. AMD drivers: ATI Catalyst™ 9.8 (ATI Radeon™ HD 48xx), prerelease driver (ATI Radeon HD 5870). Nvidia driver: GeForce 190.62. AMD client: [x86/Stream], v2.9106.513 (beta8). Nvidia client: [x86/CUDA-2.2], v2.9105.512 (beta8). 14 GPU Acceleration in Technical Applications Tomographic Reconstruction • Alain Bonissent, Centre de Physique des Particules de Marseille • Reporting 42-60x* speedups • This image: 7 minutes in optimized C++; 10 seconds in Brook+ EDA Simulation • ACCIT • Currently beta testing applications and reporting >10x speedup** Seismic Processing • Brown Deer • Achieving 120x speedup vs CPU on 3D 2nd order finite-difference time-domain (FDTD) seismic processing algorithm Options Trading • Scotia Capital: • Reported a 28x speedup over a quad-core CPU. Neural Networks • Neurala • Developing Neurala Technology Platform for advanced brain-based machine learning applications • Report achieving 10-200x speedups on biologically inspired neural models 15 Embedded Stream Computing Products 16 ATI Radeon™ E4690 Embedded GPU Performance Gaming and Computation • 384 GLOPS Peak • 320 stream processors @ 600 MHz • Bigger & Wider Memory • 128-bit memory interface • 512 MB GDDR3 in the package • Available as a Chip or MXM Module • MXM 3.0 Type A • 5 Year Planned Longevity of Supply* * Availability subject to change without notice. 17 17 ATI Radeon™E4870 MXM Module Top Performance for Stream Computing • GLOPS Peak single precision • Up to 240 GFLOPS double-precision • 800 stream processors • 115 GB/s Memory Interface • 256 bits wide • 1 GB GDDR5 on the module • MXM Module • MXM 3.0 Type II • Limited availability • Targeted for engineering demos • See an AMD Embedded Sales Representative 18 18 ONLY AMD! 19 Disclaimer and Attribution DISCLAIMER The information presented in this document is for informational purposes only and may contain technical inaccuracies, omissions and typographical errors. The information contained herein is subject to change and may be rendered inaccurate for many reasons, including but not limited to product and roadmap changes, component and motherboard version changes, new model and/or product releases, product differences between differing manufacturers, software changes, BIOS flashes, firmware upgrades, or the like. AMD assumes no obligation to update or otherwise correct or revise this information. However, AMD reserves the right to revise this information and to make changes from time to time to the content hereof without obligation of AMD to notify any person of such revisions or changes. AMD MAKES NO REPRESENTATIONS OR WARRANTIES WITH RESPECT TO THE CONTENTS HEREOF AND ASSUMES NO RESPONSIBILITY FOR ANY INACCURACIES, ERRORS OR OMISSIONS THAT MAY APPEAR IN THIS INFORMATION. AMD SPECIFICALLY DISCLAIMS ANY IMPLIED WARRANTIES OF MERCHANTABILITY OR FITNESS FOR ANY PARTICULAR PURPOSE. IN NO EVENT WILL AMD BE LIABLE TO ANY PERSON FOR ANY DIRECT, INDIRECT, SPECIAL OR OTHER CONSEQUENTIAL DAMAGES ARISING FROM THE USE OF ANY INFORMATION CONTAINED HEREIN, EVEN IF AMD IS EXPRESSLY ADVISED OF THE POSSIBILITY OF SUCH DAMAGES. ATTRIBUTION © 2010 Advanced Micro Devices, Inc. All rights reserved. AMD, the AMD Arrow logo, ATI, the ATI logo, Radeon and combinations thereof are trademarks of Advanced Micro Devices, Inc. OpenCL is a trademark of Apple Inc. used under license to the Khronos Group Inc. Other names are for informational purposes only and may be trademarks of their respective owners. CAUTIONARY STATEMENT This release contains forward-looking statements concerning revenue and other expectations which are made pursuant to the safe harbor provisions of the Private Securities Litigation Reform Act of 1995. Forward-looking statements are commonly identified by words such as “would,” “may,” “expects,” “believes,” “plans,” “intends,” “projects,” and other terms with similar meaning. Investors are cautioned that the forward-looking statements in this release are based on current beliefs, assumptions and expectations, speak only as of the date of this release and involve risks and uncertainties that could cause actual results to differ materially from current expectations. 20.
Recommended publications
  • AMD Firestream™ 9350 GPU Compute Accelerator
    AMD FireStream™ 9350 GPU Compute Accelerator High Density Server GPU Acceleration The AMD FireStream™ 9350 GPU compute accelerator card offers industry- leading performance-per-watt in a highly dense, single-slot form factor. This AMD FireStream 9350 low-profile GPU card is designed for scalable high performance computing > Heterogeneous computing that (HPC) systems that require density and power efficiency. The AMD FireStream leverages AMD GPUs and x86 CPUs 9350 is ideal for a wide range of HPC applications across several industries > High performance per watt at 2.9 including Finance, Energy (Oil and Gas), Geosciences, Life Sciences, GFLOPS / Watt Manufacturing (CAE, CFD, etc.), Defense, and more. > Industry’s most dense server GPU card 1 Utilizing the multi-billion-transistor ASICs developed for AMD Radeon™ graphics > Low profile and passively cooled cards, AMD’s FireStream 9350 cards provide maximum performance-per-slot > Massively parallel, programmable GPU and are designed to meet demanding performance and reliability requirements architecture of HPC systems that can scale to thousands of nodes. AMD FireStream 9350 > Open standard OpenCL™ and cards include a single DisplayPort output. DirectCompute2 AMD FireStream 9350 cards can be used in scalable servers, blades, PCIe® > 2.0 TFLOPS, Single-Precision Peak chassis, and are available from many leading server OEMs and HPC solution > 400 GFLOPS, Double-Precision Peak providers. > 2GB GDDR5 memory Priced competitively, AMD FireStream GPUs offer unparalleled performance > Industry’s only Single-Slot PCIe® Accelerator and value for high performance computing. > PCI Express® 2.1 Compliant Note: For workstation and server > AMD Accelerated Parallel Processing installations that require active (APP) Technology SDK with OpenCL3 cooling, AMD FirePro™ Professional > 3-year planned availability; 3-year Graphics boards are software- limited warranty compatible and offer comparable features.
    [Show full text]
  • ATI Radeon™ HD 4870 Computation Highlights
    AMD Entering the Golden Age of Heterogeneous Computing Michael Mantor Senior GPU Compute Architect / Fellow AMD Graphics Product Group [email protected] 1 The 4 Pillars of massively parallel compute offload •Performance M’Moore’s Law Î 2x < 18 Month s Frequency\Power\Complexity Wall •Power Parallel Î Opportunity for growth •Price • Programming Models GPU is the first successful massively parallel COMMODITY architecture with a programming model that managgped to tame 1000’s of parallel threads in hardware to perform useful work efficiently 2 Quick recap of where we are – Perf, Power, Price ATI Radeon™ HD 4850 4x Performance/w and Performance/mm² in a year ATI Radeon™ X1800 XT ATI Radeon™ HD 3850 ATI Radeon™ HD 2900 XT ATI Radeon™ X1900 XTX ATI Radeon™ X1950 PRO 3 Source of GigaFLOPS per watt: maximum theoretical performance divided by maximum board power. Source of GigaFLOPS per $: maximum theoretical performance divided by price as reported on www.buy.com as of 9/24/08 ATI Radeon™HD 4850 Designed to Perform in Single Slot SP Compute Power 1.0 T-FLOPS DP Compute Power 200 G-FLOPS Core Clock Speed 625 Mhz Stream Processors 800 Memory Type GDDR3 Memory Capacity 512 MB Max Board Power 110W Memory Bandwidth 64 GB/Sec 4 ATI Radeon™HD 4870 First Graphics with GDDR5 SP Compute Power 1.2 T-FLOPS DP Compute Power 240 G-FLOPS Core Clock Speed 750 Mhz Stream Processors 800 Memory Type GDDR5 3.6Gbps Memory Capacity 512 MB Max Board Power 160 W Memory Bandwidth 115.2 GB/Sec 5 ATI Radeon™HD 4870 X2 Incredible Balance of Performance,,, Power, Price
    [Show full text]
  • Download Radeon Grpadhics Catalylist and Driver AMD Catalyst Graphics Driver 13.4 X64
    download radeon grpadhics catalylist and driver AMD Catalyst Graphics Driver 13.4 x64. Fixes : - With the Radeon HD7790 graphics card, corruption may be seen in the game Tomb Raider with TressFX enabled - Texture flickering may be observed in DirectX 9.0c enabled applications - Corruption observed in Tomb Raider with TressFX enabled for AMD Crossfire and single GPU system configurations - Corruption observed on objects and textures in the game, Call of Duty – Black Ops 2 - Corruption observed in the game, Alan Wake (DirectX 9 Mode) when attempting to change in-game settings with Stereoscopic 3D using Tridef - Corruption observed on textures in the game, Battle Field: Bad Company 2 with forced anti-aliasing enabled - Adobe Photoshop CS6 experiences screen flicker under Windows 8 based system - Horizontal line corruption may be observed with forced anti-aliasing enabled in the game, Dishonored - System hang may occur when scrolling through the game selection menu in the game, DiRT 3 - F1 2012 may crash to desktop on systems supporting AMD PowerXpress (Enduro) Technology - Support added for AMD Radeon HD7790 and AMD Radeon HD 7990 - Catalyst Driver optimizations to improve performance in Far Cry 3, Crysis 3, and 3DMark - Significantly improves latency performance in Skyrim, Boderlands 2, Guild Wars 2, Tomb Raider and Hitman Absolution. Performance gains seen with the entire AMD Radeon HD 7000 Series for the following: - Batman Arkham City (1920x1200): Performance improvements up to 5% - Borderlands 2 (2560x1600): Performance improvements
    [Show full text]
  • AMD APP SDK Developer Release Notes
    AMD APP SDK v3.0 Beta Developer Release Notes 1 What’s New in AMD APP SDK v3.0 Beta 1.1 New features in AMD APP SDK v3.0 Beta AMD APP SDK v3.0 Beta includes the following new features: OpenCL 2.0: There are 20 samples that demonstrate various features of OpenCL 2.0 such as Shared Virtual Memory, Platform Atomics, Device-side Enqueue, Pipes, New workgroup built-in functions, Program Scope Variables, Generic Address Space, and OpenCL 2.0 image features. For the complete list of the samples, see the AMD APP SDK Samples Release Notes (AMD_APP_SDK_Release_Notes_Samples.pdf) document. Support for Bolt 1.3 library. 6 additional samples that demonstrate various APIs in the Bolt C++ AMP library. One new sample that demonstrates the consumption of SPIR 1.2 binary. Enhancements and bug fixes in several samples. A lightweight installer that supports the following features: Customized online installation Ability to download the full installer for install and distribution 1.2 New features for AMD CodeXL version 1.6 The following new features in AMD CodeXL version 1.6 provide the following improvements to the developer experience: GPU Profiler support for OpenCL 2.0 API-level debugging for OpenCL 2.0 Power Profiling For information about CodeXL and about how to use CodeXL to gather performance data about your OpenCL application, such as application traces and timeline views, see the CodeXL home page. Developer Release Notes 1 of 4 2 Important Notes OpenCL 2.0 runtime support is limited to 64-bit applications running on 64-bit Windows and Linux operating systems only.
    [Show full text]
  • The Compute Shader
    AMD Entering the Golden Age of Heterogeneous Computing Michael Mantor Senior GPU Compute Architect / Fellow AMD Graphics Product Group [email protected] 1 The 4 Pillars of massively parallel compute offload •Performance M’Moore’s Law Î 2x < 18 Month s Frequency\Power\Complexity Wall •Power Parallel Î Opportunity for growth •Price • Programming Models GPU is the first successful massively parallel COMMODITY architecture with a programming model that managgped to tame 1000’s of parallel threads in hardware to perform useful work efficiently 2 Quick recap of where we are – Perf, Power, Price ATI Radeon™ HD 4850 4x Performance/w and Performance/mm² in a year ATI Radeon™ X1800 XT ATI Radeon™ HD 3850 ATI Radeon™ HD 2900 XT ATI Radeon™ X1900 XTX ATI Radeon™ X1950 PRO 3 Source of GigaFLOPS per watt: maximum theoretical performance divided by maximum board power. Source of GigaFLOPS per $: maximum theoretical performance divided by price as reported on www.buy.com as of 9/24/08 ATI Radeon™HD 4850 Designed to Perform in Single Slot SP Compute Power 1.0 T-FLOPS DP Compute Power 200 G-FLOPS Core Clock Speed 625 Mhz Stream Processors 800 Memory Type GDDR3 Memory Capacity 512 MB Max Board Power 110W Memory Bandwidth 64 GB/Sec 4 ATI Radeon™HD 4870 First Graphics with GDDR5 SP Compute Power 1.2 T-FLOPS DP Compute Power 240 G-FLOPS Core Clock Speed 750 Mhz Stream Processors 800 Memory Type GDDR5 3.6Gbps Memory Capacity 512 MB Max Board Power 160 W Memory Bandwidth 115.2 GB/Sec 5 ATI Radeon™HD 4870 X2 Incredible Balance of Performance,,, Power, Price
    [Show full text]
  • Catalyst™ Software Suite Version 9.5 Release Notes
    Catalyst™ Software Suite Version 9.5 Release Notes This release note provides information on the latest posting of AMD’s industry leading software suite, Catalyst™. This particular software suite updates both the AMD Display Driver, and the Catalyst™ Control Center. This unified driver has been further enhanced to provide the highest level of power, performance, and reliability. The AMD Catalyst™ software suite is the ultimate in performance and stability. For exclusive Catalyst™ updates follow Catalyst Maker on Twitter. This release note provides information on the following: z Web Content z AMD Product Support z Operating Systems Supported z New Features z Performance Improvements z Resolved Issues for the Windows Vista Operating System z Resolved Issues for the Windows XP Operating System z Resolved Issues for the Windows 7 Operating System z Known Issues Under the Windows Vista Operating System z Known Issues Under the Windows XP Operating System z Known Issues Under the Windows 7 Operating System z Installing the Catalyst™ Vista Software Driver z Catalyst™ Crew Driver Feedback ATI Catalyst™ Release Note Version 9.5 1 Web Content The Catalyst™ Software Suite 9.5 contains the following: z Radeon™ display driver 8.612 z HydraVision™ for both Windows XP and Vista z HydraVision™ Basic Edition (Windows XP only) z WDM Driver Install Bundle z Southbridge/IXP Driver z Catalyst™ Control Center Version 8.612 Caution: The Catalyst™ software driver and the Catalyst™ Control Center can be downloaded independently of each other. However, for maximum stability and performance AMD recommends that both components be updated from the same Catalyst™ release.
    [Show full text]
  • AMD APP SDK V2.9.1 Developer Release Notes
    AMD APP SDK v2.9.1 Developer Release Notes 1 What’s New in AMD APP SDK v2.9.1 1.1 New features in AMD APP SDK v2.9.1 AMD APP SDK v2.9.1 includes the following new features: The AMD APP SDK can now be installed on Linux with root as well as non-root permissions. The installation logs display additional details about the status of the installation. Several samples have been enhanced. All the samples are now compatible with the gcc version, 4.8.1. The APARAPI samples have been updated to work with the latest APARAPI (OpenCL trunk) library. The OpenCV-CL samples now work with the OpenCV 2.4.9 library. The AMD APP SDK Bolt samples have been updated to work with the Bolt 1.2 library. For information about the features released in AMD APP SDK v2.9, see the AMD APP SDK v2.9 Developer Release Notes document. 1.2 Key features supported in the AMD Catalyst 14.4 driver Support for Kaveri graphics AMD Radeon R5/R6/R7 Support for AMD Radeon R9 295X Full support for OpenGL 4.4 Support for the following operating systems: Windows 8.1 (32- and 64-bit versions) Windows 7 (32- and 64-bit versions with SP1 or higher) 1.3 New features for AMD CodeXL version 1.4 The following new features for AMD CodeXL version 1.4 provide the following improvements to the developer experience: Microsoft Visual Studio 2013 CodeXL Extension New Shader Analyzer Support for Volcanic Islands family GPUs: debugging, profiling and analyzing.
    [Show full text]
  • Lewis University Dr. James Girard Summer Undergraduate Research Program 2021 Faculty Mentor - Project Application
    Lewis University Dr. James Girard Summer Undergraduate Research Program 2021 Faculty Mentor - Project Application Exploring the Use of High-level Parallel Abstractions and Parallel Computing for Functional and Gate-Level Simulation Acceleration Dr. Lucien Ngalamou Department of Engineering, Computing and Mathematical Sciences Abstract System-on-Chip (SoC) complexity growth has multiplied non-stop, and time-to- market pressure has driven demand for innovation in simulation performance. Logic simulation is the primary method to verify the correctness of such systems. Logic simulation is used heavily to verify the functional correctness of a design for a broad range of abstraction levels. In mainstream industry verification methodologies, typical setups coordinate the validation e↵ort of a complex digital system by distributing logic simulation tasks among vast server farms for months at a time. Yet, the performance of logic simulation is not sufficient to satisfy the demand, leading to incomplete validation processes, escaped functional bugs, and continuous pressure on the EDA1 industry to develop faster simulation solutions. In this research, we will explore a solution that uses high-level parallel abstractions and parallel computing to boost the performance of logic simulation. 1Electronic Design Automation 1 1 Project Description 1.1 Introduction and Background SoC complexity is increasing rapidly, driven by demands in the mobile market, and in- creasingly by the fast-growth of assisted- and autonomous-driving applications. SoC teams utilize many verification technologies to address their complexity and time-to-market chal- lenges; however, logic simulation continues to be the foundation for all verification flows, and continues to account for more than 90% [10] of all verification workloads.
    [Show full text]
  • AMD APP SDK V2.8.1 Developer Release Notes
    AMD APP SDK v2.8.1 Developer Release Notes 1 What’s New in AMD APP SDK v2.8.1 1.1 New features in AMD APP SDK v2.8.1 AMD APP SDK v2.8.1 includes the following new features: Bolt: With the launch of Bolt 1.0, several new samples have been added to demonstrate the use of the features of Bolt 1.0. These features showcase the use of valuable Bolt APIs such as scan, sort, reduce and transform. Other new samples highlight the ease of porting from STL and the performance benefits achieved over equivalent STL implementations. Other samples demonstrate the different fallback options in Bolt 1.0 when no GPU is available. These options include a fallback to multicore-CPU if TBB libraries are installed, or falling all the way back to serial-CPU if needed to ensure your code runs correctly on any platform. OpenCV: AMD has been working closely with the OpenCV open source community to add heterogeneous acceleration capability to the world’s most popular computer vision library. These changes are already integrated into OpenCV and are readily available for developers who want to improve the performance and efficiency of their computer vision applications. The new samples illustrate these improvements and highlight how simple it is to include them in your application. For information on the latest OpenCV enhancements, see Harris’ blog. GCN: AMD recently launched a new Graphics Core Next (GCN) Architecture on several AMD products. GCN is based on a scalar architecture versus the VLIW vector architecture of prior generations, so carefully hand-tuned vectorization to optimize hardware utilization is no longer needed.
    [Show full text]
  • Introduction Hardware Acceleration Philosophy Popular Accelerators In
    Special Purpose Accelerators Special Purpose Accelerators Introduction Recap: General purpose processors excel at various jobs, but are no Theme: Towards Reconfigurable High-Performance Computing mathftch for acce lera tors w hen dea ling w ith spec ilidtialized tas ks Lecture 4 Objectives: Platforms II: Special Purpose Accelerators Define the role and purpose of modern accelerators Provide information about General Purpose GPU computing Andrzej Nowak Contents: CERN openlab (Geneva, Switzerland) Hardware accelerators GPUs and general purpose computing on GPUs Related hardware and software technologies Inverted CERN School of Computing, 3-5 March 2008 1 iCSC2008, Andrzej Nowak, CERN openlab 2 iCSC2008, Andrzej Nowak, CERN openlab Special Purpose Accelerators Special Purpose Accelerators Hardware acceleration philosophy Popular accelerators in general Floating point units Old CPUs were really slow Embedded CPUs often don’t have a hardware FPU 1980’s PCs – the FPU was an optional add on, separate sockets for the 8087 coprocessor Video and image processing MPEG decoders DV decoders HD decoders Digital signal processing (including audio) Sound Blaster Live and friends 3 iCSC2008, Andrzej Nowak, CERN openlab 4 iCSC2008, Andrzej Nowak, CERN openlab Towards Reconfigurable High-Performance Computing Lecture 4 iCSC 2008 3-5 March 2008, CERN Special Purpose Accelerators 1 Special Purpose Accelerators Special Purpose Accelerators Mainstream accelerators today Integrated FPUs Realtime graphics GiGaming car ds Gaming physics
    [Show full text]
  • Amd Firestream 9170 Gpgpu Implementation of a Brook+
    AMD FIRESTREAM 9170 GPGPU IMPLEMENTATION OF A BROOK+ PARALLEL ALGORITHM TO SEARCH FOR COUNTEREXAMPLES TO BEAL’S CONJECTURE _______________ A Thesis Presented to the Faculty of San Diego State University _______________ In Partial Fulfillment of the Requirements for the Degree Master of Science in Computer Science _______________ by Nirav Shah Summer 2010 iii Copyright © 2010 by Nirav Shah All Rights Reserved iv DEDICATION I dedicate my work to my parents for their support and encouragement. v ABSTRACT OF THE THESIS AMD Firestream 9170 GPGPU Implementation of a Brook+ Parallel Algorithm to Search for Counterexamples to Beal’s Conjecture by Nirav Shah Master of Science in Computer Science San Diego State University, 2010 Beal’s Conjecture states that if Ax + By = Cz for integers A,B,C > 0 and integers x,y,z > 2, then A, B, and C must share a common factor. Norvig and others have shown that the conjecture holds for A,B,C,x,y,z < 1000, but the truth of the general conjecture remains unresolved. Extending the search for a counterexample to significantly greater values of the conjecture’s six integer parameters is a task ideally suited to the use of an SIMD parallel algorithm implemented on a GPGPU platform. This thesis project encompassed the design, coding, and testing of such an algorithm, implemented in the Brook+ language for execution on the AMD Firestream 9170 GPGPU. vi TABLE OF CONTENTS PAGE ABSTRACT ...............................................................................................................................v LIST
    [Show full text]
  • AMD Firepro™ W9000 Graphics Accelerator
    AMD FirePro™ W9000 Graphics Accelerator User Guide Part Number: 52015_enu_1.0 ii © 2012 Advanced Micro Devices Inc. All rights reserved. The contents of this document are provided in connection with Advanced Micro Devices, Inc. (“AMD”) products. AMD makes no representations or warranties with respect to the accuracy or completeness of the contents of this publication and reserves the right to discontinue or make changes to products, specifications, product descriptions or documentation at any time without notice. The information contained herein may be of a preliminary or advance nature. No license, whether express, implied, arising by estoppel or otherwise, to any intellectual property rights is granted by this publication. Except as set forth in AMD's Standard Terms and Conditions of Sale, AMD assumes no liability whatsoever, and disclaims any express or implied warranty, relating to its products including, but not limited to, the implied warranty of merchantability, fitness for a particular purpose, or infringement of any intellectual property right. AMD's products are not designed, intended, authorized or warranted for use as components in systems intended for surgical implant into the body, or in other applications intended to support or sustain life, or in any other application in which the failure of AMD's product could create a situation where personal injury, death, or severe property or environmental damage may occur. AMD reserves the right to discontinue or make changes to its products at any time without notice. USE OF THIS PRODUCT IN ANY MANNER THAT COMPLIES WITH THE MPEG-2 STANDARD IS EXPRESSLY PROHIBITED WITHOUT A LICENSE UNDER APPLICABLE PATENTS IN THE MPEG-2 PATENT PORTFOLIO, WHICH LICENSE IS AVAILABLE FROM MPEG LA, L.L.C., 6312 S.
    [Show full text]