Comparing the Use of Amazon Dynamodb and Apache Hbase for Nosql AWS Whitepaper Comparing the Use of Amazon Dynamodb and Apache Hbase for Nosql AWS Whitepaper

Comparing the Use of Amazon Dynamodb and Apache Hbase for Nosql AWS Whitepaper Comparing the Use of Amazon Dynamodb and Apache Hbase for Nosql AWS Whitepaper

Comparing the Use of Amazon DynamoDB and Apache HBase for NoSQL AWS Whitepaper Comparing the Use of Amazon DynamoDB and Apache HBase for NoSQL AWS Whitepaper Comparing the Use of Amazon DynamoDB and Apache HBase for NoSQL: AWS Whitepaper Copyright © Amazon Web Services, Inc. and/or its affiliates. All rights reserved. Amazon's trademarks and trade dress may not be used in connection with any product or service that is not Amazon's, in any manner that is likely to cause confusion among customers, or in any manner that disparages or discredits Amazon. All other trademarks not owned by Amazon are the property of their respective owners, who may or may not be affiliated with, connected to, or sponsored by Amazon. Comparing the Use of Amazon DynamoDB and Apache HBase for NoSQL AWS Whitepaper Table of Contents Abstract ............................................................................................................................................ 1 Abstract .................................................................................................................................... 1 Introduction ...................................................................................................................................... 2 Amazon DynamoDB Overview ............................................................................................................. 3 Apache HBase Overview ..................................................................................................................... 4 Apache HBase Deployment Options ..................................................................................................... 5 Managed Apache HBase on Amazon EMR (Amazon S3 Storage Mode) ............................................... 5 Managed Apache HBase on Amazon EMR (HDFS Storage Mode) ....................................................... 5 Self-Managed Apache HBase Deployment Model on Amazon EC2 ..................................................... 6 Feature Summary ............................................................................................................................... 7 Use Cases .......................................................................................................................................... 9 Data Models .................................................................................................................................... 10 Data Types .............................................................................................................................. 14 Indexing .................................................................................................................................. 15 Data Processing ............................................................................................................................... 19 Throughput Model .................................................................................................................... 19 Consistency Model .................................................................................................................... 20 Transaction Model .................................................................................................................... 20 Table Operations ...................................................................................................................... 21 Architecture ..................................................................................................................................... 22 Amazon DynamoDB Architecture Overview .................................................................................. 22 Apache HBase Architecture Overview .......................................................................................... 22 Apache HBase on Amazon EMR Architecture Overview .......................................................... 23 Managed Apache HBase on Amazon EMR (Amazon S3 Storage Mode) ..................................... 23 Partitioning ............................................................................................................................. 23 Performance Optimizations ............................................................................................................... 25 Amazon DynamoDB Performance Considerations ......................................................................... 25 On-demand Mode – No Capacity Planning .......................................................................... 25 Provisioned Throughput Considerations .............................................................................. 25 Read Performance Considerations ...................................................................................... 26 Primary Key Design Considerations ..................................................................................... 26 Local Secondary Index Considerations ................................................................................. 26 Global Secondary Index Considerations ............................................................................... 27 Apache HBase Performance Considerations ................................................................................. 27 Memory Considerations ..................................................................................................... 27 Apache HBase Configurations ............................................................................................ 27 Schema Design ................................................................................................................ 28 Client API Considerations .................................................................................................. 28 Compression Techniques ................................................................................................... 28 Apache HBase on Amazon EMR (HDFS Mode) ...................................................................... 29 Apache HBase on Amazon EMR (Amazon S3 Storage Mode) ................................................... 29 Conclusion ....................................................................................................................................... 31 Contributors .................................................................................................................................... 32 Further Reading ............................................................................................................................... 33 Document revisions .......................................................................................................................... 34 Notices ............................................................................................................................................ 35 iii Comparing the Use of Amazon DynamoDB and Apache HBase for NoSQL AWS Whitepaper Abstract Comparing the Use of Amazon DynamoDB and Apache HBase for NoSQL Publication date: January 29, 2020 (Document revisions (p. 34)) Abstract One challenge that architects and developers face today is how to process large volumes of data in a timely, cost effective, and reliable manner. There are several NoSQL solutions in the market and choosing the most appropriate one for your particular use case can be difficult. This paper compares two popular NoSQL data stores—Amazon DynamoDB, a fully managed NoSQL cloud database service, and Apache HBase, an open-source, column-oriented, distributed big data store. Both Amazon DynamoDB and Apache HBase are available in the Amazon Web Services (AWS) Cloud. 1 Comparing the Use of Amazon DynamoDB and Apache HBase for NoSQL AWS Whitepaper Introduction The AWS Cloud accelerates big data analytics. With access to instant scalability and elasticity on AWS, you can focus on analytics instead of infrastructure. Whether you are indexing large data sets, analyzing massive amounts of scientific data, or processing clickstream logs, AWS provides a range of big data products and services that you can leverage for virtually any data-intensive project. There is a wide adoption of NoSQL databases in the growing industry of big data and real-time web applications. Amazon DynamoDB and Apache HBase are examples of NoSQL databases, which are highly optimized to yield significant performance benefits over a traditional relational database management system (RDBMS). Both Amazon DynamoDB and Apache HBase can process large volumes of data with high performance and throughput. Amazon DynamoDB provides a fast, fully managed NoSQL database service. It lets you offload operating and scaling a highly available, distributed database cluster. Apache HBase is an open-source, column- oriented, distributed big data store that runs on the Apache Hadoop framework and is typically deployed on top of the Hadoop Distributed File System (HDFS), which provides a scalable, persistent, storage layer. In the AWS Cloud, you can choose to deploy Apache HBase on Amazon Elastic Compute Cloud (Amazon EC2) and manage it yourself. Alternatively, you can leverage Apache HBase as a managed service on Amazon EMR, a fully managed, hosted Hadoop framework on top of Amazon EC2. With Apache HBase on Amazon EMR, you can use Amazon Simple Storage Service (Amazon S3) as a data store using the EMR File System (EMRFS), an implementation of HDFS that all Amazon EMR clusters use for reading and writing regular files from Amazon EMR directly to Amazon S3. The following figure shows the relationship between

View Full Text

Details

  • File Type
    pdf
  • Upload Time
    -
  • Content Languages
    English
  • Upload User
    Anonymous/Not logged-in
  • File Pages
    38 Page
  • File Size
    -

Download

Channel Download Status
Express Download Enable

Copyright

We respect the copyrights and intellectual property rights of all users. All uploaded documents are either original works of the uploader or authorized works of the rightful owners.

  • Not to be reproduced or distributed without explicit permission.
  • Not used for commercial purposes outside of approved use cases.
  • Not used to infringe on the rights of the original creators.
  • If you believe any content infringes your copyright, please contact us immediately.

Support

For help with questions, suggestions, or problems, please contact us