Learning Scrapy
Total Page:16
File Type:pdf, Size:1020Kb
www.allitebooks.com Learning Scrapy Learn the art of efficient web scraping and crawling with Python Dimitrios Kouzis-Loukas BIRMINGHAM - MUMBAI www.allitebooks.com Learning Scrapy Copyright © 2016 Packt Publishing All rights reserved. No part of this book may be reproduced, stored in a retrieval system, or transmitted in any form or by any means, without the prior written permission of the publisher, except in the case of brief quotations embedded in critical articles or reviews. Every effort has been made in the preparation of this book to ensure the accuracy of the information presented. However, the information contained in this book is sold without warranty, either express or implied. Neither the author, nor Packt Publishing, and its dealers and distributors will be held liable for any damages caused or alleged to be caused directly or indirectly by this book. Packt Publishing has endeavored to provide trademark information about all of the companies and products mentioned in this book by the appropriate use of capitals. However, Packt Publishing cannot guarantee the accuracy of this information. First published: January 2016 Production reference: 1220116 Published by Packt Publishing Ltd. Livery Place 35 Livery Street Birmingham B3 2PB, UK. ISBN 978-1-78439-978-8 www.packtpub.com www.allitebooks.com Credits Author Project Coordinator Dimitrios Kouzis-Loukas Nidhi Joshi Reviewer Proofreader Lazar Telebak Safis Editing Commissioning Editor Indexer Akram Hussain Monica Ajmera Mehta Acquisition Editor Graphics Subho Gupta Disha Haria Content Development Editor Production Coordinator Kirti Patil Nilesh R. Mohite Technical Editor Cover Work Siddhesh Ghadi Nilesh R. Mohite Copy Editor Priyanka Ravi www.allitebooks.com About the Author Dimitrios Kouzis-Loukas has over fifteen years experience as a topnotch software developer. He uses his acquired knowledge and expertise to teach a wide range of audiences how to write great software, as well. He studied and mastered several disciplines, including mathematics, physics, and microelectronics. His thorough understanding of these subjects helped him raise his standards beyond the scope of "pragmatic solutions." He knows that true solutions should be as certain as the laws of physics, as robust as ECC memories, and as universal as mathematics. Dimitrios now develops distributed, low-latency, highly-availability systems using the latest datacenter technologies. He is language agnostic, yet has a slight preference for Python, C++, and Java. A firm believer in open source software and hardware, he hopes that his contributions will benefit individual communities as well as all of humanity. www.allitebooks.com About the Reviewer Lazar Telebak is a freelance web developer specializing in web scraping, crawling, and indexing web pages using Python libraries/frameworks. He has worked mostly on projects that deal with automation and website scraping, crawling, and exporting data to various formats, including CSV, JSON, XML, and TXT, and databases such as MongoDB, SQLAlchemy, and Postgres. He also has experience in frontend technologies and the languages: HTML, CSS, JS, and jQuery. www.allitebooks.com www.PacktPub.com Support files, eBooks, discount offers, and more For support files and downloads related to your book, please visit www.PacktPub.com. Did you know that Packt offers eBook versions of every book published, with PDF and ePub files available? You can upgrade to the eBook version at www.PacktPub. com and as a print book customer, you are entitled to a discount on the eBook copy. Get in touch with us at [email protected] for more details. At www.PacktPub.com, you can also read a collection of free technical articles, sign up for a range of free newsletters and receive exclusive discounts and offers on Packt books and eBooks. TM https://www2.packtpub.com/books/subscription/packtlib Do you need instant solutions to your IT questions? PacktLib is Packt's online digital book library. Here, you can search, access, and read Packt's entire library of books. Why subscribe? • Fully searchable across every book published by Packt • Copy and paste, print, and bookmark content • On demand and accessible via a web browser Free access for Packt account holders If you have an account with Packt at www.PacktPub.com, you can use this to access PacktLib today and view 9 entirely free books. Simply use your login credentials for immediate access. www.allitebooks.com Table of Contents Preface vii Chapter 1: Introducing Scrapy 1 Hello Scrapy 1 More reasons to love Scrapy 2 About this book: aim and usage 3 The importance of mastering automated data scraping 4 Developing robust, quality applications, and providing realistic schedules 5 Developing quality minimum viable products quickly 5 Scraping gives you scale; Google couldn't use forms 6 Discovering and integrating into your ecosystem 7 Being a good citizen in a world full of spiders 8 What Scrapy is not 8 Summary 9 Chapter 2: Understanding HTML and XPath 11 HTML, the DOM tree representation, and the XPath 11 The URL 12 The HTML document 12 The tree representation 14 What you see on the screen 15 Selecting HTML elements with XPath 16 Useful XPath expressions 17 Using Chrome to get XPath expressions 20 Examples of common tasks 21 Anticipating changes 22 Summary 23 Chapter 3: Basic Crawling 25 Installing Scrapy 26 MacOS 26 [ i ] www.allitebooks.com Table of Contents Windows 27 Linux 27 Ubuntu or Debian Linux 28 Red Hat or CentOS Linux 28 From the latest source 28 Upgrading Scrapy 29 Vagrant: this book's official way to run examples 29 UR2IM – the fundamental scraping process 31 The URL 32 The request and the response 33 The Items 34 A Scrapy project 40 Defining items 41 Writing spiders 42 Populating an item 46 Saving to files 47 Cleaning up – item loaders and housekeeping fields 49 Creating contracts 53 Extracting more URLs 55 Two-direction crawling with a spider 58 Two-direction crawling with a CrawlSpider 61 Summary 62 Chapter 4: From Scrapy to a Mobile App 63 Choosing a mobile application framework 63 Creating a database and a collection 64 Populating the database with Scrapy 66 Creating a mobile application 69 Creating a database access service 70 Setting up the user interface 70 Mapping data to the User Interface 72 Mappings between database fields and User Interface controls 73 Testing, sharing, and exporting your mobile app 74 Summary 75 Chapter 5: Quick Spider Recipes 77 A spider that logs in 78 A spider that uses JSON APIs and AJAX pages 84 Passing arguments between responses 87 A 30-times faster property spider 88 A spider that crawls based on an Excel file 92 Summary 96 [ ii ] www.allitebooks.com Table of Contents Chapter 6: Deploying to Scrapinghub 97 Signing up, signing in, and starting a project 98 Deploying our spiders and scheduling runs 100 Accessing our items 102 Scheduling recurring crawls 104 Summary 104 Chapter 7: Configuration and Management 105 Using Scrapy settings 106 Essential settings 107 Analysis 107 Logging 108 Stats 108 Telnet 108 Performance 110 Stopping crawls early 111 HTTP caching and working offline 111 Example 2 – working offline by using the cache 111 Crawling style 112 Feeds 113 Downloading media 114 Other media 114 Amazon Web Services 115 Using proxies and crawlers 116 Example 4 – using proxies and Crawlera's clever proxy 116 Further settings 117 Project-related settings 118 Extending Scrapy settings 118 Fine-tuning downloading 119 Autothrottle extension settings 119 Memory UsageExtension settings 119 Logging and debugging 120 Summary 120 Chapter 8: Programming Scrapy 121 Scrapy is a Twisted application 122 Deferreds and deferred chains 124 Understanding Twisted and nonblocking I/O – a Python tale 127 Overview of Scrapy architecture 134 Example 1 - a very simple pipeline 137 Signals 138 Example 2 - an extension that measures throughput and latencies 140 [ iii ] www.allitebooks.com Table of Contents Extending beyond middlewares 144 Summary 146 Chapter 9: Pipeline Recipes 147 Using REST APIs 148 Using treq 148 A pipeline that writes to Elasticsearch 148 A pipeline that geocodes using the Google Geocoding API 151 Enabling geoindexing on Elasticsearch 158 Interfacing databases with standard Python clients 159 A pipeline that writes to MySQL 159 Interfacing services using Twisted-specific clients 163 A pipeline that reads/writes to Redis 163 Interfacing CPU-intensive, blocking, or legacy functionality 167 A pipeline that performs CPU-intensive or blocking operations 167 A pipeline that uses binaries or scripts 170 Summary 173 Chapter 10: Understanding Scrapy's Performance 175 Scrapy's engine – an intuitive approach 176 Cascading queuing systems 177 Identifying the bottleneck 178 Scrapy's performance model 179 Getting component utilization using telnet 180 Our benchmark system 182 The standard performance model 185 Solving performance problems 187 Case #1 – saturated CPU 188 Case #2 – blocking code 189 Case #3 – "garbage" on the downloader 191 Case #4 – overflow due to many or large responses 194 Case #5 – overflow due to limited/excessive item concurrency 195 Case #6 – the downloader doesn't have enough to do 197 Troubleshooting flow 199 Summary 200 Chapter 11: Distributed Crawling with Scrapyd and Real-Time Analytics 201 How does the title of a property affect the price? 202 Scrapyd 202 Overview of our distributed system 205 Changes to our spider and middleware 207 Sharded-index crawling 207 [ iv ] Table of Contents Batching crawl URLs 209 Getting start URLs from settings 214 Deploy your project to scrapyd servers