Proceedings of 8th International Research Conference, KDU, Published November 2015 A Comparative Study on Web Scraping SCM de S Sirisuriya Department of Computer Science, Faculty of Computing, General Sir John Kotelawala Defence University, Ratmalana, Sri Lanka
[email protected] Abstract— The World Wide Web contains all kinds of I. INTRODUCTION information of different origins; some of those are social, Data is an essential part of any research, either it can be financial, security and academic. Most people access academic, marketing or scientific (SysNucleus, n.d.). information through internet for educational purposes. People might want to collect and analyse data from Information on the web is available in different formats multiple websites. The different websites which belongs and through different access interfaces. Therefore, to the specific category displays information in different indexing or semantic processing of the data through formats. Even with a single website you may not be able to see all the data at once. The data may be spanned websites could be cumbersome. Web Scraping is the across multiple pages under various sections. Most technique which aims to address this issue. Web scraping websites do not allow to save a copy of the data, is used to transform unstructured data on the web into displayed in their web sites to your local storage (Penman structured data that can be stored and analysed in a et al., 2009). The only option is to manually copy and central local database or spreadsheet. There are various paste the data shown by the website to a local file in your web scraping techniques including Traditional copy-and- computer.