Tutorial: An example of statistical data analysis using the R environment for statistical computing D G Rossiter Version 1.4; May 6, 2017 Subsoil vs. topsoil clay, by zone Regression Residuals vs. Fitted Values, subsoil clay % 128 80 15 138 ● 17119 137 1 ● 139 70 2 ● 3 10 ● 4 ● 60 ● ● 5 50 0 Slopes: Residual 40 zone 1 : 0.834 Subsoil clay % Subsoil clay ● ● zone 2 : 0.739 zone 3 : 0.564 −5 30 zone 4 : 1.081 overall: 0.829 −10 20 81 −15 10 145 10 20 30 40 50 60 70 80 20 30 40 50 60 70 Topsoil clay % Fitted GLS 2nd−order trend surface, subsoil clay % 340000 335000 330000 N 325000 320000 315000 660000 670000 680000 690000 700000 E Copyright © D G Rossiter 2008 { 2010, 2014, 2017 All rights reserved. Repro- duction and dissemination of the work as a whole (not parts) freely permitted if this original copyright notice is included. Sale or placement on a web site where payment must be made to access this document is strictly prohibited. To adapt or translate please contact the author (
[email protected]). Contents 1 Introduction1 2 Example Data Set2 2.1 Loading the dataset...........................3 2.2 A normalized database structure*...................5 3 Research questions8 4 Univariarte Analysis9 4.1 Univariarte Exploratory Data Analysis................9 4.2 Point estimation; inference of the mean............... 14 4.3 Answers.................................. 15 5 Bivariate correlation and regression 16 5.1 Conceptual issues in correlation and regression........... 16 5.2 Bivariate Exploratory Data Analysis................. 18 5.3 Bivariate Correlation Analysis..................... 22 5.4 Fitting a regression line........................