# Predictor Coef StDev T P Constant X S = R-Sq = 0.0% R-Sq(adj) = 0.

Size: px
Start display at page:

Download "Predictor Coef StDev T P Constant 970667056 616256122 1.58 0.154 X 0.00293 0.06163 0.05 0.963. S = 0.5597 R-Sq = 0.0% R-Sq(adj) = 0."

Transcription

1 Statistical analysis using Microsoft Excel Microsoft Excel spreadsheets have become somewhat of a standard for data storage, at least for smaller data sets. This, along with the program often being packaged with new computers, naturally encourages its use for statistical analyses. This is unfortunate, since Excel is most decidedly not a statistical package. Here s an example of how the numerical inaccuracies in Excel can get you into trouble. Consider the following data set: Data Display Row X Y Here is Minitab output for the regression: Regression Analysis The regression equation is Y =9.71E X Predictor Coef StDev T P Constant X S = R-Sq = 0.0% R-Sq(adj) = 0.0% Analysis of Variance Source DF SS MS F P Regression c 2008, Jeffrey S. Simonoff 1

2 Residual Error Total Now, here are the values obtained when using the regression program available in the Analysis Toolpak of Microsoft Excel 2002 (the same results came from earlier versions of Excel; I will say something about Excel 2003 later): SUMMARY OUTPUT Regression Statistics Multiple R R Square Adjusted R Square Standard Error Observations 10 ANOVA df SS MS F Signif F Regression #NUM! Residual Total Coeff Standard Error t Stat P-value Intercept #NUM! X Variable #NUM! Each of the nine numbers given above is incorrect! The slope estimate has the wrong sign, the estimated standard errors of the coefficients are zero (making it impossible to construct t statistics), and the values of R 2, F and the regression sum of squares are negative! It s obvious here that the output is garbage (even Excel seems to know this, as the #NUM! s seem to imply), but what if the numbers that had come out weren t absurd just wrong? Unless Excel does better at addressing these computational problems, it cannot be considered a serious candidate for use in statistical analysis. What went wrong here? The summary statistics from Excel give us a clue: X Y Mean Mean Standard Error 0 Standard Error Median Median c 2008, Jeffrey S. Simonoff 2

3 Mode #N/A Mode Standard Deviation 0 Standard Deviation Sample Variance 0 Sample Variance Here are the corresponding values if the all X values are decreased by , and all Y values are decreased by The standard deviations and sample variances should, of course, be identical in the two cases, but they are not (the values below are correct): X Y Mean 5.5 Mean 0.61 Standard Error Standard Error Median 5.5 Median Mode #N/A Mode 0 Standard Deviation Standard Deviation Sample Variance Sample Variance Thus, simple descriptive statistics are not trustworthy either in situations where the standard deviation is small relative to the absolute level of the data. Using Excel to analyze multiple regression data brings its own problems. Consider the following data set, provided by Gary Simon: Y X1 X2 X3 X c 2008, Jeffrey S. Simonoff 3

4 There is nothing apparently unusual about these data, and they are, in fact, from an actual clinical experiment. Here is output from Excel 2002 (and earlier versions) for a regression of Y on X1, X2, X3, and X4: c 2008, Jeffrey S. Simonoff 4

5 SUMMARY OUTPUT Regression Statistics Multiple R R Square Adjusted R Square Standard Error Observations 54 ANOVA df SS MS F Significance F Regression Residual Total Coefficients Standard Error t Stat P-value Intercept #NUM! X Variable X Variable #NUM! X Variable #NUM! X Variable #NUM! Obviously there s something strange going on here: the intercept and three of the four coefficients have standard error equal to zero, with undefined p values (why Excel gives what would seem to be t statistics equal to infinity as is a different matter!). One coefficient has more sensible looking output. In any event, Excel does give a fitted regression with associated F statistic and standard error of the estimate. Unfortunately, this is all incorrect. There is no meaningful regression possible here, because the predictors are perfectly collinear (this was done inadvertently by the clinical researcher). That is, no regression model can be fit using all four predictors. Here is what happens if you try to use Minitab to fit the model: Regression Analysis * X4 is highly correlated with other X variables * X4 has been removed from the equation The regression equation is Y = X X X3 Predictor Coef StDev T P c 2008, Jeffrey S. Simonoff 5

6 Constant X X X S = R-Sq = 4.8% R-Sq(adj) = 0.0% Analysis of Variance Source DF SS MS F P Regression Residual Error Total Minitab correctly notes the perfect collinearity among the four predictors and drops one, allowing the regression to proceed. Which variable is dropped out depends on the order of the predictors given to Minitab, but all of the fitted models yield the same R 2, F, and standard error of the estimate (of these statistics, Excel only gets the R 2 right, since it mistakenly thinks that there are four predictors in the model, affecting the other calculations). This is another indication that the numerical methods used by these versions of Excel are hopelessly out of date, and cannot be trusted. These problems have been known in the statistical community for many years, going back to the earliest versions of Excel, but new versions of Excel continued to be released without them being addressed. Finally, with the release of Excel 2003, the basic algorithmic instabilities in the regression function LINEST() were addressed, and the software yields correct answers for these regression examples (as well as for the univariate statistics example). Excel 2003 also recognizes the perfect collinearity in the previous example, and gives the slope coefficient for one variable as 0 with a standard error of 0 (although it still tries to calculate a t-test, resulting in t = 65535). Unfortunately, not all of Excel s problems were fixed in the latest version. Here is another data set: X1 X c 2008, Jeffrey S. Simonoff 6

7 Let s say that these are paired data, and we are interested in whether the population mean for X1 is different from that of X2. Minitab output for a paired sample t test is as follows: Paired T-Test and Confidence Interval Paired T for X1 - X2 N Mean StDev SE Mean X X Difference % CI for mean difference: (0.09, 4.91) T-Test of mean difference = 0 (vs not = 0): T-Value = 2.34 P-Value = Here is output from Excel: t-test: Paired Two Sample for Means Variable 1 Variable 2 Mean Variance Observations Pearson Correlation 0 Hypothesized Mean Difference 0 df 9 t Stat P(T<=t) one-tail t Critical one-tail P(T<=t) two-tail t Critical two-tail c 2008, Jeffrey S. Simonoff 7

8 The output is (basically) the same, of course, as it should be. Now, let s say that the data have a couple of more observations with missing data: X1 X Obviously, these two additional observations don t provide any information about the difference between X1 and X2, so they shouldn t change the paired t test. They don t change the Minitab output, but look at the Excel output: t-test: Paired Two Sample for Means Variable 1 Variable 2 Mean Variance Observations Pearson Correlation 0 Hypothesized Mean Difference 0 df 10 t Stat P(T<=t) one-tail t Critical one-tail P(T<=t) two-tail t Critical two-tail c 2008, Jeffrey S. Simonoff 8

9 I don t know what Excel has done here, but it s certainly not right! The statistics for each variable separately (means, variances) are correct, but irrelevant. Interestingly, the results were different (but still wrong) in Excel 97, so apparently a new error was introduced in the later versions of the software, which has still not been corrected. The same results are obtained if the observations with missing data are put in the first two rows, rather than the last two. These are not the results that are obtained if the two additional observations are collapsed into one (with no missing data), which are correct: t-test: Paired Two Sample for Means Variable 1 Variable 2 Mean Variance Observations Pearson Correlation Hypothesized Mean Difference 0 df 10 t Stat P(T<=t) one-tail t Critical one-tail P(T<=t) two-tail t Critical two-tail Missing data can cause other problems in all versions of Excel. For example, if you try to perform a regression using variables with missing data (either in the predictors or target), you get the error message Regression - LINEST() function returns error. Please check input ranges again. This means that you would have to cut and paste the variables to new locations, omitting any rows with missing data yourself. Other, less catastrophic, problems come from using (any version of) Excel to do statistical analysis. Excel requires you to put all predictors in a regression in contiguous columns, requiring repeated reorganizations of the data as different models are fit. Further, the software does not provide any record of what is done, making it virtually impossible to document or duplicate what was done. In addition, you might think that what Excel calls a Normal probability plot is a normal (qq) plot of the residuals, but you d be wrong. In fact, the plot that comes out is a plot of the ordered target values y (i) versus 50(2i 1)/n (the ordered percentiles). That is, it is effectively a plot checking uniformity of the target variable (something of no interest in a regression context), and has nothing to do with c 2008, Jeffrey S. Simonoff 9

10 normality at all! You should also know that if a column of an Excel spreadsheet is too narrow (so that pound signs replace an actual entry), and you copy and paste the column into Word, the pound signs are pasted over, not the actual entry (this cannot happen when pasting from Minitab, since it will automatically convert a number to scientific notation and automatically widen a text column to be as wide as is necessary). To my way of thinking, at a bare minimum, to be considered even remotely useful any regression package must do the following (I ve left out the obvious things, such as providing accurate least squares estimates, t-tests, F-tests, R 2 values, p-values, and so on): (1) Provide an audit trail, so that changes in data and different analyses can be tracked and recorded. An alternative model (used by the packages S-Plus, R, and SAS, for example) is the construction and use of a command language, so that a data analyst can write scripts for this purpose. (2) Allow for totally flexible choices of predictors from a spreadsheet/ worksheet (i.e., nonconsecutive columns). (3) Have a large set of easy-to-use transformations. (4) Easily produce correct residual plots. (5) Easily produce columns of standard regression diagnostics (standardized residuals, leverage values, Cook s distances), and fitted values. (6) Provide variance inflation factors. (7) Provide confidence and prediction intervals for new observations. (8) Allow for weights (i.e., weighted least squares). (9) Handle categorical predictors, at least at a basic level. Items 1 to 7 are fundamentally necessary even for regression at the level of this course, and if you ever need to perform a regression analysis after you leave the course, you almost immediately need items 8 and 9, so encouraging people to use software to do regression if those things aren t provided is to my mind not a good idea. I ve left other things out that are very useful (for example, the ability to easily create subsets of the data based on conditional statements, reasonably high quality graphics [histograms, scatter plots, side-by-side boxplots, etc.], correct accounting for missing values, built-in probability distributions for the F, t, and binomial, so tests for other hypotheses can be constructed, etc.), but you get the point. The solution to all of these problems is to perform statistical analyses using the appropriate tool a good statistical package. Many such packages are available, usually with a Windows-type graphical user interface (such as Minitab), often costing \$100 or less. A remarkably powerful package, R, is free! c 2008, Jeffrey S. Simonoff 10

12 6. Knüsel, L. (2005), On the accuracy of statistical distributions in Microsoft Excel 2003, Computational Statistics and Data Analysis, 48, McCullough, B.D. (2002), Does Microsoft fix errors in Excel? Proceedings of the 2001 Joint Statistical Meetings. 8. McCullough, B.D. and Heiser, D.A. (2008), On the accuracy of statistical procedures in Microsoft Excel 2007, Computational Statistics and Data Analysis, 52, McCullough, B.D. and Wilson, B. (1999), On the accuracy of statistical procedures in Microsoft Excel 97, Computational Statistics and Data Analysis, 31, McCullough, B.D. and Wilson, B. (2002), On the accuracy of statistical procedures in Microsoft Excel 2000 and Excel XP, Computational Statistics and Data Analysis, 40, McCullough, B.D. and Wilson, B. (2005), On the accuracy of statistical procedures in Microsoft Excel 2003, Computational Statistics and Data Analysis, 49, Pottel, H. (2001), Statistical flaws in Excel, ( nhunt/pottel.pdf). 13. Rotz, W., Falk, E., Wood, D., and Mulrow, J. (2002), A comparison of random number generators used in business, Proceedings of the 2001 Joint Statistical Meetings. c 2008, Jeffrey S. Simonoff 12

### Problems With Using Microsoft Excel for Statistics

Proceedings of the Annual Meeting of the American Statistical Association, August 5-9, 2001 Problems With Using Microsoft Excel for Statistics Jonathan D. Cryer (Jon-Cryer@uiowa.edu) Department of Statistics

### Regression Analysis: A Complete Example

Regression Analysis: A Complete Example This section works out an example that includes all the topics we have discussed so far in this chapter. A complete example of regression analysis. PhotoDisc, Inc./Getty

### 3 The spreadsheet execution model and its consequences

Paper SP06 On the use of spreadsheets in statistical analysis Martin Gregory, Merck Serono, Darmstadt, Germany 1 Abstract While most of us use spreadsheets in our everyday work, usually for keeping track

### c 2015, Jeffrey S. Simonoff 1

Modeling Lowe s sales Forecasting sales is obviously of crucial importance to businesses. Revenue streams are random, of course, but in some industries general economic factors would be expected to have

### Statistical flaws in Excel

Statistical flaws in Excel Hans Pottel Innogenetics NV, Technologiepark 6, 9052 Zwijnaarde, Belgium Introduction In 1980, our ENBIS Newsletter Editor, Tony Greenfield, published a paper on Statistical

### Chapter 7: Simple linear regression Learning Objectives

Chapter 7: Simple linear regression Learning Objectives Reading: Section 7.1 of OpenIntro Statistics Video: Correlation vs. causation, YouTube (2:19) Video: Intro to Linear Regression, YouTube (5:18) -

### Bill Burton Albert Einstein College of Medicine william.burton@einstein.yu.edu April 28, 2014 EERS: Managing the Tension Between Rigor and Resources 1

Bill Burton Albert Einstein College of Medicine william.burton@einstein.yu.edu April 28, 2014 EERS: Managing the Tension Between Rigor and Resources 1 Calculate counts, means, and standard deviations Produce

### Data Analysis Tools. Tools for Summarizing Data

Data Analysis Tools This section of the notes is meant to introduce you to many of the tools that are provided by Excel under the Tools/Data Analysis menu item. If your computer does not have that tool

### 1. What is the critical value for this 95% confidence interval? CV = z.025 = invnorm(0.025) = 1.96

1 Final Review 2 Review 2.1 CI 1-propZint Scenario 1 A TV manufacturer claims in its warranty brochure that in the past not more than 10 percent of its TV sets needed any repair during the first two years

### EXCEL Analysis TookPak [Statistical Analysis] 1. First of all, check to make sure that the Analysis ToolPak is installed. Here is how you do it:

EXCEL Analysis TookPak [Statistical Analysis] 1 First of all, check to make sure that the Analysis ToolPak is installed. Here is how you do it: a. From the Tools menu, choose Add-Ins b. Make sure Analysis

### MULTIPLE LINEAR REGRESSION ANALYSIS USING MICROSOFT EXCEL. by Michael L. Orlov Chemistry Department, Oregon State University (1996)

MULTIPLE LINEAR REGRESSION ANALYSIS USING MICROSOFT EXCEL by Michael L. Orlov Chemistry Department, Oregon State University (1996) INTRODUCTION In modern science, regression analysis is a necessary part

### When to use Excel. When NOT to use Excel 9/24/2014

Analyzing Quantitative Assessment Data with Excel October 2, 2014 Jeremy Penn, Ph.D. Director When to use Excel You want to quickly summarize or analyze your assessment data You want to create basic visual

### 17. SIMPLE LINEAR REGRESSION II

17. SIMPLE LINEAR REGRESSION II The Model In linear regression analysis, we assume that the relationship between X and Y is linear. This does not mean, however, that Y can be perfectly predicted from X.

### Guide to Microsoft Excel for calculations, statistics, and plotting data

Page 1/47 Guide to Microsoft Excel for calculations, statistics, and plotting data Topic Page A. Writing equations and text 2 1. Writing equations with mathematical operations 2 2. Writing equations with

Bowerman, O'Connell, Aitken Schermer, & Adcock, Business Statistics in Practice, Canadian edition Online Learning Centre Technology Step-by-Step - Excel Microsoft Excel is a spreadsheet software application

### business statistics using Excel OXFORD UNIVERSITY PRESS Glyn Davis & Branko Pecar

business statistics using Excel Glyn Davis & Branko Pecar OXFORD UNIVERSITY PRESS Detailed contents Introduction to Microsoft Excel 2003 Overview Learning Objectives 1.1 Introduction to Microsoft Excel

### TRINITY COLLEGE. Faculty of Engineering, Mathematics and Science. School of Computer Science & Statistics

UNIVERSITY OF DUBLIN TRINITY COLLEGE Faculty of Engineering, Mathematics and Science School of Computer Science & Statistics BA (Mod) Enter Course Title Trinity Term 2013 Junior/Senior Sophister ST7002

### KSTAT MINI-MANUAL. Decision Sciences 434 Kellogg Graduate School of Management

KSTAT MINI-MANUAL Decision Sciences 434 Kellogg Graduate School of Management Kstat is a set of macros added to Excel and it will enable you to do the statistics required for this course very easily. To

### Chapter 13 Introduction to Linear Regression and Correlation Analysis

Chapter 3 Student Lecture Notes 3- Chapter 3 Introduction to Linear Regression and Correlation Analsis Fall 2006 Fundamentals of Business Statistics Chapter Goals To understand the methods for displaing

### 2. Simple Linear Regression

Research methods - II 3 2. Simple Linear Regression Simple linear regression is a technique in parametric statistics that is commonly used for analyzing mean response of a variable Y which changes according

### 1.1. Simple Regression in Excel (Excel 2010).

.. Simple Regression in Excel (Excel 200). To get the Data Analysis tool, first click on File > Options > Add-Ins > Go > Select Data Analysis Toolpack & Toolpack VBA. Data Analysis is now available under

### How To Run Statistical Tests in Excel

How To Run Statistical Tests in Excel Microsoft Excel is your best tool for storing and manipulating data, calculating basic descriptive statistics such as means and standard deviations, and conducting

### Using Excel for Statistics Tips and Warnings

Using Excel for Statistics Tips and Warnings November 2000 University of Reading Statistical Services Centre Biometrics Advisory and Support Service to DFID Contents 1. Introduction 3 1.1 Data Entry and

### Statistical Functions in Excel

Statistical Functions in Excel There are many statistical functions in Excel. Moreover, there are other functions that are not specified as statistical functions that are helpful in some statistical analyses.

### Final Exam Practice Problem Answers

Final Exam Practice Problem Answers The following data set consists of data gathered from 77 popular breakfast cereals. The variables in the data set are as follows: Brand: The brand name of the cereal

### Multiple Regression in SPSS This example shows you how to perform multiple regression. The basic command is regression : linear.

Multiple Regression in SPSS This example shows you how to perform multiple regression. The basic command is regression : linear. In the main dialog box, input the dependent variable and several predictors.

### Getting Correct Results from PROC REG

Getting Correct Results from PROC REG Nathaniel Derby, Statis Pro Data Analytics, Seattle, WA ABSTRACT PROC REG, SAS s implementation of linear regression, is often used to fit a line without checking

### NCSS Statistical Software Principal Components Regression. In ordinary least squares, the regression coefficients are estimated using the formula ( )

Chapter 340 Principal Components Regression Introduction is a technique for analyzing multiple regression data that suffer from multicollinearity. When multicollinearity occurs, least squares estimates

### 11. Analysis of Case-control Studies Logistic Regression

Research methods II 113 11. Analysis of Case-control Studies Logistic Regression This chapter builds upon and further develops the concepts and strategies described in Ch.6 of Mother and Child Health:

### Using Excel s Analysis ToolPak Add-In

Using Excel s Analysis ToolPak Add-In S. Christian Albright, September 2013 Introduction This document illustrates the use of Excel s Analysis ToolPak add-in for data analysis. The document is aimed at

### An introduction to using Microsoft Excel for quantitative data analysis

Contents An introduction to using Microsoft Excel for quantitative data analysis 1 Introduction... 1 2 Why use Excel?... 2 3 Quantitative data analysis tools in Excel... 3 4 Entering your data... 6 5 Preparing

### 1. The parameters to be estimated in the simple linear regression model Y=α+βx+ε ε~n(0,σ) are: a) α, β, σ b) α, β, ε c) a, b, s d) ε, 0, σ

STA 3024 Practice Problems Exam 2 NOTE: These are just Practice Problems. This is NOT meant to look just like the test, and it is NOT the only thing that you should study. Make sure you know all the material

### Premaster Statistics Tutorial 4 Full solutions

Premaster Statistics Tutorial 4 Full solutions Regression analysis Q1 (based on Doane & Seward, 4/E, 12.7) a. Interpret the slope of the fitted regression = 125,000 + 150. b. What is the prediction for

### Answer: C. The strength of a correlation does not change if units change by a linear transformation such as: Fahrenheit = 32 + (5/9) * Centigrade

Statistics Quiz Correlation and Regression -- ANSWERS 1. Temperature and air pollution are known to be correlated. We collect data from two laboratories, in Boston and Montreal. Boston makes their measurements

### Business Statistics. Successful completion of Introductory and/or Intermediate Algebra courses is recommended before taking Business Statistics.

Business Course Text Bowerman, Bruce L., Richard T. O'Connell, J. B. Orris, and Dawn C. Porter. Essentials of Business, 2nd edition, McGraw-Hill/Irwin, 2008, ISBN: 978-0-07-331988-9. Required Computing

### 1/27/2013. PSY 512: Advanced Statistics for Psychological and Behavioral Research 2

PSY 512: Advanced Statistics for Psychological and Behavioral Research 2 Introduce moderated multiple regression Continuous predictor continuous predictor Continuous predictor categorical predictor Understand

### How To Test For Significance On A Data Set

Non-Parametric Univariate Tests: 1 Sample Sign Test 1 1 SAMPLE SIGN TEST A non-parametric equivalent of the 1 SAMPLE T-TEST. ASSUMPTIONS: Data is non-normally distributed, even after log transforming.

### Using R for Linear Regression

Using R for Linear Regression In the following handout words and symbols in bold are R functions and words and symbols in italics are entries supplied by the user; underlined words and symbols are optional

### DEPARTMENT OF PSYCHOLOGY UNIVERSITY OF LANCASTER MSC IN PSYCHOLOGICAL RESEARCH METHODS ANALYSING AND INTERPRETING DATA 2 PART 1 WEEK 9

DEPARTMENT OF PSYCHOLOGY UNIVERSITY OF LANCASTER MSC IN PSYCHOLOGICAL RESEARCH METHODS ANALYSING AND INTERPRETING DATA 2 PART 1 WEEK 9 Analysis of covariance and multiple regression So far in this course,

### Exercise 1.12 (Pg. 22-23)

Individuals: The objects that are described by a set of data. They may be people, animals, things, etc. (Also referred to as Cases or Records) Variables: The characteristics recorded about each individual.

### The importance of graphing the data: Anscombe s regression examples

The importance of graphing the data: Anscombe s regression examples Bruce Weaver Northern Health Research Conference Nipissing University, North Bay May 30-31, 2008 B. Weaver, NHRC 2008 1 The Objective

### " Y. Notation and Equations for Regression Lecture 11/4. Notation:

Notation: Notation and Equations for Regression Lecture 11/4 m: The number of predictor variables in a regression Xi: One of multiple predictor variables. The subscript i represents any number from 1 through

### Factors affecting online sales

Factors affecting online sales Table of contents Summary... 1 Research questions... 1 The dataset... 2 Descriptive statistics: The exploratory stage... 3 Confidence intervals... 4 Hypothesis tests... 4

### The Dummy s Guide to Data Analysis Using SPSS

The Dummy s Guide to Data Analysis Using SPSS Mathematics 57 Scripps College Amy Gamble April, 2001 Amy Gamble 4/30/01 All Rights Rerserved TABLE OF CONTENTS PAGE Helpful Hints for All Tests...1 Tests

### Analysing Questionnaires using Minitab (for SPSS queries contact -) Graham.Currell@uwe.ac.uk

Analysing Questionnaires using Minitab (for SPSS queries contact -) Graham.Currell@uwe.ac.uk Structure As a starting point it is useful to consider a basic questionnaire as containing three main sections:

### Good luck! BUSINESS STATISTICS FINAL EXAM INSTRUCTIONS. Name:

Glo bal Leadership M BA BUSINESS STATISTICS FINAL EXAM Name: INSTRUCTIONS 1. Do not open this exam until instructed to do so. 2. Be sure to fill in your name before starting the exam. 3. You have two hours

### Elementary Statistics Sample Exam #3

Elementary Statistics Sample Exam #3 Instructions. No books or telephones. Only the supplied calculators are allowed. The exam is worth 100 points. 1. A chi square goodness of fit test is considered to

### Assignment objectives:

Assignment objectives: Regression Pivot table Exercise #1- Simple Linear Regression Often the relationship between two variables, Y and X, can be adequately represented by a simple linear equation of the

Table of Contents Introduction to Minitab...2 Example 1 One-Way ANOVA...3 Determining Sample Size in One-way ANOVA...8 Example 2 Two-factor Factorial Design...9 Example 3: Randomized Complete Block Design...14

CHAPTER 9 ADD-INS: ENHANCING EXCEL This chapter discusses the following topics: WHAT CAN AN ADD-IN DO? WHY USE AN ADD-IN (AND NOT JUST EXCEL MACROS/PROGRAMS)? ADD INS INSTALLED WITH EXCEL OTHER ADD-INS

### Module 5: Multiple Regression Analysis

Using Statistical Data Using to Make Statistical Decisions: Data Multiple to Make Regression Decisions Analysis Page 1 Module 5: Multiple Regression Analysis Tom Ilvento, University of Delaware, College

### Course Text. Required Computing Software. Course Description. Course Objectives. StraighterLine. Business Statistics

Course Text Business Statistics Lind, Douglas A., Marchal, William A. and Samuel A. Wathen. Basic Statistics for Business and Economics, 7th edition, McGraw-Hill/Irwin, 2010, ISBN: 9780077384470 [This

### 4. Multiple Regression in Practice

30 Multiple Regression in Practice 4. Multiple Regression in Practice The preceding chapters have helped define the broad principles on which regression analysis is based. What features one should look

### Statistics Review PSY379

Statistics Review PSY379 Basic concepts Measurement scales Populations vs. samples Continuous vs. discrete variable Independent vs. dependent variable Descriptive vs. inferential stats Common analyses

### Simple Regression Theory II 2010 Samuel L. Baker

SIMPLE REGRESSION THEORY II 1 Simple Regression Theory II 2010 Samuel L. Baker Assessing how good the regression equation is likely to be Assignment 1A gets into drawing inferences about how close the

### Simple Predictive Analytics Curtis Seare

Using Excel to Solve Business Problems: Simple Predictive Analytics Curtis Seare Copyright: Vault Analytics July 2010 Contents Section I: Background Information Why use Predictive Analytics? How to use

### Regression step-by-step using Microsoft Excel

Step 1: Regression step-by-step using Microsoft Excel Notes prepared by Pamela Peterson Drake, James Madison University Type the data into the spreadsheet The example used throughout this How to is a regression

### Simple Linear Regression Inference

Simple Linear Regression Inference 1 Inference requirements The Normality assumption of the stochastic term e is needed for inference even if it is not a OLS requirement. Therefore we have: Interpretation

### Multiple Linear Regression

Multiple Linear Regression A regression with two or more explanatory variables is called a multiple regression. Rather than modeling the mean response as a straight line, as in simple regression, it is

### Software for Intro Stats: Is Excel an Option?

Software for Intro Stats: Is Excel an Option? Roger L. Berger Arizona State University August 2006 New Researchers Conference Seattle, WA 1 OUTLINE 1. Course description 2. Survey 3. Why use Excel? 4.

### Directions for using SPSS

Directions for using SPSS Table of Contents Connecting and Working with Files 1. Accessing SPSS... 2 2. Transferring Files to N:\drive or your computer... 3 3. Importing Data from Another File Format...

### 2013 MBA Jump Start Program. Statistics Module Part 3

2013 MBA Jump Start Program Module 1: Statistics Thomas Gilbert Part 3 Statistics Module Part 3 Hypothesis Testing (Inference) Regressions 2 1 Making an Investment Decision A researcher in your firm just

### Using Excel for Statistical Analysis

2010 Using Excel for Statistical Analysis Microsoft Excel is spreadsheet software that is used to store information in columns and rows, which can then be organized and/or processed. Excel is a powerful

### EXCEL Tutorial: How to use EXCEL for Graphs and Calculations.

EXCEL Tutorial: How to use EXCEL for Graphs and Calculations. Excel is powerful tool and can make your life easier if you are proficient in using it. You will need to use Excel to complete most of your

### Univariate Regression

Univariate Regression Correlation and Regression The regression line summarizes the linear relationship between 2 variables Correlation coefficient, r, measures strength of relationship: the closer r is

### GLM I An Introduction to Generalized Linear Models

GLM I An Introduction to Generalized Linear Models CAS Ratemaking and Product Management Seminar March 2009 Presented by: Tanya D. Havlicek, Actuarial Assistant 0 ANTITRUST Notice The Casualty Actuarial

### An SPSS companion book. Basic Practice of Statistics

An SPSS companion book to Basic Practice of Statistics SPSS is owned by IBM. 6 th Edition. Basic Practice of Statistics 6 th Edition by David S. Moore, William I. Notz, Michael A. Flinger. Published by

### LAB 4 INSTRUCTIONS CONFIDENCE INTERVALS AND HYPOTHESIS TESTING

LAB 4 INSTRUCTIONS CONFIDENCE INTERVALS AND HYPOTHESIS TESTING In this lab you will explore the concept of a confidence interval and hypothesis testing through a simulation problem in engineering setting.

### Introduction to Statistical Computing in Microsoft Excel By Hector D. Flores; hflores@rice.edu, and Dr. J.A. Dobelman

Introduction to Statistical Computing in Microsoft Excel By Hector D. Flores; hflores@rice.edu, and Dr. J.A. Dobelman Statistics lab will be mainly focused on applying what you have learned in class with

### Chapter 2 Simple Comparative Experiments Solutions

Solutions from Montgomery, D. C. () Design and Analysis of Experiments, Wiley, NY Chapter Simple Comparative Experiments Solutions - The breaking strength of a fiber is required to be at least 5 psi. Past

### Analysis of Variance. MINITAB User s Guide 2 3-1

3 Analysis of Variance Analysis of Variance Overview, 3-2 One-Way Analysis of Variance, 3-5 Two-Way Analysis of Variance, 3-11 Analysis of Means, 3-13 Overview of Balanced ANOVA and GLM, 3-18 Balanced

### Figure 1. An embedded chart on a worksheet.

8. Excel Charts and Analysis ToolPak Charts, also known as graphs, have been an integral part of spreadsheets since the early days of Lotus 1-2-3. Charting features have improved significantly over the

### Using Excel for Data Manipulation and Statistical Analysis: How-to s and Cautions

2010 Using Excel for Data Manipulation and Statistical Analysis: How-to s and Cautions This document describes how to perform some basic statistical procedures in Microsoft Excel. Microsoft Excel is spreadsheet

### MGT 267 PROJECT. Forecasting the United States Retail Sales of the Pharmacies and Drug Stores. Done by: Shunwei Wang & Mohammad Zainal

MGT 267 PROJECT Forecasting the United States Retail Sales of the Pharmacies and Drug Stores Done by: Shunwei Wang & Mohammad Zainal Dec. 2002 The retail sale (Million) ABSTRACT The present study aims

### Testing Group Differences using T-tests, ANOVA, and Nonparametric Measures

Testing Group Differences using T-tests, ANOVA, and Nonparametric Measures Jamie DeCoster Department of Psychology University of Alabama 348 Gordon Palmer Hall Box 870348 Tuscaloosa, AL 35487-0348 Phone:

### SPSS Tests for Versions 9 to 13

SPSS Tests for Versions 9 to 13 Chapter 2 Descriptive Statistic (including median) Choose Analyze Descriptive statistics Frequencies... Click on variable(s) then press to move to into Variable(s): list

### STATISTICAL ANALYSIS WITH EXCEL COURSE OUTLINE

STATISTICAL ANALYSIS WITH EXCEL COURSE OUTLINE Perhaps Microsoft has taken pains to hide some of the most powerful tools in Excel. These add-ins tools work on top of Excel, extending its power and abilities

: Table of Contents... 1 Overview of Model... 1 Dispersion... 2 Parameterization... 3 Sigma-Restricted Model... 3 Overparameterized Model... 4 Reference Coding... 4 Model Summary (Summary Tab)... 5 Summary

### 5. Linear Regression

5. Linear Regression Outline.................................................................... 2 Simple linear regression 3 Linear model............................................................. 4

### Traditional Conjoint Analysis with Excel

hapter 8 Traditional onjoint nalysis with Excel traditional conjoint analysis may be thought of as a multiple regression problem. The respondent s ratings for the product concepts are observations on the

### Statistical Models in R

Statistical Models in R Some Examples Steven Buechler Department of Mathematics 276B Hurley Hall; 1-6233 Fall, 2007 Outline Statistical Models Linear Models in R Regression Regression analysis is the appropriate

Business Valuation Review Regression Analysis in Valuation Engagements By: George B. Hawkins, ASA, CFA Introduction Business valuation is as much as art as it is science. Sage advice, however, quantitative

### Using Microsoft Excel for Probability and Statistics

Introduction Using Microsoft Excel for Probability and Despite having been set up with the business user in mind, Microsoft Excel is rather poor at handling precisely those aspects of statistics which

### Statistics with Ms Excel

Statistics with Ms Excel Simple Statistics with Excel and Minitab Elementary Concepts in Statistics Multiple Regression ANOVA Elementary Concepts in Statistics Overview of Elementary Concepts in Statistics.

### One-Way Analysis of Variance (ANOVA) Example Problem

One-Way Analysis of Variance (ANOVA) Example Problem Introduction Analysis of Variance (ANOVA) is a hypothesis-testing technique used to test the equality of two or more population (or treatment) means

### Biology statistics made simple using Excel

Millar Biology statistics made simple using Excel Biology statistics made simple using Excel Neil Millar Spreadsheet programs such as Microsoft Excel can transform the use of statistics in A-level science

### The Volatility Index Stefan Iacono University System of Maryland Foundation

1 The Volatility Index Stefan Iacono University System of Maryland Foundation 28 May, 2014 Mr. Joe Rinaldi 2 The Volatility Index Introduction The CBOE s VIX, often called the market fear gauge, measures

### Difference of Means and ANOVA Problems

Difference of Means and Problems Dr. Tom Ilvento FREC 408 Accounting Firm Study An accounting firm specializes in auditing the financial records of large firm It is interested in evaluating its fee structure,particularly

### Module 6: Introduction to Time Series Forecasting

Using Statistical Data to Make Decisions Module 6: Introduction to Time Series Forecasting Titus Awokuse and Tom Ilvento, University of Delaware, College of Agriculture and Natural Resources, Food and

### 5 Analysis of Variance models, complex linear models and Random effects models

5 Analysis of Variance models, complex linear models and Random effects models In this chapter we will show any of the theoretical background of the analysis. The focus is to train the set up of ANOVA

### Simple Methods and Procedures Used in Forecasting

Simple Methods and Procedures Used in Forecasting The project prepared by : Sven Gingelmaier Michael Richter Under direction of the Maria Jadamus-Hacura What Is Forecasting? Prediction of future events

### Descriptive Statistics

Descriptive Statistics Primer Descriptive statistics Central tendency Variation Relative position Relationships Calculating descriptive statistics Descriptive Statistics Purpose to describe or summarize

### EXPLORATORY DATA ANALYSIS

CHAPTER 3 EXPLORATORY DATA ANALYSIS HYPOTHESIS TESTING VERSUS EXPLORATORY DATA ANALYSIS GETTING TO KNOW THE DATA SET DEALING WITH CORRELATED VARIABLES EXPLORING CATEGORICAL VARIABLES USING EDA TO UNCOVER

### Stepwise Regression. Chapter 311. Introduction. Variable Selection Procedures. Forward (Step-Up) Selection

Chapter 311 Introduction Often, theory and experience give only general direction as to which of a pool of candidate variables (including transformed variables) should be included in the regression model.

### Analyzing Research Data Using Excel

Analyzing Research Data Using Excel Fraser Health Authority, 2012 The Fraser Health Authority ( FH ) authorizes the use, reproduction and/or modification of this publication for purposes other than commercial

### DESCRIPTIVE STATISTICS. The purpose of statistics is to condense raw data to make it easier to answer specific questions; test hypotheses.

DESCRIPTIVE STATISTICS The purpose of statistics is to condense raw data to make it easier to answer specific questions; test hypotheses. DESCRIPTIVE VS. INFERENTIAL STATISTICS Descriptive To organize,

### Section Format Day Begin End Building Rm# Instructor. 001 Lecture Tue 6:45 PM 8:40 PM Silver 401 Ballerini

NEW YORK UNIVERSITY ROBERT F. WAGNER GRADUATE SCHOOL OF PUBLIC SERVICE Course Syllabus Spring 2016 Statistical Methods for Public, Nonprofit, and Health Management Section Format Day Begin End Building

### Case Study Call Centre Hypothesis Testing

is often thought of as an advanced Six Sigma tool but it is a very useful technique with many applications and in many cases it can be quite simple to use. Hypothesis tests are used to make comparisons