# MSwM examples. Jose A. Sanchez-Espigares, Alberto Lopez-Moreno Dept. of Statistics and Operations Research UPC-BarcelonaTech.

Save this PDF as:

Size: px
Start display at page:

Download "MSwM examples. Jose A. Sanchez-Espigares, Alberto Lopez-Moreno Dept. of Statistics and Operations Research UPC-BarcelonaTech."

## Transcription

1 MSwM examples Jose A. Sanchez-Espigares, Alberto Lopez-Moreno Dept. of Statistics and Operations Research UPC-BarcelonaTech February 24, 2014 Abstract Two examples are described to illustrate the use of the MSwM package. First, a simulated dataset is modeled in detail. Next, Markov Switching Models are fitted to a real dataset with a discrete response variable. The main methods and graphical representations are used to validate different approaches to model these datasets. 1 Simulated Example The example data is a simulated data set to show how msmfit can detect the presence of two different regimes: one in which the response variable is highly correlated and other in which the response only depends on an exogenous variable x. The autocorrelated observations are in the intervals 1 to 100, 151 to 180 and 251 to 300. The real models for each regime are: y t = { 8 + 2x t + ε (1) t ɛ (1) t N(0, 1) t = 101 : 150, 181 : y t 1 + ε (2) t ɛ (2) t N(0, 0.5) t = 1 : 100, 151 : 180, 251 : 300 > data(example) The plot in Fig.1 shows that in the intervals where does not exist autocorrelation the response variable y has a similar behaviour as the covariant x. A linear model is fitted to study how the covariate x explains the variable response y. > mod=lm(y~x,example) > summary(mod) Call: lm(formula = y ~ x, data = example) Residuals: Min 1Q Median 3Q Max

2 > plot(ts(example)) ts(example) y x Time Figure 1: Simulated data. The y variable is the response variable and there are two periods in which this depends on the x covariate Coefficients: Estimate Std. Error t value Pr(> t ) (Intercept) < 2e-16 *** x *** --- Signif. codes: 0 Ś***Š Ś**Š 0.01 Ś*Š 0.05 Ś.Š 0.1 Ś Š 1 Residual standard error: on 298 degrees of freedom Multiple R-squared: , Adjusted R-squared: F-statistic: on 1 and 298 DF, p-value: The covariate is really significant but the data behaviour is very bad explained by the model. The plot of the linear model residuals in Fig.1 indicates that their autocorrelation is significant. The diagnostics plots for the residuals (Fig.2) confirm that they does not seem to be white noise and that they have autocorrelation. Next, a Autoregressive Markov Switching Model (MSM-AR) is fitted to the data. The order for the autoregressive part is set to one. In order to indicate that all the parameters can be different in both periods, the switch-

3 Normal Q Q Plot Series resid(mod) Sample Quantiles ACF Theoretical Quantiles Figure 2: Normal Probability plot and Autocorrelation Function of the residuals from the linear model ing parameter (sw) is set to a vector with four components with value equal to TRUE. The last value when fitting a linear model is referred to the residual standard deviation. There are some options to control the estimation process, like a logical parameter to indicate whether parallelization of the process is done or not. > mod.mswm=msmfit(mod,k=2,p=1,sw=c(t,t,t,t),control=list(parallel=f)) > summary(mod.mswm) Markov Switching Model Call: msmfit(object = mod, k = 2, sw = c(t, T, T, T), p = 1, control = list(parallel = F)) AIC BIC loglik Coefficients: Regime Estimate Std. Error t value Pr(> t ) (Intercept)(S) ** x(s) y_1(s) < 2.2e-16 *** --- Signif. codes: 0 Ś***Š Ś**Š 0.01 Ś*Š 0.05 Ś.Š 0.1 Ś Š 1 Residual standard error: Multiple R-squared:

4 Standardized Residuals: Min Q1 Med Q3 Max Regime Estimate Std. Error t value Pr(> t ) (Intercept)(S) < 2.2e-16 *** x(s) e-09 *** y_1(s) Signif. codes: 0 Ś***Š Ś**Š 0.01 Ś*Š 0.05 Ś.Š 0.1 Ś Š 1 Residual standard error: Multiple R-squared: Standardized Residuals: Min Q1 Med Q3 Max Transition probabilities: Regime 1 Regime 2 Regime Regime The model mod.mswm has a regime where the covariant x is very significant and in the other regime the autocorrelation variable is very significant too. In both, the R-squared have high values. Finally, the transition probabilities matrix has high values which indicate that is difficult to change from on regime to the other. The model detect perfectly the periods of each state. The residuals look like to be white noise and they fit to the Normal Distribution. Moreover, the autocorrelation has disappeared. > plotprob(mod.mswm,which=1)

5 Residuals Regime 1 Regime 2 Conditional Order Figure 3: Residuals form the Autoregressive MSM model Regime 1 Smoothed Probabilities Filtered Probabilities Regime 2 Smoothed Probabilities Filtered Probabilities Figure 5: Filtered and Smoothed Probabilities for both regimes in the MSM-AR model

6 Normal Q Q Plot Pooled Residuals Theoretical Quantiles Sample Quantiles ACF ACF of Residuals Partial ACF PACF of Residuals ACF ACF of Square Residuals Partial ACF PACF of Square Residuals Figure 4: Normal Probability plot and Autocorrelation Function of the residuals from the Complete MSwM model. They are obtained by using the plotdiag method > plotprob(mod.mswm,which=2) Regime 1 y vs. Smooth Probabilities y Figure 6: Response variable indicating which observations are associated to regime 1 The graphics show that the periods for each regime have been detected perfectly.

7 > plotreg(mod.mswm,expl="x") Regime1 y with Smooth Probabilities x with Smooth Probabilities Figure 7: Relationship between x and y locating the two regimes 2 Daily Traffic Casualties by car accidents in Spain The traffic data (Fig.8) contains the daily number of deaths in traffic accidents in Spain during the year 2010, the average daily temperature and the daily sum of precipitations. The interest of this data is to study the relation between the number of deaths with the climate conditions. We illustrate the use of a Generalized Markov Switching Model in this case because there exists a different behaviour between the variables during weekends and working days. To avoid a long example, the explanations of how the functions work and repeated results are skipped. In this example, the response variable is a counting variable. For this reason we fit a Poisson Generalized Linear Model. > model=glm(ndead~temp+prec,traffic,family="poisson") > summary(model) Call:

8 Traffic Prec Temp NDead Time Figure 8: Traffic data: Daily traffic casualties in Spain and climate variables glm(formula = NDead ~ Temp + Prec, family = "poisson", data = traffic) Deviance Residuals: Min 1Q Median 3Q Max Coefficients: Estimate Std. Error z value Pr(> z ) (Intercept) < 2e-16 *** Temp e-08 *** Prec * --- Signif. codes: 0 Ś***Š Ś**Š 0.01 Ś*Š 0.05 Ś.Š 0.1 Ś Š 1 (Dispersion parameter for poisson family taken to be 1) Null deviance: on 364 degrees of freedom Residual deviance: on 362 degrees of freedom AIC: Number of Fisher Scoring iterations: 5

9 In the next step, the Markov Switching Model is fitted using msmfit. To fit a Generalized Markov Switching Model, the family parameter has to be included. Moreover, the glm s don t have the standard deviation parameter and, because of this, the sw parameter doesn t contain its switching parameter. > m1=msmfit(model,k=2,sw=c(t,t,t),family="poisson",control=list(parallel=f)) > summary(m1) Markov Switching Model Call: msmfit(object = model, k = 2, sw = c(t, T, T), family = "poisson", control = list(parallel = F)) AIC BIC loglik Coefficients: Regime Estimate Std. Error t value Pr(> t ) (Intercept)(S) e-05 *** Temp(S) *** Prec(S) Signif. codes: 0 Ś***Š Ś**Š 0.01 Ś*Š 0.05 Ś.Š 0.1 Ś Š 1 Regime Estimate Std. Error t value Pr(> t ) (Intercept)(S) < 2e-16 *** Temp(S) * Prec(S) * --- Signif. codes: 0 Ś***Š Ś**Š 0.01 Ś*Š 0.05 Ś.Š 0.1 Ś Š 1 Transition probabilities: Regime 1 Regime 2 Regime Regime Both states have significant covariates, but the precipitation covariate is only significant in one of the two. > intervals(m1)

10 Aproximate intervals for the coefficients. Level= 0.95 (Intercept): Lower Estimation Upper Regime Regime Temp: Lower Estimation Upper Regime Regime Prec: Lower Estimation Upper Regime e Regime e Residuals Regime 1 Regime 2 Conditional Order Figure 9: Residuals form the Autoregressive MSM model The Pearson residuals from Fig. 9 are calculated from an object of class MSM.glm because the model is an extension of a General Linear Model. The residuals have the classical structure of white noise. The residuals aren t au-

11 Normal Q Q Plot Pooled Residuals Theoretical Quantiles Sample Quantiles ACF ACF of Residuals Partial ACF PACF of Residuals ACF ACF of Square Residuals Partial ACF PACF of Square Residuals Figure 10: Normal Probability plot and Autocorrelation Function of the residuals from the Autoregressive MSwM model. They are obtained by using the plotdiag method tocorrelated but they don t fit very well to a Normal Distribution. However, normality of the Pearson residuals is not a critical condition for generalized linear model validation. > plotprob(m1,which=2)

12 Regime 1 NDead NDead vs. Smooth Probabilities 0.0 Figure 11: Response variable indicating which observations are associated to regime 1 Using the function plotprob we can see how the regimes are distributed in shorts periods because the bigger one contains basically working days.

### Generalized Linear Models

Generalized Linear Models We have previously worked with regression models where the response variable is quantitative and normally distributed. Now we turn our attention to two types of models where the

### Régression logistique : introduction

Chapitre 16 Introduction à la statistique avec R Régression logistique : introduction Une variable à expliquer binaire Expliquer un risque suicidaire élevé en prison par La durée de la peine L existence

### Multiple Linear Regression

Multiple Linear Regression A regression with two or more explanatory variables is called a multiple regression. Rather than modeling the mean response as a straight line, as in simple regression, it is

### Regression in ANOVA. James H. Steiger. Department of Psychology and Human Development Vanderbilt University

Regression in ANOVA James H. Steiger Department of Psychology and Human Development Vanderbilt University James H. Steiger (Vanderbilt University) 1 / 30 Regression in ANOVA 1 Introduction 2 Basic Linear

### Logistic Regression (a type of Generalized Linear Model)

Logistic Regression (a type of Generalized Linear Model) 1/36 Today Review of GLMs Logistic Regression 2/36 How do we find patterns in data? We begin with a model of how the world works We use our knowledge

### 12-1 Multiple Linear Regression Models

12-1.1 Introduction Many applications of regression analysis involve situations in which there are more than one regressor variable. A regression model that contains more than one regressor variable is

### Statistical Models in R

Statistical Models in R Some Examples Steven Buechler Department of Mathematics 276B Hurley Hall; 1-6233 Fall, 2007 Outline Statistical Models Linear Models in R Regression Regression analysis is the appropriate

### 5. Linear Regression

5. Linear Regression Outline.................................................................... 2 Simple linear regression 3 Linear model............................................................. 4

### E(y i ) = x T i β. yield of the refined product as a percentage of crude specific gravity vapour pressure ASTM 10% point ASTM end point in degrees F

Random and Mixed Effects Models (Ch. 10) Random effects models are very useful when the observations are sampled in a highly structured way. The basic idea is that the error associated with any linear,

### A Handbook of Statistical Analyses Using R. Brian S. Everitt and Torsten Hothorn

A Handbook of Statistical Analyses Using R Brian S. Everitt and Torsten Hothorn CHAPTER 6 Logistic Regression and Generalised Linear Models: Blood Screening, Women s Role in Society, and Colonic Polyps

### DEPARTMENT OF PSYCHOLOGY UNIVERSITY OF LANCASTER MSC IN PSYCHOLOGICAL RESEARCH METHODS ANALYSING AND INTERPRETING DATA 2 PART 1 WEEK 9

DEPARTMENT OF PSYCHOLOGY UNIVERSITY OF LANCASTER MSC IN PSYCHOLOGICAL RESEARCH METHODS ANALYSING AND INTERPRETING DATA 2 PART 1 WEEK 9 Analysis of covariance and multiple regression So far in this course,

### Multivariate Logistic Regression

1 Multivariate Logistic Regression As in univariate logistic regression, let π(x) represent the probability of an event that depends on p covariates or independent variables. Then, using an inv.logit formulation

### Basic Statistics and Data Analysis for Health Researchers from Foreign Countries

Basic Statistics and Data Analysis for Health Researchers from Foreign Countries Volkert Siersma siersma@sund.ku.dk The Research Unit for General Practice in Copenhagen Dias 1 Content Quantifying association

### We extended the additive model in two variables to the interaction model by adding a third term to the equation.

Quadratic Models We extended the additive model in two variables to the interaction model by adding a third term to the equation. Similarly, we can extend the linear model in one variable to the quadratic

### EDUCATION AND VOCABULARY MULTIPLE REGRESSION IN ACTION

EDUCATION AND VOCABULARY MULTIPLE REGRESSION IN ACTION EDUCATION AND VOCABULARY 5-10 hours of input weekly is enough to pick up a new language (Schiff & Myers, 1988). Dutch children spend 5.5 hours/day

### Quantitative Understanding in Biology Module II: Model Parameter Estimation Lecture I: Linear Correlation and Regression

Quantitative Understanding in Biology Module II: Model Parameter Estimation Lecture I: Linear Correlation and Regression Correlation Linear correlation and linear regression are often confused, mostly

### Time Series Analysis

Time Series 1 April 9, 2013 Time Series Analysis This chapter presents an introduction to the branch of statistics known as time series analysis. Often the data we collect in environmental studies is collected

### Unit 2 Regression and Correlation Practice Problems. SOLUTIONS Version R

PubHlth 640. Regression and Correlation Page 1 of 0 Unit Regression and Correlation Practice Problems SOLUTIONS Version R 1. A regression analysis of measurements of a dependent variable Y on an independent

### Correlation and Simple Linear Regression

Correlation and Simple Linear Regression We are often interested in studying the relationship among variables to determine whether they are associated with one another. When we think that changes in a

### Lecture 8: Gamma regression

Lecture 8: Gamma regression Claudia Czado TU München c (Claudia Czado, TU Munich) ZFS/IMS Göttingen 2004 0 Overview Models with constant coefficient of variation Gamma regression: estimation and testing

### Stat 411/511 ANOVA & REGRESSION. Charlotte Wickham. stat511.cwick.co.nz. Nov 31st 2015

Stat 411/511 ANOVA & REGRESSION Nov 31st 2015 Charlotte Wickham stat511.cwick.co.nz This week Today: Lack of fit F-test Weds: Review email me topics, otherwise I ll go over some of last year s final exam

### Lecture 6: Poisson regression

Lecture 6: Poisson regression Claudia Czado TU München c (Claudia Czado, TU Munich) ZFS/IMS Göttingen 2004 0 Overview Introduction EDA for Poisson regression Estimation and testing in Poisson regression

### N-Way Analysis of Variance

N-Way Analysis of Variance 1 Introduction A good example when to use a n-way ANOVA is for a factorial design. A factorial design is an efficient way to conduct an experiment. Each observation has data

### Chapter 7: Simple linear regression Learning Objectives

Chapter 7: Simple linear regression Learning Objectives Reading: Section 7.1 of OpenIntro Statistics Video: Correlation vs. causation, YouTube (2:19) Video: Intro to Linear Regression, YouTube (5:18) -

### SAS Software to Fit the Generalized Linear Model

SAS Software to Fit the Generalized Linear Model Gordon Johnston, SAS Institute Inc., Cary, NC Abstract In recent years, the class of generalized linear models has gained popularity as a statistical modeling

### Threshold Autoregressive Models in Finance: A Comparative Approach

University of Wollongong Research Online Applied Statistics Education and Research Collaboration (ASEARC) - Conference Papers Faculty of Informatics 2011 Threshold Autoregressive Models in Finance: A Comparative

### Psychology 205: Research Methods in Psychology

Psychology 205: Research Methods in Psychology Using R to analyze the data for study 2 Department of Psychology Northwestern University Evanston, Illinois USA November, 2012 1 / 38 Outline 1 Getting ready

### Time-Series Regression and Generalized Least Squares in R

Time-Series Regression and Generalized Least Squares in R An Appendix to An R Companion to Applied Regression, Second Edition John Fox & Sanford Weisberg last revision: 11 November 2010 Abstract Generalized

### Paired Differences and Regression

Paired Differences and Regression Students sometimes have difficulty distinguishing between paired data and independent samples when comparing two means. One can return to this topic after covering simple

### data on Down's syndrome

DATA a; INFILE 'downs.dat' ; INPUT AgeL AgeU BirthOrd Cases Births ; MidAge = (AgeL + AgeU)/2 ; Rate = 1000*Cases/Births; LogRate = Log( (Cases+0.5)/Births ); LogDenom = Log(Births); age_c = MidAge - 30;

### Comparing Nested Models

Comparing Nested Models ST 430/514 Two models are nested if one model contains all the terms of the other, and at least one additional term. The larger model is the complete (or full) model, and the smaller

### Regression, least squares

Regression, least squares Joe Felsenstein Department of Genome Sciences and Department of Biology Regression, least squares p.1/24 Fitting a straight line X Two distinct cases: The X values are chosen

### Exercise Page 1 of 32

Exercise 10.1 (a) Plot wages versus LOS. Describe the relationship. There is one woman with relatively high wages for her length of service. Circle this point and do not use it in the rest of this exercise.

### ANOVA. February 12, 2015

ANOVA February 12, 2015 1 ANOVA models Last time, we discussed the use of categorical variables in multivariate regression. Often, these are encoded as indicator columns in the design matrix. In [1]: %%R

### Applied Statistics. J. Blanchet and J. Wadsworth. Institute of Mathematics, Analysis, and Applications EPF Lausanne

Applied Statistics J. Blanchet and J. Wadsworth Institute of Mathematics, Analysis, and Applications EPF Lausanne An MSc Course for Applied Mathematicians, Fall 2012 Outline 1 Model Comparison 2 Model

### where b is the slope of the line and a is the intercept i.e. where the line cuts the y axis.

Least Squares Introduction We have mentioned that one should not always conclude that because two variables are correlated that one variable is causing the other to behave a certain way. However, sometimes

### Using R for Linear Regression

Using R for Linear Regression In the following handout words and symbols in bold are R functions and words and symbols in italics are entries supplied by the user; underlined words and symbols are optional

### Time Series Analysis with R - Part I. Walter Zucchini, Oleg Nenadić

Time Series Analysis with R - Part I Walter Zucchini, Oleg Nenadić Contents 1 Getting started 2 1.1 Downloading and Installing R.................... 2 1.2 Data Preparation and Import in R.................

### All Models are wrong, but some are useful.

1 Goodness of Fit in Logistic Regression As in linear regression, goodness of fit in logistic regression attempts to get at how well a model fits the data. It is usually applied after a final model has

### Stock Price Forecasting Using Information from Yahoo Finance and Google Trend

Stock Price Forecasting Using Information from Yahoo Finance and Google Trend Selene Yue Xu (UC Berkeley) Abstract: Stock price forecasting is a popular and important topic in financial and academic studies.

### Logistic regression (with R)

Logistic regression (with R) Christopher Manning 4 November 2007 1 Theory We can transform the output of a linear regression to be suitable for probabilities by using a logit link function on the lhs as

### Lab 13: Logistic Regression

Lab 13: Logistic Regression Spam Emails Today we will be working with a corpus of emails received by a single gmail account over the first three months of 2012. Just like any other email address this account

### Stat 5303 (Oehlert): Tukey One Degree of Freedom 1

Stat 5303 (Oehlert): Tukey One Degree of Freedom 1 > catch

### SCHOOL OF MATHEMATICS AND STATISTICS

RESTRICTED OPEN BOOK EXAMINATION (Not to be removed from the examination hall) Data provided: Statistics Tables by H.R. Neave MAS5052 SCHOOL OF MATHEMATICS AND STATISTICS Basic Statistics Spring Semester

### TIME SERIES ANALYSIS

TIME SERIES ANALYSIS Ramasubramanian V. I.A.S.R.I., Library Avenue, New Delhi- 110 012 ram_stat@yahoo.co.in 1. Introduction A Time Series (TS) is a sequence of observations ordered in time. Mostly these

### MIXED MODEL ANALYSIS USING R

Research Methods Group MIXED MODEL ANALYSIS USING R Using Case Study 4 from the BIOMETRICS & RESEARCH METHODS TEACHING RESOURCE BY Stephen Mbunzi & Sonal Nagda www.ilri.org/rmg www.worldagroforestrycentre.org/rmg

### Vector Time Series Model Representations and Analysis with XploRe

0-1 Vector Time Series Model Representations and Analysis with plore Julius Mungo CASE - Center for Applied Statistics and Economics Humboldt-Universität zu Berlin mungo@wiwi.hu-berlin.de plore MulTi Motivation

### Statistical Models in R

Statistical Models in R Some Examples Steven Buechler Department of Mathematics 276B Hurley Hall; 1-6233 Fall, 2007 Outline Statistical Models Structure of models in R Model Assessment (Part IA) Anova

### Cov(x, y) V ar(x)v ar(y)

Simple linear regression Systematic components: β 0 + β 1 x i Stochastic component : error term ε Y i = β 0 + β 1 x i + ε i ; i = 1,..., n E(Y X) = β 0 + β 1 x the central parameter is the slope parameter

### MGT 267 PROJECT. Forecasting the United States Retail Sales of the Pharmacies and Drug Stores. Done by: Shunwei Wang & Mohammad Zainal

MGT 267 PROJECT Forecasting the United States Retail Sales of the Pharmacies and Drug Stores Done by: Shunwei Wang & Mohammad Zainal Dec. 2002 The retail sale (Million) ABSTRACT The present study aims

### Lecture 5 Hypothesis Testing in Multiple Linear Regression

Lecture 5 Hypothesis Testing in Multiple Linear Regression BIOST 515 January 20, 2004 Types of tests 1 Overall test Test for addition of a single variable Test for addition of a group of variables Overall

### Forecasting Analytics. Group members: - Arpita - Kapil - Kaushik - Ridhima - Ushhan

Forecasting Analytics Group members: - Arpita - Kapil - Kaushik - Ridhima - Ushhan Business Problem Forecast daily sales of dairy products (excluding milk) to make a good prediction of future demand, and

### Basic Statistical and Modeling Procedures Using SAS

Basic Statistical and Modeling Procedures Using SAS One-Sample Tests The statistical procedures illustrated in this handout use two datasets. The first, Pulse, has information collected in a classroom

### The importance of graphing the data: Anscombe s regression examples

The importance of graphing the data: Anscombe s regression examples Bruce Weaver Northern Health Research Conference Nipissing University, North Bay May 30-31, 2008 B. Weaver, NHRC 2008 1 The Objective

### Simple Linear Regression Inference

Simple Linear Regression Inference 1 Inference requirements The Normality assumption of the stochastic term e is needed for inference even if it is not a OLS requirement. Therefore we have: Interpretation

: Table of Contents... 1 Overview of Model... 1 Dispersion... 2 Parameterization... 3 Sigma-Restricted Model... 3 Overparameterized Model... 4 Reference Coding... 4 Model Summary (Summary Tab)... 5 Summary

### STATISTICS: AN INTRODUCTION USING R. By M.J. Crawley. Exercises 4. REGRESSION

STATISTICS: AN INTRODUCTION USING R By M.J. Crawley Exercises 4. REGRESSION Regression is the statistical model we use when the explanatory variable is continuous. If the explanatory variables were categorical

### NCSS Statistical Software Principal Components Regression. In ordinary least squares, the regression coefficients are estimated using the formula ( )

Chapter 340 Principal Components Regression Introduction is a technique for analyzing multiple regression data that suffer from multicollinearity. When multicollinearity occurs, least squares estimates

### Introduction to Generalized Linear Models

to Generalized Linear Models Heather Turner ESRC National Centre for Research Methods, UK and Department of Statistics University of Warwick, UK WU, 2008 04 22-24 Copyright c Heather Turner, 2008 to Generalized

### Extended control charts

Extended control charts The control chart types listed below are recommended as alternative and additional tools to the Shewhart control charts. When compared with classical charts, they have some advantages

### Lucky vs. Unlucky Teams in Sports

Lucky vs. Unlucky Teams in Sports Introduction Assuming gambling odds give true probabilities, one can classify a team as having been lucky or unlucky so far. Do results of matches between lucky and unlucky

### Regression in SPSS. Workshop offered by the Mississippi Center for Supercomputing Research and the UM Office of Information Technology

Regression in SPSS Workshop offered by the Mississippi Center for Supercomputing Research and the UM Office of Information Technology John P. Bentley Department of Pharmacy Administration University of

### MISSING DATA TECHNIQUES WITH SAS. IDRE Statistical Consulting Group

MISSING DATA TECHNIQUES WITH SAS IDRE Statistical Consulting Group ROAD MAP FOR TODAY To discuss: 1. Commonly used techniques for handling missing data, focusing on multiple imputation 2. Issues that could

### 1. What is the critical value for this 95% confidence interval? CV = z.025 = invnorm(0.025) = 1.96

1 Final Review 2 Review 2.1 CI 1-propZint Scenario 1 A TV manufacturer claims in its warranty brochure that in the past not more than 10 percent of its TV sets needed any repair during the first two years

### GLM I An Introduction to Generalized Linear Models

GLM I An Introduction to Generalized Linear Models CAS Ratemaking and Product Management Seminar March 2009 Presented by: Tanya D. Havlicek, Actuarial Assistant 0 ANTITRUST Notice The Casualty Actuarial

### Random effects and nested models with SAS

Random effects and nested models with SAS /************* classical2.sas ********************* Three levels of factor A, four levels of B Both fixed Both random A fixed, B random B nested within A ***************************************************/

### Forecasting of Paddy Production in Sri Lanka: A Time Series Analysis using ARIMA Model

Tropical Agricultural Research Vol. 24 (): 2-3 (22) Forecasting of Paddy Production in Sri Lanka: A Time Series Analysis using ARIMA Model V. Sivapathasundaram * and C. Bogahawatte Postgraduate Institute

### Regression step-by-step using Microsoft Excel

Step 1: Regression step-by-step using Microsoft Excel Notes prepared by Pamela Peterson Drake, James Madison University Type the data into the spreadsheet The example used throughout this How to is a regression

### Chapter 5 Analysis of variance SPSS Analysis of variance

Chapter 5 Analysis of variance SPSS Analysis of variance Data file used: gss.sav How to get there: Analyze Compare Means One-way ANOVA To test the null hypothesis that several population means are equal,

### Discrete with many values are often treated as continuous: Minnesota math scores, years of education.

Lecture 9 1. Continuous and categorical predictors 2. Indicator variables 3. General linear model: child s IQ example 4. Proc GLM and Proc Reg 5. ANOVA example: Rat diets 6. Factorial design: 2 2 7. Interactions

### Lets suppose we rolled a six-sided die 150 times and recorded the number of times each outcome (1-6) occured. The data is

In this lab we will look at how R can eliminate most of the annoying calculations involved in (a) using Chi-Squared tests to check for homogeneity in two-way tables of catagorical data and (b) computing

### Didacticiel - Études de cas

1 Topic Regression analysis with LazStats (OpenStat). LazStat 1 is a statistical software which is developed by Bill Miller, the father of OpenStat, a wellknow tool by statisticians since many years. These

### Seminar paper Statistics

Seminar paper Statistics The seminar paper must contain: - the title page - the characterization of the data (origin, reason why you have chosen this analysis,...) - the list of the data (in the table)

### 2. Regression and Correlation. Simple Linear Regression Software: R

2. Regression and Correlation Simple Linear Regression Software: R Create txt file from SAS data set data _null_; file 'C:\Documents and Settings\sphlab\Desktop\slr1.txt'; set temp; put input day:date7.

### Statistiek II. John Nerbonne. March 24, 2010. Information Science, Groningen Slides improved a lot by Harmut Fitz, Groningen!

Information Science, Groningen j.nerbonne@rug.nl Slides improved a lot by Harmut Fitz, Groningen! March 24, 2010 Correlation and regression We often wish to compare two different variables Examples: compare

### Chapter 13 Introduction to Linear Regression and Correlation Analysis

Chapter 3 Student Lecture Notes 3- Chapter 3 Introduction to Linear Regression and Correlation Analsis Fall 2006 Fundamentals of Business Statistics Chapter Goals To understand the methods for displaing

### Technology Step-by-Step Using StatCrunch

Technology Step-by-Step Using StatCrunch Section 1.3 Simple Random Sampling 1. Select Data, highlight Simulate Data, then highlight Discrete Uniform. 2. Fill in the following window with the appropriate

J. Stuart Hunter Princeton University Copyright 2007, SAS Institute Inc. All rights reserved. JMP INNOVATOR S SUMMIT Time Series Approaches to Quality Prediction for Production and the Arts of Charts Traverse

### Testing for Lack of Fit

Chapter 6 Testing for Lack of Fit How can we tell if a model fits the data? If the model is correct then ˆσ 2 should be an unbiased estimate of σ 2. If we have a model which is not complex enough to fit

### Yiming Peng, Department of Statistics. February 12, 2013

Regression Analysis Using JMP Yiming Peng, Department of Statistics February 12, 2013 2 Presentation and Data http://www.lisa.stat.vt.edu Short Courses Regression Analysis Using JMP Download Data to Desktop

### Univariate Regression

Univariate Regression Correlation and Regression The regression line summarizes the linear relationship between 2 variables Correlation coefficient, r, measures strength of relationship: the closer r is

### BIOSTAT640 R-Solution for HW7 Logistic Minming Li & Steele H. Valenzuela Mar.10, 2016

BIOSTAT640 R-Solution for HW7 Logistic Minming Li & Steele H. Valenzuela Mar.10, 2016 Required Libraries If there is an error in loading the libraries, you must first install the packaged library (install.packages("package

### Business Statistics. Lecture 8: More Hypothesis Testing

Business Statistics Lecture 8: More Hypothesis Testing 1 Goals for this Lecture Review of t-tests Additional hypothesis tests Two-sample tests Paired tests 2 The Basic Idea of Hypothesis Testing Start

### Institute of Actuaries of India Subject CT3 Probability and Mathematical Statistics

Institute of Actuaries of India Subject CT3 Probability and Mathematical Statistics For 2015 Examinations Aim The aim of the Probability and Mathematical Statistics subject is to provide a grounding in

### A Multiplicative Seasonal Box-Jenkins Model to Nigerian Stock Prices

A Multiplicative Seasonal Box-Jenkins Model to Nigerian Stock Prices Ette Harrison Etuk Department of Mathematics/Computer Science, Rivers State University of Science and Technology, Nigeria Email: ettetuk@yahoo.com

### Part 2: Analysis of Relationship Between Two Variables

Part 2: Analysis of Relationship Between Two Variables Linear Regression Linear correlation Significance Tests Multiple regression Linear Regression Y = a X + b Dependent Variable Independent Variable

### Lab 5 Linear Regression with Within-subject Correlation. Goals: Data: Use the pig data which is in wide format:

Lab 5 Linear Regression with Within-subject Correlation Goals: Data: Fit linear regression models that account for within-subject correlation using Stata. Compare weighted least square, GEE, and random

### Lecture 11: Confidence intervals and model comparison for linear regression; analysis of variance

Lecture 11: Confidence intervals and model comparison for linear regression; analysis of variance 14 November 2007 1 Confidence intervals and hypothesis testing for linear regression Just as there was

### Business Statistics. Successful completion of Introductory and/or Intermediate Algebra courses is recommended before taking Business Statistics.

Business Course Text Bowerman, Bruce L., Richard T. O'Connell, J. B. Orris, and Dawn C. Porter. Essentials of Business, 2nd edition, McGraw-Hill/Irwin, 2008, ISBN: 978-0-07-331988-9. Required Computing

### Week 5: Multiple Linear Regression

BUS41100 Applied Regression Analysis Week 5: Multiple Linear Regression Parameter estimation and inference, forecasting, diagnostics, dummy variables Robert B. Gramacy The University of Chicago Booth School

### Classical Hypothesis Testing in R. R can do all of the common analyses that are available in SPSS, including:

Classical Hypothesis Testing in R R can do all of the common analyses that are available in SPSS, including: Classical Hypothesis Testing in R R can do all of the common analyses that are available in

### Simple Linear Regression

Inference for Regression Simple Linear Regression IPS Chapter 10.1 2009 W.H. Freeman and Company Objectives (IPS Chapter 10.1) Simple linear regression Statistical model for linear regression Estimating

### SELF-TEST: SIMPLE REGRESSION

ECO 22000 McRAE SELF-TEST: SIMPLE REGRESSION Note: Those questions indicated with an (N) are unlikely to appear in this form on an in-class examination, but you should be able to describe the procedures

### Math 141. Lecture 24: Model Comparisons and The F-test. Albyn Jones 1. 1 Library jones/courses/141

Math 141 Lecture 24: Model Comparisons and The F-test Albyn Jones 1 1 Library 304 jones@reed.edu www.people.reed.edu/ jones/courses/141 Nested Models Two linear models are Nested if one (the restricted

### CORRELATION AND SIMPLE REGRESSION ANALYSIS USING SAS IN DAIRY SCIENCE

CORRELATION AND SIMPLE REGRESSION ANALYSIS USING SAS IN DAIRY SCIENCE A. K. Gupta, Vipul Sharma and M. Manoj NDRI, Karnal-132001 When analyzing farm records, simple descriptive statistics can reveal a

### Lecture 5: Correlation and Linear Regression

Lecture 5: Correlation and Linear Regression 3.5. (Pearson) correlation coefficient The correlation coefficient measures the strength of the linear relationship between two variables. The correlation is

### Sydney Roberts Predicting Age Group Swimmers 50 Freestyle Time 1. 1. Introduction p. 2. 2. Statistical Methods Used p. 5. 3. 10 and under Males p.

Sydney Roberts Predicting Age Group Swimmers 50 Freestyle Time 1 Table of Contents 1. Introduction p. 2 2. Statistical Methods Used p. 5 3. 10 and under Males p. 8 4. 11 and up Males p. 10 5. 10 and under

Adatelemzés II. [SST35] Idősorok 0. Lőw András low.andras@gmail.com 2011. október 26. Lőw A (low.andras@gmail.com Adatelemzés II. [SST35] 2011. október 26. 1 / 24 Vázlat 1 Honnan tudja a summary, hogy