# Missing Data & How to Deal: An overview of missing data. Melissa Humphries Population Research Center

Save this PDF as:

Size: px
Start display at page:

Download "Missing Data & How to Deal: An overview of missing data. Melissa Humphries Population Research Center"

## Transcription

1 Missing Data & How to Deal: An overview of missing data Melissa Humphries Population Research Center

2 Goals Discuss ways to evaluate and understand missing data Discuss common missing data methods Know the advantages and disadvantages of common methods Review useful commands in Stata for missing data

3 General Steps for Analysis with Missing Data 1. Identify patterns/reasons for missing and recode correctly 2. Understand distribution of missing data 3. Decide on best method of analysis

4 Step One: Understand your data Attrition due to social/natural processes Example: School graduation, dropout, death Skip pattern in survey Example: Certain questions only asked to respondents who indicate they are married Intentional missing as part of data collection process Random data collection issues Respondent refusal/non-response

5 Find information from survey (codebook, questionnaire) Identify skip patterns and/or sampling strategy from documentation

6 Recode for analysis: mvdecode command Mvdecode How stata reads missing Tip.># s Nmissing npresent

7 Recode for analysis: mvdecode command Mvdecode How stata reads missing Tip.># s Nmissing npresent Note: Stata reads missing (.) as a value greater than any number.

8 Analyze missing data patterns: misstable command

9 Step Two: Missing data Mechanism (or probability distribution of missingness) Consider the probability of missingness Are certain groups more likely to have missing values? Example: Respondents in service occupations less likely to report income Are certain responses more likely to be missing? Example: Respondents with high income less likely to report income Certain analysis methods assume a certain probability distribution

10 Missing Data Mechanisms Missing Completely at Random (MCAR) Missing value (y) neither depends on x nor y Example: some survey questions asked of a simple random sample of original sample Missing at Random (MAR) Missing value (y) depends on x, but not y Example: Respondents in service occupations less likely to report income Missing not at Random (NMAR) The probability of a missing value depends on the variable that is missing Example: Respondents with high income less likely to report income

11 Exploring missing data mechanisms Can t be 100% sure about probability of missing (since we don t actually know the missing values) Could test for MCAR (t-tests) but not totally accurate Many missing data methods assume MCAR or MAR but our data often are MNAR Some methods specifically for MNAR Selection model (Heckman) Pattern mixture models

12 Good News!! Some MAR analysis methods using MNAR data are still pretty good. May be another measured variable that indirectly can predict the probability of missingness Example: those with higher incomes are less likely to report income BUT we have a variable for years of education and/or number of investments ML and MI are often unbiased with NMAR data even though assume data is MAR See Schafer & Graham 2002

13 Step 3: Deal with missing data Use what you know about Why data is missing Distribution of missing data Decide on the best analysis strategy to yield the least biased estimates Deletion Methods Listwise deletion, pairwise deletion Single Imputation Methods Mean/mode substitution, dummy variable method, single regression Model-Based Methods Maximum Likelihood, Multiple imputation

14 Deletion Methods Listwise deletion AKA complete case analysis Pairwise deletion

15 Listwise Deletion (Complete Case Analysis) Only analyze cases with available data on each variable Advantages: Simplicity Comparability across analyses Disadvantages: Reduces statistical power (because lowers n) Doesn t use all information Estimates may be biased if data not MCAR* Gender 8 th grade math test score F 45. M. 99 F F F F M 95. M F F th grade math score *NOTE: List-wise deletion often produces unbiased regression slope estimates as long as missingness is not a function of outcome variable.

16 Application in Stata Any analysis including multiple variables automatically applies listwise deletion.

17 Pairwise deletion (Available Case Analysis) Analysis with all cases in which the variables of interest are present. Advantage: Keeps as many cases as possible for each analysis Uses all information possible with each analysis Disadvantage: Can t compare analyses because sample different each time

18 Single imputation methods Mean/Mode substitution Dummy variable control Conditional mean substitution

19 Mean/Mode Substitution Replace missing value with sample mean or mode Run analyses as if all complete cases Advantages: Can use complete case analysis methods Disadvantages: Reduces variability Weakens covariance and correlation estimates in the data (because ignores relationship between variables)

20 th grade math test score imputed 12th grade math test score (mean sub)

21 Dummy variable adjustment Create an indicator for missing value (1=value is missing for observation; 0=value is observed for observation) Impute missing values to a constant (such as the mean) Include missing indicator in regression Advantage: Uses all available information about missing observation Disadvantage: Results in biased estimates Not theoretically driven NOTE: Results not biased if value is missing because of a legitimate skip

22 Regression Imputation Replaces missing values with predicted score from a regression equation. Advantage: Uses information from observed data Disadvantages: Overestimates model fit and correlation estimates Weakens variance

23 th grade math test score imputed 12th grade math test score (single regression)

24 Model-based methods Maximum Likelihood Multiple imputation

25 Model-based Methods: Maximum Likelihood Estimation Identifies the set of parameter values that produces the highest log-likelihood. ML estimate: value that is most likely to have resulted in the observed data Conceptually, process the same with or without missing data Advantages: Uses full information (both complete cases and incomplete cases) to calculate log likelihood Unbiased parameter estimates with MCAR/MAR data Disadvantages SEs biased downward can be adjusted by using observed information matrix

26 Multiple Imputation 1. Impute: Data is filled in with imputed values using specified regression model This step is repeated m times, resulting in a separate dataset each time. 2. Analyze: Analyses performed within each dataset 3. Pool: Results pooled into one estimate Advantages: Variability more accurate with multiple imputations for each missing value Considers variability due to sampling AND variability due to imputation Disadvantages: Cumbersome coding Room for error when specifying models

27 Multiple Imputation Process 1. Impute 2. Analyze 3. Pool Dataset with Missing Values Final Estimates Imputed Datasets Analysis results of each dataset

28 Multiple Imputation: Stata & SAS SAS: Proc mi Stata: ice (imputation using chained equations) & mim (analysis with multiply imputed dataset) mi commands mi set mi register mi impute mi estimate NOTE: the ice command is the only chained equation method until Stata12. Chained equations can be used as an option of mi impute since Stata12.

29 ice & mim ice: Imputation using chained equations Series of equations predicting one variable at a time Creates as many datasets as desired mim: prefix used before analysis that performs analyses across datasets and pools estimates

30 ice command 1. Impute 2. Analyze 3. Pool Dataset with Missing Values Final Estimates Imputed Datasets Analysis results of each dataset

31 mim command 1. Impute 2. Analyze 3. Pool Dataset with Missing Values Final Estimates Imputed Datasets Analysis results of each dataset

32 ice female lm latino black asian other F1PARED AGE1 intact bymirt ESL2 ALG2OH acgpa ac_engall hardwtr Lksch MAE10 RAE10 hilep midw south public catholic colltype aceng_esl Lksch_ESL, /// saving(imputed2) m(5) cmd (Lksch:ologit) Variable Command Prediction equation female [No missing data in estimation sample] lm [No missing data in estimation sample] latino [No missing data in estimation sample] black [No missing data in estimation sample] ALG2OH logit female lm latino black asian other F1PARED AGE1 intact bymirt ESL2 acgpa ac_engall hardwtr Lksch MAE10 RAE10 hilep midw south public catholic colltype aceng_esl Lksch_ESL acgpa regress female lm latino black asian other F1PARED AGE1 intact bymirt ESL2 ALG2OH ac_engall hardwtr Lksch MAE10 RAE10 hilep midw south public catholic colltype aceng_esl Lksch_ESL ac_engall regress female lm latino black asian other F1PARED AGE1 intact bymirt ESL2 ALG2OH acgpa hardwtr Lksch MAE10 RAE10 hilep midw south public catholic colltype Lksch_ESL hardwtr logit female lm latino black asian other F1PARED AGE1 intact bymirt ESL2 ALG2OH acgpa ac_engall Lksch MAE10 RAE10 hilep midw south public catholic colltype aceng_esl Lksch_ESL Lksch ologit female lm latino black asian other F1PARED AGE1 intact bymirt ESL2 ALG2OH acgpa ac_engall hardwtr MAE10 RAE10 hilep midw south public catholic colltype aceng_esl MAE10 regress female lm latino black asian other F1PARED AGE1 intact bymirt ESL2 ALG2OH acgpa ac_engall hardwtr Lksch RAE10 hilep midw south public catholic colltype aceng_esl Lksch_ESL RAE10 regress female lm latino black asian other F1PARED AGE1 intact bymirt ESL2 ALG2OH acgpa ac_engall hardwtr Lksch MAE10 hilep midw south public catholic colltype aceng_esl Lksch_ESL hilep logit female lm latino black asian other F1PARED AGE1 intact bymirt ESL2 ALG2OH acgpa ac_engall hardwtr Lksch MAE10 RAE10 midw south public catholic colltype aceng_esl Lksch_ESL Imputing file imputed2.dta saved

33 mim, storebv: svy: mlogit colltype ESL2 lm female latino black asian other F1PARED lowinc AGE1 intact bymirt ALG2OH acgpa Lksch, b(0) Multiple-imputation estimates (svy: mlogit) Imputations = 5 Survey: Multinomial logistic regression Minimum obs = Minimum dof = colltype Coef. Std. Err. t P> t [95% Conf. Int.] FMI ESL lm female latino black asian other F1PARED lowinc AGE intact bymirt ALG2OH acgpa Lksch _cons ESL lm female latino black asian other F1PARED lowinc AGE intact bymirt ALG2OH acgpa Lksch _cons

34 mi commands Included in Stata 11 Includes univariate multiple imputation (impute only one variable) Multivariate imputation probably more useful for our data Specific order: mi set mi register mi impute mi estimate

35 mi set and mi register commands 1. Impute 2. Analyze 3. Pool Dataset with Missing Values Final Estimates Imputed Datasets Analysis results of each dataset

36 mi impute command 1. Impute 2. Analyze 3. Pool Dataset with Missing Values Final Estimates Imputed Datasets Analysis results of each dataset

37 mi estimate command 1. Impute 2. Analyze 3. Pool Dataset with Missing Values Final Estimates Imputed Datasets Analysis results of each dataset

38 *******set data to be multiply imputed (can set to 'wide' format also) mi set flong *******register variables as "imputed" (variables with missing data that you want imputed) or "regular" mi register imputed readtest8 worked mathtest8 mi register regular sex race *******describing data mi describe *******setting seed so results are replicable set seed 8945 *******imputing using chained equations using ols regression for predicting read and math test using mlogit to predict worked mi impute chained (regress) readtest8 mathtest8 (mlogit) worked=sex i.race, add(10) ********check new imputed dataset mi describe *******estimating model using imputed values mi estimate:regress mathtest12 mathtest8 sex race

39

40

41 Dataset after imputation

42

43 Notes and help with mi in stata LOTS of options Can specify exactly how you want imputed Can specify the model appropriately (ex. Using svy command) mi impute mvn (multivariate normal regression) also useful Help mi is useful Also, UCLA has great website about ice and mi

44 General Tips Try a few methods: often if result in similar estimates, can put as a footnote to support method Some don t impute dependent variable But would still use to impute independent variables

45 References Allison, Paul D Missing Data. Sage University Papers Series on Quantitative Applications in the Social Sciences. Thousand Oaks: Sage. Enders, Craig Applied Missing Data Analysis. Guilford Press: New York. Little, Roderick J., Donald Rubin Statistical Analysis with Missing Data. John Wiley & Sons, Inc: Hoboken. Schafer, Joseph L., John W. Graham Missing Data: Our View of the State of the Art. Psychological Methods.

### Missing Data: Part 1 What to Do? Carol B. Thompson Johns Hopkins Biostatistics Center SON Brown Bag 3/20/13

Missing Data: Part 1 What to Do? Carol B. Thompson Johns Hopkins Biostatistics Center SON Brown Bag 3/20/13 Overview Missingness and impact on statistical analysis Missing data assumptions/mechanisms Conventional

### MISSING DATA TECHNIQUES WITH SAS. IDRE Statistical Consulting Group

MISSING DATA TECHNIQUES WITH SAS IDRE Statistical Consulting Group ROAD MAP FOR TODAY To discuss: 1. Commonly used techniques for handling missing data, focusing on multiple imputation 2. Issues that could

### A Basic Introduction to Missing Data

John Fox Sociology 740 Winter 2014 Outline Why Missing Data Arise Why Missing Data Arise Global or unit non-response. In a survey, certain respondents may be unreachable or may refuse to participate. Item

### APPLIED MISSING DATA ANALYSIS

APPLIED MISSING DATA ANALYSIS Craig K. Enders Series Editor's Note by Todd D. little THE GUILFORD PRESS New York London Contents 1 An Introduction to Missing Data 1 1.1 Introduction 1 1.2 Chapter Overview

### Handling missing data in Stata a whirlwind tour

Handling missing data in Stata a whirlwind tour 2012 Italian Stata Users Group Meeting Jonathan Bartlett www.missingdata.org.uk 20th September 2012 1/55 Outline The problem of missing data and a principled

### Challenges in Longitudinal Data Analysis: Baseline Adjustment, Missing Data, and Drop-out

Challenges in Longitudinal Data Analysis: Baseline Adjustment, Missing Data, and Drop-out Sandra Taylor, Ph.D. IDDRC BBRD Core 23 April 2014 Objectives Baseline Adjustment Introduce approaches Guidance

### A REVIEW OF CURRENT SOFTWARE FOR HANDLING MISSING DATA

123 Kwantitatieve Methoden (1999), 62, 123-138. A REVIEW OF CURRENT SOFTWARE FOR HANDLING MISSING DATA Joop J. Hox 1 ABSTRACT. When we deal with a large data set with missing data, we have to undertake

### Problem of Missing Data

VASA Mission of VA Statisticians Association (VASA) Promote & disseminate statistical methodological research relevant to VA studies; Facilitate communication & collaboration among VA-affiliated statisticians;

### Missing Data: Patterns, Mechanisms & Prevention. Edith de Leeuw

Missing Data: Patterns, Mechanisms & Prevention Edith de Leeuw Thema middag Nonresponse en Missing Data, Universiteit Groningen, 30 Maart 2006 Item-Nonresponse Pattern General pattern: various variables

### Missing Data. A Typology Of Missing Data. Missing At Random Or Not Missing At Random

[Leeuw, Edith D. de, and Joop Hox. (2008). Missing Data. Encyclopedia of Survey Research Methods. Retrieved from http://sage-ereference.com/survey/article_n298.html] Missing Data An important indicator

### Statistical modelling with missing data using multiple imputation. Session 4: Sensitivity Analysis after Multiple Imputation

Statistical modelling with missing data using multiple imputation Session 4: Sensitivity Analysis after Multiple Imputation James Carpenter London School of Hygiene & Tropical Medicine Email: james.carpenter@lshtm.ac.uk

### Analyzing Structural Equation Models With Missing Data

Analyzing Structural Equation Models With Missing Data Craig Enders* Arizona State University cenders@asu.edu based on Enders, C. K. (006). Analyzing structural equation models with missing data. In G.

### Missing Data in Longitudinal Studies: To Impute or not to Impute? Robert Platt, PhD McGill University

Missing Data in Longitudinal Studies: To Impute or not to Impute? Robert Platt, PhD McGill University 1 Outline Missing data definitions Longitudinal data specific issues Methods Simple methods Multiple

### Missing Data Dr Eleni Matechou

1 Statistical Methods Principles Missing Data Dr Eleni Matechou matechou@stats.ox.ac.uk References: R.J.A. Little and D.B. Rubin 2nd edition Statistical Analysis with Missing Data J.L. Schafer and J.W.

### Dealing with Missing Data

Res. Lett. Inf. Math. Sci. (2002) 3, 153-160 Available online at http://www.massey.ac.nz/~wwiims/research/letters/ Dealing with Missing Data Judi Scheffer I.I.M.S. Quad A, Massey University, P.O. Box 102904

### Imputing Attendance Data in a Longitudinal Multilevel Panel Data Set

Imputing Attendance Data in a Longitudinal Multilevel Panel Data Set April 2015 SHORT REPORT Baby FACES 2009 This page is left blank for double-sided printing. Imputing Attendance Data in a Longitudinal

### Data Cleaning and Missing Data Analysis

Data Cleaning and Missing Data Analysis Dan Merson vagabond@psu.edu India McHale imm120@psu.edu April 13, 2010 Overview Introduction to SACS What do we mean by Data Cleaning and why do we do it? The SACS

### Handling missing data in large data sets. Agostino Di Ciaccio Dept. of Statistics University of Rome La Sapienza

Handling missing data in large data sets Agostino Di Ciaccio Dept. of Statistics University of Rome La Sapienza The problem Often in official statistics we have large data sets with many variables and

### Review of the Methods for Handling Missing Data in. Longitudinal Data Analysis

Int. Journal of Math. Analysis, Vol. 5, 2011, no. 1, 1-13 Review of the Methods for Handling Missing Data in Longitudinal Data Analysis Michikazu Nakai and Weiming Ke Department of Mathematics and Statistics

### Module 14: Missing Data Stata Practical

Module 14: Missing Data Stata Practical Jonathan Bartlett & James Carpenter London School of Hygiene & Tropical Medicine www.missingdata.org.uk Supported by ESRC grant RES 189-25-0103 and MRC grant G0900724

### Best Practices for Missing Data Management in Counseling Psychology

Journal of Counseling Psychology 2010 American Psychological Association 2010, Vol. 57, No. 1, 1 10 0022-0167/10/\$12.00 DOI: 10.1037/a0018082 Best Practices for Missing Data Management in Counseling Psychology

### Missing Data. Katyn & Elena

Missing Data Katyn & Elena What to do with Missing Data Standard is complete case analysis/listwise dele;on ie. Delete cases with missing data so only complete cases are le> Two other popular op;ons: Mul;ple

### Multiple Imputation for Missing Data: A Cautionary Tale

Multiple Imputation for Missing Data: A Cautionary Tale Paul D. Allison University of Pennsylvania Address correspondence to Paul D. Allison, Sociology Department, University of Pennsylvania, 3718 Locust

### Dealing with missing data: Key assumptions and methods for applied analysis

Technical Report No. 4 May 6, 2013 Dealing with missing data: Key assumptions and methods for applied analysis Marina Soley-Bori msoley@bu.edu This paper was published in fulfillment of the requirements

### Handling attrition and non-response in longitudinal data

Longitudinal and Life Course Studies 2009 Volume 1 Issue 1 Pp 63-72 Handling attrition and non-response in longitudinal data Harvey Goldstein University of Bristol Correspondence. Professor H. Goldstein

### Missing data and net survival analysis Bernard Rachet

Workshop on Flexible Models for Longitudinal and Survival Data with Applications in Biostatistics Warwick, 27-29 July 2015 Missing data and net survival analysis Bernard Rachet General context Population-based,

### Modern Methods for Missing Data

Modern Methods for Missing Data Paul D. Allison, Ph.D. Statistical Horizons LLC www.statisticalhorizons.com 1 Introduction Missing data problems are nearly universal in statistical practice. Last 25 years

### Missing Data Part 1: Overview, Traditional Methods Page 1

Missing Data Part 1: Overview, Traditional Methods Richard Williams, University of Notre Dame, http://www3.nd.edu/~rwilliam/ Last revised January 17, 2015 This discussion borrows heavily from: Applied

### Missing Data. Paul D. Allison INTRODUCTION

4 Missing Data Paul D. Allison INTRODUCTION Missing data are ubiquitous in psychological research. By missing data, I mean data that are missing for some (but not all) variables and for some (but not all)

### IBM SPSS Missing Values 22

IBM SPSS Missing Values 22 Note Before using this information and the product it supports, read the information in Notices on page 23. Product Information This edition applies to version 22, release 0,

### In almost any research you perform, there is the potential for missing or

SIX DEALING WITH MISSING OR INCOMPLETE DATA Debunking the Myth of Emptiness In almost any research you perform, there is the potential for missing or incomplete data. Missing data can occur for many reasons:

### An introduction to modern missing data analyses

Journal of School Psychology 48 (2010) 5 37 An introduction to modern missing data analyses Amanda N. Baraldi, Craig K. Enders Arizona State University, United States Received 19 October 2009; accepted

### A Review of Methods for Missing Data

Educational Research and Evaluation 1380-3611/01/0704-353\$16.00 2001, Vol. 7, No. 4, pp. 353±383 # Swets & Zeitlinger A Review of Methods for Missing Data Therese D. Pigott Loyola University Chicago, Wilmette,

### SPSS TRAINING SESSION 3 ADVANCED TOPICS (PASW STATISTICS 17.0) Sun Li Centre for Academic Computing lsun@smu.edu.sg

SPSS TRAINING SESSION 3 ADVANCED TOPICS (PASW STATISTICS 17.0) Sun Li Centre for Academic Computing lsun@smu.edu.sg IN SPSS SESSION 2, WE HAVE LEARNT: Elementary Data Analysis Group Comparison & One-way

### Introduction to mixed model and missing data issues in longitudinal studies

Introduction to mixed model and missing data issues in longitudinal studies Hélène Jacqmin-Gadda INSERM, U897, Bordeaux, France Inserm workshop, St Raphael Outline of the talk I Introduction Mixed models

### Title: Categorical Data Imputation Using Non-Parametric or Semi-Parametric Imputation Methods

Masters by Coursework and Research Report Mathematical Statistics School of Statistics and Actuarial Science Title: Categorical Data Imputation Using Non-Parametric or Semi-Parametric Imputation Methods

### Approaches for Addressing Missing Data in Statistical Analyses of Female and Male Adolescent Fertility 1

1 Approaches for Addressing Missing Data in Statistical Analyses of Female and Male Adolescent Fertility 1 Eugenia Conde Texas A&M University and Dudley L. Poston, Jr. Texas A&M University 1 2 Introduction

### Auxiliary Variables in Mixture Modeling: 3-Step Approaches Using Mplus

Auxiliary Variables in Mixture Modeling: 3-Step Approaches Using Mplus Tihomir Asparouhov and Bengt Muthén Mplus Web Notes: No. 15 Version 8, August 5, 2014 1 Abstract This paper discusses alternatives

### 2. Making example missing-value datasets: MCAR, MAR, and MNAR

Lecture 20 1. Types of missing values 2. Making example missing-value datasets: MCAR, MAR, and MNAR 3. Common methods for missing data 4. Compare results on example MCAR, MAR, MNAR data 1 Missing Data

### Analysis of Longitudinal Data with Missing Values.

Analysis of Longitudinal Data with Missing Values. Methods and Applications in Medical Statistics. Ingrid Garli Dragset Master of Science in Physics and Mathematics Submission date: June 2009 Supervisor:

### Overview. Longitudinal Data Variation and Correlation Different Approaches. Linear Mixed Models Generalized Linear Mixed Models

Overview 1 Introduction Longitudinal Data Variation and Correlation Different Approaches 2 Mixed Models Linear Mixed Models Generalized Linear Mixed Models 3 Marginal Models Linear Models Generalized Linear

### Imputation and Analysis. Peter Fayers

Missing Data in Palliative Care Research Imputation and Analysis Peter Fayers Department of Public Health University of Aberdeen NTNU Det medisinske fakultet Missing data Missing data is a major problem

### When to Use Which Statistical Test

When to Use Which Statistical Test Rachel Lovell, Ph.D., Senior Research Associate Begun Center for Violence Prevention Research and Education Jack, Joseph, and Morton Mandel School of Applied Social Sciences

### Missing Data Techniques for Structural Equation Modeling

Journal of Abnormal Psychology Copyright 2003 by the American Psychological Association, Inc. 2003, Vol. 112, No. 4, 545 557 0021-843X/03/\$12.00 DOI: 10.1037/0021-843X.112.4.545 Missing Data Techniques

### How To Use A Monte Carlo Study To Decide On Sample Size and Determine Power

How To Use A Monte Carlo Study To Decide On Sample Size and Determine Power Linda K. Muthén Muthén & Muthén 11965 Venice Blvd., Suite 407 Los Angeles, CA 90066 Telephone: (310) 391-9971 Fax: (310) 391-8971

### Nonrandomly Missing Data in Multiple Regression Analysis: An Empirical Comparison of Ten Missing Data Treatments

Brockmeier, Kromrey, & Hogarty Nonrandomly Missing Data in Multiple Regression Analysis: An Empirical Comparison of Ten s Lantry L. Brockmeier Jeffrey D. Kromrey Kristine Y. Hogarty Florida A & M University

### Dealing with Missing Data

Dealing with Missing Data Roch Giorgi email: roch.giorgi@univ-amu.fr UMR 912 SESSTIM, Aix Marseille Université / INSERM / IRD, Marseille, France BioSTIC, APHM, Hôpital Timone, Marseille, France January

### Craig K. Enders Arizona State University Department of Psychology craig.enders@asu.edu

Craig K. Enders Arizona State University Department of Psychology craig.enders@asu.edu Topic Page Missing Data Patterns And Missing Data Mechanisms 1 Traditional Missing Data Techniques 7 Maximum Likelihood

### Failure to take the sampling scheme into account can lead to inaccurate point estimates and/or flawed estimates of the standard errors.

Analyzing Complex Survey Data: Some key issues to be aware of Richard Williams, University of Notre Dame, http://www3.nd.edu/~rwilliam/ Last revised January 24, 2015 Rather than repeat material that is

### IBM SPSS Missing Values 20

IBM SPSS Missing Values 20 Note: Before using this information and the product it supports, read the general information under Notices on p. 87. This edition applies to IBM SPSS Statistics 20 and to all

### Comparison of Imputation Methods in the Survey of Income and Program Participation

Comparison of Imputation Methods in the Survey of Income and Program Participation Sarah McMillan U.S. Census Bureau, 4600 Silver Hill Rd, Washington, DC 20233 Any views expressed are those of the author

### Marginal Effects for Continuous Variables Richard Williams, University of Notre Dame, http://www3.nd.edu/~rwilliam/ Last revised February 21, 2015

Marginal Effects for Continuous Variables Richard Williams, University of Notre Dame, http://www3.nd.edu/~rwilliam/ Last revised February 21, 2015 References: Long 1997, Long and Freese 2003 & 2006 & 2014,

### Imputing Missing Data using SAS

ABSTRACT Paper 3295-2015 Imputing Missing Data using SAS Christopher Yim, California Polytechnic State University, San Luis Obispo Missing data is an unfortunate reality of statistics. However, there are

### MISSING DATA IMPUTATION IN CARDIAC DATA SET (SURVIVAL PROGNOSIS)

MISSING DATA IMPUTATION IN CARDIAC DATA SET (SURVIVAL PROGNOSIS) R.KAVITHA KUMAR Department of Computer Science and Engineering Pondicherry Engineering College, Pudhucherry, India DR. R.M.CHADRASEKAR Professor,

### CHOOSING APPROPRIATE METHODS FOR MISSING DATA IN MEDICAL RESEARCH: A DECISION ALGORITHM ON METHODS FOR MISSING DATA

CHOOSING APPROPRIATE METHODS FOR MISSING DATA IN MEDICAL RESEARCH: A DECISION ALGORITHM ON METHODS FOR MISSING DATA Hatice UENAL Institute of Epidemiology and Medical Biometry, Ulm University, Germany

### Advances in Missing Data Methods and Implications for Educational Research. Chao-Ying Joanne Peng, Indiana University-Bloomington

Advances in Missing Data 1 Running head: MISSING DATA METHODS Advances in Missing Data Methods and Implications for Educational Research Chao-Ying Joanne Peng, Indiana University-Bloomington Michael Harwell,

### Psychology 594L Quantitative Behavioral Methods Missing Data T/R 3:40 5:00 Lago W-0272

Psychology 594L Quantitative Behavioral Methods Missing Data T/R 3:40 5:00 Lago W-0272 Instructor Todd Abraham Email abrahamt@iastate.edu Office W-053 Lagomarcino Hall Phone 515-294-4948 Office Hours (1/12-3/27)

### A Guide to Imputing Missing Data with Stata Revision: 1.4

A Guide to Imputing Missing Data with Stata Revision: 1.4 Mark Lunt December 6, 2011 Contents 1 Introduction 3 2 Installing Packages 4 3 How big is the problem? 5 4 First steps in imputation 5 5 Imputation

### PROC LOGISTIC: Traps for the unwary Peter L. Flom, Independent statistical consultant, New York, NY

PROC LOGISTIC: Traps for the unwary Peter L. Flom, Independent statistical consultant, New York, NY ABSTRACT Keywords: Logistic. INTRODUCTION This paper covers some gotchas in SAS R PROC LOGISTIC. A gotcha

### Econ 371 Problem Set #3 Answer Sheet

Econ 371 Problem Set #3 Answer Sheet 4.1 In this question, you are told that a OLS regression analysis of third grade test scores as a function of class size yields the following estimated model. T estscore

### Models for Longitudinal and Clustered Data

Models for Longitudinal and Clustered Data Germán Rodríguez December 9, 2008, revised December 6, 2012 1 Introduction The most important assumption we have made in this course is that the observations

### Item Imputation Without Specifying Scale Structure

Original Article Item Imputation Without Specifying Scale Structure Stef van Buuren TNO Quality of Life, Leiden, The Netherlands University of Utrecht, The Netherlands Abstract. Imputation of incomplete

ALLISON 1 TABLE OF CONTENTS 1. INTRODUCTION... 3 2. ASSUMPTIONS... 6 MISSING COMPLETELY AT RANDOM (MCAR)... 6 MISSING AT RANDOM (MAR)... 7 IGNORABLE... 8 NONIGNORABLE... 8 3. CONVENTIONAL METHODS... 10

### Re-analysis using Inverse Probability Weighting and Multiple Imputation of Data from the Southampton Women s Survey

Re-analysis using Inverse Probability Weighting and Multiple Imputation of Data from the Southampton Women s Survey MRC Biostatistics Unit Institute of Public Health Forvie Site Robinson Way Cambridge

### in press, Communication Methods and Measures Effective Tool for Handling Missing Data Teresa A. Myers George Mason University

HOTDECK: An SPSS Tool for Handling Missing Data 1 in press, Communication Methods and Measures Goodbye, Listwise Deletion: Presenting Hot Deck Imputation as an Easy and Effective Tool for Handling Missing

### ECON 142 SKETCH OF SOLUTIONS FOR APPLIED EXERCISE #2

University of California, Berkeley Prof. Ken Chay Department of Economics Fall Semester, 005 ECON 14 SKETCH OF SOLUTIONS FOR APPLIED EXERCISE # Question 1: a. Below are the scatter plots of hourly wages

### Visualization of missing values using the R-package VIM

Institut f. Statistik u. Wahrscheinlichkeitstheorie 040 Wien, Wiedner Hauptstr. 8-0/07 AUSTRIA http://www.statistik.tuwien.ac.at Visualization of missing values using the R-package VIM M. Templ and P.

### Missing data in randomized controlled trials (RCTs) can

EVALUATION TECHNICAL ASSISTANCE BRIEF for OAH & ACYF Teenage Pregnancy Prevention Grantees May 2013 Brief 3 Coping with Missing Data in Randomized Controlled Trials Missing data in randomized controlled

### Using Stata 9 & Higher for OLS Regression Richard Williams, University of Notre Dame, http://www3.nd.edu/~rwilliam/ Last revised January 8, 2015

Using Stata 9 & Higher for OLS Regression Richard Williams, University of Notre Dame, http://www3.nd.edu/~rwilliam/ Last revised January 8, 2015 Introduction. This handout shows you how Stata can be used

### Introduction to Multilevel Modeling Using HLM 6. By ATS Statistical Consulting Group

Introduction to Multilevel Modeling Using HLM 6 By ATS Statistical Consulting Group Multilevel data structure Students nested within schools Children nested within families Respondents nested within interviewers

### From this it is not clear what sort of variable that insure is so list the first 10 observations.

MNL in Stata We have data on the type of health insurance available to 616 psychologically depressed subjects in the United States (Tarlov et al. 1989, JAMA; Wells et al. 1989, JAMA). The insurance is

### Lecture 15. Endogeneity & Instrumental Variable Estimation

Lecture 15. Endogeneity & Instrumental Variable Estimation Saw that measurement error (on right hand side) means that OLS will be biased (biased toward zero) Potential solution to endogeneity instrumental

### From the help desk: hurdle models

The Stata Journal (2003) 3, Number 2, pp. 178 184 From the help desk: hurdle models Allen McDowell Stata Corporation Abstract. This article demonstrates that, although there is no command in Stata for

### PATTERN MIXTURE MODELS FOR MISSING DATA. Mike Kenward. London School of Hygiene and Tropical Medicine. Talk at the University of Turku,

PATTERN MIXTURE MODELS FOR MISSING DATA Mike Kenward London School of Hygiene and Tropical Medicine Talk at the University of Turku, April 10th 2012 1 / 90 CONTENTS 1 Examples 2 Modelling Incomplete Data

### Analyzing Intervention Effects: Multilevel & Other Approaches. Simplest Intervention Design. Better Design: Have Pretest

Analyzing Intervention Effects: Multilevel & Other Approaches Joop Hox Methodology & Statistics, Utrecht Simplest Intervention Design R X Y E Random assignment Experimental + Control group Analysis: t

### Linda K. Muthén Bengt Muthén. Copyright 2008 Muthén & Muthén www.statmodel.com. Table Of Contents

Mplus Short Courses Topic 2 Regression Analysis, Eploratory Factor Analysis, Confirmatory Factor Analysis, And Structural Equation Modeling For Categorical, Censored, And Count Outcomes Linda K. Muthén

### Chapter 5: Analysis of The National Education Longitudinal Study (NELS:88)

Chapter 5: Analysis of The National Education Longitudinal Study (NELS:88) Introduction The National Educational Longitudinal Survey (NELS:88) followed students from 8 th grade in 1988 to 10 th grade in

### Call Centre Forecasting. A comprehensive analysis of missing data, extreme values, holiday influences and different forecasting methods

A comprehensive analysis of missing data, extreme values, holiday influences and different forecasting methods Master Thesis Econometrics and Operations Research Master Operations Research and Management

### M. Ehren N. Shackleton. Institute of Education, University of London. June 2014. Grant number: 511490-2010-LLP-NL-KA1-KA1SCR

Impact of school inspections on teaching and learning in primary and secondary education in the Netherlands; Technical report ISI-TL project year 1-3 data M. Ehren N. Shackleton Institute of Education,

### REGRESSION LINES IN STATA

REGRESSION LINES IN STATA THOMAS ELLIOTT 1. Introduction to Regression Regression analysis is about eploring linear relationships between a dependent variable and one or more independent variables. Regression

### How Do We Test Multiple Regression Coefficients?

How Do We Test Multiple Regression Coefficients? Suppose you have constructed a multiple linear regression model and you have a specific hypothesis to test which involves more than one regression coefficient.

### Rethinking the Cultural Context of Schooling Decisions in Disadvantaged Neighborhoods: From Deviant Subculture to Cultural Heterogeneity

Rethinking the Cultural Context of Schooling Decisions in Disadvantaged Neighborhoods: From Deviant Subculture to Cultural Heterogeneity Sociology of Education David J. Harding, University of Michigan

### Sensitivity Analysis in Multiple Imputation for Missing Data

Paper SAS270-2014 Sensitivity Analysis in Multiple Imputation for Missing Data Yang Yuan, SAS Institute Inc. ABSTRACT Multiple imputation, a popular strategy for dealing with missing values, usually assumes

### Survey Data Analysis in Stata

Survey Data Analysis in Stata Jeff Pitblado Associate Director, Statistical Software StataCorp LP Stata Conference DC 2009 J. Pitblado (StataCorp) Survey Data Analysis DC 2009 1 / 44 Outline 1 Types of

### How to set the main menu of STATA to default factory settings standards

University of Pretoria Data analysis for evaluation studies Examples in STATA version 11 List of data sets b1.dta (To be created by students in class) fp1.xls (To be provided to students) fp1.txt (To be

### Introduction to Stata

Introduction to Stata September 23, 2014 Stata is one of a few statistical analysis programs that social scientists use. Stata is in the mid-range of how easy it is to use. Other options include SPSS,

### CHAPTER 12 EXAMPLES: MONTE CARLO SIMULATION STUDIES

Examples: Monte Carlo Simulation Studies CHAPTER 12 EXAMPLES: MONTE CARLO SIMULATION STUDIES Monte Carlo simulation studies are often used for methodological investigations of the performance of statistical

### NCEE 2009-0049. What to Do When Data Are Missing in Group Randomized Controlled Trials

NCEE 2009-0049 What to Do When Data Are Missing in Group Randomized Controlled Trials What to Do When Data Are Missing in Group Randomized Controlled Trials October 2009 Michael J. Puma Chesapeake Research

### Power Calculation Using the Online Variance Almanac (Web VA): A User s Guide

Power Calculation Using the Online Variance Almanac (Web VA): A User s Guide Larry V. Hedges & E.C. Hedberg This research was supported by the National Science Foundation under Award Nos. 0129365 and 0815295.

### Introduction to Quantitative Methods

Introduction to Quantitative Methods October 15, 2009 Contents 1 Definition of Key Terms 2 2 Descriptive Statistics 3 2.1 Frequency Tables......................... 4 2.2 Measures of Central Tendencies.................

### Using Medical Research Data to Motivate Methodology Development among Undergraduates in SIBS Pittsburgh

Using Medical Research Data to Motivate Methodology Development among Undergraduates in SIBS Pittsburgh Megan Marron and Abdus Wahed Graduate School of Public Health Outline My Experience Motivation for

### MODEL I: DRINK REGRESSED ON GPA & MALE, WITHOUT CENTERING

Interpreting Interaction Effects; Interaction Effects and Centering Richard Williams, University of Notre Dame, http://www3.nd.edu/~rwilliam/ Last revised February 20, 2015 Models with interaction effects

### Copyright 2010 The Guilford Press. Series Editor s Note

This is a chapter excerpt from Guilford Publications. Applied Missing Data Analysis, by Craig K. Enders. Copyright 2010. Series Editor s Note Missing data are a real bane to researchers across all social

### Software Cost Estimation with Incomplete Data

890 IEEE TRANSACTIONS ON SOFTWARE ENGINEERING, VOL. 27, NO. 10, OCTOBER 2001 Software Cost Estimation with Incomplete Data Kevin Strike, Khaled El Emam, and Nazim Madhavji AbstractÐThe construction of

### 4.1 Exploratory Analysis: Once the data is collected and entered, the first question is: "What do the data look like?"

Data Analysis Plan The appropriate methods of data analysis are determined by your data types and variables of interest, the actual distribution of the variables, and the number of cases. Different analyses

### Survey Data Analysis in Stata

Survey Data Analysis in Stata Jeff Pitblado Associate Director, Statistical Software StataCorp LP 2009 Canadian Stata Users Group Meeting Outline 1 Types of data 2 2 Survey data characteristics 4 2.1 Single

### Lecture 10: Logistical Regression II Multinomial Data. Prof. Sharyn O Halloran Sustainable Development U9611 Econometrics II

Lecture 10: Logistical Regression II Multinomial Data Prof. Sharyn O Halloran Sustainable Development U9611 Econometrics II Logit vs. Probit Review Use with a dichotomous dependent variable Need a link

### Using multilevel models to analyze couple and family treatment data: Basic and advanced issues. David C. Atkins. Travis Research Institute

Multilevel analyses 1 RUNNING HEAD: Multilevel analyses Using multilevel models to analyze couple and family treatment data: Basic and advanced issues David C. Atkins Travis Research Institute Fuller Graduate