Linda K. Muthén Bengt Muthén. Copyright 2008 Muthén & Muthén Table Of Contents

Size: px
Start display at page:

Download "Linda K. Muthén Bengt Muthén. Copyright 2008 Muthén & Muthén www.statmodel.com. Table Of Contents"

Transcription

1 Mplus Short Courses Topic 2 Regression Analysis, Eploratory Factor Analysis, Confirmatory Factor Analysis, And Structural Equation Modeling For Categorical, Censored, And Count Outcomes Linda K. Muthén Bengt Muthén Copyright 28 Muthén & Muthén Table Of Contents General Latent Variable Modeling Framework Analysis With Categorical Observed And Latent Variables Categorical Observed Variables Logit And Probit Regression British Coal Miner Eample Logistic Regression And Adjusted Odds Ratios Latent Response Variable Formulation Versus Probability Curve Formulation Ordered Polytomous Regression Alcohol Consumption Eample Unordered Polytomous Regression Censored Regression Count Regression Poisson Regression Negative Binomial Regression Path Analysis With Categorical Outcomes Occupational Destination Eample

2 Table Of Contents (Continued) Categorical Observed And Continuous Latent Variables Item Response Theory Eploratory Factor Analysis Practical Issues CFA With Covariates Antisocial Behavior Eample Multiple Group Analysis With Categorical Outcomes Technical Issues For Weighted Least Squares Estimation References Inefficient dissemination of statistical methods: Many good methods contributions from biostatistics, psychometrics, etc are underutilized in practice Fragmented presentation of methods: Technical descriptions in many different journals Many different pieces of limited software Mplus: Integration of methods in one framework Easy to use: Simple, non-technical language, graphics Powerful: General modeling capabilities Mplus versions V: November 998 V3: March 24 V5: November 27 Mplus Background V2: February 2 V4: February 26 Mplus team: Linda & Bengt Muthén, Thuy Nguyen, Tihomir Asparouhov, Michelle Conn, Jean Maninger 4 2

3 Statistical Analysis With Latent Variables A General Modeling Framework Statistical Concepts Captured By Latent Variables Continuous Latent Variables Measurement errors Factors Random effects Frailties, liabilities Variance components Missing data Categorical Latent Variables Latent classes Clusters Finite mitures Missing data 5 Statistical Analysis With Latent Variables A General Modeling Framework (Continued) Models That Use Latent Variables Continuous Latent Variables Factor analysis models Structural equation models Growth curve models Multilevel models Categorical Latent Variables Latent class models Miture models Discrete-time survival models Missing data models Mplus integrates the statistical concepts captured by latent variables into a general modeling framework that includes not only all of the models listed above but also combinations and etensions of these models. 6 3

4 General Latent Variable Modeling Framework Observed variables background variables (no model structure) y continuous and censored outcome variables u categorical (dichotomous, ordinal, nominal) and count outcome variables Latent variables f continuous variables c interactions among f s categorical variables multiple c s 7 Several programs in one Eploratory factor analysis Structural equation modeling Item response theory analysis Latent class analysis Latent transition analysis Survival analysis Growth modeling Multilevel analysis Comple survey data analysis Monte Carlo simulation Mplus Fully integrated in the general latent variable framework 8 4

5 Overview Of Mplus Courses Topic. March 8, 28, Johns Hopkins University: Introductory - advanced factor analysis and structural equation modeling with continuous outcomes Topic 2. March 9, 28, Johns Hopkins University: Introductory - advanced regression analysis, IRT, factor analysis and structural equation modeling with categorical, censored, and count outcomes Topic 3. August 2, 28, Johns Hopkins University: Introductory and intermediate growth modeling Topic 4. August 2, 28, Johns Hopkins University: Advanced growth modeling, survival analysis, and missing data analysis 9 Overview Of Mplus Courses (Continued) Topic 5. November, 28, University of Michigan, Ann Arbor: Categorical latent variable modeling with crosssectional data Topic 6. November, 28, University of Michigan, Ann Arbor: Categorical latent variable modeling with longitudinal data Topic 7. March 7, 29, Johns Hopkins University: Multilevel modeling of cross-sectional data Topic 8. March 8, 29, Johns Hopkins University: Multilevel modeling of longitudinal data 5

6 Analysis With Categorical Observed And Latent Variables Categorical Variable Modeling Categorical observed variables Categorical observed variables, continuous latent variables Categorical observed variables, categorical latent variables 2 6

7 Categorical Observed Variables 3 Two Eamples Alcohol Dependence And Gender In The NLSY Female Male n Not Dep Dep Prop Odds (Prop/(-Prop)) Odds Ratio =.79/.59 = 3.9 Eample wording: Males are three times more likely than females to be alcohol dependent. Colds And Vitamin C n No Cold Cold Prop Odds Placebo Vitamin C

8 Categorical Outcomes: Probability Concepts Probabilities: Joint: P (u, ) Marginal: P (u) Conditional: P (u ) Joint Female Alcohol Eample Conditional Not Dep.47 Dep.3 Male.43.8 Marginal.9. Distributions: Bernoulli: u = /; E(u) = π Binomial: sum or prop. (u = ), E(prop.) = π, V(prop.) = π( π)/n, π = prop Multinomial (#parameters = #cells ) Independent multinomial (product multinomial) Poisson Categorical Outcomes: Probability Concepts (Continued) u = u = Cross-product ratio (odds ratio): = π π π = / π π = π π π / ( ππ) = π / π P(u =, = ) / P(u =, = ) / P(u =, = ) / P(u =, = ) Tests: Log odds ratio (appro. normal) Test of proportions (appro. normal) Pearson χ 2 = Σ(O E) 2 / E (e.g. independence) Likelihood Ratio χ 2 = 2 Σ Olog(O / E ) 6 8

9 Further Readings On Categorical Variable Analysis Agresti, A. (22). Categorical data analysis. Second edition. New York: John Wiley & Sons. Agresti, A. (996). An introduction to categorical data analysis. New York: Wiley. Hosmer, D. W. & Lemeshow, S. (2). Applied logistic regression. Second edition. New York: John Wiley & Sons. Long, S. (997). Regression models for categorical and limited dependent variables. Thousand Oaks: Sage. 7 Logit And Probit Regression Dichotomous outcome Adjusted log odds Ordered, polytomous outcome Unordered, polytomous outcome Multivariate categorical outcomes 8 9

10 Logs Logarithmic Function Logistic Distribution Function e log P(u = ) Logit Logit [P(u = )] Logistic Density Density u * 9 Binary Outcome: Logistic Regression The logistic function P(u = ) = F ( + )=. + e ( + ) Logistic distribution function Logistic density F ( + ) F ( + ) + + Logistic score Logistic density: δ F / δ z = F( F) = f (z;, π 2 /3) 2

11 Binary Outcome: Probit Regression Probit regression considers P (u = ) = Φ ( + ), (6) where Φ is the standard normal distribution function. Using the inverse normal function Φ -, gives a linear probit equation Φ - [P(u = )] = +. (6) Normal distribution function Normal density Φ ( + ) Φ ( + ) + + z score 2 Interpreting Logit And Probit Coefficients Sign and significance Odds and odds ratios Probabilities 22

12 2 23 Logistic Regression And Log Odds Odds (u = ) = P(u = )/ P(u = ) = P(u = ) / ( P(u = )). The logistic function gives a log odds linear in, + + = + + ) ( / log ) ( ) ( e e [ ] e ) ( log + = = = ) ( ) ( ) ( * log e e e logit = log [odds (u = )] = log [P(u = ) / ( P(u = ))] ) ( ) ( - e u P + + = = 24 Logistic Regression And Log Odds (Continued) logit = log odds = + When changes one unit, the logit (log odds) changes units When changes one unit, the odds changes units e

13 British Coal Miner Data Have you eperienced breathlessness? Proportion yes Age 25 Plot Of Sample Logits Logit Age Sample logit = log [proportion / ( proportion)] 26 3

14 British Coal Miner Data (Continued) Age () N N Yes Proportion Yes OLS Estimated Probability Logit Estimated Probability Probit Estimated Probability ,952,79 2,3 2, ,393 2,9, , , SOURCE: Ashford & Sowden (97), Muthén (993) 2 Logit model: χ LRT (7) = 7.3 (p >.) Probit model: χ 2 LRT (7) = Coal Miner Data u w

15 Mplus Input For Categorical Outcomes Specifying dependent variables as categorical use the CATEGORICAL option CATEGORICAL ARE u u2 u3; Thresholds used instead of intercepts only different in sign Referring to thresholds in the model use $ number added to a variable name the number of thresholds is equal to the number of categories minus u$ refers to threshold of u u$2 refers to threshold 2 of u 29 Mplus Input For Categorical Outcomes (Continued) u2$ refers to threshold of u2 u2$2 refers to threshold 2 of u2 u2$3 refers to threshold 3 of u2 u3$ refers to threshold of u3 Referring to scale factors use { } to refer to scale factors u2 u3}; 3 5

16 Input For Logistic Regression Of Coal Miner Data TITLE: DATA: VARIABLE: DEFINE: ANALYSIS: MODEL: OUTPUT: Logistic regression of coal miner data FILE = coalminer.dat; NAMES = u w; CATEGORICAL = u; FREQWEIGHT = w; = /; ESTIMATOR = ML; u ON ; TECH SAMPSTAT STANDARDIZED; 3 Input For Probit Regression Of Coal Miner Data TITLE: DATA: VARIABLE: DEFINE: MODEL: OUTPUT: Probit regression of coal miner data FILE = coalminer.dat; NAMES = u w; CATEGORICAL = u; FREQWEIGHT = w; = /; u ON ; TECH SAMPSTAT STANDARDIZED; 32 6

17 Output Ecerpts Logistic Regression Of Coal Miner Data Model Results Estimates S.E. Est./S.E. Std StdYX U ON X Thresholds U$ Odds: e.25 = 2.79 As increases unit ( years), the odds of breathlessness increases Estimated Logistic Regression Probabilities For Coal Miner Data P ( u = ) = + e L where L = For = 6.2 (age 62) L = =.29 P( u = age 62) = + e,.29 =

18 Output Ecerpts Probit Regression Of Coal Miner Data Model Results Estimates S.E. Est./S.E. Std StdYX U ON X Thresholds U$ R-Square Observed Variable U Residual Variance. R-Square Estimated Probit Regression Probabilities For Coal Miner Data P (u = = 62) = Φ ( + ) = Φ (τ ) = Φ ( τ + ). Φ ( * 6.2) = Φ (.834).427 Note: logit probit * c where c = π 2 / 3 =

19 Categorical Outcomes: Logit And Probit Regression With One Binary And One Continuous X P(u =, 2 ) = F[ ], (22) P(u =, 2 ) = - P[u =, 2 ], where F[z] is either the standard normal (Φ[z]) or logistic (/[ + e -z ]) distribution function. Eample: Lung cancer and smoking among coal miners u lung cancer (u = ) or not (u = ) smoker ( = ), non-smoker ( = ) 2 years spent in coal mine 37 Categorical Outcomes: Logit And Probit Regression With One Binary And One Continuous X P(u =, 2 ) = F [ ], (22) P( u =, 2 ) = Probit / Logit = = =

20 Logistic Regression And Adjusted Odds Ratios Binary u variable regression on a binary variable and a continuous 2 variable: P (u =, 2 ) = - (, (62) e 2 2 ) which implies log odds = logit [P (u =, 2 )] = (63) This gives log odds{ = } = logit [P (u = =, 2 )] = + 2 2, (64) and log odds{ = } = logit [P (u = =, 2 )] = (65) 39 Logistic Regression And Adjusted Odds Ratios (Continued) The log odds ratio for u and adjusted for 2 is odds log OR = log [ ] = log odds log odds = (66) odds so that OR = ep ( ), constant for all values of 2. If an interaction term for and 2 is introduced, the constancy of the OR no longer holds. Eample wording: The odds of lung cancer adjusted for years is OR times higher for smokers than for nonsmokers The odds ratio adjusted for years is OR 4 2

21 Analysis Of NLSY Data: Odds Ratios For Alcohol Dependence And Gender Adjusting for Age First Started Drinking (n=976) Observed Frequencies, Proportions, and Odds Ratios Frequency Proportion Dependent Age st Female Male Female Male OR 2 or < or > Analysis Of NLSY Data: Odds Ratios For Alcohol Dependence And Gender (Continued) Estimated Probabilities and Odds Ratios Age st 2 or < or > Logit Female Male OR Female Male OR Probit Logit model: χ p (2) = 54.2 Probit model: χ 2 p (2) =

22 Analysis Of NLSY Data: Odds Ratios For Alcohol Dependence And Gender (Continued) Dependence on Gender and Age First Started Drinking Unstd. Coeff. Logit Regression s.e. t Std. Unstd. Coeff. Probit Regression s.e. t Std. Unstd. Coeff Rescaled To Logit Intercept Male Age st R OR = e.98 = 2.66 logit probit * c where c = π 2 / 3 =.8 43 NELS 88 Table 2.2 Odds ratios of eighth-grade students in 988 performing below basic levels of reading and mathematics in 988 and dropping out of school, 988 to 99, by basic demographics Variable Below basic mathematics Below basic reading Dropped out Se Female vs. male.8*.73**.92 Race ethnicity Asian vs. white Hispanic vs. white Black vs. white Native American vs. white ** 2.23** 2.43**.42** 2.29** 2.64** 3.5**.59 2.** 2.23** 2.5** Socioeconomic status Low vs. middle High vs. middle.9**.46**.9**.4** 3.95**.39* SOURCE: U.S. Department of Education, National Center for Education Statistics, National Education Longitudinal Study of 988 (NELS:88), Base Year and First Follow-Up surveys

23 NELS 88 Table 2.3 Adjusted odds ratios of eighth-grade students in 988 performing below basic levels of reading and mathematics in 988 and dropping out of school, 988 to 99, by basic demographics Variable Below basic mathematics Below basic reading Dropped out Se Female vs. male.77**.7**.86 Race ethnicity Asian vs. white Hispanic vs. white Black vs. white Native American vs. white.84.6**.77** 2.2**.46**.74** 2.9** 2.87** Socioeconomic status Low vs. middle High vs. middle.68**.49**.66**.44** 3.74**.4* 45 Latent Response Variable Formulation Versus Probability Curve Formulation Probability curve formulation in the binary u case: P (u = ) = F ( + ), (67) where F is the standard normal or logistic distribution function. Latent response variable formulation defines a threshold τ on a continuous u * variable so that u = is observed when u * eceeds τ while otherwise u = is observed, where δ ~ N (, V (δ)). u * = γ + δ, u = (68) u = τ u* 46 23

CHAPTER 12 EXAMPLES: MONTE CARLO SIMULATION STUDIES

CHAPTER 12 EXAMPLES: MONTE CARLO SIMULATION STUDIES Examples: Monte Carlo Simulation Studies CHAPTER 12 EXAMPLES: MONTE CARLO SIMULATION STUDIES Monte Carlo simulation studies are often used for methodological investigations of the performance of statistical

More information

CHAPTER 9 EXAMPLES: MULTILEVEL MODELING WITH COMPLEX SURVEY DATA

CHAPTER 9 EXAMPLES: MULTILEVEL MODELING WITH COMPLEX SURVEY DATA Examples: Multilevel Modeling With Complex Survey Data CHAPTER 9 EXAMPLES: MULTILEVEL MODELING WITH COMPLEX SURVEY DATA Complex survey data refers to data obtained by stratification, cluster sampling and/or

More information

The Latent Variable Growth Model In Practice. Individual Development Over Time

The Latent Variable Growth Model In Practice. Individual Development Over Time The Latent Variable Growth Model In Practice 37 Individual Development Over Time y i = 1 i = 2 i = 3 t = 1 t = 2 t = 3 t = 4 ε 1 ε 2 ε 3 ε 4 y 1 y 2 y 3 y 4 x η 0 η 1 (1) y ti = η 0i + η 1i x t + ε ti

More information

13. Poisson Regression Analysis

13. Poisson Regression Analysis 136 Poisson Regression Analysis 13. Poisson Regression Analysis We have so far considered situations where the outcome variable is numeric and Normally distributed, or binary. In clinical work one often

More information

CHAPTER 3 EXAMPLES: REGRESSION AND PATH ANALYSIS

CHAPTER 3 EXAMPLES: REGRESSION AND PATH ANALYSIS Examples: Regression And Path Analysis CHAPTER 3 EXAMPLES: REGRESSION AND PATH ANALYSIS Regression analysis with univariate or multivariate dependent variables is a standard procedure for modeling relationships

More information

Latent Class Regression Part II

Latent Class Regression Part II This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike License. Your use of this material constitutes acceptance of that license and the conditions of use of materials on this

More information

Automated Statistical Modeling for Data Mining David Stephenson 1

Automated Statistical Modeling for Data Mining David Stephenson 1 Automated Statistical Modeling for Data Mining David Stephenson 1 Abstract. We seek to bridge the gap between basic statistical data mining tools and advanced statistical analysis software that requires

More information

Ordinal Regression. Chapter

Ordinal Regression. Chapter Ordinal Regression Chapter 4 Many variables of interest are ordinal. That is, you can rank the values, but the real distance between categories is unknown. Diseases are graded on scales from least severe

More information

11. Analysis of Case-control Studies Logistic Regression

11. Analysis of Case-control Studies Logistic Regression Research methods II 113 11. Analysis of Case-control Studies Logistic Regression This chapter builds upon and further develops the concepts and strategies described in Ch.6 of Mother and Child Health:

More information

Models for Count Data With Overdispersion

Models for Count Data With Overdispersion Models for Count Data With Overdispersion Germán Rodríguez November 6, 2013 Abstract This addendum to the WWS 509 notes covers extra-poisson variation and the negative binomial model, with brief appearances

More information

It is important to bear in mind that one of the first three subscripts is redundant since k = i -j +3.

It is important to bear in mind that one of the first three subscripts is redundant since k = i -j +3. IDENTIFICATION AND ESTIMATION OF AGE, PERIOD AND COHORT EFFECTS IN THE ANALYSIS OF DISCRETE ARCHIVAL DATA Stephen E. Fienberg, University of Minnesota William M. Mason, University of Michigan 1. INTRODUCTION

More information

Multivariate Logistic Regression

Multivariate Logistic Regression 1 Multivariate Logistic Regression As in univariate logistic regression, let π(x) represent the probability of an event that depends on p covariates or independent variables. Then, using an inv.logit formulation

More information

Categorical Data Analysis

Categorical Data Analysis Richard L. Scheaffer University of Florida The reference material and many examples for this section are based on Chapter 8, Analyzing Association Between Categorical Variables, from Statistical Methods

More information

Calculating the Probability of Returning a Loan with Binary Probability Models

Calculating the Probability of Returning a Loan with Binary Probability Models Calculating the Probability of Returning a Loan with Binary Probability Models Associate Professor PhD Julian VASILEV (e-mail: vasilev@ue-varna.bg) Varna University of Economics, Bulgaria ABSTRACT The

More information

Generalized Linear Models

Generalized Linear Models Generalized Linear Models We have previously worked with regression models where the response variable is quantitative and normally distributed. Now we turn our attention to two types of models where the

More information

STATISTICA Formula Guide: Logistic Regression. Table of Contents

STATISTICA Formula Guide: Logistic Regression. Table of Contents : Table of Contents... 1 Overview of Model... 1 Dispersion... 2 Parameterization... 3 Sigma-Restricted Model... 3 Overparameterized Model... 4 Reference Coding... 4 Model Summary (Summary Tab)... 5 Summary

More information

Multinomial and Ordinal Logistic Regression

Multinomial and Ordinal Logistic Regression Multinomial and Ordinal Logistic Regression ME104: Linear Regression Analysis Kenneth Benoit August 22, 2012 Regression with categorical dependent variables When the dependent variable is categorical,

More information

Logistic regression modeling the probability of success

Logistic regression modeling the probability of success Logistic regression modeling the probability of success Regression models are usually thought of as only being appropriate for target variables that are continuous Is there any situation where we might

More information

Overview Classes. 12-3 Logistic regression (5) 19-3 Building and applying logistic regression (6) 26-3 Generalizations of logistic regression (7)

Overview Classes. 12-3 Logistic regression (5) 19-3 Building and applying logistic regression (6) 26-3 Generalizations of logistic regression (7) Overview Classes 12-3 Logistic regression (5) 19-3 Building and applying logistic regression (6) 26-3 Generalizations of logistic regression (7) 2-4 Loglinear models (8) 5-4 15-17 hrs; 5B02 Building and

More information

VI. Introduction to Logistic Regression

VI. Introduction to Logistic Regression VI. Introduction to Logistic Regression We turn our attention now to the topic of modeling a categorical outcome as a function of (possibly) several factors. The framework of generalized linear models

More information

Multiple Choice Models II

Multiple Choice Models II Multiple Choice Models II Laura Magazzini University of Verona laura.magazzini@univr.it http://dse.univr.it/magazzini Laura Magazzini (@univr.it) Multiple Choice Models II 1 / 28 Categorical data Categorical

More information

SHORT COURSE ON MPLUS Getting Started with Mplus HANDOUT

SHORT COURSE ON MPLUS Getting Started with Mplus HANDOUT SHORT COURSE ON MPLUS Getting Started with Mplus HANDOUT Instructor: Cathy Zimmer 962-0516, cathy_zimmer@unc.edu INTRODUCTION a) Who am I? Who are you? b) Overview of Course i) Mplus capabilities ii) The

More information

The Proportional Odds Model for Assessing Rater Agreement with Multiple Modalities

The Proportional Odds Model for Assessing Rater Agreement with Multiple Modalities The Proportional Odds Model for Assessing Rater Agreement with Multiple Modalities Elizabeth Garrett-Mayer, PhD Assistant Professor Sidney Kimmel Comprehensive Cancer Center Johns Hopkins University 1

More information

SAS Software to Fit the Generalized Linear Model

SAS Software to Fit the Generalized Linear Model SAS Software to Fit the Generalized Linear Model Gordon Johnston, SAS Institute Inc., Cary, NC Abstract In recent years, the class of generalized linear models has gained popularity as a statistical modeling

More information

I L L I N O I S UNIVERSITY OF ILLINOIS AT URBANA-CHAMPAIGN

I L L I N O I S UNIVERSITY OF ILLINOIS AT URBANA-CHAMPAIGN Beckman HLM Reading Group: Questions, Answers and Examples Carolyn J. Anderson Department of Educational Psychology I L L I N O I S UNIVERSITY OF ILLINOIS AT URBANA-CHAMPAIGN Linear Algebra Slide 1 of

More information

Generalized Linear Models. Today: definition of GLM, maximum likelihood estimation. Involves choice of a link function (systematic component)

Generalized Linear Models. Today: definition of GLM, maximum likelihood estimation. Involves choice of a link function (systematic component) Generalized Linear Models Last time: definition of exponential family, derivation of mean and variance (memorize) Today: definition of GLM, maximum likelihood estimation Include predictors x i through

More information

Multinomial Logistic Regression

Multinomial Logistic Regression Multinomial Logistic Regression Dr. Jon Starkweather and Dr. Amanda Kay Moske Multinomial logistic regression is used to predict categorical placement in or the probability of category membership on a

More information

Basic Statistical and Modeling Procedures Using SAS

Basic Statistical and Modeling Procedures Using SAS Basic Statistical and Modeling Procedures Using SAS One-Sample Tests The statistical procedures illustrated in this handout use two datasets. The first, Pulse, has information collected in a classroom

More information

When to Use Which Statistical Test

When to Use Which Statistical Test When to Use Which Statistical Test Rachel Lovell, Ph.D., Senior Research Associate Begun Center for Violence Prevention Research and Education Jack, Joseph, and Morton Mandel School of Applied Social Sciences

More information

VI. The Investigation of the Determinants of Bicycling in Colorado

VI. The Investigation of the Determinants of Bicycling in Colorado VI. The Investigation of the Determinants of Bicycling in Colorado Using the data described earlier in this report, statistical analyses are performed to identify the factors that influence the propensity

More information

CHAPTER 8 EXAMPLES: MIXTURE MODELING WITH LONGITUDINAL DATA

CHAPTER 8 EXAMPLES: MIXTURE MODELING WITH LONGITUDINAL DATA Examples: Mixture Modeling With Longitudinal Data CHAPTER 8 EXAMPLES: MIXTURE MODELING WITH LONGITUDINAL DATA Mixture modeling refers to modeling with categorical latent variables that represent subpopulations

More information

Unit 12 Logistic Regression Supplementary Chapter 14 in IPS On CD (Chap 16, 5th ed.)

Unit 12 Logistic Regression Supplementary Chapter 14 in IPS On CD (Chap 16, 5th ed.) Unit 12 Logistic Regression Supplementary Chapter 14 in IPS On CD (Chap 16, 5th ed.) Logistic regression generalizes methods for 2-way tables Adds capability studying several predictors, but Limited to

More information

Example: Credit card default, we may be more interested in predicting the probabilty of a default than classifying individuals as default or not.

Example: Credit card default, we may be more interested in predicting the probabilty of a default than classifying individuals as default or not. Statistical Learning: Chapter 4 Classification 4.1 Introduction Supervised learning with a categorical (Qualitative) response Notation: - Feature vector X, - qualitative response Y, taking values in C

More information

Multiple logistic regression analysis of cigarette use among high school students

Multiple logistic regression analysis of cigarette use among high school students Multiple logistic regression analysis of cigarette use among high school students ABSTRACT Joseph Adwere-Boamah Alliant International University A binary logistic regression analysis was performed to predict

More information

Examples of Using R for Modeling Ordinal Data

Examples of Using R for Modeling Ordinal Data Examples of Using R for Modeling Ordinal Data Alan Agresti Department of Statistics, University of Florida Supplement for the book Analysis of Ordinal Categorical Data, 2nd ed., 2010 (Wiley), abbreviated

More information

Additional sources Compilation of sources: http://lrs.ed.uiuc.edu/tseportal/datacollectionmethodologies/jin-tselink/tselink.htm

Additional sources Compilation of sources: http://lrs.ed.uiuc.edu/tseportal/datacollectionmethodologies/jin-tselink/tselink.htm Mgt 540 Research Methods Data Analysis 1 Additional sources Compilation of sources: http://lrs.ed.uiuc.edu/tseportal/datacollectionmethodologies/jin-tselink/tselink.htm http://web.utk.edu/~dap/random/order/start.htm

More information

Analysis of Microdata

Analysis of Microdata Rainer Winkelmann Stefan Boes Analysis of Microdata With 38 Figures and 41 Tables 4y Springer Contents 1 Introduction 1 1.1 What Are Microdata? 1 1.2 Types of Microdata 4 1.2.1 Qualitative Data 4 1.2.2

More information

Predicting Successful Completion of the Nursing Program: An Analysis of Prerequisites and Demographic Variables

Predicting Successful Completion of the Nursing Program: An Analysis of Prerequisites and Demographic Variables Predicting Successful Completion of the Nursing Program: An Analysis of Prerequisites and Demographic Variables Introduction In the summer of 2002, a research study commissioned by the Center for Student

More information

The Probit Link Function in Generalized Linear Models for Data Mining Applications

The Probit Link Function in Generalized Linear Models for Data Mining Applications Journal of Modern Applied Statistical Methods Copyright 2013 JMASM, Inc. May 2013, Vol. 12, No. 1, 164-169 1538 9472/13/$95.00 The Probit Link Function in Generalized Linear Models for Data Mining Applications

More information

Auxiliary Variables in Mixture Modeling: 3-Step Approaches Using Mplus

Auxiliary Variables in Mixture Modeling: 3-Step Approaches Using Mplus Auxiliary Variables in Mixture Modeling: 3-Step Approaches Using Mplus Tihomir Asparouhov and Bengt Muthén Mplus Web Notes: No. 15 Version 8, August 5, 2014 1 Abstract This paper discusses alternatives

More information

HLM software has been one of the leading statistical packages for hierarchical

HLM software has been one of the leading statistical packages for hierarchical Introductory Guide to HLM With HLM 7 Software 3 G. David Garson HLM software has been one of the leading statistical packages for hierarchical linear modeling due to the pioneering work of Stephen Raudenbush

More information

Using the Delta Method to Construct Confidence Intervals for Predicted Probabilities, Rates, and Discrete Changes

Using the Delta Method to Construct Confidence Intervals for Predicted Probabilities, Rates, and Discrete Changes Using the Delta Method to Construct Confidence Intervals for Predicted Probabilities, Rates, Discrete Changes JunXuJ.ScottLong Indiana University August 22, 2005 The paper provides technical details on

More information

Aileen Murphy, Department of Economics, UCC, Ireland. WORKING PAPER SERIES 07-10

Aileen Murphy, Department of Economics, UCC, Ireland. WORKING PAPER SERIES 07-10 AN ECONOMETRIC ANALYSIS OF SMOKING BEHAVIOUR IN IRELAND Aileen Murphy, Department of Economics, UCC, Ireland. DEPARTMENT OF ECONOMICS WORKING PAPER SERIES 07-10 1 AN ECONOMETRIC ANALYSIS OF SMOKING BEHAVIOUR

More information

Master programme in Statistics

Master programme in Statistics Master programme in Statistics Björn Holmquist 1 1 Department of Statistics Lund University Cramérsällskapets årskonferens, 2010-03-25 Master programme Vad är ett Master programme? Breddmaster vs Djupmaster

More information

PROC LOGISTIC: Traps for the unwary Peter L. Flom, Independent statistical consultant, New York, NY

PROC LOGISTIC: Traps for the unwary Peter L. Flom, Independent statistical consultant, New York, NY PROC LOGISTIC: Traps for the unwary Peter L. Flom, Independent statistical consultant, New York, NY ABSTRACT Keywords: Logistic. INTRODUCTION This paper covers some gotchas in SAS R PROC LOGISTIC. A gotcha

More information

Simple Linear Regression Inference

Simple Linear Regression Inference Simple Linear Regression Inference 1 Inference requirements The Normality assumption of the stochastic term e is needed for inference even if it is not a OLS requirement. Therefore we have: Interpretation

More information

Longitudinal Data Analysis. Wiley Series in Probability and Statistics

Longitudinal Data Analysis. Wiley Series in Probability and Statistics Brochure More information from http://www.researchandmarkets.com/reports/2172736/ Longitudinal Data Analysis. Wiley Series in Probability and Statistics Description: Longitudinal data analysis for biomedical

More information

Department of Epidemiology and Public Health Miller School of Medicine University of Miami

Department of Epidemiology and Public Health Miller School of Medicine University of Miami Department of Epidemiology and Public Health Miller School of Medicine University of Miami BST 630 (3 Credit Hours) Longitudinal and Multilevel Data Wednesday-Friday 9:00 10:15PM Course Location: CRB 995

More information

SPSS TRAINING SESSION 3 ADVANCED TOPICS (PASW STATISTICS 17.0) Sun Li Centre for Academic Computing lsun@smu.edu.sg

SPSS TRAINING SESSION 3 ADVANCED TOPICS (PASW STATISTICS 17.0) Sun Li Centre for Academic Computing lsun@smu.edu.sg SPSS TRAINING SESSION 3 ADVANCED TOPICS (PASW STATISTICS 17.0) Sun Li Centre for Academic Computing lsun@smu.edu.sg IN SPSS SESSION 2, WE HAVE LEARNT: Elementary Data Analysis Group Comparison & One-way

More information

Handling missing data in Stata a whirlwind tour

Handling missing data in Stata a whirlwind tour Handling missing data in Stata a whirlwind tour 2012 Italian Stata Users Group Meeting Jonathan Bartlett www.missingdata.org.uk 20th September 2012 1/55 Outline The problem of missing data and a principled

More information

Structural Equation Models for Comparing Dependent Means and Proportions. Jason T. Newsom

Structural Equation Models for Comparing Dependent Means and Proportions. Jason T. Newsom Structural Equation Models for Comparing Dependent Means and Proportions Jason T. Newsom How to Do a Paired t-test with Structural Equation Modeling Jason T. Newsom Overview Rationale Structural equation

More information

Logistic Regression. http://faculty.chass.ncsu.edu/garson/pa765/logistic.htm#sigtests

Logistic Regression. http://faculty.chass.ncsu.edu/garson/pa765/logistic.htm#sigtests Logistic Regression http://faculty.chass.ncsu.edu/garson/pa765/logistic.htm#sigtests Overview Binary (or binomial) logistic regression is a form of regression which is used when the dependent is a dichotomy

More information

Calculating Effect-Sizes

Calculating Effect-Sizes Calculating Effect-Sizes David B. Wilson, PhD George Mason University August 2011 The Heart and Soul of Meta-analysis: The Effect Size Meta-analysis shifts focus from statistical significance to the direction

More information

Multinomial and ordinal logistic regression using PROC LOGISTIC Peter L. Flom National Development and Research Institutes, Inc

Multinomial and ordinal logistic regression using PROC LOGISTIC Peter L. Flom National Development and Research Institutes, Inc ABSTRACT Multinomial and ordinal logistic regression using PROC LOGISTIC Peter L. Flom National Development and Research Institutes, Inc Logistic regression may be useful when we are trying to model a

More information

Service courses for graduate students in degree programs other than the MS or PhD programs in Biostatistics.

Service courses for graduate students in degree programs other than the MS or PhD programs in Biostatistics. Course Catalog In order to be assured that all prerequisites are met, students must acquire a permission number from the education coordinator prior to enrolling in any Biostatistics course. Courses are

More information

Logistic Regression (1/24/13)

Logistic Regression (1/24/13) STA63/CBB540: Statistical methods in computational biology Logistic Regression (/24/3) Lecturer: Barbara Engelhardt Scribe: Dinesh Manandhar Introduction Logistic regression is model for regression used

More information

A Basic Guide to Modeling Techniques for All Direct Marketing Challenges

A Basic Guide to Modeling Techniques for All Direct Marketing Challenges A Basic Guide to Modeling Techniques for All Direct Marketing Challenges Allison Cornia Database Marketing Manager Microsoft Corporation C. Olivia Rud Executive Vice President Data Square, LLC Overview

More information

5. Ordinal regression: cumulative categories proportional odds. 6. Ordinal regression: comparison to single reference generalized logits

5. Ordinal regression: cumulative categories proportional odds. 6. Ordinal regression: comparison to single reference generalized logits Lecture 23 1. Logistic regression with binary response 2. Proc Logistic and its surprises 3. quadratic model 4. Hosmer-Lemeshow test for lack of fit 5. Ordinal regression: cumulative categories proportional

More information

Interpretation of Somers D under four simple models

Interpretation of Somers D under four simple models Interpretation of Somers D under four simple models Roger B. Newson 03 September, 04 Introduction Somers D is an ordinal measure of association introduced by Somers (96)[9]. It can be defined in terms

More information

Lecture #2 Overview. Basic IRT Concepts, Models, and Assumptions. Lecture #2 ICPSR Item Response Theory Workshop

Lecture #2 Overview. Basic IRT Concepts, Models, and Assumptions. Lecture #2 ICPSR Item Response Theory Workshop Basic IRT Concepts, Models, and Assumptions Lecture #2 ICPSR Item Response Theory Workshop Lecture #2: 1of 64 Lecture #2 Overview Background of IRT and how it differs from CFA Creating a scale An introduction

More information

Nominal and ordinal logistic regression

Nominal and ordinal logistic regression Nominal and ordinal logistic regression April 26 Nominal and ordinal logistic regression Our goal for today is to briefly go over ways to extend the logistic regression model to the case where the outcome

More information

Yew May Martin Maureen Maclachlan Tom Karmel Higher Education Division, Department of Education, Training and Youth Affairs.

Yew May Martin Maureen Maclachlan Tom Karmel Higher Education Division, Department of Education, Training and Youth Affairs. How is Australia s Higher Education Performing? An analysis of completion rates of a cohort of Australian Post Graduate Research Students in the 1990s. Yew May Martin Maureen Maclachlan Tom Karmel Higher

More information

Mplus Short Courses Topic 7. Multilevel Modeling With Latent Variables Using Mplus: Cross-Sectional Analysis

Mplus Short Courses Topic 7. Multilevel Modeling With Latent Variables Using Mplus: Cross-Sectional Analysis Mplus Short Courses Topic 7 Multilevel Modeling With Latent Variables Using Mplus: Cross-Sectional Analysis Linda K. Muthén Bengt Muthén Copyright 2011 Muthén & Muthén www.statmodel.com 03/29/2011 1 Table

More information

Module 4 - Multiple Logistic Regression

Module 4 - Multiple Logistic Regression Module 4 - Multiple Logistic Regression Objectives Understand the principles and theory underlying logistic regression Understand proportions, probabilities, odds, odds ratios, logits and exponents Be

More information

Statistical Analysis of Life Insurance Policy Termination and Survivorship

Statistical Analysis of Life Insurance Policy Termination and Survivorship Statistical Analysis of Life Insurance Policy Termination and Survivorship Emiliano A. Valdez, PhD, FSA Michigan State University joint work with J. Vadiveloo and U. Dias Session ES82 (Statistics in Actuarial

More information

INTRODUCTORY STATISTICS

INTRODUCTORY STATISTICS INTRODUCTORY STATISTICS FIFTH EDITION Thomas H. Wonnacott University of Western Ontario Ronald J. Wonnacott University of Western Ontario WILEY JOHN WILEY & SONS New York Chichester Brisbane Toronto Singapore

More information

Sun Li Centre for Academic Computing lsun@smu.edu.sg

Sun Li Centre for Academic Computing lsun@smu.edu.sg Sun Li Centre for Academic Computing lsun@smu.edu.sg Elementary Data Analysis Group Comparison & One-way ANOVA Non-parametric Tests Correlations General Linear Regression Logistic Models Binary Logistic

More information

Students' Opinion about Universities: The Faculty of Economics and Political Science (Case Study)

Students' Opinion about Universities: The Faculty of Economics and Political Science (Case Study) Cairo University Faculty of Economics and Political Science Statistics Department English Section Students' Opinion about Universities: The Faculty of Economics and Political Science (Case Study) Prepared

More information

CREDIT SCORING MODEL APPLICATIONS:

CREDIT SCORING MODEL APPLICATIONS: Örebro University Örebro University School of Business Master in Applied Statistics Thomas Laitila Sune Karlsson May, 2014 CREDIT SCORING MODEL APPLICATIONS: TESTING MULTINOMIAL TARGETS Gabriela De Rossi

More information

Rethinking the Cultural Context of Schooling Decisions in Disadvantaged Neighborhoods: From Deviant Subculture to Cultural Heterogeneity

Rethinking the Cultural Context of Schooling Decisions in Disadvantaged Neighborhoods: From Deviant Subculture to Cultural Heterogeneity Rethinking the Cultural Context of Schooling Decisions in Disadvantaged Neighborhoods: From Deviant Subculture to Cultural Heterogeneity Sociology of Education David J. Harding, University of Michigan

More information

When to Use a Particular Statistical Test

When to Use a Particular Statistical Test When to Use a Particular Statistical Test Central Tendency Univariate Descriptive Mode the most commonly occurring value 6 people with ages 21, 22, 21, 23, 19, 21 - mode = 21 Median the center value the

More information

Discussion Section 4 ECON 139/239 2010 Summer Term II

Discussion Section 4 ECON 139/239 2010 Summer Term II Discussion Section 4 ECON 139/239 2010 Summer Term II 1. Let s use the CollegeDistance.csv data again. (a) An education advocacy group argues that, on average, a person s educational attainment would increase

More information

Module 14: Missing Data Stata Practical

Module 14: Missing Data Stata Practical Module 14: Missing Data Stata Practical Jonathan Bartlett & James Carpenter London School of Hygiene & Tropical Medicine www.missingdata.org.uk Supported by ESRC grant RES 189-25-0103 and MRC grant G0900724

More information

XPost: Excel Workbooks for the Post-estimation Interpretation of Regression Models for Categorical Dependent Variables

XPost: Excel Workbooks for the Post-estimation Interpretation of Regression Models for Categorical Dependent Variables XPost: Excel Workbooks for the Post-estimation Interpretation of Regression Models for Categorical Dependent Variables Contents Simon Cheng hscheng@indiana.edu php.indiana.edu/~hscheng/ J. Scott Long jslong@indiana.edu

More information

Lecture 25. December 19, 2007. Department of Biostatistics Johns Hopkins Bloomberg School of Public Health Johns Hopkins University.

Lecture 25. December 19, 2007. Department of Biostatistics Johns Hopkins Bloomberg School of Public Health Johns Hopkins University. This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike License. Your use of this material constitutes acceptance of that license and the conditions of use of materials on this

More information

Free Trial - BIRT Analytics - IAAs

Free Trial - BIRT Analytics - IAAs Free Trial - BIRT Analytics - IAAs 11. Predict Customer Gender Once we log in to BIRT Analytics Free Trial we would see that we have some predefined advanced analysis ready to be used. Those saved analysis

More information

From the help desk: hurdle models

From the help desk: hurdle models The Stata Journal (2003) 3, Number 2, pp. 178 184 From the help desk: hurdle models Allen McDowell Stata Corporation Abstract. This article demonstrates that, although there is no command in Stata for

More information

Virtual Parental Involvement: The Role of the Internet in Parent-School Communications

Virtual Parental Involvement: The Role of the Internet in Parent-School Communications Virtual Parental Involvement: The Role of the Internet in Parent-School Communications Suzanne M. Bouffard Department of Psychology, Duke University & Harvard Family Research Project, Harvard Graduate

More information

Developing Risk Adjustment Techniques Using the SAS@ System for Assessing Health Care Quality in the lmsystem@

Developing Risk Adjustment Techniques Using the SAS@ System for Assessing Health Care Quality in the lmsystem@ Developing Risk Adjustment Techniques Using the SAS@ System for Assessing Health Care Quality in the lmsystem@ Yanchun Xu, Andrius Kubilius Joint Commission on Accreditation of Healthcare Organizations,

More information

Data Mining: An Overview of Methods and Technologies for Increasing Profits in Direct Marketing. C. Olivia Rud, VP, Fleet Bank

Data Mining: An Overview of Methods and Technologies for Increasing Profits in Direct Marketing. C. Olivia Rud, VP, Fleet Bank Data Mining: An Overview of Methods and Technologies for Increasing Profits in Direct Marketing C. Olivia Rud, VP, Fleet Bank ABSTRACT Data Mining is a new term for the common practice of searching through

More information

Comparison of Estimation Methods for Complex Survey Data Analysis

Comparison of Estimation Methods for Complex Survey Data Analysis Comparison of Estimation Methods for Complex Survey Data Analysis Tihomir Asparouhov 1 Muthen & Muthen Bengt Muthen 2 UCLA 1 Tihomir Asparouhov, Muthen & Muthen, 3463 Stoner Ave. Los Angeles, CA 90066.

More information

Tips for surviving the analysis of survival data. Philip Twumasi-Ankrah, PhD

Tips for surviving the analysis of survival data. Philip Twumasi-Ankrah, PhD Tips for surviving the analysis of survival data Philip Twumasi-Ankrah, PhD Big picture In medical research and many other areas of research, we often confront continuous, ordinal or dichotomous outcomes

More information

Introduction to Hypothesis Testing. Point estimation and confidence intervals are useful statistical inference procedures.

Introduction to Hypothesis Testing. Point estimation and confidence intervals are useful statistical inference procedures. Introduction to Hypothesis Testing Point estimation and confidence intervals are useful statistical inference procedures. Another type of inference is used frequently used concerns tests of hypotheses.

More information

STA 4273H: Statistical Machine Learning

STA 4273H: Statistical Machine Learning STA 4273H: Statistical Machine Learning Russ Salakhutdinov Department of Statistics! rsalakhu@utstat.toronto.edu! http://www.cs.toronto.edu/~rsalakhu/ Lecture 6 Three Approaches to Classification Construct

More information

Cross Validation techniques in R: A brief overview of some methods, packages, and functions for assessing prediction models.

Cross Validation techniques in R: A brief overview of some methods, packages, and functions for assessing prediction models. Cross Validation techniques in R: A brief overview of some methods, packages, and functions for assessing prediction models. Dr. Jon Starkweather, Research and Statistical Support consultant This month

More information

Introduction to Data Analysis in Hierarchical Linear Models

Introduction to Data Analysis in Hierarchical Linear Models Introduction to Data Analysis in Hierarchical Linear Models April 20, 2007 Noah Shamosh & Frank Farach Social Sciences StatLab Yale University Scope & Prerequisites Strong applied emphasis Focus on HLM

More information

Institute of Actuaries of India Subject CT3 Probability and Mathematical Statistics

Institute of Actuaries of India Subject CT3 Probability and Mathematical Statistics Institute of Actuaries of India Subject CT3 Probability and Mathematical Statistics For 2015 Examinations Aim The aim of the Probability and Mathematical Statistics subject is to provide a grounding in

More information

ANNUITY LAPSE RATE MODELING: TOBIT OR NOT TOBIT? 1. INTRODUCTION

ANNUITY LAPSE RATE MODELING: TOBIT OR NOT TOBIT? 1. INTRODUCTION ANNUITY LAPSE RATE MODELING: TOBIT OR NOT TOBIT? SAMUEL H. COX AND YIJIA LIN ABSTRACT. We devise an approach, using tobit models for modeling annuity lapse rates. The approach is based on data provided

More information

Introduction to Longitudinal Data Analysis

Introduction to Longitudinal Data Analysis Introduction to Longitudinal Data Analysis Longitudinal Data Analysis Workshop Section 1 University of Georgia: Institute for Interdisciplinary Research in Education and Human Development Section 1: Introduction

More information

Exploring Relationships using SPSS inferential statistics (Part II) Dwayne Devonish

Exploring Relationships using SPSS inferential statistics (Part II) Dwayne Devonish Exploring Relationships using SPSS inferential statistics (Part II) Dwayne Devonish Reminder: Types of Variables Categorical Variables Based on qualitative type variables. Gender, Ethnicity, religious

More information

240ST014 - Data Analysis of Transport and Logistics

240ST014 - Data Analysis of Transport and Logistics Coordinating unit: Teaching unit: Academic year: Degree: ECTS credits: 2015 240 - ETSEIB - Barcelona School of Industrial Engineering 715 - EIO - Department of Statistics and Operations Research MASTER'S

More information

Simple Predictive Analytics Curtis Seare

Simple Predictive Analytics Curtis Seare Using Excel to Solve Business Problems: Simple Predictive Analytics Curtis Seare Copyright: Vault Analytics July 2010 Contents Section I: Background Information Why use Predictive Analytics? How to use

More information

MS 2007-0081-RR Booil Jo. Supplemental Materials (to be posted on the Web)

MS 2007-0081-RR Booil Jo. Supplemental Materials (to be posted on the Web) MS 2007-0081-RR Booil Jo Supplemental Materials (to be posted on the Web) Table 2 Mplus Input title: Table 2 Monte Carlo simulation using externally generated data. One level CACE analysis based on eqs.

More information

Statistical Models in R

Statistical Models in R Statistical Models in R Some Examples Steven Buechler Department of Mathematics 276B Hurley Hall; 1-6233 Fall, 2007 Outline Statistical Models Structure of models in R Model Assessment (Part IA) Anova

More information

Applied Multiple Regression/Correlation Analysis for the Behavioral Sciences

Applied Multiple Regression/Correlation Analysis for the Behavioral Sciences Applied Multiple Regression/Correlation Analysis for the Behavioral Sciences Third Edition Jacob Cohen (deceased) New York University Patricia Cohen New York State Psychiatric Institute and Columbia University

More information

2015 TUHH Online Summer School: Overview of Statistical and Path Modeling Analyses

2015 TUHH Online Summer School: Overview of Statistical and Path Modeling Analyses : Overview of Statistical and Path Modeling Analyses Prof. Dr. Christian M. Ringle (Hamburg Univ. of Tech., TUHH) Prof. Dr. Jӧrg Henseler (University of Twente) Dr. Geoffrey Hubona (The Georgia R School)

More information

EXPANDING THE EVIDENCE BASE IN OUTCOMES RESEARCH: USING LINKED ELECTRONIC MEDICAL RECORDS (EMR) AND CLAIMS DATA

EXPANDING THE EVIDENCE BASE IN OUTCOMES RESEARCH: USING LINKED ELECTRONIC MEDICAL RECORDS (EMR) AND CLAIMS DATA EXPANDING THE EVIDENCE BASE IN OUTCOMES RESEARCH: USING LINKED ELECTRONIC MEDICAL RECORDS (EMR) AND CLAIMS DATA A CASE STUDY EXAMINING RISK FACTORS AND COSTS OF UNCONTROLLED HYPERTENSION ISPOR 2013 WORKSHOP

More information

Raul Cruz-Cano, HLTH653 Spring 2013

Raul Cruz-Cano, HLTH653 Spring 2013 Multilevel Modeling-Logistic Schedule 3/18/2013 = Spring Break 3/25/2013 = Longitudinal Analysis 4/1/2013 = Midterm (Exercises 1-5, not Longitudinal) Introduction Just as with linear regression, logistic

More information

Regression 3: Logistic Regression

Regression 3: Logistic Regression Regression 3: Logistic Regression Marco Baroni Practical Statistics in R Outline Logistic regression Logistic regression in R Outline Logistic regression Introduction The model Looking at and comparing

More information

The Chi-Square Test. STAT E-50 Introduction to Statistics

The Chi-Square Test. STAT E-50 Introduction to Statistics STAT -50 Introduction to Statistics The Chi-Square Test The Chi-square test is a nonparametric test that is used to compare experimental results with theoretical models. That is, we will be comparing observed

More information