SUGI 29 Statistics and Data Analysis
|
|
|
- Charlene Newton
- 10 years ago
- Views:
Transcription
1 Paper Head of the CLASS: Impress your colleagues with a superior understanding of the CLASS statement in PROC LOGISTIC Michelle L. Pritchard and David J. Pasta Ovation Research Group, San Francisco, CA ABSTRACT Data analysts and statistical programmers in a variety of disciplines use PROC LOGISTIC to fit logistic and proportional odds models to binary and ordinal responses. These models frequently contain categorical explanatory variables, and programmers must use the CLASS statement to identify which variables are categorical predictors. SAS Version 8 allows analysts more power and flexibility when handling categorized effects, specifically by allowing you to specify the method of parameterization and the reference category. But, be warned, this new set of tools should be used with caution! The default parameterization may not furnish what you would expect if you are used to the default provided by the CLASS statement in PROC GLM. The selected parameterization method has a profound effect on how CONTRAST statements are specified and in the interpretation of the parameter estimates. Also, the syntax is new and does take some getting used to. This paper provides a series of examples to allow programmers to become comfortable using the CLASS statement to specify and test desired hypotheses in PROC LOGISTIC. INTRODUCTION Beginning in SAS Version 8 (SAS version 8.02 for Windows 2000 will be used throughout this paper), the LOGISTIC procedure enables programmers to specify a full-rank parameterization (with the choice of effect, reference, polynomial, or orthogonal polynomial coding), or a less than full-rank parameterization (as in the GLM procedure). This is a big step forward from the days of doing your own coding of dummy variables, like you would for regression. Let s begin by stepping through the new syntax. The GLM parameterization is only available as a global option (i.e., for all variables in the class statement), but the full rank parameterizations can be specified either globally or for individual variables. Global parameterization is specified by the PARAM= <EFFECT GLM ORTHPOLY POLYNOMIAL REFERENCE> option after the forward slash (/) in the CLASS statement, as follows: class <classvar1> <classvar2>/param = glm; title3"example 1: PARAM = GLM Global Option"; class <classvar1> <classvar2>/param = ref; title3"example 2: PARAM = REF Global Option"; Alternatively, parameterization for individual variables can be specified in parentheses immediately following the variable in the CLASS statement, as follows: class <classvar1> (param = ref) <classvar2> (param = effect); title3"example 3: Individual Variable Parameterization Options"; As an extra bit of flexibility, the REF= option in the CLASS statement determines the reference level for the EFFECT and REFERENCE coding. As with the PARAM= option, the REF= option can be coded globally or separately for each classification variable. To program the reference level for each variable individually, use the REF = option in parentheses immediately following the variable in the CLASS statement, or use the keyword FIRST or LAST (FIRST
2 designates the first ordered category as the reference and LAST designates the last ordered category as the reference). Note that unless you use FIRST or LAST, the reference level should be placed in either single or double quotation marks, as follows: class <classvar1> (param = ref ref = <refcat> ) <classvar2> (param = effect ref = last); title3"example 4: Individual Variable Parameterization Options with Individual Specification of Reference Category"; Alternatively, the reference category can be applied to all variables in the CLASS statement by using the keyword FIRST or LAST in the REF= option after the forward slash (/) in the CLASS statement. class <classvar1> (param = ref) <classvar2> (param = effect)/ref = first; title3"example 5: Individual Variable Parameterization Options with Global Specification of Reference Category"; Noteworthy notes: 1. If you neglect to specify a coding method and/or reference category, then the default parameterization method is PARAM=EFFECT and the default reference category is the last ordered category (REF=LAST). 2. Individual parameterization trumps the global option, unless the global option is GLM. 3. REF= is only valid when PARAM=EFFECT or PARAM=REF. 4. If a format has been applied to the classification variables, then either the formatted value (or FIRST or LAST) must be used when specifying the reference category. If the keyword FIRST or LAST is used in this situation, then it pertains to the first or last ordered formatted category. 5. We recommend limiting the classification variables to no more than one parameterization method, unless there are specific multiple hypotheses that you want to test and you are confident in what you are doing and you won t be showing the output to anyone else and you are sure you ll remember later why you did what you did and you get the idea. DATA Now that we are familiar with the new PARAM= and REF= syntax, let s consider a model with a dichotomous response and three classification variables. The variables involved are distributed as follows: Cumulative Cumulative response Frequency Percent Frequency Percent Cumulative Cumulative female Frequency Percent Frequency Percent 0. Male Female Cumulative Cumulative smoker Frequency Percent Frequency Percent 1. Never Past Current Cumulative Cumulative agecat Frequency Percent Frequency Percent
3 female smoker agecat response Frequency Percent MODEL A: DEFAULT PARAMETERIZATION (EFFECT) class female smoker agecat ; model response = female smoker agecat; title3"model A: Logistic regression with three categorical predictors and default options PARAM=EFFECT and REF=LAST"; In Model A, the method of parameterization is not specified, so the default EFFECT parameterization will be used. (Also, by default the last ordered category will be used as the reference category.) The columns of the design matrix are created based on the EFFECT coding scheme, as follows: For each variable, c-1 columns comprise the design matrix, where c is the number of levels of the classification variable. For the nonreference levels, the columns represent group membership (1=member, 0 =non-member), and for the reference level, the row is populated with a 1. In our Model A, the reference level is the last ordered category, so the design matrix is: 3
4 Class Level Information Design Variables Class Value female smoker agecat Using EFFECT coding, the beta estimates are estimating the difference in the effect of each nonreference level compared to the average effect over all levels. Analysis of Maximum Likelihood Estimates Standard Wald Parameter DF Estimate Error Chi-Square Pr > ChiSq Intercept female smoker smoker agecat agecat agecat Therefore, the odds of having response = 1 for never vs. current smokers is: exp( )/exp[-( )] = Point 95% Wald Effect Estimate Confidence Limits female 0 vs smoker 1 vs smoker 2 vs agecat 1 vs agecat 2 vs agecat 3 vs Odds Ratio Estimates MODEL B: GLM PARAMETERIZATION class female smoker agecat/param = glm ; model response = female smoker agecat; title3"model B: Logistic regression with three categorical predictors and PARAM = GLM option"; In Model B, the GLM parameterization has been specified. With GLM parameterization, the programmer does not have control over the reference category using the REF= option; instead, SAS includes variables as long as they are linearly independent from those that went before. This has the effect of making the last ordered category the omitted (reference) category. The columns of the design matrix are created based on the GLM coding scheme, as follows: For each variable, c columns comprise the design matrix, where c is the number of levels of the classification variable. For all levels, the columns represent group membership (1=member, 0 =non-member). The design matrix for Model B is: 4
5 Class Level Information Design Variables Class Value female smoker agecat Using GLM coding, the beta estimates are estimating the difference in the effect of each nonreference level compared to the reference (last) level. Analysis of Maximum Likelihood Estimates Standard Wald Parameter DF Estimate Error Chi-Square Pr > ChiSq intercept female female smoker smoker smoker agecat agecat agecat agecat Therefore, the odds of having response = 1 for never vs. current smokers is: exp( ) = Point 95% Wald Effect Estimate Confidence Limits female 0 vs smoker 1 vs smoker 2 vs agecat 1 vs agecat 2 vs agecat 3 vs Odds Ratio Estimates One of the reasons this parameterization may be desirable is that it corresponds to the behavior of GLM and MIXED. That parameterization is familiar to long-time users of those PROCs. One disadvantage is that in order to control which category is omitted (treated as the reference category), it is necessary to order the categories so the category to be omitted is last. MODEL C: REFERENCE PARAMETERIZATION The REFERENCE parameterization is similar to GLM except you actually omit one of the categories (so the resulting design matrix is full rank) and you have explicit control over which category is to be omitted; it need not be the last category. class female smoker agecat/param = ref; model response = female smoker agecat; title3"model C: Logistic regression with three categorical predictors and PARAM = REF option"; 5
6 In Model C, the REFERENCE parameterization has been specified. (Also, by default the last ordered category will be used as the reference category.) The columns of the design matrix are created based on the REFERENCE coding scheme, as follows: For each variable, c-1 columns comprise the design matrix, where c is the number of levels of the classification variable. For the nonreference levels, the columns represent group membership (1=member, 0 =non-member), and for the reference level, the row is populated with a 0. In our Model C, the reference level is the last ordered category, so the design matrix is: Design Variables Class Value female smoker agecat Class Level Information Using REFERENCE coding, the beta estimates are estimating the difference in the effect of each nonreference level compared to the effect of the reference level. Analysis of Maximum Likelihood Estimates Standard Wald Parameter DF Estimate Error Chi-Square Pr > ChiSq Intercept female smoker smoker agecat agecat agecat Therefore, as with the GLM parameterization, the odds of having response = 1 for never vs. current smokers is: exp( ) = Point 95% Wald Effect Estimate Confidence Limits female 0 vs smoker 1 vs smoker 2 vs agecat 1 vs agecat 2 vs agecat 3 vs Odds Ratio Estimates MODEL D: REFERENCE PARAMETERIZATION WITH REFERENCE CATEGORY SPECIFICATION In this model we use the reference parameterization but specify REF=FIRST. This changes the design matrix and the odds ratios are inverted for binary variables. class female smoker agecat/ param = ref ref = first ; model response = female smoker agecat; title3"model D: Logistic regression with three categorical predictors and PARAM=REF and REF=first"; 6
7 Class Level Information Design Variables Class Value female smoker agecat Analysis of Maximum Likelihood Estimates Standard Wald Parameter DF Estimate Error Chi-Square Pr > ChiSq Intercept female smoker smoker agecat agecat agecat Therefore, the odds of having response = 1 for current vs. never smokers is: exp(0.1337) = Point 95% Wald Effect Estimate Confidence Limits female 1 vs smoker 2 vs smoker 3 vs agecat 2 vs agecat 3 vs agecat 4 vs Odds Ratio Estimates MODEL E: DEFAULT WITH INTERACTIONS One thing that analysts new to the CLASS statement sometimes find mysterious is that when the specify interactions sometimes SAS changes the order of the variables. They specified female*smoker, for example, but in the output they see smoker*female. What is SAS doing? It turns out it s paying attention to the order of the variables in the CLASS statement. class female smoker agecat; model response = female smoker agecat female*smoker female*agecat smoker*agecat; title3"model E1: Logistic regression with three categorical predictors plus two-way interactions, default options"; class smoker agecat female; model response = female smoker agecat female*smoker female*agecat smoker*agecat; title3"model E2: Logistic regression with three categorical predictors plus two-way interactions, default options"; 7
8 MODEL F: PARAM=GLM WITH INTERACTIONS class female smoker agecat/param = glm; model response = female smoker agecat female*smoker female*agecat smoker*agecat; title3"model F: Logistic regression with three categorical predictors plus two-way interactions, with PARAM = GLM option"; HYPOTHESIS TESTING WITH CONTRAST STATEMENTS The CONTRAST statement in PROC LOGISTIC works much like the corresponding statement in other procedures such as GLM and MIXED. It allows you to specify one or more linear combinations of the parameters to test against zero. You can test each of the linear combinations against zero or specify that several linear combinations be tested against zero simultaneously (a multiple-degrees-of-freedom test). In order to correctly specify the CONTRAST statement, you need to pay close attention to the parameterization. For example, consider the AGECAT variable as parameterized in Model A, the default EFFECT coding. To test group 1 against group 2, the contrast statement would be contrast '1 vs. 2' agecat / e; The 0 at the end is not necessary but is a good reminder that there are three parameters for the AGECAT variable. The E option is extremely useful: it requests that the linear combination be displayed, giving you a chance to make sure the coefficients are lined up with the proper parameters. We strongly recomment using the E option when testing new constrast statements. Testing that groups 1, 2, and 3 are simultaneously equal requires a two degree of freedom test: contrast '1 = 2 = 3' agecat 1 1 0, agecat / e; or, equivalently, contrast '1 = 2 = 3' agecat 1 1 0, agecat / e; or contrast '1 = 2 = 3' agecat 1 0 1, agecat / e; all of which test the same hypothesis, that the first three AGECAT groups are equal. Because the fourth AGECAT group is coded as the negation of the first three, contrasts involving that group look rather different. If you label the parameters for the first three groups A1, A2, and A3, then the parameter for the fourth group is A1- A2-A3. Thus if you want to compare group 1 to group 4, you want to test (A1)-(-A1-A2-A3)=0, or 2*A1+A2+A3=0: contrast '1 vs. 4' agecat / e; It is clear how useful it is to label each contrast it would be easy to forget what this contrast is designed to test! The construction of CONTRAST statements can be very tricky in the presence of many interactions. To be sure you are testing the hypothesis you want to test, use the E option and write down the resulting equation separately. It may be worth seeking advice from a statistician for especially complex models. CONCLUSION The different model parameterizations allowed in PROC LOGISTIC are a powerful tool for specifying models and testing hypotheses. With that power comes responsibility responsibility to be sure the results you get are what you intended. To make sure that happens, remember to: 1. pay attention to the order of the levels of the class variables 2. pay attention to the order of the variables in the CLASS statement 3. pay attention to the design matrix 4. understand the beta estimates and hypothesis tests 5. understand the syntax and defaults, and what options are and are not available for each method of parameterization 6. be aware of how formatted data are treated CONTACT INFORMATION The authors welcome questions and comments. Please direct inquiries to: Michelle L. Pritchard David J. Pasta Ovation Research Group Ovation Research Group 120 Howard St., Ste Howard St., Ste. 600 San Francisco, CA San Francisco, CA [email protected] [email protected] (310) (415)
ln(p/(1-p)) = α +β*age35plus, where p is the probability or odds of drinking
Dummy Coding for Dummies Kathryn Martin, Maternal, Child and Adolescent Health Program, California Department of Public Health ABSTRACT There are a number of ways to incorporate categorical variables into
Statistics and Data Analysis
NESUG 27 PRO LOGISTI: The Logistics ehind Interpreting ategorical Variable Effects Taylor Lewis, U.S. Office of Personnel Management, Washington, D STRT The goal of this paper is to demystify how SS models
Cool Tools for PROC LOGISTIC
Cool Tools for PROC LOGISTIC Paul D. Allison Statistical Horizons LLC and the University of Pennsylvania March 2013 www.statisticalhorizons.com 1 New Features in LOGISTIC ODDSRATIO statement EFFECTPLOT
PROC LOGISTIC: Traps for the unwary Peter L. Flom, Independent statistical consultant, New York, NY
PROC LOGISTIC: Traps for the unwary Peter L. Flom, Independent statistical consultant, New York, NY ABSTRACT Keywords: Logistic. INTRODUCTION This paper covers some gotchas in SAS R PROC LOGISTIC. A gotcha
SAS Software to Fit the Generalized Linear Model
SAS Software to Fit the Generalized Linear Model Gordon Johnston, SAS Institute Inc., Cary, NC Abstract In recent years, the class of generalized linear models has gained popularity as a statistical modeling
Chapter 39 The LOGISTIC Procedure. Chapter Table of Contents
Chapter 39 The LOGISTIC Procedure Chapter Table of Contents OVERVIEW...1903 GETTING STARTED...1906 SYNTAX...1910 PROCLOGISTICStatement...1910 BYStatement...1912 CLASSStatement...1913 CONTRAST Statement.....1916
11. Analysis of Case-control Studies Logistic Regression
Research methods II 113 11. Analysis of Case-control Studies Logistic Regression This chapter builds upon and further develops the concepts and strategies described in Ch.6 of Mother and Child Health:
STATISTICA Formula Guide: Logistic Regression. Table of Contents
: Table of Contents... 1 Overview of Model... 1 Dispersion... 2 Parameterization... 3 Sigma-Restricted Model... 3 Overparameterized Model... 4 Reference Coding... 4 Model Summary (Summary Tab)... 5 Summary
Multinomial and Ordinal Logistic Regression
Multinomial and Ordinal Logistic Regression ME104: Linear Regression Analysis Kenneth Benoit August 22, 2012 Regression with categorical dependent variables When the dependent variable is categorical,
Is it statistically significant? The chi-square test
UAS Conference Series 2013/14 Is it statistically significant? The chi-square test Dr Gosia Turner Student Data Management and Analysis 14 September 2010 Page 1 Why chi-square? Tests whether two categorical
Chapter 29 The GENMOD Procedure. Chapter Table of Contents
Chapter 29 The GENMOD Procedure Chapter Table of Contents OVERVIEW...1365 WhatisaGeneralizedLinearModel?...1366 ExamplesofGeneralizedLinearModels...1367 TheGENMODProcedure...1368 GETTING STARTED...1370
An Introduction to Statistical Tests for the SAS Programmer Sara Beck, Fred Hutchinson Cancer Research Center, Seattle, WA
ABSTRACT An Introduction to Statistical Tests for the SAS Programmer Sara Beck, Fred Hutchinson Cancer Research Center, Seattle, WA Often SAS Programmers find themselves in situations where performing
Analysis of Survey Data Using the SAS SURVEY Procedures: A Primer
Analysis of Survey Data Using the SAS SURVEY Procedures: A Primer Patricia A. Berglund, Institute for Social Research - University of Michigan Wisconsin and Illinois SAS User s Group June 25, 2014 1 Overview
Using An Ordered Logistic Regression Model with SAS Vartanian: SW 541
Using An Ordered Logistic Regression Model with SAS Vartanian: SW 541 libname in1 >c:\=; Data first; Set in1.extract; A=1; PROC LOGIST OUTEST=DD MAXITER=100 ORDER=DATA; OUTPUT OUT=CC XBETA=XB P=PROB; MODEL
ABSTRACT INTRODUCTION
Paper SP03-2009 Illustrative Logistic Regression Examples using PROC LOGISTIC: New Features in SAS/STAT 9.2 Robert G. Downer, Grand Valley State University, Allendale, MI Patrick J. Richardson, Van Andel
Statistics in Retail Finance. Chapter 2: Statistical models of default
Statistics in Retail Finance 1 Overview > We consider how to build statistical models of default, or delinquency, and how such models are traditionally used for credit application scoring and decision
Logistic Regression. http://faculty.chass.ncsu.edu/garson/pa765/logistic.htm#sigtests
Logistic Regression http://faculty.chass.ncsu.edu/garson/pa765/logistic.htm#sigtests Overview Binary (or binomial) logistic regression is a form of regression which is used when the dependent is a dichotomy
VI. Introduction to Logistic Regression
VI. Introduction to Logistic Regression We turn our attention now to the topic of modeling a categorical outcome as a function of (possibly) several factors. The framework of generalized linear models
Ordinal Regression. Chapter
Ordinal Regression Chapter 4 Many variables of interest are ordinal. That is, you can rank the values, but the real distance between categories is unknown. Diseases are graded on scales from least severe
Binary Logistic Regression
Binary Logistic Regression Main Effects Model Logistic regression will accept quantitative, binary or categorical predictors and will code the latter two in various ways. Here s a simple model including
Logistic (RLOGIST) Example #1
Logistic (RLOGIST) Example #1 SUDAAN Statements and Results Illustrated EFFECTS RFORMAT, RLABEL REFLEVEL EXP option on MODEL statement Hosmer-Lemeshow Test Input Data Set(s): BRFWGT.SAS7bdat Example Using
Basic Statistical and Modeling Procedures Using SAS
Basic Statistical and Modeling Procedures Using SAS One-Sample Tests The statistical procedures illustrated in this handout use two datasets. The first, Pulse, has information collected in a classroom
This can dilute the significance of a departure from the null hypothesis. We can focus the test on departures of a particular form.
One-Degree-of-Freedom Tests Test for group occasion interactions has (number of groups 1) number of occasions 1) degrees of freedom. This can dilute the significance of a departure from the null hypothesis.
Improved Interaction Interpretation: Application of the EFFECTPLOT statement and other useful features in PROC LOGISTIC
Paper AA08-2013 Improved Interaction Interpretation: Application of the EFFECTPLOT statement and other useful features in PROC LOGISTIC Robert G. Downer, Grand Valley State University, Allendale, MI ABSTRACT
Generalized Linear Models
Generalized Linear Models We have previously worked with regression models where the response variable is quantitative and normally distributed. Now we turn our attention to two types of models where the
Overview Classes. 12-3 Logistic regression (5) 19-3 Building and applying logistic regression (6) 26-3 Generalizations of logistic regression (7)
Overview Classes 12-3 Logistic regression (5) 19-3 Building and applying logistic regression (6) 26-3 Generalizations of logistic regression (7) 2-4 Loglinear models (8) 5-4 15-17 hrs; 5B02 Building and
Developing Risk Adjustment Techniques Using the SAS@ System for Assessing Health Care Quality in the lmsystem@
Developing Risk Adjustment Techniques Using the SAS@ System for Assessing Health Care Quality in the lmsystem@ Yanchun Xu, Andrius Kubilius Joint Commission on Accreditation of Healthcare Organizations,
How to set the main menu of STATA to default factory settings standards
University of Pretoria Data analysis for evaluation studies Examples in STATA version 11 List of data sets b1.dta (To be created by students in class) fp1.xls (To be provided to students) fp1.txt (To be
Students' Opinion about Universities: The Faculty of Economics and Political Science (Case Study)
Cairo University Faculty of Economics and Political Science Statistics Department English Section Students' Opinion about Universities: The Faculty of Economics and Political Science (Case Study) Prepared
Two Correlated Proportions (McNemar Test)
Chapter 50 Two Correlated Proportions (Mcemar Test) Introduction This procedure computes confidence intervals and hypothesis tests for the comparison of the marginal frequencies of two factors (each with
Free Trial - BIRT Analytics - IAAs
Free Trial - BIRT Analytics - IAAs 11. Predict Customer Gender Once we log in to BIRT Analytics Free Trial we would see that we have some predefined advanced analysis ready to be used. Those saved analysis
Simple Linear Regression, Scatterplots, and Bivariate Correlation
1 Simple Linear Regression, Scatterplots, and Bivariate Correlation This section covers procedures for testing the association between two continuous variables using the SPSS Regression and Correlate analyses.
Interaction effects and group comparisons Richard Williams, University of Notre Dame, http://www3.nd.edu/~rwilliam/ Last revised February 20, 2015
Interaction effects and group comparisons Richard Williams, University of Notre Dame, http://www3.nd.edu/~rwilliam/ Last revised February 20, 2015 Note: This handout assumes you understand factor variables,
Yew May Martin Maureen Maclachlan Tom Karmel Higher Education Division, Department of Education, Training and Youth Affairs.
How is Australia s Higher Education Performing? An analysis of completion rates of a cohort of Australian Post Graduate Research Students in the 1990s. Yew May Martin Maureen Maclachlan Tom Karmel Higher
Credit Risk Analysis Using Logistic Regression Modeling
Credit Risk Analysis Using Logistic Regression Modeling Introduction A loan officer at a bank wants to be able to identify characteristics that are indicative of people who are likely to default on loans,
Logistic (RLOGIST) Example #3
Logistic (RLOGIST) Example #3 SUDAAN Statements and Results Illustrated PREDMARG (predicted marginal proportion) CONDMARG (conditional marginal proportion) PRED_EFF pairwise comparison COND_EFF pairwise
Data Mining: An Overview of Methods and Technologies for Increasing Profits in Direct Marketing. C. Olivia Rud, VP, Fleet Bank
Data Mining: An Overview of Methods and Technologies for Increasing Profits in Direct Marketing C. Olivia Rud, VP, Fleet Bank ABSTRACT Data Mining is a new term for the common practice of searching through
Multinomial and ordinal logistic regression using PROC LOGISTIC Peter L. Flom National Development and Research Institutes, Inc
ABSTRACT Multinomial and ordinal logistic regression using PROC LOGISTIC Peter L. Flom National Development and Research Institutes, Inc Logistic regression may be useful when we are trying to model a
International Statistical Institute, 56th Session, 2007: Phil Everson
Teaching Regression using American Football Scores Everson, Phil Swarthmore College Department of Mathematics and Statistics 5 College Avenue Swarthmore, PA198, USA E-mail: [email protected] 1. Introduction
Modeling Lifetime Value in the Insurance Industry
Modeling Lifetime Value in the Insurance Industry C. Olivia Parr Rud, Executive Vice President, Data Square, LLC ABSTRACT Acquisition modeling for direct mail insurance has the unique challenge of targeting
MULTIPLE REGRESSION WITH CATEGORICAL DATA
DEPARTMENT OF POLITICAL SCIENCE AND INTERNATIONAL RELATIONS Posc/Uapp 86 MULTIPLE REGRESSION WITH CATEGORICAL DATA I. AGENDA: A. Multiple regression with categorical variables. Coding schemes. Interpreting
Guido s Guide to PROC FREQ A Tutorial for Beginners Using the SAS System Joseph J. Guido, University of Rochester Medical Center, Rochester, NY
Guido s Guide to PROC FREQ A Tutorial for Beginners Using the SAS System Joseph J. Guido, University of Rochester Medical Center, Rochester, NY ABSTRACT PROC FREQ is an essential procedure within BASE
Categorical Data Analysis
Richard L. Scheaffer University of Florida The reference material and many examples for this section are based on Chapter 8, Analyzing Association Between Categorical Variables, from Statistical Methods
Lecture 14: GLM Estimation and Logistic Regression
Lecture 14: GLM Estimation and Logistic Regression Dipankar Bandyopadhyay, Ph.D. BMTRY 711: Analysis of Categorical Data Spring 2011 Division of Biostatistics and Epidemiology Medical University of South
Gamma Distribution Fitting
Chapter 552 Gamma Distribution Fitting Introduction This module fits the gamma probability distributions to a complete or censored set of individual or grouped data values. It outputs various statistics
ANALYSING LIKERT SCALE/TYPE DATA, ORDINAL LOGISTIC REGRESSION EXAMPLE IN R.
ANALYSING LIKERT SCALE/TYPE DATA, ORDINAL LOGISTIC REGRESSION EXAMPLE IN R. 1. Motivation. Likert items are used to measure respondents attitudes to a particular question or statement. One must recall
Simple Predictive Analytics Curtis Seare
Using Excel to Solve Business Problems: Simple Predictive Analytics Curtis Seare Copyright: Vault Analytics July 2010 Contents Section I: Background Information Why use Predictive Analytics? How to use
Descriptive Statistics
Descriptive Statistics Primer Descriptive statistics Central tendency Variation Relative position Relationships Calculating descriptive statistics Descriptive Statistics Purpose to describe or summarize
Segmentation For Insurance Payments Michael Sherlock, Transcontinental Direct, Warminster, PA
Segmentation For Insurance Payments Michael Sherlock, Transcontinental Direct, Warminster, PA ABSTRACT An online insurance agency has built a base of names that responded to different offers from various
Charles Secolsky County College of Morris. Sathasivam 'Kris' Krishnan The Richard Stockton College of New Jersey
Using logistic regression for validating or invalidating initial statewide cut-off scores on basic skills placement tests at the community college level Abstract Charles Secolsky County College of Morris
GLM I An Introduction to Generalized Linear Models
GLM I An Introduction to Generalized Linear Models CAS Ratemaking and Product Management Seminar March 2009 Presented by: Tanya D. Havlicek, Actuarial Assistant 0 ANTITRUST Notice The Casualty Actuarial
Elements of statistics (MATH0487-1)
Elements of statistics (MATH0487-1) Prof. Dr. Dr. K. Van Steen University of Liège, Belgium December 10, 2012 Introduction to Statistics Basic Probability Revisited Sampling Exploratory Data Analysis -
1 Theory: The General Linear Model
QMIN GLM Theory - 1.1 1 Theory: The General Linear Model 1.1 Introduction Before digital computers, statistics textbooks spoke of three procedures regression, the analysis of variance (ANOVA), and the
MISSING DATA TECHNIQUES WITH SAS. IDRE Statistical Consulting Group
MISSING DATA TECHNIQUES WITH SAS IDRE Statistical Consulting Group ROAD MAP FOR TODAY To discuss: 1. Commonly used techniques for handling missing data, focusing on multiple imputation 2. Issues that could
Consider a study in which. How many subjects? The importance of sample size calculations. An insignificant effect: two possibilities.
Consider a study in which How many subjects? The importance of sample size calculations Office of Research Protections Brown Bag Series KB Boomer, Ph.D. Director, [email protected] A researcher conducts
Paper D10 2009. Ranking Predictors in Logistic Regression. Doug Thompson, Assurant Health, Milwaukee, WI
Paper D10 2009 Ranking Predictors in Logistic Regression Doug Thompson, Assurant Health, Milwaukee, WI ABSTRACT There is little consensus on how best to rank predictors in logistic regression. This paper
Regression 3: Logistic Regression
Regression 3: Logistic Regression Marco Baroni Practical Statistics in R Outline Logistic regression Logistic regression in R Outline Logistic regression Introduction The model Looking at and comparing
Auxiliary Variables in Mixture Modeling: 3-Step Approaches Using Mplus
Auxiliary Variables in Mixture Modeling: 3-Step Approaches Using Mplus Tihomir Asparouhov and Bengt Muthén Mplus Web Notes: No. 15 Version 8, August 5, 2014 1 Abstract This paper discusses alternatives
Unit 12 Logistic Regression Supplementary Chapter 14 in IPS On CD (Chap 16, 5th ed.)
Unit 12 Logistic Regression Supplementary Chapter 14 in IPS On CD (Chap 16, 5th ed.) Logistic regression generalizes methods for 2-way tables Adds capability studying several predictors, but Limited to
Abbas S. Tavakoli, DrPH, MPH, ME 1 ; Nikki R. Wooten, PhD, LISW-CP 2,3, Jordan Brittingham, MSPH 4
1 Paper 1680-2016 Using GENMOD to Analyze Correlated Data on Military System Beneficiaries Receiving Inpatient Behavioral Care in South Carolina Care Systems Abbas S. Tavakoli, DrPH, MPH, ME 1 ; Nikki
Model Fitting in PROC GENMOD Jean G. Orelien, Analytical Sciences, Inc.
Paper 264-26 Model Fitting in PROC GENMOD Jean G. Orelien, Analytical Sciences, Inc. Abstract: There are several procedures in the SAS System for statistical modeling. Most statisticians who use the SAS
Lecture 19: Conditional Logistic Regression
Lecture 19: Conditional Logistic Regression Dipankar Bandyopadhyay, Ph.D. BMTRY 711: Analysis of Categorical Data Spring 2011 Division of Biostatistics and Epidemiology Medical University of South Carolina
Chapter 5 Analysis of variance SPSS Analysis of variance
Chapter 5 Analysis of variance SPSS Analysis of variance Data file used: gss.sav How to get there: Analyze Compare Means One-way ANOVA To test the null hypothesis that several population means are equal,
HYPOTHESIS TESTING: CONFIDENCE INTERVALS, T-TESTS, ANOVAS, AND REGRESSION
HYPOTHESIS TESTING: CONFIDENCE INTERVALS, T-TESTS, ANOVAS, AND REGRESSION HOD 2990 10 November 2010 Lecture Background This is a lightning speed summary of introductory statistical methods for senior undergraduate
Additional sources Compilation of sources: http://lrs.ed.uiuc.edu/tseportal/datacollectionmethodologies/jin-tselink/tselink.htm
Mgt 540 Research Methods Data Analysis 1 Additional sources Compilation of sources: http://lrs.ed.uiuc.edu/tseportal/datacollectionmethodologies/jin-tselink/tselink.htm http://web.utk.edu/~dap/random/order/start.htm
Directions for using SPSS
Directions for using SPSS Table of Contents Connecting and Working with Files 1. Accessing SPSS... 2 2. Transferring Files to N:\drive or your computer... 3 3. Importing Data from Another File Format...
How to Get More Value from Your Survey Data
Technical report How to Get More Value from Your Survey Data Discover four advanced analysis techniques that make survey research more effective Table of contents Introduction..............................................................2
This chapter will demonstrate how to perform multiple linear regression with IBM SPSS
CHAPTER 7B Multiple Regression: Statistical Methods Using IBM SPSS This chapter will demonstrate how to perform multiple linear regression with IBM SPSS first using the standard method and then using the
SPSS Explore procedure
SPSS Explore procedure One useful function in SPSS is the Explore procedure, which will produce histograms, boxplots, stem-and-leaf plots and extensive descriptive statistics. To run the Explore procedure,
Weight of Evidence Module
Formula Guide The purpose of the Weight of Evidence (WoE) module is to provide flexible tools to recode the values in continuous and categorical predictor variables into discrete categories automatically,
Logistic (RLOGIST) Example #7
Logistic (RLOGIST) Example #7 SUDAAN Statements and Results Illustrated EFFECTS UNITS option EXP option SUBPOPX REFLEVEL Input Data Set(s): SAMADULTED.SAS7bdat Example Using 2006 NHIS data, determine for
ASSIGNMENT 4 PREDICTIVE MODELING AND GAINS CHARTS
DATABASE MARKETING Fall 2015, max 24 credits Dead line 15.10. ASSIGNMENT 4 PREDICTIVE MODELING AND GAINS CHARTS PART A Gains chart with excel Prepare a gains chart from the data in \\work\courses\e\27\e20100\ass4b.xls.
CREDIT SCORING MODEL APPLICATIONS:
Örebro University Örebro University School of Business Master in Applied Statistics Thomas Laitila Sune Karlsson May, 2014 CREDIT SCORING MODEL APPLICATIONS: TESTING MULTINOMIAL TARGETS Gabriela De Rossi
SUMAN DUVVURU STAT 567 PROJECT REPORT
SUMAN DUVVURU STAT 567 PROJECT REPORT SURVIVAL ANALYSIS OF HEROIN ADDICTS Background and introduction: Current illicit drug use among teens is continuing to increase in many countries around the world.
New SAS Procedures for Analysis of Sample Survey Data
New SAS Procedures for Analysis of Sample Survey Data Anthony An and Donna Watts, SAS Institute Inc, Cary, NC Abstract Researchers use sample surveys to obtain information on a wide variety of issues Many
IBM SPSS Direct Marketing 23
IBM SPSS Direct Marketing 23 Note Before using this information and the product it supports, read the information in Notices on page 25. Product Information This edition applies to version 23, release
Multivariate Logistic Regression
1 Multivariate Logistic Regression As in univariate logistic regression, let π(x) represent the probability of an event that depends on p covariates or independent variables. Then, using an inv.logit formulation
ABSTRACT INTRODUCTION STUDY DESCRIPTION
ABSTRACT Paper 1675-2014 Validating Self-Reported Survey Measures Using SAS Sarah A. Lyons MS, Kimberly A. Kaphingst ScD, Melody S. Goodman PhD Washington University School of Medicine Researchers often
Predicting Successful Completion of the Nursing Program: An Analysis of Prerequisites and Demographic Variables
Predicting Successful Completion of the Nursing Program: An Analysis of Prerequisites and Demographic Variables Introduction In the summer of 2002, a research study commissioned by the Center for Student
NCSS Statistical Software Principal Components Regression. In ordinary least squares, the regression coefficients are estimated using the formula ( )
Chapter 340 Principal Components Regression Introduction is a technique for analyzing multiple regression data that suffer from multicollinearity. When multicollinearity occurs, least squares estimates
MULTIPLE REGRESSION EXAMPLE
MULTIPLE REGRESSION EXAMPLE For a sample of n = 166 college students, the following variables were measured: Y = height X 1 = mother s height ( momheight ) X 2 = father s height ( dadheight ) X 3 = 1 if
Recommend Continued CPS Monitoring. 63 (a) 17 (b) 10 (c) 90. 35 (d) 20 (e) 25 (f) 80. Totals/Marginal 98 37 35 170
Work Sheet 2: Calculating a Chi Square Table 1: Substance Abuse Level by ation Total/Marginal 63 (a) 17 (b) 10 (c) 90 35 (d) 20 (e) 25 (f) 80 Totals/Marginal 98 37 35 170 Step 1: Label Your Table. Label
I L L I N O I S UNIVERSITY OF ILLINOIS AT URBANA-CHAMPAIGN
Beckman HLM Reading Group: Questions, Answers and Examples Carolyn J. Anderson Department of Educational Psychology I L L I N O I S UNIVERSITY OF ILLINOIS AT URBANA-CHAMPAIGN Linear Algebra Slide 1 of
LOGISTIC REGRESSION ANALYSIS
LOGISTIC REGRESSION ANALYSIS C. Mitchell Dayton Department of Measurement, Statistics & Evaluation Room 1230D Benjamin Building University of Maryland September 1992 1. Introduction and Model Logistic
A Primer on Mathematical Statistics and Univariate Distributions; The Normal Distribution; The GLM with the Normal Distribution
A Primer on Mathematical Statistics and Univariate Distributions; The Normal Distribution; The GLM with the Normal Distribution PSYC 943 (930): Fundamentals of Multivariate Modeling Lecture 4: September
USING LOGISTIC REGRESSION TO PREDICT CUSTOMER RETENTION. Andrew H. Karp Sierra Information Services, Inc. San Francisco, California USA
USING LOGISTIC REGRESSION TO PREDICT CUSTOMER RETENTION Andrew H. Karp Sierra Information Services, Inc. San Francisco, California USA Logistic regression is an increasingly popular statistical technique
A Handbook of Statistical Analyses Using R. Brian S. Everitt and Torsten Hothorn
A Handbook of Statistical Analyses Using R Brian S. Everitt and Torsten Hothorn CHAPTER 6 Logistic Regression and Generalised Linear Models: Blood Screening, Women s Role in Society, and Colonic Polyps
Beginning Tutorials. PROC FREQ: It s More Than Counts Richard Severino, The Queen s Medical Center, Honolulu, HI OVERVIEW.
Paper 69-25 PROC FREQ: It s More Than Counts Richard Severino, The Queen s Medical Center, Honolulu, HI ABSTRACT The FREQ procedure can be used for more than just obtaining a simple frequency distribution
SP10 From GLM to GLIMMIX-Which Model to Choose? Patricia B. Cerrito, University of Louisville, Louisville, KY
SP10 From GLM to GLIMMIX-Which Model to Choose? Patricia B. Cerrito, University of Louisville, Louisville, KY ABSTRACT The purpose of this paper is to investigate several SAS procedures that are used in
Data Mining and Data Warehousing. Henryk Maciejewski. Data Mining Predictive modelling: regression
Data Mining and Data Warehousing Henryk Maciejewski Data Mining Predictive modelling: regression Algorithms for Predictive Modelling Contents Regression Classification Auxiliary topics: Estimation of prediction
HLM software has been one of the leading statistical packages for hierarchical
Introductory Guide to HLM With HLM 7 Software 3 G. David Garson HLM software has been one of the leading statistical packages for hierarchical linear modeling due to the pioneering work of Stephen Raudenbush
" Y. Notation and Equations for Regression Lecture 11/4. Notation:
Notation: Notation and Equations for Regression Lecture 11/4 m: The number of predictor variables in a regression Xi: One of multiple predictor variables. The subscript i represents any number from 1 through
THE INFLUENCE OF MARKETING INTELLIGENCE ON PERFORMANCES OF ROMANIAN RETAILERS. Adrian MICU 1 Angela-Eliza MICU 2 Nicoleta CRISTACHE 3 Edit LUKACS 4
THE INFLUENCE OF MARKETING INTELLIGENCE ON PERFORMANCES OF ROMANIAN RETAILERS Adrian MICU 1 Angela-Eliza MICU 2 Nicoleta CRISTACHE 3 Edit LUKACS 4 ABSTRACT The paper was dedicated to the assessment of
Chapter 7: Simple linear regression Learning Objectives
Chapter 7: Simple linear regression Learning Objectives Reading: Section 7.1 of OpenIntro Statistics Video: Correlation vs. causation, YouTube (2:19) Video: Intro to Linear Regression, YouTube (5:18) -
Module 4 - Multiple Logistic Regression
Module 4 - Multiple Logistic Regression Objectives Understand the principles and theory underlying logistic regression Understand proportions, probabilities, odds, odds ratios, logits and exponents Be
LOGIT AND PROBIT ANALYSIS
LOGIT AND PROBIT ANALYSIS A.K. Vasisht I.A.S.R.I., Library Avenue, New Delhi 110 012 [email protected] In dummy regression variable models, it is assumed implicitly that the dependent variable Y
EARLY VS. LATE ENROLLERS: DOES ENROLLMENT PROCRASTINATION AFFECT ACADEMIC SUCCESS? 2007-08
EARLY VS. LATE ENROLLERS: DOES ENROLLMENT PROCRASTINATION AFFECT ACADEMIC SUCCESS? 2007-08 PURPOSE Matthew Wetstein, Alyssa Nguyen & Brianna Hays The purpose of the present study was to identify specific
The Dummy s Guide to Data Analysis Using SPSS
The Dummy s Guide to Data Analysis Using SPSS Mathematics 57 Scripps College Amy Gamble April, 2001 Amy Gamble 4/30/01 All Rights Rerserved TABLE OF CONTENTS PAGE Helpful Hints for All Tests...1 Tests
