Showing posts with label linear models. Show all posts
Showing posts with label linear models. Show all posts

Saturday, July 29, 2023

On LLMs and LPMs: Does the LL in LLM Stand for Linear Literalism?

 I've blogged in the past about what I call linear literalism and fundamentalist econometrics. And I've blogged a bit about linear probability models (LPMs). Recently I have had some concerns about people outsourcing their thinking to LLMs and the use of these tools like Dunning-Kruger-as-a-Service (DKaaS) where the critical thinking and actual learning starts and stops with prompt engineering and a response. Out of curiosity I asked ChatGPT about the appropriateness of using linear probability models. Although the overall response was thoughtful about thinking more carefully about causality, it still gave the canned 'thou shalt not'  theoretically correct fundamentalist response. My prompt could have been more sophisticated, but I tried to prompt from a user's prospective, someone who may not be as familiar with applied statistics work, or who may have even read my blog and wanted to question something about the use of LPMs and may not be thinking about the tradeoffs or who may be unfamiliar with the social norms and practices related to their use.  As has been noted before on this blog, in applied work, there is no consensus among practitioners that nonlinear models (like logistic regression) are 'better' than LMPs when estimating treatment effects. If anything this illustrates at best, a response from an LLM about applied econometric analysis could be just as good as having another expert in the room, but an experienced practitioner understands that experts often disagree, and that disagreement comes with a lot of nuance, and is often as much the result of social norms and practices as theory. Perhaps someone could take the fundamentalist response from this prompt and do their analysis and solve their problem and there is no harm at the end of the day. But there is danger in fundamentalism, if this leads them to ignore great work and potential learnings derived from LPMs, or prevents them from getting more actionable and interpretable results vs. stumbling through the mental gymnastics sometimes required to interpret things like odds ratios vs. marginal effects! 

Note: As I refined my prompt I got slightly better answers, but not one that really captures the nuance I would like to have. Again....it's like arguing with an expert that knows the theory but may not be as familiar with all the nuances of different disciplines and the practical implications of the tradeoffs involved. I have to say - to give a response close to what I would expect from another practitioner that maybe outside the field of economics, is impressive! 

While it did a great job articulating the pros and cons of LPMs vs logistic regression or other models, I think these responses would leave the casual reader with the impression that they should be wary of relying on LPMs to estimate treatment effects in most situations. And as noted above that could be an issue depending on the situation.  Because I disagree with the tone of the response doesn't mean chatGPT is a bad tool, there are great textbooks and econometricians that I think would certainly endorse these responses! Its the missing nuance around these differences in opinion that concern me and while its not super critical in this situation, I'd be concerned about others. 

And I have to disagree that some of the 'cons' listed about LPMs are really as big of an issue that a reader might think if they take this response very seriously. Most of the listed 'cons' can often be addressed easily and in my experience despite all the alleged 'cons' LPMs generally lead to the same decision you would make when properly interpreting and understanding results from a logistic regression, while providing all the very practical benefits you get from an LPM.  The nuance is, the 'cons' in theory don't always have implications that matter in any practical sense. As Angrist and Pischke state: "there is a lot to be said for sticking to a linear regression function as compared to a fairly arbitrary choice of a non-linear one!"

Prompt: Is it OK to use a linear probability model to measure causal effects?

Answer:

Using a linear probability model to measure causal effects can be problematic and is generally not recommended. A linear probability model (LPM) is a linear regression model where the dependent variable is a binary outcome, typically coded as 0 or 1. In this model, the predicted values can be outside the [0, 1] range, which can lead to biased estimates and incorrect interpretations.

There are several issues with using a linear probability model for causal effects:

1. Heteroscedasticity: The variance of the error term in the linear probability model is not constant across all values of the independent variables, violating one of the assumptions of linear regression. This can lead to inefficient estimates and incorrect standard errors.

2. Predicted probabilities outside [0, 1]: Since linear regression does not constrain the predicted values to be between 0 and 1, it can produce probabilities that fall outside this range, which is not meaningful for binary outcomes.

3. Non-linearity: The relationship between the independent variables and the probability of the binary outcome is often non-linear. Using a linear model may not capture this relationship accurately.

To address these issues, researchers typically use alternative methods for estimating causal effects with binary outcomes. The most common approaches are logistic regression and probit regression. These models are specifically designed for binary outcomes and provide estimates that are valid probabilities between 0 and 1.

Additionally, for estimating causal effects, researchers often use methods such as randomized controlled trials (RCTs) or quasi-experimental designs (e.g., difference-in-differences, matching methods, instrumental variables) to establish causality more convincingly.

In summary, while linear probability models might be used for exploratory analysis, they are generally not suitable for measuring causal effects with binary outcomes. Logistic regression or probit regression should be preferred when dealing with binary dependent variables, and careful consideration of study design and potential confounders is essential for establishing causality.

Additional Reading:

Love It or Logit, or: Man, People *Really* Care About Binary Dependent Variables. Marc Bellemare. http://marcfbellemare.com/wordpress/9024

Whether to probit or to probe it: in defense of the Linear Probability Model. World Bank Blogs. Jed Friedman. https://blogs.worldbank.org/impactevaluations/whether-to-probit-or-to-probe-it-in-defense-of-the-linear-probability-model 

Sunday, February 17, 2019

Was It Meant to Be? OR Sometimes Playing Match Maker Can Be a Bad Idea: Matching with Difference-in-Differences

Previously I discussed the unique aspects of modeling claims and addressing those with generalized linear models. I followed that with a discussion of the challenges of using difference-in-differences in the context of GLM models and some ways to deal with this. In this post I want to dig into into what some folks are debating in terms of issues related to combining matching with DID. Laura Hatfield covers it well on twitter:

Link: https://twitter.com/laura_tastic/status/1022890688525029376

Also, they picked up on this it at the incidental economist and gave a good summary of the key papers here.

You can find citations for the relevant papers below. I won't plagerize what both Laura and the folks at the Incidental Economist have already explained very well. But, at a risk of oversimplifying the big picture I'll try to summarize a bit. Matching in a few special cases can improve the precision of the estimate in a DID framework,  and occasionally reduces bias. Remember, that  matching on pre-period observables is not necessary for the validity of difference in difference models. There are  cases when the treatment group is in fact determined by pre-period outcome levels. In these cases matching is necessary. At other times, if not careful, matching in DID introduces risks for regression to the mean…what Laura Hatfield describes as a ‘bounce back’ effect in the post period that can generate or inflate treatment effects when they do not really exist.

Both the previous discussion on DID in a GLM context and combining matching with DID indicate the risks involved in just plug and play causal inference and the challenges of bridging the gap between theory and application.


References:

Daw, J. R. and Hatfield, L. A. (2018), Matching and Regression to the Mean in Difference‐in‐Differences Analysis. Health Serv Res, 53: 4138-4156. doi:10.1111/1475-6773.12993

Daw, J. R. and Hatfield, L. A. (2018), Matching in Difference‐in‐Differences: between a Rock and a Hard Place. Health Serv Res, 53: 4111-4117. doi:10.1111/1475-6773.13017



Thursday, January 24, 2019

Modeling Claims with Linear vs. Non-Linear Difference-in-Difference Models

Previously I have discussed the issues with modeling claims costs. Typically medical claims exhibit non-negative highly skewed values with high zero mass and heterskedasticity. The most commonly suggested approach to addressing these distributional concerns in the literature call for the use of non-linear GLM models.  However, as previously discussed (see here and here) there are challenges with using difference-in-difference models in the context of GLM models. So once again, the gap between theory and application presents challenges, tradeoffs, and compromises that need to be made by the applied econometrician.

In the past I have written about the accepted (although controversial in some circles) practice of leveraging linear probability models to estimate marginal effects in applied work when outcomes are dichotomous. But what about doing this in the context of claims analysis? In my original post regarding the challenges of using difference-in-differences with claims I speculated:

"So as Angrist and Pischke might ask, what is an applied guy to do? One approach even in the context of skewed distributions with high mass points (as is common in the healthcare econometrics space) is to specify a linear model. For count outcomes (utilization like ER visits or hospital admissions are often dichotomized and modeled by logit or probit models) you can just use a linear probability model. For skewed distributions with heavy mass points, dichotomization with a LPM may also be an attractive alternative."

 I have found that this advice is pretty consistent with the social norms and practices in the field.

In their analysis of the ACA Cantor, et al (2012) leverage linear probability models for difference-in-differences for healthcare utilization stating:

"Linear probability models are fit to produce coefficients that are direct estimates of the relevant policy impacts and are easily interpreted as percentage point changes in coverage outcomes. This approach has been applied in earlier evaluations of insurance market reforms (Buchmueller and DiNardo 2002; Monheit and Steinberg Schone 2004;  Levine, McKnight, and Heep 2011;  Monheit et al. 2011). It also avoids complications associated with estimation and interpretation of multiple interaction terms and their standard errors in logit or probit models (Ai and Norton 2003)."

Jhamb et al (2015) use LPMs for dichotomous outcomes as well as OLS models for counts in a DID framework.

Interestingly, Deb and Norton (2018) discuss an approach to address the challenges of DID in a GLM framework head on:

"Puhani argued, using the potential outcomes framework, that the treatment effect on the treated in the difference-in-difference regression equals the expected value of the dependent variable for the treatment group in the post period with treatment compared with the hypothetical expected value of the dependent variable for the treatment group in the post period if they had not received treatment. In nonlinear models, the treatment effect on the treated equals the difference in two predicted values. It always has the same sign as the coefficient on the interaction term. Because we estimate many nonlinear models using a difference-in-differences study design, we report the treatment effect on the treated in all tables of results."

In presenting their results they compare their GLM based approach to results from linear models of healthcare expenditures. While they argue the differences are substantial in supporting their approach, I did not find the OLS estimate (-$323.4) to be practically different from the second part (conditional on positive) of the two part GLM model (-$321.4), although the combined results from the two part model had large practical differences from OLS. It does not appear they compared a two-part GLM to a two-part linear model (which could be problematic if the first part OLS model gave probabilities greater than 1 or less than zero). In their paper they cited a number of authors using linear difference-in-differences to model claims you will find below.

See the references below for a number of examples (including those cited above).

Related: Linear Literalism and Fundamentalist Econometrics

References:

Cantor JC, Monheit AC, DeLia D, Lloyd K. Early impact of the Affordable Care Act on health insurance coverage of young adults. Health Serv Res. 2012;47(5):1773-90.

Modeling Health Care Expenditures and Use
Partha Deb and Edward C. Norton
Annual Review of Public Health 2018 39:1, 489-505

Buchmueller T, DiNardo J. “Did Community Rating Induce an Adverse Selection Death Spiral? Evidence from New York, Pennsylvania and Connecticut” American Economic Review. 2002;92(1):280–94.

Monheit AC, Cantor JC, DeLia D, Belloff D. “How Have State Policies to Expand Dependent Coverage Affected the Health Insurance Status of Young Adults?” Health Services Research. 2011;46(1 Pt 2):251–67

Amuedo-Dorantes C, Yaya ME. 2016. The impact of the ACA’s extension of coverage to dependents on young adults’ access to care and prescription drugs. South. Econ. J. 83:25–44

Barbaresco S, Courtemanche CJ, Qi Y. 2015. Impacts of the Affordable Care Act dependent coverage provision on health-related outcomes of young adults. J. Health Econ. 40:54–68

Jhamb J, Dave D, Colman G. 2015. The Patient Protection and Affordable Care Act and the utilization of health care services among young adults. Int. J. Health Econ. Dev. 1:8–25

Sommers BD, Buchmueller T, Decker SL, Carey C, Kronick R. 2013. The Affordable Care Act has led
to significant gains in health insurance and access to care for young adults. Health Aff. 32:165–74




Friday, June 12, 2015

Linear Literalism & Fundamentalist Econometrics

Your tweet stream may have included the recently trending article by Angrist and Pischke entitled: "Why econometrics teaching needs an overhaul". Here is an excerpt:

“In addition to its more up-to-date contents, our book renews the econometrics canon by abandoning the childish literalism of the legacy approach to econometric instruction. In this spirit, we eschew the notion that regression is tied to a literal linear model. Regression describes differences in averages whether or not these averages fit a linear equation. This is a universal property – one that is reliably true – and we don’t intimidate readers with descriptions of the punishments to be meted out for the failure of classical assumptions. Our regression discussion begins by challenging readers to ask themselves, first, what the target causal effect is, and, second, by asking, ‘what is the regression you want’? In other words, what would you like to hold fixed when trying to regress-out an average causal effect?”

They are referencing their text Mastering Metrics, which I highly recommend.  In their other text, Mostly Harmles Econometrics, they also state:

"In fact, the validity of linear regression as an empirical tool does not turn on linearity either...The statement that regression approximates the CEF lines up with our view of empirical work as an effort to describe the essential features of statistical relationships, without necessarily trying to pin them down exactly."  - Mostly Harmless Econometrics, p. 26 & 29

On their MHE blog a reader asks about their pedagogy which focuses on 'best linear projection' vs the traditional BLUE criteria to which they respond:

"our undergrad econometrics training (like most people’s) focused on the sampling distribution of OLS. Hence you were tortured with the Gauss-Markov Thm, which says that OLS is a Best Linear Unbiased Estimator (BLUE). MHE and MM are largely unconcerned with such things. Rather, we try to give our students a clear understanding of what regression means. To that end, we introduce regression as the best linear approximation to whatever conditional expectation fn. (CEF) motivates your empirical work – this is the BLP property you mention, which is a regression feature unrelated to samples. (MM also emphasises our interpretation of regression as a form of “automated matching”)." read more...

The notion of using regression as a means of making like comparisons has also been echoed by Andrew Gelman:

http://andrewgelman.com/2013/01/understanding-regression-models-and-regression-coefficients/

"It's all about comparisons, nothing about how a variable "responds to change." Why? Because, in its most basic form, regression tells you nothing at all about change. It's a structured way of computing average comparisons in data."

Linear literalism or fundamentalist undergraduate econometrics (being tortured with BLUE as A&P might put it) can have long term consequences for students. I think this has caused harms that I encounter from time to time even among more seasoned practitioners and even graduate degree holders. This isn't too different from what Leo Brieman described as a 'statistical straight jacket' that can arbitrarily limit fruitful empirical work. Overly clinical concerns with linearity, heteroskedasticity, and multicollinearity might crowd out more important concerns around causality and prediction.

Heteroskedasticity

We could simplify this as a notion of non-constant variance. As Angrist and Pischke note:

"Our view of regression as an approximation to the CEF makes heteroskedasticity seem natural. If the CEF is nonlinear....the residuals will be larger, on average, at values of X where the fit is poorer...as an empirical matter, heteroskedasticity may matter little" -Ch 3, p.46-47 MHE

 Of course the concern is correct standard errors and valid inference, which can be addressed via heteroskedasticity corrected standard errors. But I am afraid some students, after taking a traditional econometrics course, may terminate all thought processes after a cookbook test hints of its existence.

Multicollinearity

When covariates are highly correlated, it may be difficult to parse out the independent information about each variable and  lead to inflated standard errors. Again this is a phenomena related to inference, not prediction. Even in professional and academic settings when I have presented or attended other presentations related to forecasting or predictive analtyics you will get the occasional criticism or self aggrandizing questions about multicollinearity being a concern.

"Multicollinearity has a very different impact if your goal is prediction from when your goal is estimation. When predicting, multicollinearity is not really a problem provided the values of your predictors lie within the hyper-​​region of the predictors used when estimating the model."-  Statist. Sci.  Volume 25, Number 3 (2010), 289-310.

Undue criticisms and literalism related to multicollinearity often results from a failure to recognize the differences between goals related to explaining vs. predicting.

Paul Allison offers some additional advice on when not to worry about multicolinearity. I have highlighted a couple points of interest here.

Dave Giles provides great context around Arthur S. Goldberger's parody of multicollinearity referencing 'micronumerosity.'

Kennedy has a similar discussion:

"The worth of an econometrics textbook tends to be inversely related to the technical material devoted to  multicollinearity" - Williams, R. Economic Record 68, 80-1. (1992).  via Kennedy, A Guide to Econometrics (6th edition).

This kind of linear fundamentalist paradigm can lead students and practitioners to adopt more complicated methods than necessary or abandon promising empirical work altogether, become overly critical and dismissive of other important work done by others, or completely miss more important questions related to selection bias and identification and unobserved heterogeneity and endogeneity.

Some of this also is a the result of the huge gap between theoretical and applied econometrics.

See also:
Marc Bellemare discusses a similar vein of literalism that is averse to linear probability models here.

Linear Probability Models

Regression as an empirical tool
Quasi-Experimental Design Roundup

Saturday, June 28, 2014

Linear Probability Models for Skewed Distributions with High Mass Points

There are a lot of methods discussed in the literature related to modeling skewed distributions with high mass points including log transformations, two part models,  GLM etc. In some previous posts I have discussed linear probability models in the context of causal inference.  I've also discussed the use of quantile regression as a strategy to model highly skewed continuous and count data. Mullahy (2009) alludes to the use of quantile regression as well:

"Such concerns should translate into empirical strategies that target the high-end parameters of particular interest, e.g. models for Prob(y ≥ k | x) or quantile regression models"

The focus on high end parameters  using linear probability models is mentioned in Angrist and Pischke (2009) :

"COP [conditional-on-positive] effects are sometimes motivated by a researcher's sense that when the outcome distribution has a mass point-that is, when it piles up on a particular value, such as zero-or has a heavily skewed distribution, or both, then an analysis of effects on averages misses something. Analysis of effects on averages indeed miss some things, such as changes in the probability of specific values or a shift in quantiles away from the median. But why not look at these distribution effects directly? Distribution outcomes include the likelihood that annual medical expenditures exceed zero, 100 dollars, 200 dollars, and so on. In other words, put 1[Yi > c] for different choices of c on the left hand side of the regression of interest...the idea of looking directly at distribution effects with linear probability models is illustrated by Angrist (2001),...Alternatively, if quantiles provide a focal point, we can use quantile regressions to model them."

References:

Mostly Harmless Econometrics. Angrist and Pischke. 2009

Angrist, J.D. Estimation of Limited Dependent Variable Models With Dummy Endogenous Regressors: Simple Strategies for Empirical Practice. Journal of Business & Economic Statistics January 2001, Vol. 19, No. 1.

ECONOMETRIC MODELING OF HEALTH CARE COSTS AND EXPENDITURES: A SURVEY OF ANALTICAL ISSUES AND RELATED POLICY CONSIDERATIONS
John Mullahy Univ. of Wisconsin-Madison
January 2009


Saturday, January 4, 2014

The Oregon Medicaid Experiment and Linear Probability Models

I just recently discussed the methodology used in some recent papers analyzing the Oregon Medicaid expansion (see: http://econometricsense.blogspot.com/2014/01/the-oregon-medicaid-experiment-applied.html ). This was one of the papers:

"The Oregon Experiment--Effects of Medicaid on Clinical Outcomes," by Katherine Baicker, et al. New England Journal of Medicine, 2013; 368:1713-1722. http://www.nejm.org/doi/full/10.1056/NEJMsa1212321 

If you read the supplementary appendix you will find the following:

In all of our ITT estimates and in our subsequent instrumental variable estimates (see below), we fit linear models even though a number of our outcomes are binary. Because we are interested in the difference in conditional means for the treatments and controls, linear probability models would pose no concerns in the absence of covariates or in fully saturated models (Angrist 2001, Angrist and Pischke 2009). Our models are not fully saturated, however, so it is possible that results could be affected by this functional form choice, especially for outcomes with very low or very high mean probability. We therefore explore the sensitivity of our results to an alternate specification using logistic regression and calculating average marginal effects for all binary outcomes, and are reassured that the results look very similar (see Table S15a-d below). 

You will find a similar methodology in the more recent article in Science previously discussed. This weaves well with some of my past posts:

Linear Regression and Analysis of Variance with Binary Dependent Variables

Regression as an Empirical Tool (matching and linear probability models)


Tuesday, September 10, 2013

Regression as an Empirical Tool


Linear regression is a powerful empirical tool for the social sciences.  Its robustness is often underrated, while at other times its use and interpretation is mischaracterized.  Andrew Gelman and Agrist and Pischke are two great sources for learning about regression in an applied context.

I particularly like Gelman's comment here:
"It's all about comparisons, nothing about how a variable "responds to change." Why? Because, in its most basic form, regression tells you nothing at all about change. It's a structured way of computing average comparisons in data."

This 'computing average comparisons of data' interpretation is why regression works as sort of a matching estimator as Angrist and Pischke  argue. 

“Our view is that regression can be motivated as a particular sort of weighted matching estimator, and therefore the differences between regression and matching estimates are unlikely to be of major
empirical importance” (Chapter 3 p. 70)

 In further discussion Gelman goes on to say, I think in a very appropriate interpretation:

“They're saying (Angrist and Pischke ) that regression, like matching, is a way of comparing-like-with-like in estimating a comparison. This point seems commonplace from a statistical standpoint but may be news to some economists who might think that regression relies on the linear model being true.”

This brings up a very important point, one also in-line with Angrist and Pischke regarding the use of regression as an empirical tool in the social sciences:

"In fact, the validity of linear regression as an empirical tool does not turn on linearity either...The statement that regression approximates the CEF lines up with our view of empirical work as an effort to describe the essential features of statistical relationships, without necessarily trying to pin them down exactly."  - Mostly Harmless Econometrics, p. 26 & 29

Regression users can seem at odds with each other at times. On one extreme they can get caught up in making very clinical assumptions  about linearity (see somewhat related discussions related to linear probability models  here and here) then on the other hand, take robustness to extremes by failing to consider at times questions of unobserved heterogeneity, endogeneity, selection bias, and identification.

Cellini(2008) discusses this issue in an analysis of the impact of financial aid on college enrollment:

“While simple ordinary least squares estimates of the impact of aid on college-going can reveal a correlation between financial aid policies and enrollment, these estimates are likely to suffer from omitted variable bias due to self-selection, potentially overestimating or underestimating the causal impact of these policies on enrollment.”

“The discussion above has outlined several methods for addressing the problem of omitted variable bias in financial aid research… proxy variable, fixed effects, and difference in- differences approaches are becoming quite common. Indeed, these approaches have replaced basic multivariate regression as the new standard for education research in the economics literature”

 This is where quasi-experimental methods come in to play.


References

 Stephanie Riegg Cellini. Causal Inference and Omitted Variable Bias in Financial Aid Research: Assessing Solutions The Review of Higher Education Spring 2008, Volume 31, No. 3, pp. 329–354

Saturday, January 5, 2013

Interpreting Regression Coefficients

Nice discussion on regression here:(one of my favorite blogs)

http://andrewgelman.com/2013/01/understanding-regression-models-and-regression-coefficients/

I particularly like Gelman's comment:

"It's all about comparisons, nothing about how a variable "responds to change." Why? Because, in its most basic form, regression tells you nothing at all about change. It's a structured way of computing average comparisons in data."

This 'computing average comparisons of data' interpretation is why regression works as sort of a matching estimator as Angrist and Pischke  argue, and we've all including Dr. Gelman discussed before here: 

"Think of the world of difference between using a regression model for prediction and using one for estimating a parameter with a causal interpretation, for example, the effect of class size on school children's test scores. With prediction, we don't need our relationship to be causal, but we do need to be concerned with the relation between our training and our test set. If we have reason to think that our future test set may differ from our past training set in unknown ways, nothing, including cross-validation, will save us. When estimating the causal parameter, we do need to ask whether the children were randomly assigned to classes of different sizes, and if not, we need to find a way to deal with possible selection bias. If we have not measured suitable covariates on our children, we may not be able to adjust for any bias."

If we are talking about the following specification: 

E[Y­­­i|ci=1] - E[Y­­­i|ci=0] =E[Y1i-Y0i|ci=1]  +{ E[Y0i|ci=1] - E[Y0i|ci=0]}

Observed effect = treatment effect on the treated + {selection bias}

 I think that framework is the most useful for characterizing and understanding selection bias. I could be missing something but I don't see how the block quote from Terry above is really inconsistent with the potential outcome framework of causal inference, unless maybe you completely refuse to think of regression as a matching estimator. I think he does a good job pointing out what most people don't see as different applications of regression. As Dr. Gelman says, inference may be a special case of prediction, but when I here this distinction I can't help but think of this comment from Greene: 

 "It remains an interesting question for research whether fitting well or obtaining good parameter estimates is a preferable estimation criterion. Evidently, they need not be the same thing."
  p. 686 Greene,  Econometric Analysis 5th ed

Friday, July 27, 2012

Empirical Work in The Social Sciences- from Mostly Harmless Econometrics



"In fact, the validity of linear regression as an empirical tool does not turn on linearity either...The statement that regression approximates the CEF lines up with our view of empirical work as an effort to describe the essential features of statistical relationships, without necessarily trying to pin them down exactly."  - Mostly Harmless Econometrics, p. 26 & 29

I really like Mostly Harmless Econometrics. I started reading it some time ago. I have had several formal courses in econometrics, mathematical statistics, and experimental design, and have spent a lot of time in the pages of Golberger's A Course in Econometrics., Greene's Econometric Analysis, and Kennedy's A Guide to Econometrics, as well as Hastie, Tibshirani, and Friedman's Elements of Statistical Learning: Data Mining, Inference, and Prediction, but Angrist and Pishke's book really speaks to me in the every day empirical work that I find myself caught up in. The other textbooks are great, maybe essential for a student, but Mostly Harmless Econometrics is a must have for the practitioner. MHE is not a substitute for a solid econometrics background, in fact, it would not have made sense to me without it, but it wouldn't have as much meaning either without some prior experience. I find myself re-reading sections because on the job data challenges continue to make this book more relevant every day. The other textbooks get you started with the theory (again very important). They are great for 'highway' use, when your are coasting on the smooth surfaces of textbook ideals. MHE gives you that push you need when you get bogged down on the very muddy roads of real life empirical work.  Its the off-road backwoods survival manual for practitioners.



Wednesday, August 31, 2011

Linear Regression and Analysis of Variance with a Binary Dependent Variable

(see also my posts related to Logistic Regression

If for instance Y is dichotomous or binary, Y = { 1 if ‘yes’  0 if ‘no’}, would  you consider it valid to do an analysis of variance or fit a linear regression model?

We might not think so based on traditional assumptions, because besides assuming that Y is continuous….

1)      ANOVA / linear regression both work under the assumption of a uniform (homoskedastic) error term ‘e’

2)      For a dichotomous y, the expected value E(y|X)  =  ‘probability’ and may be continuous, but the  error terms follow a binomial distribution with mean ‘p’ and a variance that is a function of the mean, which is inherently heteroskedastic , var(error) ~ n*p*(1-p)   

3)      Therefore the assumption of uniform variance is violated, so the F test and other tests based on standard errors based on this assumption are questionable

Some people will argue that violations of this assumption may only matter by degree, and under certain conditions ANOVA and linear regression using least squares is OK with a binary dependent variable (Lunney 1971, D'Agostino 1971,Astin & Dey 1993, Angrist & Pischke 2008).

I recall from mathematical statistics and econometrics, the lectures (and test questions) related to properties of estimators including efficiency, consistency, unbiasedness etc. We also spoke some about robustness, but in the theoretical work, its not so easy to 'prove' robustness as it is to show that a certain estimator is unbiased or consistent. In this way, I think a lot of times, robustness to assumptions isn't given a lot of credence by students or practitioners. ( However, Angrist and Pischke in their book 'Mostly Harmless Economics' spend a lot of time discussing ideas related to the robustness of the least squares estimator).  Robustness to assumptions related to the distribution of the error terms (under OLS / ANOVA are discussed in the following:

LITERATURE RELATED TO REGRESSION AND ANOVA  WITH A DICHOTOMOUS DEPENDENT VARIABLE

ANALYSIS OF VARIANCE CONTEXT--------------------------------------------

A Second Look at Analysis of Variance on Dichotomous Data
Author(s): Ralph B. D'Agostino
Source: Journal of Educational Measurement, Vol. 8, No. 4 (Winter, 1971), pp. 327-333

“probably a safe rule of thumb for deciding when the I x J ANOVA techniques may be used on dichotomous data with equal sample sizes in each cell is; the sample proportions for the cells should lie between .25 and .75 and there should be at least 20 degrees offreedom for error. This rule combines Lunney's results along with standard rules (Snedecor & Cochran, 1967, p. 494). The reason for the .25 and .75 lies in the fact that for this range there is little change between the within cell variances, p(l - p), and so there is a sufficient homogeneity of variances. Given that this rule is satisfied a standard procedure for analysis is the ANOVA on the original data. In this range (.25 to .75) it is doubtful if any alternative valid procedure, such as an analysis of transformed data, would lead to different conclusions”

Using Analysis of Variance with a Dichotomous Dependent Variable: An Empirical Study
Author(s): Gerald H. Lunney
Source: Journal of Educational Measurement, Vol. 7, No. 4 (Winter, 1970), pp. 263-269

“The findings show the analysis of variance to be an appropriate statistical technique for analyzing dichotomous data in fixed effects models where cell frequencies are equal under the following conditions: (a) the proportion of responses in the smaller response category is equal to or greater than .2 and there are at least 20 degrees of freedom for error, or (b) the proportion of responses in the smaller response category is less than .2 and there are at least 40 degrees of freedom for error.”

REGRESSION CONTEXT---------------------------------------------------

Dey ,Eric L. and Alexander W. Astin. Statistical Alternatives For Studying College Student Retention: A Comparative Analysis of Logit, Probit, and Linear Regression. Research in Higher Education, Vol. 34, No. 5. 1993. link http://www.jstor.org/stable/40196112  

"These results indicate that despite the theoretical advantages offered by logistic regression and probit analysis, there is little practical difference between either of these two techniques and more traditional linear regression. While this may not always be the case, these and other analyses show that for variables that are moderately distributed (say, within a .75/. 25 split; for example, see Cleary and Angel, 1984) there is little practical difference in obtained results upon which to make a decision about one technique or another, especially in large samples."

Angrist, Joshua D. & Jörn-Steffen Pischke. Mostly Harmless Econometrics: An Empiricist's Companion. Princeton University Press. NJ. 2008.

"While a nonlinear model may fit the CEF (population conditional expectation function) for LDVs (limited dependent variables) more closely than a linear model, when it comes to marginal effects, this probably matters little"

BOTH CONTEXTS---------------------------------------------

The Analysis of Relationships Involving Dichotomous Dependent Variables
Author(s): Paul D. Cleary and Ronald AngelSource: Journal of Health and Social Behavior, Vol. 25, No. 3 (Sep., 1984), pp. 334-348

"If the researcher wishes to estimate the probability of an outcome as a linear function and, A. If the sample size is moderately large and the dependent variable is not too skewed (.25 < p < .75), then OLS regression or ordinary ANOVA is adequate."

Wednesday, April 6, 2011

Topics Related to Linear Models


While attending a session at this year’s SAS Global Forum, I was reminded of several topics that I have not covered in previous posts, but should have.

Fixed and Random Effects Models and Panel Data (see also Mixed, Fixed and Random Effects)

Notes: Panel data, or repeated measures data, is characterized by multiple observations on individuals over time.  A consequence of this is that measurements on the same individual will likely be more correlated than measurements for (or between) different individuals.  Measurements taken closer together over time will also likely be more correlated than measurements taken at greater intervals. As a result, assumptions from OLS regression regarding independence and homogenous variance will likely be violated.  As noted during the session I attended, in SAS, PROC MIXED appropriately handles within subject and time dependent correlations and the covariance structure associated with repeated measures.

Let’s structure the model as follows: Yit = Xit + Ait +Uit

Where Ait is a unobserved individual effect. There are two assumptions that we can work under when estimating this model,

1)Random Effects (RE): Assumes that Ait is independent of X . X may be considered a  predetermined, or other fixed effect. Ait is a random effect.

Estimation: Feasible Generalized Least Squares-  Bfgls = (X’W-1X)-1X’W-1y
i.e V(e) != σ2  I  which would be the case under OLS

2)Fixed Effects (FE): Assumes Ait is not independent of X.

Estimation: differencing or subtracting the respective means from X and Y  and running OLS on the adjusted data produces the estimate of B.

You could also subtract the lagged version of each variable X and Y from itself respectively to move the time-correlated component, or you could remove Ait through dummy variable regression.

Fixed and Random Effects in Analysis of Variance (in general)

In my previous posts related to AOV and Mixed Models,  I did not discuss these in the context of  AOV.

Typically fixed effects are repeatable factors that are set by the experimenter,  often the ‘treatments.’  Random effects are effects that are selected randomly from a population, often ‘blocks.’

Type I and III Sums of Squares

Type I Sums of Squares: often referred to as sequential sums of squares, this is the SS for each effect adjusted for all effects that appear earlier in the model.

Example:
                                   
Y1 = X1B1                       

Type I SS = SSR(Y1) = SS(x1)

Y2 = X1B1 + X2B2

Type I SS = SSR(Y2) – SSR(Y1) = SS(X2)

Type III sums of squares are adjusted for every X in the full model.

Example:

Y1 = B1X1 + B2X2
Y2 = B1X1

Type III SS = SSR(Y1) – SSR(Y2) = SS(X1)

Y3 = B2X2

Type III SS = SSR(Y1) – SSR(Y3) = SS(X2)

Means vs. LS Means

Means = overall mean for a treatment or factor level

LS Means = within group means adjusted for other effects in the model

Type I SS are used to test differences between means while type III SS are used to test differences between LS Means.