"As Philip Dawid once said "a causal model is just an ambitious associational model". A carefully-considered regression model, with an appropriate set of potential confounders (possibly identified using a causal diagram – see below) measured and included as covariates, is the most appropriate causal model in many simple settings."
http://csm.lshtm.ac.uk/themes/causal-inference/
To paraphrase Angrist and Pischke:
To the extent that the population CEF that it is estimating is causal, so is linear regression. (And that includes LPMs)
An attempt to make sense of econometrics, biostatistics, machine learning, experimental design, bioinformatics, ....
Sunday, March 30, 2014
Tuesday, March 25, 2014
Institutional Research Presentations at SAS Global Forum
I'm not attending #SASGF14, but some of my colleagues in higher ed are. Here is what they are doing. If you are not attending global forum, or can't make their talks, I encourage you to check out their papers via the online proceedings once they are posted.
Tuesday March 25
Paper 1448 - From Providing Support to Driving Decisions: Improving the Value of Institutional Research For almost two decades, Western Kentucky University's Office of Institutional Research (WKU-IR) has used SAS® to help shape the future of the institution by providing faculty and administrators with information they can use to make a difference in the lives of their students. This presentation provides specific examples of how WKU-IR has shaped the policies and practices of our institution and discusses how WKU-IR moved from a support unit to a key strategic partner. In addition, the presentation covers the following topics: How the WKU Office of Institutional Research developed over time; Why WKU abandoned reactive reporting for a more accurate, convenient system using SAS® Enterprise Intelligence Suite for Education; How WKU shifted from investigating what happened to predicting outcomes using SAS® Enterprise Miner™ and SAS® Text Miner; How the office keeps the system relevant and utilized by key decision makers; What the office has accomplished and key plans for the future.
Paper 1638 - Institutional Research: Serving University Deans and Department Heads Administrators at Western Kentucky University rely on the Institutional Research department to perform detailed statistical analyses to deepen the understanding of issues associated with enrollment management, student and faculty performance, and overall program operations. This paper presents several instances of analyses performed for the university to help it identify and recruit suitable candidates, uncover root causes in grade and enrollment trends, evaluate faculty effectiveness, and assess the impact of student characteristics, programs, or student activities on retention and graduation rates. The paper briefly discusses the data infrastructure created and used by Institutional Research. For each analysis performed, it reviews the SAS® program and key components of the SAS code involved. The studies presented include the use of SAS® Enterprise Miner™ to create a retention model incorporating dozens of student background variables. It shows an examination of grade trends in the same courses taught by different faculty and subsequent student behavior and success, providing insights into the nuances and subtleties of evaluating faculty performance. Another analysis uncovers the possible influence of fraternities and sororities in freshmen algebra courses. Two investigations explore the impact of programs on student retention and graduation rates. Each example and its findings illustrate how Institutional Research can support the administration of university operations. The target audience is any SAS professional interested in learning more about Institutional Research in higher education and how SAS software is used by an Institutional Research department to serve its organization.
Monday March 24
Paper 1689 - Simple ODS Tips to Get RWI (Really Wonderful Information) SAS® continues to expand and improve its reporting capability. With new SAS® 9.4 enhancements in ODS (Output Delivery System), the opportunity to create stunning reports has expanded even further. If you are charged with creating relevant, informative, easy-to-read reports for clients or administrators, then the ODS Report Writing Interface, ODS LAYOUT enhancements, and the new ODSTEXT procedure are important tools to use. These tools allow you to create reports in a smart, eye-catching format that can be turned around quite quickly and programmed to provide optimum flexibility. How many times have you worked hours to tweak and fine-tune a report directly in Microsoft Excel, Microsoft Word, Microsoft Power Point or some other similar software only to be asked for a “quick update”, which would then take hours to recreate because you are manually transferring data? Do you ever dread receiving the compliment, “This is really wonderful information!!!!” because you know it will be followed by “Can you run this for EVERY region?” Well, dread no more, because when you harness the power of SAS® ODS, you can create first-rate, flexible, fabulous reports! Join me as I share with you two real-world examples of ODS capabilities using (1) a marketing piece I designed to help the president of our university spotlight county- and region-specific data as he recruited across the state and (2) our academic program review form, a multi-page report that outputs to Word so that program coordinators can add personalized commentary to support their program’s effectiveness.
Tuesday March 25
Paper 1448 - From Providing Support to Driving Decisions: Improving the Value of Institutional Research For almost two decades, Western Kentucky University's Office of Institutional Research (WKU-IR) has used SAS® to help shape the future of the institution by providing faculty and administrators with information they can use to make a difference in the lives of their students. This presentation provides specific examples of how WKU-IR has shaped the policies and practices of our institution and discusses how WKU-IR moved from a support unit to a key strategic partner. In addition, the presentation covers the following topics: How the WKU Office of Institutional Research developed over time; Why WKU abandoned reactive reporting for a more accurate, convenient system using SAS® Enterprise Intelligence Suite for Education; How WKU shifted from investigating what happened to predicting outcomes using SAS® Enterprise Miner™ and SAS® Text Miner; How the office keeps the system relevant and utilized by key decision makers; What the office has accomplished and key plans for the future.
Paper 1638 - Institutional Research: Serving University Deans and Department Heads Administrators at Western Kentucky University rely on the Institutional Research department to perform detailed statistical analyses to deepen the understanding of issues associated with enrollment management, student and faculty performance, and overall program operations. This paper presents several instances of analyses performed for the university to help it identify and recruit suitable candidates, uncover root causes in grade and enrollment trends, evaluate faculty effectiveness, and assess the impact of student characteristics, programs, or student activities on retention and graduation rates. The paper briefly discusses the data infrastructure created and used by Institutional Research. For each analysis performed, it reviews the SAS® program and key components of the SAS code involved. The studies presented include the use of SAS® Enterprise Miner™ to create a retention model incorporating dozens of student background variables. It shows an examination of grade trends in the same courses taught by different faculty and subsequent student behavior and success, providing insights into the nuances and subtleties of evaluating faculty performance. Another analysis uncovers the possible influence of fraternities and sororities in freshmen algebra courses. Two investigations explore the impact of programs on student retention and graduation rates. Each example and its findings illustrate how Institutional Research can support the administration of university operations. The target audience is any SAS professional interested in learning more about Institutional Research in higher education and how SAS software is used by an Institutional Research department to serve its organization.
Monday March 24
Paper 1689 - Simple ODS Tips to Get RWI (Really Wonderful Information) SAS® continues to expand and improve its reporting capability. With new SAS® 9.4 enhancements in ODS (Output Delivery System), the opportunity to create stunning reports has expanded even further. If you are charged with creating relevant, informative, easy-to-read reports for clients or administrators, then the ODS Report Writing Interface, ODS LAYOUT enhancements, and the new ODSTEXT procedure are important tools to use. These tools allow you to create reports in a smart, eye-catching format that can be turned around quite quickly and programmed to provide optimum flexibility. How many times have you worked hours to tweak and fine-tune a report directly in Microsoft Excel, Microsoft Word, Microsoft Power Point or some other similar software only to be asked for a “quick update”, which would then take hours to recreate because you are manually transferring data? Do you ever dread receiving the compliment, “This is really wonderful information!!!!” because you know it will be followed by “Can you run this for EVERY region?” Well, dread no more, because when you harness the power of SAS® ODS, you can create first-rate, flexible, fabulous reports! Join me as I share with you two real-world examples of ODS capabilities using (1) a marketing piece I designed to help the president of our university spotlight county- and region-specific data as he recruited across the state and (2) our academic program review form, a multi-page report that outputs to Word so that program coordinators can add personalized commentary to support their program’s effectiveness.
Saturday, March 22, 2014
Quantile Regression with Count Data
I stumbled upon this paper recently:
Reforming health care: Evidence from quantile regressions for counts
Rainer Winkelmann
Journal of Health Economics 25 (2006) 131–145
"Basically, the approach transforms the discrete data problem into a continuous data problem by adding a random uniform variable to each count. The quantile regression functions of the transformed variable can then be estimated using standard quantile regression software. To interpret the results, one can compare the freely estimated quantile functions to those implied by the respective Poisson or negative binomial estimates in order to detect excess sensitivity in specific parts of the distribution, such as the lower or upper tails."
See also:
Machado, J.A.F. and Santos Silva, J.M.C. (2005), Quantiles for Counts, Journal of the American Statistical Association, vol. 100, no. 472, pp. 1226-1237.
R:
http://www.inside-r.org/packages/cran/lqmm/docs/lqm.counts
STATA:
http://ideas.repec.org/c/boc/bocode/s456714.html
Reforming health care: Evidence from quantile regressions for counts
Rainer Winkelmann
Journal of Health Economics 25 (2006) 131–145
"Basically, the approach transforms the discrete data problem into a continuous data problem by adding a random uniform variable to each count. The quantile regression functions of the transformed variable can then be estimated using standard quantile regression software. To interpret the results, one can compare the freely estimated quantile functions to those implied by the respective Poisson or negative binomial estimates in order to detect excess sensitivity in specific parts of the distribution, such as the lower or upper tails."
See also:
Machado, J.A.F. and Santos Silva, J.M.C. (2005), Quantiles for Counts, Journal of the American Statistical Association, vol. 100, no. 472, pp. 1226-1237.
R:
http://www.inside-r.org/packages/cran/lqmm/docs/lqm.counts
STATA:
http://ideas.repec.org/c/boc/bocode/s456714.html
Friday, January 17, 2014
Propensity Score Matching Meets Survival Analysis
In one of my earlier posts regarding propensity score applications in higher ed research, a reader asked in the comment section about using propensity score methods in the context of survival analysis. Ironically, just a few days prior, I was having a similar discussion with another higher education researcher. Unfortunately, I have not been able to answer their questions adequately, but I think this is an interesting topic. I've recently located a few articles that deal with this. Unfortunately I have not had a chance to read through them but thought they may be of interest. In the least I now have them bookmarked for future reference. Hopefully they address some of these issues.
Effect of
radiation therapy on survival in surgically resected retroperitoneal sarcoma: a
propensity score-adjusted SEER analysis
Ann Oncol (2012)
A. H. Choi1,
J. S. Barnholtz-Sloan2 and
J. A. Kim3*
Propensity score
methods were used to perform survival analysis in patients who received
radiation matched with patients who underwent surgery alone...Propensity
scoring (309 matched pairs) and survival analysis using Kaplan–Meier methods
demonstrated no difference between propensity score-matched patients receiving
radiation therapy and those who did not (P = 0.35).
Propensity score applied to survival data
analysis through proportional hazards models: a Monte Carlo study.
Pharm Stat. 2012 Mar 12. doi: 10.1002/pst.537.
Gayat E, Resche-Rigon M, Mary JY, Porcher R.
A Monte Carlo
simulation study was used to compare the performance of several survival models
to estimate both marginal and conditional treatment effects. The impact of
accounting or not for pairing when analysing propensity-score-matched survival
data was assessed. In addition, the influence of unmeasured confounders was
investigated....Our study showed that propensity scores applied to survival
data can lead to unbiased estimation of both marginal and conditional treatment
effect, when marginal and adjusted Cox models are used. In all cases, it is
necessary to account for pairing when analysing propensity-score-matched data,
using a robust estimator of the variance.
The performance of
different propensity score methods for estimating marginal hazard ratios.
Stat Med. 2013 Jul 20;32(16):2837-49. doi: 10.1002/sim.5705. Epub 2012 Dec 12. Austin
PC.
...in biomedical
research, time-to-event outcomes occur frequently. There is a paucity of research
into the performance of different propensity score methods for estimating the
effect of treatment on time-to-event outcomes....We conducted an extensive
series of Monte Carlo simulations to examine the performance of propensity
score matching (1:1 greedy nearest-neighbor matching within propensity score
calipers), stratification on the propensity score, inverse probability of
treatment weighting (IPTW) using the propensity score, and covariate adjustment
using the propensity score to estimate marginal hazard ratios. We found that
both propensity score matching and IPTW using the propensity score allow for
the estimation of marginal hazard ratios with minimal bias. Of these two
approaches, IPTW using the propensity score resulted in estimates with lower mean
squared error when estimating the effect of treatment in the treated.
Stratification on the propensity score and covariate adjustment using the
propensity score result in biased estimation of both marginal and conditional
hazard ratios. Applied researchers are encouraged to use propensity score
matching and IPTW using the propensity score when estimating the relative
effect of treatment on time-to-event outcomes.
Tuesday, January 14, 2014
Analytics vs. Causal Inference
When I think of analytics, I primarily think about what Leo Breiman referred to as an 'algorithmic' approach to analysis. Breiman states:"There are two cultures in the use of statistical modeling to reach conclusions from data. One "assumes that the data are generated by a given stochastic data model" the other"uses algorithmic models and treats the data mechanism as unknown."
To take an example from higher education, lets look at a hypothetical I have proposed before. Assume we want to evaluate the causal effects of a summer camp designed to prepare high school graduates for their first year of college. If we are interested in making inferences about the causal impact of the camp on retention (this fits under the stochastic data modeling culture) we realize that the impact of the camp program itself (which is the causal effect of interest) as well as academic potential and motivation are confounded. Students that attend the camp may also be very motivated and academically strong, and likely to have high retention rates regardless of their attendance. We proceed with our analysis using some form of experimental or quasi-experimental design to attempt to statistically estimate or 'identify' the causal effects related to camp. We are concerned with things like standard errors, confidence intervals, statistical significance etc. We may use the results to evaluate the effectiveness of the camp program and improve resource allocation.
On the other hand, we may simply be interested in building a model that gives a measure of the probability of retention for first year students. We may want to take those results and segment the student population into strata based on their risk profile, and tailor programs to help with their academic success an improve resource allocation. The variable indicating 'camp' attendance among others might be a good predictor and aid in producing the probability estimates and well calibrated stratifications. We might accomplish this with logistic regression, decision trees, neural networks, random forests or gradient boosting or some other machine learning algorithm. We are concerned with things like predictive accuracy, discrimination, sensitivity, specificity, true positives, false positives, ranking, and calibration. We are not so concerned with p-values, statistical significance, standard errors etc.
Both approaches are data driven, and one is not more meritorious than the other. People often adopt one culture or paradigm and impugn the other. As Brieman states:
"Approaching problems by looking for a data model imposes an apriori straight jacket that restricts the ability of statisticians to deal with a wide range of statistical problems."
In a past Campus Technology article "The Predictive Analytics Framework Moves Forward" the following comments depict this schism that sometimes occurs between cultures:
"For some in education what we’re doing might seem a bit heretical at first--we’ve all been warned in research methods classes that data snooping is bad! But the newer technologies and the sophisticated analyses that those technologies have enabled have helped us to move away from looking askance at pattern recognition. That may take a while in some research circles, but in decision-making circles it’s clear that pattern recognition techniques are making a real difference in terms of the ways numerous enterprises in many industries are approaching their work"
In fact, both paradigms can often be complimentary in application. We might first develop an algorithm as indicated in the second approach and based on those results design a program or intervention targeted at students with a certain risk profile. We might then step into the world of inferential statistics and attempt to evaluate the effectiveness of this program.
In summary, both are going to utilize some sort of data driven model development, and as the popular parahrase of statistican George E.P. Box goes, all models are wrong, but some are useful. Regardless of the paradigm you are working in, the key is to to produce something useful for solving the problem at hand.
References:
The Predictive Analytics Reporting Framework Moves Forward
A Q & A with WCET Executive Director Ellen Wagner on the PAR Framework. Mary Grush 01/18/12 Campus Technology. http://campustechnology.com/articles/2012/01/18/the-predictive-analytics-reporting-framework-moves-forward.aspx
'Statistical Modeling: The Two Cultures' by L. Breiman (Statistical Science
2001, Vol. 16, No. 3, 199–231)
See also:
Predictive Analytics in Higher Ed http://econometricsense.blogspot.com/2012/01/predictive-analytics-in-higher-ed.html
Culture War: Classical Statistics vs. Machine Learning: http://econometricsense.blogspot.com/2011/01/classical-statistics-vs-machine.html
Economists as Data Scientists http://econometricsense.blogspot.com/2012/10/economists-as-data-scientists.html
To take an example from higher education, lets look at a hypothetical I have proposed before. Assume we want to evaluate the causal effects of a summer camp designed to prepare high school graduates for their first year of college. If we are interested in making inferences about the causal impact of the camp on retention (this fits under the stochastic data modeling culture) we realize that the impact of the camp program itself (which is the causal effect of interest) as well as academic potential and motivation are confounded. Students that attend the camp may also be very motivated and academically strong, and likely to have high retention rates regardless of their attendance. We proceed with our analysis using some form of experimental or quasi-experimental design to attempt to statistically estimate or 'identify' the causal effects related to camp. We are concerned with things like standard errors, confidence intervals, statistical significance etc. We may use the results to evaluate the effectiveness of the camp program and improve resource allocation.
On the other hand, we may simply be interested in building a model that gives a measure of the probability of retention for first year students. We may want to take those results and segment the student population into strata based on their risk profile, and tailor programs to help with their academic success an improve resource allocation. The variable indicating 'camp' attendance among others might be a good predictor and aid in producing the probability estimates and well calibrated stratifications. We might accomplish this with logistic regression, decision trees, neural networks, random forests or gradient boosting or some other machine learning algorithm. We are concerned with things like predictive accuracy, discrimination, sensitivity, specificity, true positives, false positives, ranking, and calibration. We are not so concerned with p-values, statistical significance, standard errors etc.
Both approaches are data driven, and one is not more meritorious than the other. People often adopt one culture or paradigm and impugn the other. As Brieman states:
"Approaching problems by looking for a data model imposes an apriori straight jacket that restricts the ability of statisticians to deal with a wide range of statistical problems."
In a past Campus Technology article "The Predictive Analytics Framework Moves Forward" the following comments depict this schism that sometimes occurs between cultures:
"For some in education what we’re doing might seem a bit heretical at first--we’ve all been warned in research methods classes that data snooping is bad! But the newer technologies and the sophisticated analyses that those technologies have enabled have helped us to move away from looking askance at pattern recognition. That may take a while in some research circles, but in decision-making circles it’s clear that pattern recognition techniques are making a real difference in terms of the ways numerous enterprises in many industries are approaching their work"
In fact, both paradigms can often be complimentary in application. We might first develop an algorithm as indicated in the second approach and based on those results design a program or intervention targeted at students with a certain risk profile. We might then step into the world of inferential statistics and attempt to evaluate the effectiveness of this program.
In summary, both are going to utilize some sort of data driven model development, and as the popular parahrase of statistican George E.P. Box goes, all models are wrong, but some are useful. Regardless of the paradigm you are working in, the key is to to produce something useful for solving the problem at hand.
References:
The Predictive Analytics Reporting Framework Moves Forward
A Q & A with WCET Executive Director Ellen Wagner on the PAR Framework. Mary Grush 01/18/12 Campus Technology. http://campustechnology.com/articles/2012/01/18/the-predictive-analytics-reporting-framework-moves-forward.aspx
'Statistical Modeling: The Two Cultures' by L. Breiman (Statistical Science
2001, Vol. 16, No. 3, 199–231)
See also:
Predictive Analytics in Higher Ed http://econometricsense.blogspot.com/2012/01/predictive-analytics-in-higher-ed.html
Culture War: Classical Statistics vs. Machine Learning: http://econometricsense.blogspot.com/2011/01/classical-statistics-vs-machine.html
Economists as Data Scientists http://econometricsense.blogspot.com/2012/10/economists-as-data-scientists.html
Friday, January 10, 2014
Do You Realy Have a 'Population' and Does It Matter?
They have taken up this question over at the Mostly Harmless Econometrics blog:
http://www.mostlyharmlesseconometrics.com/2013/11/pop-quiz-2/
See also: Is Inference Valid for Population Data
http://www.mostlyharmlesseconometrics.com/2013/11/pop-quiz-2/
See also: Is Inference Valid for Population Data
Saturday, January 4, 2014
The Oregon Medicaid Experiment and Linear Probability Models
I just recently discussed the methodology used in some recent papers analyzing the Oregon Medicaid expansion (see: http://econometricsense.blogspot.com/2014/01/the-oregon-medicaid-experiment-applied.html ). This was one of the papers:
"The Oregon Experiment--Effects of Medicaid on Clinical Outcomes," by Katherine Baicker, et al. New England Journal of Medicine, 2013; 368:1713-1722. http://www.nejm.org/doi/full/10.1056/NEJMsa1212321
If you read the supplementary appendix you will find the following:
In all of our ITT estimates and in our subsequent instrumental variable estimates (see below), we fit linear models even though a number of our outcomes are binary. Because we are interested in the difference in conditional means for the treatments and controls, linear probability models would pose no concerns in the absence of covariates or in fully saturated models (Angrist 2001, Angrist and Pischke 2009). Our models are not fully saturated, however, so it is possible that results could be affected by this functional form choice, especially for outcomes with very low or very high mean probability. We therefore explore the sensitivity of our results to an alternate specification using logistic regression and calculating average marginal effects for all binary outcomes, and are reassured that the results look very similar (see Table S15a-d below).
You will find a similar methodology in the more recent article in Science previously discussed. This weaves well with some of my past posts:
Linear Regression and Analysis of Variance with Binary Dependent Variables
Regression as an Empirical Tool (matching and linear probability models)
"The Oregon Experiment--Effects of Medicaid on Clinical Outcomes," by Katherine Baicker, et al. New England Journal of Medicine, 2013; 368:1713-1722. http://www.nejm.org/doi/full/10.1056/NEJMsa1212321
If you read the supplementary appendix you will find the following:
In all of our ITT estimates and in our subsequent instrumental variable estimates (see below), we fit linear models even though a number of our outcomes are binary. Because we are interested in the difference in conditional means for the treatments and controls, linear probability models would pose no concerns in the absence of covariates or in fully saturated models (Angrist 2001, Angrist and Pischke 2009). Our models are not fully saturated, however, so it is possible that results could be affected by this functional form choice, especially for outcomes with very low or very high mean probability. We therefore explore the sensitivity of our results to an alternate specification using logistic regression and calculating average marginal effects for all binary outcomes, and are reassured that the results look very similar (see Table S15a-d below).
You will find a similar methodology in the more recent article in Science previously discussed. This weaves well with some of my past posts:
Linear Regression and Analysis of Variance with Binary Dependent Variables
Regression as an Empirical Tool (matching and linear probability models)
Subscribe to:
Posts (Atom)