Monday, December 23, 2013

An Actuarial Application of Social Network Analysis

While the article from The Future Actuary linked to below doesn't really mention social network analysis (SNA), it speaks to insurance products designed to protect you from loss, personal liability, or related litigation issues involving social media. I'm sure there has already been plenty of work done in this area and if I get the chance I'll provide some updates. I would think knowledge of SNA would be critical i.e what is the relationship between centrality measures and network structure and risk? How can this help price products designed to protect individuals and businesses?

"As a social media entrepreneur, Matt Mullenweg put it: “we’re only one click away from a billion people.” While these technological advancements have made ‘getting connected’ one of the top priorities for employers and individuals, it doesn’t come without its perils. The increasing risks of legality/operation, security, and defamation are also just one click away. Maybe it’s time for social media risk management plans to take off. "

Tuesday, December 17, 2013

Econometrics and Big Data

Thanks to Dave Giles at Econometrics Beat for pointing out Hal Varian's paper. Big Data: New Tricks for Econometrics

"I believe that these methods have a lot to offer and should be more widely known and used by economists. In fact, my standard advice to graduate students these days is “go to the computer science department and take a class in machine learning.” There have been very fruitful collaborations between computer scientists and statisticians in the last decade or so, and I expect collaborations between computer scientists and econometricians will also be productive in the future."

See also Economists as Data Scientists.

The Convergence of Big Data, Genomics, Biotechnology and Seed Choice

Some recent discussion including big data found on my applied economics blog EconomicSense:

-Big Data, Genomics, and Seed Choice
-AgriTalk Discussion

Also, Blake Hurst's American article:

"Tyler Cowen argues that we’re about to see an even wider disparity in incomes between the 10 to 15 percent of the population that can relate well to computers and the vast majority of us who will deliver services to the computer-savvy class. Farming may be one of the first industries to explore the validity of Cowen’s thesis. All of us involved in agriculture will soon have to decide whether we want to occupy the nostalgic niche providing artisanal beets and heritage pork to Cowen’s 10 percent, or whether we’ll roll the dice on surviving the transition to a data-driven agriculture."

Wednesday, November 6, 2013

Kentucky Association of Institutional Research (KAIR) Presentations

I'll be giving the following two talks at KAIR this year.
  
A Guide to Quasi-Experimental Designs

Quasi-experimental designs including propensity score methods, instrumental variables, regression discontinuity, and difference-in-difference estimators offer an inferentially rigorous alternative for program evaluation. In this guide, I begin with an introduction to the potential outcomes framework for rigorously characterizing selection bias and follow with discussions of quasi-experimental methods that may be useful to practitioners involved in program evaluation.

Paper: http://works.bepress.com/matt_bogard/24/
Slides: http://works.bepress.com/matt_bogard/27/

A Data Driven Analytic Strategy for Increasing Yield and Retention

As many Universities face the constraints of declining enrollment demographics, pressure from state governments for increased student success, as well as declining revenues, the costs of utilizing anecdotal evidence and intuition based on ‘gut’ feelings to make time and resource allocation decisions become significant. This presentation describes several data mining algorithms that could be used to develop a model to score university students based on their probability of enrollment and retention early in the enrollment funnel so that staff and administrators can work to recruit students that not only have an average or better chance of enrolling but also succeeding once they enroll. The use of SAS® Enterprise Miner and SAS® EBI server are discussed as an example of a business intelligence platform that can be used for implementing the models and delivering flexible web based reports to end users.

Paper: http://works.bepress.com/matt_bogard/26/
Slides:http://works.bepress.com/matt_bogard/28/




Tuesday, September 10, 2013

Causal Inference and Quasi-experimental Design Roundup


I’ve dedicated several posts recently to the subject of quasi-experimental designs and causal inference. I’ve tried to organize the following related links for a bigger picture.

First off, regression is often mischaracterized by a clinical view of  assumptions related to linearity. However, as Angrist and Pischke state:

"In fact, the validity of linear regression as an empirical tool does not turn on linearity either...The statement that regression approximates the CEF lines up with our view of empirical work as an effort to describe the essential features of statistical relationships, without necessarily trying to pin them down exactly."  - Mostly Harmless Econometrics, p. 26 & 29

For more discussion see:


In Mostly Harmless Econometrics, not only is linear regression rigorously developed and discussed, but quasi-experimental designs are given very heavy emphasis. As discussed in Cellini (2008):

… proxy variable, fixed effects, and difference in- differences approaches are becoming quite common. Indeed, these approaches have replaced basic multivariate regression as the new standard for education research in the economics literature”

The links below attempt to highlight at least in a heuristic sense, these methods and the issues they attempt to address:








Time Series Methods:


Regression as an Empirical Tool


Linear regression is a powerful empirical tool for the social sciences.  Its robustness is often underrated, while at other times its use and interpretation is mischaracterized.  Andrew Gelman and Agrist and Pischke are two great sources for learning about regression in an applied context.

I particularly like Gelman's comment here:
"It's all about comparisons, nothing about how a variable "responds to change." Why? Because, in its most basic form, regression tells you nothing at all about change. It's a structured way of computing average comparisons in data."

This 'computing average comparisons of data' interpretation is why regression works as sort of a matching estimator as Angrist and Pischke  argue. 

“Our view is that regression can be motivated as a particular sort of weighted matching estimator, and therefore the differences between regression and matching estimates are unlikely to be of major
empirical importance” (Chapter 3 p. 70)

 In further discussion Gelman goes on to say, I think in a very appropriate interpretation:

“They're saying (Angrist and Pischke ) that regression, like matching, is a way of comparing-like-with-like in estimating a comparison. This point seems commonplace from a statistical standpoint but may be news to some economists who might think that regression relies on the linear model being true.”

This brings up a very important point, one also in-line with Angrist and Pischke regarding the use of regression as an empirical tool in the social sciences:

"In fact, the validity of linear regression as an empirical tool does not turn on linearity either...The statement that regression approximates the CEF lines up with our view of empirical work as an effort to describe the essential features of statistical relationships, without necessarily trying to pin them down exactly."  - Mostly Harmless Econometrics, p. 26 & 29

Regression users can seem at odds with each other at times. On one extreme they can get caught up in making very clinical assumptions  about linearity (see somewhat related discussions related to linear probability models  here and here) then on the other hand, take robustness to extremes by failing to consider at times questions of unobserved heterogeneity, endogeneity, selection bias, and identification.

Cellini(2008) discusses this issue in an analysis of the impact of financial aid on college enrollment:

“While simple ordinary least squares estimates of the impact of aid on college-going can reveal a correlation between financial aid policies and enrollment, these estimates are likely to suffer from omitted variable bias due to self-selection, potentially overestimating or underestimating the causal impact of these policies on enrollment.”

“The discussion above has outlined several methods for addressing the problem of omitted variable bias in financial aid research… proxy variable, fixed effects, and difference in- differences approaches are becoming quite common. Indeed, these approaches have replaced basic multivariate regression as the new standard for education research in the economics literature”

 This is where quasi-experimental methods come in to play.


References

 Stephanie Riegg Cellini. Causal Inference and Omitted Variable Bias in Financial Aid Research: Assessing Solutions The Review of Higher Education Spring 2008, Volume 31, No. 3, pp. 329–354

Friday, September 6, 2013

IPTW Regression


An alternative to direct matching or matching on propensity scores involves the use of the inverse of propensity scores in a weighted regression framework (Horvitz and Thompson (1952), known as inverse probability of treatment weighted (IPTW) regression where:

 IPTW regression (with weights specified as above) specifically estimates the average treatment effect (ATE) (Astin, 2011):

ATE = E[Y1i-Y0i] 
Inverse probability of treatment weighting (IPTW) uses weights derived from the propensity scores  to create a pseudo population such that the distribution of covariates in the population are independent of treatment assignment.  (Astin,2011). This is an appeal to the CIA and Rosenbaum and Rubin’s propensity score theorem discussed before.  The weighting scheme essentially ‘weights up’ control units to look like treatment units (Stuart,2011). 



References

Austin, P.(2011). An Introduction to Propensity Score Methods for Reducing the Effects of Confounding in Observational Studies.  Multivariate Behav Research.May; 46(3): 399–424.

Horvitz D. G. & Thompson D. J.(1952) . A Generalization of Sampling Without Replacement From a Finite Universe. Journal of the American Statistical Association, Vol. 47, No. 260 (Dec., 1952), pp. 663- 685

Maciejewski, M. L. & Brookhart, M.A. (2011). Propensity score workshop . Retrieved January 19,2013. Website: http://ahrqplexnet.sharepointspace.com/Webinars/PS_webinar_followup.pdf

Stuart ,E.(2011).Propensity score methods for estimating causal effects: The why, when, and how . Johns Hopkins Bloomberg School of Public Health. Department of Mental Health. Department of      Biostatistics. Retrieved January 19,2013. Website:  www.biostat.jhsph.edu/estuart