Friday, July 27, 2012

Empirical Work in The Social Sciences- from Mostly Harmless Econometrics



"In fact, the validity of linear regression as an empirical tool does not turn on linearity either...The statement that regression approximates the CEF lines up with our view of empirical work as an effort to describe the essential features of statistical relationships, without necessarily trying to pin them down exactly."  - Mostly Harmless Econometrics, p. 26 & 29

I really like Mostly Harmless Econometrics. I started reading it some time ago. I have had several formal courses in econometrics, mathematical statistics, and experimental design, and have spent a lot of time in the pages of Golberger's A Course in Econometrics., Greene's Econometric Analysis, and Kennedy's A Guide to Econometrics, as well as Hastie, Tibshirani, and Friedman's Elements of Statistical Learning: Data Mining, Inference, and Prediction, but Angrist and Pishke's book really speaks to me in the every day empirical work that I find myself caught up in. The other textbooks are great, maybe essential for a student, but Mostly Harmless Econometrics is a must have for the practitioner. MHE is not a substitute for a solid econometrics background, in fact, it would not have made sense to me without it, but it wouldn't have as much meaning either without some prior experience. I find myself re-reading sections because on the job data challenges continue to make this book more relevant every day. The other textbooks get you started with the theory (again very important). They are great for 'highway' use, when your are coasting on the smooth surfaces of textbook ideals. MHE gives you that push you need when you get bogged down on the very muddy roads of real life empirical work.  Its the off-road backwoods survival manual for practitioners.



Sunday, June 24, 2012

How can applied econometrics and analytics benefit your business?



“We’re rapidly entering a world where everything can be monitored and measured,” said Erik Brynjolfsson, an economist and director of the Massachusetts Institute of Technology’s Center for Digital Business. “But the big problem is going to be the ability of humans to use, analyze and make sense of the data.”

“I.B.M., seeing an opportunity in data-hunting services, created a Business Analytics and Optimization Services group in April. The unit will tap the expertise of the more than 200 mathematicians, statisticians and other data analysts in its research labs — but that number is not enough. I.B.M. plans to retrain or hire 4,000 more analysts across the company.” – From ‘For Today’s Graduate, Just One Word: Statistics’ – NYT, Steve Lohr, Aug 2009.

"Some companies have built their very business on their ability to collect, analyze, and act on data" – ‘Competing on Analytics’. Harvard Bus.Review Jan 2006.

"The success of companies like Google and Amazon has encouraged a whole generation of business leaders to try and replicate their data-driven processes, and left them searching for data scientists." – Jobs for Data Scientists Explode Across The Market - NYT, July 20,2011.

Many businesses make some sort of use of their customer and market data. For some, it’s just a matter of storing and accessing customer data for record keeping and transactional purposes. Others like Google, Netflix, I.B.M. , or the Oakland Athletics make data analysis and analytics a major part of their business model. 

By analyzing past business records, data mining and analytics can help identify patterns that can support decisions that are more cost effective and efficient.  This is the specialty of what has contemporarily been dubbed the data scientist.


What is a data scientist?

“What sets data scientists apart from other data workers, including data analysts, is their ability to create logic behind the data that leads to business decisions. "Data scientists extract data, formulate models and apply quantitative analysis in a proactive manner" -Laura Kelley, Vice President, Modis.

"They can suck data out of a server log, a telecom billing file, or the alternator on a locomotive, and figure out what the heck is going on with it. They create new products and services for customers. They can also interface with carbon-based lifeforms — senior executives, product managers, CTOs, and CIOs. You need them." - Can You Live Without a Data Scientist, Harvard Business Review.
“At least as important [as big data technologies] are the people with the skill set (and the mind-set) to put them to good use…what data scientists do is make discoveries while swimming in data… the dominant trait [of data scientists] is intense curiosity—a desire to go beneath the surface of a problem, find the questions at its heart, and distill them into a very clear set of hypotheses that can be tested. This often entails the associative thinking that characterizes the most creative scientists in any field….perhaps it’s becoming clear that the word ‘scientist’ fits this emerging role”Tom Davenport and D.J. Patil
The Data Science Venn Diagram created by Drew Conway helps in defining data science and the role of a data scientist.



Data scientists not only have expertise in applied research and statistics, but they are comfortable doing the programming (a.k.a. hacking) necessary to shape data into a form suitable for analysis. In addition, data scientists have practical knowledge and expertise in the field or industry they work in.

Do you need billions of dollars and a special division staffed with expensive PhD’s to create value from your business data?

No matter the size of your business or organization, data mining and analytic expertise can help create value from your data. The size of your staff and the extent of resources devoted to data science and analytics may depend on the nature of your organization. However, if you are a small business, you don’t necessarily have to make the same kind of investments as Google or Netflix.  Competing on analytics is as much a mindset as it is an investment, and you may be able to accomplish a lot with current staff or local talent.

Can you afford to have a ‘data scientist’ on staff?  Is it difficult to find people with this skillset?

Tapping data science talent may not be as difficult as you think. You don’t necessarily need an expensive PhD statistician, there are many graduates with similar degrees that possess these skills. Take for instance agricultural economics:

"The combination of quantitative training and applied work makes agricultural economics graduates an extremely well-prepared source of employees for private industry. That's why American Express has hired over 80 agricultural economists since 1990." - David Edwards, Vice President-International Risk Management, American Express from Why Study Applied/Agricultural Economics.

This is also stated quite well in a the NYT article cited above ‘For Today’s Graduate, Just One Word: Statistics’:

"Though at the fore, statisticians are only a small part of an army of experts using modern statistical techniques for data analysis. Computing and numerical skills, experts say, matter far more than degrees. So the new data sleuths come from backgrounds like economics, computer science and mathematics."

Perhaps you see a need for data driven decision making in your business and you want to tap someone with this talent, but you just don’t think you can justify someone on staff full time to do this. I would first challenge this notion. You may start with a few questions and answers, and some predictive modeling, but you will find there is always more data being generated, more questions, and a constant need to tweak and improve models and forecasts. However, modern technology makes telecommuting a very attractive opportunity.  You might find that you can contract with someone with this skill set on an adhoc or as needed basis. There’s only one way to find out how much potential value is buried in your data, and you have to start somewhere. With just a few data mining techniques you can begin to extract insight from your data that you might not otherwise achieve even after hours or years of pouring over lists and and row after row, column after column in excel.

Data Mining and Predictive Analytics

First off, there are differences between traditional statistics and data mining. See ‘Culture War: Classical Statistics vs. Machine Learning.’ What’s important is that you have someone that can identify the best tool for the task at hand, whether it’s a traditional experimental design involving analysis of variance, a forecast or time series analysis, a predictive model using logistic regression or decision trees, or one of the many other possible data mining tools available to a data scientist.  In addition, there’s more ways to gain insight from your data than just fancy models or algorithms. Data visualization allows you to transmit information to end users without the sometimes distracting  statistical terminology, complicated equations, or never ending excel sheets. 


 

Real World Applications

My most recent analytics accomplishment to this point involves the development of a predictive model that we use to identify students that have a high risk of dropping out at WKU.  Working with my team, we’ve incorporated my model metrics into our data base/reporting/decision support system so that administrators have access to these high level analytical tools for strategic decision making.  We won an honorable mention from SAS at the recent SAS Global forum for our presentation. (see here for the paper with screenshots).

In addition, with the rise of social media and the digitization of so much data, there are some analytical tools such as text mining and social network analysis that are of particular interest that might not have seemed feasible in an applied business setting just a decade ago.

Text Mining

With Twitter, Facebook, email, online forums, open response surveys, customer and reader comments on web pages and news articles etc. there is a lot of information available to companies and organizations in the form of text. Without hiring experts to read through all of the thousands of pages worth of text available and making subjective claims about its meaning, text mining allows us to take otherwise unusable 'qualitative' data and convert it into quantitative measures that we can use for various types of reporting and modeling. In 'An Intuitive Approach to Text Mining with SAS IML'  I demonstrate how the mathematical technique of singular value decomposition  (SVD) can be used to do this.

For example, with text mining techniques based on SVD you can take the following sample text from comments about finely textured beef:


 
And quantify it to classify or segment respondents (who may be customers or clients) 

 
Or use the information in a predictive model:


 
Tools like SAS Text Miner in conjunction with SAS Enterprise Miner are designed specifically to do this type of analysis on a much larger scale. I have used both of these tools on much larger document collections and obtained promising results using text topics from SVD in predictive modeling applications. R also has open source tools as well.  I’ve used R to text mine tweets related to U.S. political issues related to the debt ceiling and U.S. budget  as well as tweets related to the term ‘Factory Farm’ as depicted below:

 
Social Network Analysis

With the rise in the use of social media, data related to social networks is ripe for analysis using techniques from social network analysis and graph theory. According to International Network for Social Network Analysis, ‘Social network analysis is focused on uncovering the patterning of people's interaction’.

Social network analysis (SNA) allows us to answer questions such as who are key  actors in a network? Who are the most influential members of a network? Who seems to be acting on the peripheral? Which connections in the network are most important?  Are there key players bridging connections or information between otherwise disconnected groups? Have policies or other forces changed the overall dynamics/interaction between people in the network (i.e. has the network structure changed in any meaningful way) and does that relate to some other performance outcome or goal?

More specific applications of SNA may include Student Integration and Persistence, Business to Business Supply Chains, Seeding Strategies for Viral Marketing, and Predicting Customer Churn. The open source software R and NetDraw provide  many tools for conducting social network analysis. See  also ‘Using Twitter to Demonstrate Basic Concepts from Social Network Analysis’ as well as ‘An Introduction to Social Network Analysis Using R and Netdraw.’   As some of these examples demonstrate, measures derived from social network analysis can be very useful in predictive modeling. For a more basic example see ‘Using SNA in Predictive Modeling'.


For more information:

If you feel you can benefit from the services of a data scientist or have further questions about applied econometrics and analytics please contact me for more information or feel free to visit my blog or selected works where you can find a copy of my CV.


LinkedIn Profile:  (link)    
Selected Works Profile: http://works.bepress.com/matt_bogard/

Monday, May 28, 2012

Notes to 'Support' an Understanding of Support Vector Machines

Some time ago I wrote a short article highlighting the similarities of tools used in economics, bioinformatics, and machine learning (Mathematical Themes in Economics, Machine Learning, and Bioinformatics). In this followup I expand more on the details 'supporting' support vector machines. While not appropriate for making the kinds of inferences econometricians are so often fond of, they are powerful tools for classification. They are certainly an artifact of  the data mining culture.  While I don't cover all of the details (for instance I don't discuss kernel methods, or duality) I cover some of the basic concepts that are likely to trip up someone new to SVMs, like how a line in R2 or hyperplane in higher dimensions can be expressed using inner products  as in <w,x> + b = 0.

Notes on Support Vector Machines - link


Matt Bogard, Western Kentucky University

Abstract

 

The most basic idea of support vector machines is to find a line (or hyperplane) that separates classes of data, and use this information to classify new examples. But we don’t want just any line, we want a line that maximizes the distance between classes (we want the best line). This line turns out to be a separating hyperplane that is equidistant between the supporting hyperplanes that ‘support’ the sets that make up each distinct class. The notes that follow discuss the concepts of supporting and separating hyperplanes and inner products as they relate to support vector machines (SVMs). Using simple examples, much detail is given to the mathematical notation used to represent hyperplanes, as well as how the SVM classification works.

Suggested Citation

 

Matt Bogard. 2012. "Notes on Support Vector Machines" The Selected Works of Matt Bogard
Available at: http://works.bepress.com/matt_bogard/20

Saturday, May 5, 2012

An Intuitive Approach to Text Mining with SAS IML

Text Mining- in Plain English: Turning Text into Numbers

With Twitter, Facebook, email, online forums, open response surveys, customer and reader comments on web pages and news articles etc. there is a lot of information available to companies and organizations in the form of text. Without hiring experts to read through all of the thousands of pages worth of text available and making subjective claims about its meaning, text mining allows us to take otherwise unusable 'qualitative' data and convert it into quantitative measures that we can use for various types of reporting and modeling. In the example below I demonstrate how the mathematical technique of singular value decomposition  (SVD) can be used to do this.

Sometimes to assess statistical techniques, or to even understand them at a basic level, you need data with properties you understand. With numerical data, that typically might call for simulation. But since we are dealing with text, I just made some up that will work to clearly demonstrate the effectiveness of SVD.

Below you will see 10 hypothetical comments to the question- 'What is your take on pink slime?' Each person's response is considered a document, and I have classified each document as type 'H' or 'S' as explained below.  The goal will be to see if we can use a basic application of SVD to convert these comments into numbers and use them to predict what type of person  made which type of comments (i.e. does a comment belong in category H or S) based on clustering or some type of predictive model. This gets way beyond simply classifying comments by doing a key word search. SVD allows us to not only classify documents by the specific words they contain, but also by how similar they are.

 The purpose of this exercise is to provide intuition for text mining and the application of singular value decomposition. The text above is made up, and specifically designed to produce the results I'm after. I'm not making any claims one way or the other about how realistic this is. But suppose these are potential customers and we want to be able to distinguish between the hippies, who favor a more local nostalgic food supply from centuries past  (designated as class or type 'H') and the animal scientists, designated as type 'S.' Obviously comments of class H are critical of modern agriculture and technology.  We might want to make these distinctions for PR, marketing, lobbying,or educational outreach purposes.  

After cleaning up the text (which many software programs like SAS Enterprise Miner provide excellent tools for doing so)by eliminating parts of speech, articles, etc. we can form a term-document frequency matrix as follows:




Singular Value Decomposition

Singular Value Decomposition (SVD) is a concept from linear algebra based on the following matrix equation:

A = USV which states that a rectangular matrix A can be decomposed into 3 other matrix components:

U  consists of the orthonormal eigenvectors of AA’,  where U’U =  I (recall from linear algebra, orthogonal vectors of unit length are ‘orthonormal’)
V consists of the orthonormal eigenvectors of A’A
S  is a diagonal matrix consisting of the square root of the eigenvalues of U or V (which are equal). The values in S depict the variance of linearly independent components along each dimension similarly to the way eigenvalues depict variance explained by ‘factors’ or components in principle components analysis (PCA).

SVD provides the mathematical foundation for text mining.

A  term document matrix A can be decomposed as in:
A = USV’

If the term document matrix A were a collection of individuals’ textual responses or comments to a survey question (or Facebook post etc. ) then each individual’s response would be considered a ‘document’. The individual words they used in their response are the ‘terms.’ The term document frequency matrix consists of rows that represent each term and columns that represent each ‘document’ or individual. 

The  vectors in U can be used to score documents in the term-document matrix, as in U’A.  (this may be analogous to the way values of x are scored by eigenvectors in PCA)  As a result, a single numerical value or ‘score’ can be assigned to each person’s textual response (i.e. each document) for each SVD dimension (i.e. each kept independent vector in U). Thus the text is converted into numerical scores (via the transformation U’A) that can then be used in predictive modeling (as predictor variables) or  clustering can be utilized to cluster the individual textual responses (or ‘documents’).  So ultimately SVD converts  otherwise unuseful  text into numeric SVD scores or more interpretable clusters.  Likewise, using  the weights from the eigenvectors that comprise  V to score the terms in the term document matrix A, as in AV’ we can score  cluster like terms as ‘topics’. 

Thus SVD of A gives allows us to derive the following scores:

U’A = SVD document vectors 
AV’ = SVD term vectors

 This can be loosely demonstrated in using PROC IML in SAS. Specifying the term document frequency matrix in PROC IML and implementing SVD produces:









 
The document vectors U’A can be depicted as follows:


 
Here is the original text document with customer types/classes and the appended SVD scores. 

 As depicted above, SVD has allowed us to replace all of the text with the quantitative values for the associated SVD scores (derived from U'A). In this case, I'm representing all of the text with just two SVD dimensions. These values can then be used for clustering or entered into a predictive model.


 
It is easy to see that the different documents types (response types H vs S) cluster very well on the dimensions SVD1 and SVD2. In fact, all responses of type S have a value of SVD2 < .5.
The SVD values can also be entered into a regression or other type of predictive model.

 
Consistent with the observed relationship between SVD2 and class ‘S’ and ‘H’ clusters, we see that SVD2 is significant in the regression. In addition, it shows that higher values of SVD2 decrease the probability of being classified as document or response type ‘S’. Yes, this is OLS on a binary dependent variable, but again the purpose is to provide intuition and motivation for using text analytics for predictive modeling. 

This seemed to work ok on a small collection of documents carefully constructed to illustrate the concepts above, but tools like SAS Text Miner in conjunction with SAS Enterprise Miner are designed specifically to do this type of analysis on a much larger scale. I have used both of these tools on much larger document collections and obtained promising results using text topics from SVD in predictive modeling applications.

SAS CODE


proc iml;
       A = {0 0      0      1      0      0      0      0      0      0,
1      1      0      0      0      1      0      0      0      0,
0      0      0      0      0      1      1      1      0      2,
0      0      0      0      0      1      0      0      0      0,
0      1      1      0      0      0      0      0      0      0,
1      0      0      0      0      0      0      0      0      0,
0      0      1      0      1      0      0      0      0      0,
0      0      0      0      1      0      0      0      0      0,
0      0      0      0      1      0      0      0      0      0,
0      0      1      0      0      0      0      0      0      0,
0      0      0      0      0      1      0      1      0      1,
0      0      1      0      0      0      0      0      0      0,
0      1      1      1      0      0      0      0      0      0,
0      0      0      1      0      0      0      0      0      0,
0      0      0      0      0      0      1      1      0      0,
1      0      1      0      0      0      0      0      0      0,
0      0      0      0      0      0      1      0      0      0,
0      0      0      0      0      0      0      0      0      1,
0      0      0      0      0      1      0      0      0      0,
0      0      1      0      0      0      0      0      0      0,
1      0      0      0      1      0      0      0      0      0,
0      1      0      1      0      0      0      0      0      0,
0      1      0      0      0      1      0      0      0      0,
0      0      0      0      0      0      0      1      0      0,
0      0      0      0      0      0      0      0      0      1,
0      0      0      0      0      0      0      0      1      1,
0      0      0      0      0      0      0      0      0      1,
1      0      0      0      1      0      0      0      0      0,
0      0      0      0      0      0      1      1      1      0,
0      0      1      0      0      0      0      0      0      0,
0      0      0      0      0      0      0      0      1      0,
0      0      0      0      0      0      0      0      0      1,
0      0      0      0      0      1      0      1      0      0,
0      0      0      0      0      0      0      0      1      0
};
       print(A); /* print matrix A*/
       n = nrow(A); /* how many rows */
       p = ncol(A); /* how many columns */
       print(n);   
       print(p);
       call svd(u,s,v,A); /* sigular value decomposition of A = usv'*/
       print(u); /* independent eigenvectors of AA' */
       print(s); /*  independent eigenvectors of A'A */
       print(v); /*singular values (sqrt(eigenvalues)) of AA' or A'A */
       uTa = T(u)*A; /* document vectors*/
       avT = A*T(v); /* term vectors */
       print(uTa);
       print(avT);
    /* scoring a data set */
       ID = {1, 2, 3, 4, 5, 6, 7, 8, 9,10}; /* create a document id matrix */
       print(ID);
       docscores =ID||T(uTa); /* combine with svd results */
       print(docscores);
       /* export as a SAS data set */
       varnames = 'svd1':'svd10';
       create svd_scores from docscores[colname = varnames];
       append from docscores;
       close svd_scores;

quit;
run;

 *-------------------------------------------------------------------------------------*
 |  SCORING, CLUSTERING AND PREDICTIVE MODELING
 *-------------------------------------------------------------------------------------*;

* CREATE CLIENT DATA SET;

DATA CLIENTS;
       INPUT ID CLASS $ ;
       CARDS;
1      H
2      H
3      H
4      H
5      H
6      S
7      S
8      S
9      S
10  S
;
RUN;

* MERGE WITH SVD SCORES;

PROC SQL;
       CREATE TABLE CLIENTS_SCORED AS
       SELECT A.ID, A.CLASS, B.SVD2 AS SVD1, B.SVD3 AS SVD2
       FROM CLIENTS A LEFT JOIN SVD_SCORES B
       ON A.ID = B.SVD1;
QUIT;

PROC PRINT DATA = CLIENTS_SCORED;
RUN;


* CLUSTER/ VISUALIZE  CLIENTS BASED ON SVD SCORES;
PROC GPLOT DATA = CLIENTS_SCORED;
       PLOT SVD1*SVD2 = ID;
RUN;
QUIT;


* RECODE FOR NUMERIC DEPENDENT VAR;

DATA CLIENTS_SCORED2;
       SET CLIENTS_SCORED;
       IF CLASS = 'H' THEN Y = 0;
       ELSE Y = 1;
RUN;

* BASIC REGRESSION MODEL;

PROC REG DATA = CLIENTS_SCORED2;
       MODEL Y = SVD1 SVD2;
RUN;
QUIT;