An attempt to make sense of econometrics, biostatistics, machine learning, experimental design, bioinformatics, ....
Friday, July 27, 2012
Empirical Work in The Social Sciences- from Mostly Harmless Econometrics
"In fact, the validity of linear regression as an empirical tool does not turn on linearity either...The statement that regression approximates the CEF lines up with our view of empirical work as an effort to describe the essential features of statistical relationships, without necessarily trying to pin them down exactly." - Mostly Harmless Econometrics, p. 26 & 29
I really like Mostly Harmless Econometrics. I started reading it some time ago. I have had several formal courses in econometrics, mathematical statistics, and experimental design, and have spent a lot of time in the pages of Golberger's A Course in Econometrics., Greene's Econometric Analysis, and Kennedy's A Guide to Econometrics, as well as Hastie, Tibshirani, and Friedman's Elements of Statistical Learning: Data Mining, Inference, and Prediction, but Angrist and Pishke's book really speaks to me in the every day empirical work that I find myself caught up in. The other textbooks are great, maybe essential for a student, but Mostly Harmless Econometrics is a must have for the practitioner. MHE is not a substitute for a solid econometrics background, in fact, it would not have made sense to me without it, but it wouldn't have as much meaning either without some prior experience. I find myself re-reading sections because on the job data challenges continue to make this book more relevant every day. The other textbooks get you started with the theory (again very important). They are great for 'highway' use, when your are coasting on the smooth surfaces of textbook ideals. MHE gives you that push you need when you get bogged down on the very muddy roads of real life empirical work. Its the off-road backwoods survival manual for practitioners.
Sunday, June 24, 2012
How can applied econometrics and analytics benefit your business?
“We’re rapidly
entering a world where everything can be monitored and measured,” said Erik
Brynjolfsson, an economist and director of the Massachusetts Institute of Technology’s
Center for Digital Business. “But the big problem is going to be the ability of
humans to use, analyze and make sense of the data.”
“I.B.M., seeing an
opportunity in data-hunting services, created a Business Analytics and
Optimization Services group in April. The unit will tap the expertise of the
more than 200 mathematicians, statisticians and other data analysts in its
research labs — but that number is not enough. I.B.M. plans to retrain or hire
4,000 more analysts across the company.” – From ‘For Today’s Graduate, Just
One Word: Statistics’ – NYT,
Steve Lohr, Aug 2009.
"Some companies
have built their very business on their ability to collect, analyze, and act on
data" – ‘Competing on Analytics’. Harvard Bus.Review Jan 2006.
"The success of
companies like Google and Amazon has encouraged a whole generation of business
leaders to try and replicate their data-driven processes, and left them
searching for data scientists." – Jobs for Data Scientists Explode
Across The Market - NYT,
July 20,2011.
Many businesses make some sort of use of their customer and
market data. For some, it’s just a matter of storing and accessing customer
data for record keeping and transactional purposes. Others like Google, Netflix,
I.B.M. , or the Oakland
Athletics make data analysis and analytics a major part of their business
model.
By analyzing past business records, data mining and
analytics can help identify patterns that can support decisions that are more
cost effective and efficient. This is
the specialty of what has contemporarily been dubbed the data scientist.
What is a data
scientist?
“What sets data
scientists apart from other data workers, including data analysts, is their
ability to create logic behind the data that leads to business decisions.
"Data scientists extract data, formulate models and apply quantitative
analysis in a proactive manner" -Laura
Kelley, Vice President, Modis.
"They can suck data out of a server log, a telecom billing file, or the alternator on a locomotive, and figure out what the heck is going on with it. They create new products and services for customers. They can also interface with carbon-based lifeforms — senior executives, product managers, CTOs, and CIOs. You need them." - Can You Live Without a Data Scientist, Harvard Business Review.
“At least as important [as big data technologies] are the people with the skill set (and the mind-set) to put them to good use…what data scientists do is make discoveries while swimming in data… the dominant trait [of data scientists] is intense curiosity—a desire to go beneath the surface of a problem, find the questions at its heart, and distill them into a very clear set of hypotheses that can be tested. This often entails the associative thinking that characterizes the most creative scientists in any field….perhaps it’s becoming clear that the word ‘scientist’ fits this emerging role”—Tom Davenport and D.J. Patil
The Data Science Venn Diagram created by Drew Conway helps
in defining data science and the role of a data scientist.
Data scientists not only have expertise in applied
research and statistics, but they are comfortable doing the programming (a.k.a.
hacking) necessary to shape data into a form suitable for analysis. In addition,
data scientists have practical knowledge and expertise in the field or
industry they work in.
Do you need billions
of dollars and a special division staffed with expensive PhD’s to create value
from your business data?
No matter the size of your business or organization, data
mining and analytic expertise can help create value from your data. The size of
your staff and the extent of resources devoted to data science and analytics
may depend on the nature of your organization. However, if you are a small
business, you don’t necessarily have to make the same kind of investments as
Google or Netflix. Competing on
analytics is as much a mindset as it is an investment, and you may be able to
accomplish a lot with current staff or local talent.
Can you afford to
have a ‘data scientist’ on staff? Is it
difficult to find people with this skillset?
Tapping data science talent may not be as difficult as you
think. You don’t necessarily need an expensive PhD statistician, there are many
graduates with similar degrees that possess these skills. Take for instance
agricultural economics:
"The combination
of quantitative training and applied work makes agricultural economics
graduates an extremely well-prepared source of employees for private industry.
That's why American Express has hired over 80 agricultural economists since
1990." - David Edwards, Vice President-International Risk Management,
American Express from Why
Study Applied/Agricultural Economics.
This is also stated quite well in a the NYT article cited
above ‘For Today’s Graduate, Just One Word: Statistics’:
"Though at the fore, statisticians are only a small
part of an army of experts using modern statistical techniques for data
analysis. Computing and numerical skills, experts say, matter far more than
degrees. So the new data sleuths come from backgrounds like economics, computer
science and mathematics."
Perhaps you see a need for data driven decision making in
your business and you want to tap someone with this talent, but you just don’t
think you can justify someone on staff full time to do this. I would first
challenge this notion. You may start with a few questions and answers, and some
predictive modeling, but you will find there is always more data being
generated, more questions, and a constant need to tweak and improve models and
forecasts. However, modern technology makes telecommuting a very attractive
opportunity. You might find that you can
contract with someone with this skill set on an adhoc or as needed basis. There’s
only one way to find out how much potential value is buried in your data, and you have to start somewhere. With just a few data mining
techniques you can begin to extract insight from your data that you might not
otherwise achieve even after hours or years of pouring over lists and and row
after row, column after column in excel.
Data Mining and
Predictive Analytics
First off, there are differences between traditional
statistics and data mining. See ‘Culture
War: Classical Statistics vs. Machine Learning.’ What’s important is that
you have someone that can identify the best tool for the task at hand, whether it’s
a traditional experimental design involving analysis of variance, a forecast or
time series analysis, a predictive model using logistic regression or decision
trees, or one of the many other possible data mining tools available to a data
scientist. In addition, there’s more
ways to gain insight from your data than just fancy models or algorithms. Data
visualization allows you to transmit information to end users without the sometimes distracting statistical terminology, complicated
equations, or never ending excel sheets.
Real World
Applications
My most recent analytics accomplishment to this point
involves the development of a predictive model that we use to identify students
that have a high risk of dropping out at WKU. Working with my team, we’ve
incorporated my model metrics into our data base/reporting/decision support
system so that administrators have access to these high level analytical tools
for strategic decision making. We won an honorable mention from SAS at
the recent SAS Global forum for our presentation. (see here
for the paper with screenshots).
In addition, with the rise of social media and the
digitization of so much data, there are some analytical tools such as text mining and social network analysis that are of
particular interest that might not have seemed feasible in an applied business
setting just a decade ago.
Text Mining
With Twitter, Facebook, email, online forums, open response
surveys, customer and reader comments on web pages and news articles etc. there
is a lot of information available to companies and organizations in the form of
text. Without hiring experts to read through all of the thousands of pages
worth of text available and making subjective claims about its meaning, text
mining allows us to take otherwise unusable 'qualitative' data and convert it
into quantitative measures that we can use for various types of reporting and
modeling. In 'An Intuitive Approach to Text Mining with SAS IML' I demonstrate how the mathematical technique
of singular value decomposition (SVD) can be used to do this.
For example, with text mining techniques based on SVD you
can take the following sample text from comments about finely textured beef:
And quantify it to classify or segment respondents (who may
be customers or clients)
Or use the information in a predictive model:
Tools like SAS
Text Miner in conjunction with SAS
Enterprise Miner are designed specifically to do this type of analysis on a
much larger scale. I have used both of these tools on much larger document
collections and obtained promising results using text topics from SVD in
predictive modeling applications. R
also has open source tools as well. I’ve
used R to text mine tweets related to U.S. political issues related to the debt
ceiling and U.S. budget as well as
tweets related to the term ‘Factory
Farm’ as depicted below:
Social Network
Analysis
With the rise in the use of social media, data related to
social networks is ripe for analysis using techniques from social network
analysis and graph theory. According to International Network for Social
Network Analysis, ‘Social network analysis is focused on uncovering the
patterning of people's interaction’.
Social network analysis (SNA) allows us to answer questions
such as who are key actors in a network? Who are the most influential
members of a network? Who seems to be acting on the peripheral? Which
connections in the network are most important? Are there key players
bridging connections or information between otherwise disconnected groups? Have
policies or other forces changed the overall dynamics/interaction between
people in the network (i.e. has the network structure changed in any meaningful
way) and does that relate to some other performance outcome or goal?
More specific applications of SNA may include Student
Integration and Persistence, Business to Business Supply Chains, Seeding Strategies for Viral Marketing, and Predicting Customer Churn. The open source software R
and NetDraw
provide many tools for conducting social
network analysis. See also ‘Using Twitter to Demonstrate Basic
Concepts from Social Network Analysis’ as well as ‘An
Introduction to Social Network Analysis Using R and Netdraw.’ As some of these examples demonstrate, measures derived from social network analysis
can be very useful in predictive modeling. For a more basic example see ‘Using
SNA in Predictive Modeling'.
For more information:
If you feel you can benefit from the services of a data
scientist or have further questions about applied econometrics and analytics
please contact me for more information or feel free to visit my blog or
selected works where you can find a copy of my CV.
LinkedIn Profile: (link)
Selected Works Profile: http://works.bepress.com/matt_bogard/
Monday, May 28, 2012
Notes to 'Support' an Understanding of Support Vector Machines
Some time ago I wrote a short article highlighting the similarities of tools used in economics, bioinformatics, and machine learning (Mathematical Themes in Economics, Machine Learning, and Bioinformatics). In this followup I expand more on the details 'supporting' support vector machines. While not appropriate for making the kinds of inferences econometricians are so often fond of, they are powerful tools for classification. They are certainly an artifact of the data mining culture. While I don't cover all of the details (for instance I don't discuss kernel methods, or duality) I cover some of the basic concepts that are likely to trip up someone new to SVMs, like how a line in R2 or hyperplane in higher dimensions can be expressed using inner products as in <w,x> + b = 0.
Matt Bogard, Western Kentucky University
Available at: http://works.bepress.com/matt_bogard/20
Notes on Support Vector Machines - link
Matt Bogard, Western Kentucky University
Abstract
The most basic idea of support vector machines is to find a line (or hyperplane) that separates classes of data, and use this information to classify new examples. But we don’t want just any line, we want a line that maximizes the distance between classes (we want the best line). This line turns out to be a separating hyperplane that is equidistant between the supporting hyperplanes that ‘support’ the sets that make up each distinct class. The notes that follow discuss the concepts of supporting and separating hyperplanes and inner products as they relate to support vector machines (SVMs). Using simple examples, much detail is given to the mathematical notation used to represent hyperplanes, as well as how the SVM classification works.
Suggested Citation
Matt Bogard. 2012. "Notes on Support Vector Machines" The Selected Works of Matt Bogard
Available at: http://works.bepress.com/matt_bogard/20
Saturday, May 5, 2012
An Intuitive Approach to Text Mining with SAS IML
Text Mining- in Plain English: Turning Text into Numbers
With Twitter, Facebook, email, online forums, open response surveys, customer and reader comments on web pages and news articles etc. there is a lot of information available to companies and organizations in the form of text. Without hiring experts to read through all of the thousands of pages worth of text available and making subjective claims about its meaning, text mining allows us to take otherwise unusable 'qualitative' data and convert it into quantitative measures that we can use for various types of reporting and modeling. In the example below I demonstrate how the mathematical technique of singular value decomposition (SVD) can be used to do this.
Sometimes to assess statistical techniques, or to even understand them at a basic level, you need data with properties you understand. With numerical data, that typically might call for simulation. But since we are dealing with text, I just made some up that will work to clearly demonstrate the effectiveness of SVD.
Below you will see 10 hypothetical comments to the question- 'What is your take on pink slime?' Each person's response is considered a document, and I have classified each document as type 'H' or 'S' as explained below. The goal will be to see if we can use a basic application of SVD to convert these comments into numbers and use them to predict what type of person made which type of comments (i.e. does a comment belong in category H or S) based on clustering or some type of predictive model. This gets way beyond simply classifying comments by doing a key word search. SVD allows us to not only classify documents by the specific words they contain, but also by how similar they are.
Singular Value Decomposition
SAS CODE
With Twitter, Facebook, email, online forums, open response surveys, customer and reader comments on web pages and news articles etc. there is a lot of information available to companies and organizations in the form of text. Without hiring experts to read through all of the thousands of pages worth of text available and making subjective claims about its meaning, text mining allows us to take otherwise unusable 'qualitative' data and convert it into quantitative measures that we can use for various types of reporting and modeling. In the example below I demonstrate how the mathematical technique of singular value decomposition (SVD) can be used to do this.
Sometimes to assess statistical techniques, or to even understand them at a basic level, you need data with properties you understand. With numerical data, that typically might call for simulation. But since we are dealing with text, I just made some up that will work to clearly demonstrate the effectiveness of SVD.
Below you will see 10 hypothetical comments to the question- 'What is your take on pink slime?' Each person's response is considered a document, and I have classified each document as type 'H' or 'S' as explained below. The goal will be to see if we can use a basic application of SVD to convert these comments into numbers and use them to predict what type of person made which type of comments (i.e. does a comment belong in category H or S) based on clustering or some type of predictive model. This gets way beyond simply classifying comments by doing a key word search. SVD allows us to not only classify documents by the specific words they contain, but also by how similar they are.
The purpose of this exercise is to provide intuition for text mining and the application of singular value decomposition. The text above is made up, and specifically designed to produce the results I'm after. I'm not making any claims one way or the other about how realistic this is. But suppose these are potential customers and we want to be able to distinguish between the hippies, who favor a more local nostalgic food supply from centuries past (designated as class or type 'H') and the animal scientists, designated as type 'S.' Obviously comments of class H are critical of modern agriculture and technology. We might want to make these distinctions for PR, marketing, lobbying,or educational outreach purposes.
After cleaning up the text (which many software programs like SAS Enterprise Miner provide excellent tools for doing so)by eliminating parts of speech, articles, etc. we can form a term-document frequency matrix as follows:
After cleaning up the text (which many software programs like SAS Enterprise Miner provide excellent tools for doing so)by eliminating parts of speech, articles, etc. we can form a term-document frequency matrix as follows:
Singular Value Decomposition
Singular Value Decomposition (SVD) is a concept from linear algebra
based on the following matrix equation:
A = USV’ which states that a rectangular matrix A can
be decomposed into 3 other matrix components:
U
consists of the orthonormal eigenvectors of AA’, where U’U
= I (recall from linear algebra, orthogonal vectors of unit length
are ‘orthonormal’)
V consists
of the orthonormal eigenvectors of A’A
S
is a diagonal matrix consisting of the square root of the eigenvalues of U
or V (which are equal). The values in S depict the variance of
linearly independent components along each dimension similarly to the way
eigenvalues depict variance explained by ‘factors’ or components in principle
components analysis (PCA).
SVD provides the mathematical foundation for text mining.
A term document
matrix A can be decomposed as in:
A = USV’
If the term document matrix A were a collection of individuals’ textual responses or comments
to a survey question (or Facebook post etc. ) then each individual’s response
would be considered a ‘document’. The individual words they used in their
response are the ‘terms.’ The term document frequency matrix consists of rows
that represent each term and columns that represent each ‘document’ or
individual.
The vectors in U can be used to score documents in the
term-document matrix, as in U’A. (this may be analogous to the way values of x
are scored by eigenvectors in PCA) As a
result, a single numerical value or ‘score’ can be assigned to each person’s
textual response (i.e. each document) for each SVD dimension (i.e. each kept
independent vector in U). Thus the
text is converted into numerical scores (via the transformation U’A) that can then be used in
predictive modeling (as predictor variables) or
clustering can be utilized to cluster the individual textual responses
(or ‘documents’). So ultimately SVD
converts otherwise unuseful text into numeric SVD scores or more
interpretable clusters. Likewise,
using the weights from the eigenvectors
that comprise V to score the terms in the term document matrix A, as in AV’ we can score cluster
like terms as ‘topics’.
Thus SVD of A gives allows us to derive the following
scores:
U’A = SVD
document vectors
AV’ = SVD term
vectors
This can be loosely demonstrated in using PROC IML in SAS. Specifying the term document frequency matrix in PROC IML and implementing SVD
produces:
The document vectors U’A can be depicted as follows:
Here is the original text document with customer
types/classes and the appended SVD scores.
As depicted above, SVD has allowed us to replace all of the text with the quantitative values for the associated SVD scores (derived from U'A). In this case, I'm representing all of the text with just two SVD dimensions. These values can then be used for clustering or entered into
a predictive model.
It is easy to see that the different documents types
(response types H vs S) cluster very well on the dimensions SVD1 and SVD2. In
fact, all responses of type S have a value of SVD2 < .5.
The SVD values can also be entered into a regression or
other type of predictive model.
Consistent with the observed relationship between SVD2
and class ‘S’ and ‘H’ clusters, we see that SVD2 is significant in the regression. In
addition, it shows that higher values of SVD2 decrease the probability of being
classified as document or response type ‘S’. Yes, this is OLS on a binary dependent variable, but again the purpose is to provide intuition and motivation for using text analytics for predictive modeling.
This seemed to work ok on a small collection of documents carefully constructed to illustrate the concepts above, but tools like SAS Text Miner in conjunction with SAS Enterprise Miner are designed specifically to do this type of analysis on a much larger scale. I have used both of these tools on much larger document collections and obtained promising results using text topics from SVD in predictive modeling applications.
SAS CODE
proc
iml;
A = {0 0 0 1 0 0 0 0 0 0,
1 1 0 0 0 1 0 0 0 0,
0 0 0 0 0 1 1 1 0 2,
0 0 0 0 0 1 0 0 0 0,
0 1 1 0 0 0 0 0 0 0,
1 0 0 0 0 0 0 0 0 0,
0 0 1 0 1 0 0 0 0 0,
0 0 0 0 1 0 0 0 0 0,
0 0 0 0 1 0 0 0 0 0,
0 0 1 0 0 0 0 0 0 0,
0 0 0 0 0 1 0 1 0 1,
0 0 1 0 0 0 0 0 0 0,
0 0 0 1 0 0 0 0 0 0,
0 0 0 0 0 0 1 1 0 0,
1 0 1 0 0 0 0 0 0 0,
0 0 0 0 0 0 1 0 0 0,
0 0 0 0 0 0 0 0 0 1,
0 0 0 0 0 1 0 0 0 0,
0 0 1 0 0 0 0 0 0 0,
1 0 0 0 1 0 0 0 0 0,
0 1 0 1 0 0 0 0 0 0,
0 1 0 0 0 1 0 0 0 0,
0 0 0 0 0 0 0 1 0 0,
0 0 0 0 0 0 0 0 0 1,
0 0 0 0 0 0 0 0 1 1,
0 0 0 0 0 0 0 0 0 1,
1 0 0 0 1 0 0 0 0 0,
0 0 0 0 0 0 1 1 1 0,
0 0 1 0 0 0 0 0 0 0,
0 0 0 0 0 0 0 0 1 0,
0 0 0 0 0 0 0 0 0 1,
0 0 0 0 0 1 0 1 0 0,
0 0 0 0 0 0 0 0 1 0
};
print(A);
/* print matrix A*/
n = nrow(A); /*
how many rows */
p = ncol(A); /*
how many columns */
print(n);
print(p);
call
svd(u,s,v,A); /* sigular value decomposition of A = usv'*/
print(u);
/* independent eigenvectors of AA' */
print(s);
/* independent
eigenvectors of A'A */
print(v);
/*singular values (sqrt(eigenvalues)) of AA' or A'A */
uTa = T(u)*A; /*
document vectors*/
avT = A*T(v); /*
term vectors */
print(uTa);
print(avT);
/* scoring a data set
*/
ID = {1,
2, 3,
4, 5,
6, 7,
8, 9,10};
/* create a document id matrix */
print(ID);
docscores =ID||T(uTa); /*
combine with svd results */
print(docscores);
/* export as a SAS data
set */
varnames = 'svd1':'svd10';
create
svd_scores from docscores[colname =
varnames];
append
from docscores;
close
svd_scores;
quit;
run;
*-------------------------------------------------------------------------------------*
|
SCORING, CLUSTERING AND PREDICTIVE MODELING
*-------------------------------------------------------------------------------------*;
* CREATE CLIENT
DATA SET;
DATA
CLIENTS;
INPUT
ID CLASS $ ;
CARDS;
1 H
2 H
3 H
4 H
5 H
6 S
7 S
8 S
9 S
10 S
;
RUN;
* MERGE WITH SVD
SCORES;
PROC
SQL;
CREATE
TABLE CLIENTS_SCORED AS
SELECT
A.ID, A.CLASS, B.SVD2 AS SVD1, B.SVD3 AS
SVD2
FROM
CLIENTS A LEFT JOIN
SVD_SCORES B
ON
A.ID = B.SVD1;
QUIT;
PROC
PRINT DATA
= CLIENTS_SCORED;
RUN;
* CLUSTER/
VISUALIZE CLIENTS BASED ON SVD SCORES;
PROC
GPLOT DATA
= CLIENTS_SCORED;
PLOT
SVD1*SVD2 = ID;
RUN;
QUIT;
* RECODE FOR
NUMERIC DEPENDENT VAR;
DATA
CLIENTS_SCORED2;
SET
CLIENTS_SCORED;
IF
CLASS = 'H' THEN
Y = 0;
ELSE
Y = 1;
RUN;
* BASIC
REGRESSION MODEL;
PROC
REG DATA =
CLIENTS_SCORED2;
MODEL
Y = SVD1 SVD2;
RUN;
QUIT;
Subscribe to:
Posts (Atom)
















