Showing posts with label analytics. Show all posts
Showing posts with label analytics. Show all posts

Wednesday, May 6, 2020

Experimentation and Causal Inference: Strategy and Innovation

Knowledge is the most important resource in a firm and the essence of organizational capability, innovation, value creation, strategy, and competitive advantage. Causal knowledge is no exception.In previous posts I have discussed the value proposition of experimentation and causal inference from both mainline and behavioral economic perspectives. This series of posts has been greatly influenced by Jim Manzi's book 'Uncontrolled: The Surprising Payoff of Trial-and-Error for Business, Politics, and Society.' Midway through the book Manzi highlights three important things that experimentation and causal inference in business settings can do:

1) Precision around the tactical implementation of strategy
2) Feedback on the performance of a strategy and refinements driven by evidence
3) Achievement of organizational and strategic alignment

Manzi explains that within any corporation there are always silos and subcultures advocating competing strategies with perverse incentives and agendas in pursuit of power and control. How do we know who is right and which programs or ideas are successful, considering the many factors that could be influencing any outcome of interest?  Manzi describes any environment where the number of causes of variation are enormous as an environment that has 'high causal density.' We can claim to address this with a data driven culture, but what does that mean? How do we know what is, and isn't supported by data? Modern companies in a digital age with AI and big data are drowning in data. This makes it easy to adorn rhetoric in advanced analytical frameworks. Because data seldom speaks, anyone can speak for the data through wily data story telling.  Decision makers fail to make the distinction between just having data, and having evidence to support good decisions.

As Jim Manzi and Stefan Thomke discuss in Harvard Business Review:

"business experiments can allow companies to look beyond correlation and investigate causality....Without it, executives have only a fragmentary understanding of their businesses, and the decisions they make can easily backfire."

Without experimentation and causal inference, there is know way to connect the things we do with the value created. In complex environments with high causal density, we don't know enough about the nature and causes of human behavior, decisions, and causal paths from actions to outcomes to list them all and measure and account for them even if we could agree how to measure them. This is the nature of decision making under uncertainty. But, as R.A. Fisher taught us with his agricultural experiments, randomized tests allow us to account for all of these hidden factors (Manzi calls them hidden conditionals). Only then does our data stand a chance to speak truth. Experimentation and causal inference don't provide perfect information but they are the only means by which we can begin to say that we have data and evidence to inform the tactical implementation of our strategy as opposed to pretending that we do based on correlations alone. As economist F.A. Hayek once said:

"I prefer true but imperfect knowledge, even if it leaves much undetermined and unpredictable, to a pretense of exact knowledge that is likely to be false"

In Dual Transformation: How to Reposition Today's Business While Creating the Future authors discuss the importance of experimentation and causal inference as a way to navigate uncertainty in causally dense environments in what they refer to as transformation B:

“Whenever you innovate, you can never be sure about the assumptions on which your business rests. So, like a good scientist, you start with a hypothesis, then design and experiment. Make sure the experiment has clear objectives (why are you running it and what do you hope to learn). Even if you have no idea what the right answer is, make a prediction. Finally, execute in such a way that you can measure the prediction, such as running a so-called A/B test in which you vary a single factor."

Experiments aren't just tinkering and trying new things. While these are helpful to innovation, just tinkering and observing still leaves you speculating about what really works and is subject to all the same behavioral biases and pitfalls of big data previously discussed.

List and Gneezy address this in The Why Axis:

"Many businesses experiment and often...businesses always tinker...and try new things...the problem is that businesses rarely conduct experiments that allow a comparison between a treatment and control group...Business experiments are research investigations that give companies the opportunity to get fast and accurate data regarding important decisions."

Three things distinguish experimentation and causal inference from just tinkering:

1) Separation of signal from noise (statistical inference)
2) Connecting cause and effect  (causal inference)
3) Clear signals on business value that follows from 1 & 2 above

Having causal knowledge helps identify more informed and calculated risks vs. risks taken on the basis of gut instinct, political motivation, or overly optimistic and behaviorally biased data-driven correlational pattern finding analytics. 

Experimentation and causal inference add incremental knowledge and value to business. No single experiment is going to be a 'killer app' that by itself will generate millions in profits. But in aggregate the knowledge created by experimentation and causal inference probably offers the greatest strategic value across an enterprise compared to any other analytic method.

As discussed earlier, experimentation and causal inference creates value by helping manage the knowledge problem within firms, it's worth repeating again from List and Gneezy:

"We think that businesses that don't experiment and fail to show, through hard data, that their ideas can actually work before the company takes action - are wasting their money....every day they set suboptimal prices, place adds that do not work, or use ineffective incentive schemes for their work force, they effectively leave millions of dollars on the table."

As Luke Froeb writes in Managerial Economics, A Problem Solving Approach (3rd Edition):

"With the benefit of hindsight, it is easy to identify successful strategies (and the reasons for their success) or failed strategies (and the reason for their failures). It's much more difficult to identify successful or failed strategies before they succeed or fail."

Again from Dual Transformation:

"Explorers recognize they can't know the right answer, so they want to invest as little as possible in learning which of their hypotheses are right and which ones are wrong"

Experimentation and causal inference offer the opportunity to test strategies early on a smaller scale to get causal feedback about potential success or failure before fully committing large amounts of irrecoverable resources. They allow us to fail smarter and learn faster. Experimentation and causal inference play a central role in product development, strategy, and innovation across a range of industries and companies like Harrah's casinos, Capital One, Petco, Publix, State Farm, Kohl's, Wal-Mart, and Humana who have been leading in this area for decades in addition to new ventures like Amazon and Uber. 

"At Uber Labs, we apply behavioral science insights and methodologies to help product teams improve the Uber customer experience. One of the most exciting areas we’ve been working on is causal inference, a category of statistical methods that is commonly used in behavioral science research to understand the causes behind the results we see from experiments or observations...Teams across Uber apply causal inference methods that enable us to bring richer insights to operations analysis, product development, and other areas critical to improving the user experience on our platform." - From: Using Causal Inference to Improve the Uber User Experience (link)

Economist Joshua Angrist explains about his students that have went on to work for companies like Amazon: "when I ask them what are they up to they say...we're running experiments."

Achieving the greatest value from experimentation and causal inference requires leadership commitment.  It also demands a culture that is genuinely open to learning through a blend of trial and error, data driven decision making informed by theory and experiments, and the infrastructure necessary for implementing enough tests and iterations to generate the knowledge necessary for rapid learning and innovation. It requires business leaders, strategists, and product managers to think about what they are trying to achieve and asking causal questions to get there (vs. data scientists sitting in an ivory tower dreaming up models or experiments of their own). The result is a corporate culture that allows an organization to formulate, implement, and modify strategy faster and more tactfully than others.

See also:
Experimentation and Causal Inference: The Knowledge Problem
Experimentation and Causal Inference: A Behavioral Economics Perspective
Statistics is a Way of Thinking, Not a Box of Tools

Monday, September 30, 2019

Wicked Problems and The Role of Expertise and AI in Data Science

In 2018, an article in Science characterized the challenge of pesticide resistance as a wicked problem:

“If we are to address this recalcitrant issue of pesticide resistance, we must treat it as a “wicked problem,” in the sense that there are social, economic, and biological uncertainties and complexities interacting in ways that decrease incentives for actions aimed at mitigation.”

In graduate school, I worked on this same problem, attempting to model the social and economic systems with game theory and behavioral economics and capturing biological complexities leveraging population genetics. 

Wicked vs. Kind Environments

In data science, we also have 'wicked' learning environments in which we try to train our models. In the EconTalk podcast with Russ Roberts, Mastery, Specialization, and Range, David Epstein discusses wicked and kind learning environments:

"The way that chess works makes it what's called a kind learning environment. So, these are terms used by psychologist Robin Hogarth. And what a kind learning environment is, is one where patterns recur; ideally a situation is constrained--so, a chessboard with very rigid rules and a literal board is very constrained; and, importantly, every time you do something you get feedback that is totally obvious...you see the consequences. The consequences are completely immediate and accurate. And you adjust accordingly. And in these kinds of kind learning environments, if you are cognitively engaged you get better just by doing the activity."

"On the opposite end of the spectrum are wicked learning environments. And this is a spectrum, from kind to wicked. Wicked learning environments: often some information is hidden. Even when it isn't, feedback may be delayed. It may be infrequent. It may be nonexistent. And it maybe be partly accurate, or inaccurate in many of the cases. So, the most wicked learning environments will reinforce the wrong types of behavior."

As discussed in the podcast, many problems fall within some spectrum ranging between very kind environments like Chess to more complex environments like self driving cars or medical diagnosis. What do experts have to offer where AI/ML falls short? The type of environment determines to a great extent the scope of disruption we might be able to expect from AI applications.

The Role of Human Expertise

In Thinking Fast and Slow, Kahneman discusses two conditions for acquiring skill:

1) an environment that is sufficiently regular to be predictable
2) an opportunity to learn these regularities through prolonged practice

This sounds a lot like the 'kind' environments discussed above. Based on research by Robin Hogarth, Kahneman also makes these distinctions describing 'wicked' environments as those environments in which those with expertise are likely to learn the wrong lessons from experience. The problem is that with wicked environments, experts often default to heuristics which can lead to wrong conclusions. Even if aware of these biases, social norms often nudge experts into the wrong direction. Kahneman gives an example involving physicians:

"Generally it is considered a weakness and a sign of vulnerability for clinicians to appear unsure. Confidence is valued over uncertainty and there is a prevailing censure against disclosing uncertainty to patients...acting on pretended knowledge is often the preferred solution."

This likely explains many of the mistakes and low value care that are problematic with healthcare delivery as well as dissatisfaction with both the quality and costs of healthcare. How many of us want our physicians to pretend to know what they are talking about? On the other hand, how many people are willing to accept an answer from their physician that rhymes with "let me look this up and get back to you later." 

One advantage AI may have over experts in kind environments is as Kahneman puts it, the opportunity to learn through prolonged practice. Machine learning can handle many more training examples than a human so to speak.

Even in kind environments, an expert may swing and miss when dealing with cases where the correct decision is like a pitch straight over the plate. One reason Kahneman discusses in Thinking Fast and Slow is the idea of 'ego' depletion. This is related to the idea that mental energy can become exhausted after significant exertion. As self-control breaks down, its easy to default to heuristics and biases that can lead to decisions that look like careless mistakes. This would certainly apply to physicians given the number of stories we hear about burnout in the profession. 

The solution seems to be what polymath economist Tyler Cowen suggested several years ago in the econtalk podcast discussion he had about his book Average is Over with Russ Roberts:

"I would stress much more that humans can always complement robots. I'm not saying every human will be good at this. That's a big part of the problem. But a large number of humans will work very effectively with robots and become far more productive, and this will be one of the driving forces behind that inequality."

Imagine the clinical situation where a physician's 'ego' is substantially depleted from a difficult case. They could then lean on AI to prevent mistakes treating more routine decisions that follow. Or perhaps leveraging AI tools, a clinician could conserve additional mental energy throughout the day so that they are less likely to default to heuristics when they encounter more complex issues. The way this synergy materializes is uncertain, but it will certainly continue to involve substantial expertise on the part of many professionals going forward. Together human expertise and AI might have the greatest chance tackling the most wicked problems.

References:

Wicked evolution: Can we address the sociobiological dilemma of pesticide resistance? | Science  https://science.sciencemag.org/content/360/6390/728.full

Thinking Fast and Slow. Daniel Kahneman. 2011

EconTalk:David Epstein on Mastery, Specialization, and Range
https://www.econtalk.org/david-epstein-on-mastery-specialization-and-range/

EconTalk: Tyler Cowen on Inequality, the Future, and Average is Over
https://www.econtalk.org/tyler-cowen-on-inequality-the-future-and-average-is-over/

Tuesday, June 6, 2017

Professional Science Master's Degree Programs in Biotechnology and Management

As an undergraduate I always had an interest in biotechnology and molecular genetics. However, lab work did not particularly appeal to me. I also recognized early on that science does not occur in a vacuum- its subject to social, political, economic, and financial forces. This drew me to the field of economics, specifically public choice theory.

When it came time for graduate school I was still torn. I really wasn't interested in an MBA and didn't really have the background to work in a lab or do field work in genetic research. I really liked economics. The combination of mathematically precise theories (microeconomics/game theory) and empirically sound methods (econometrics) provided a powerful framework for applied problem solving.

I had two advisers make recommendations that got me thinking outside the box. One suggested ultimately I would find a niche that combined both economics and genetics. The other suggested I look at programs like the Bioscience Management program that was being offered at the time at George Mason University (now Bioinformatics Management). While there were not a lot of programs like that being offered at the time, the Agriculture Department at Western Kentucky University provided enough flexibility in their masters program to include courses in biostatistics, genetics, and applied economics. I was able to work on research projects analyzing consumer perceptions of biotechnology and biotech trait resistance management using tools from econometrics, game theory, and population genetics.  Additionally I took courses in applied economics and finance from both the Department of Agriculture and College of Business where I was exposed to tools related to investment analysis, options pricing, and analysis and valuation of biotech companies as well as the impacts of technological change and biotechnology on food and economic development.

With this combination of quantitative training and applied work I have been able to leverage SAS, R, and Python to solve a number of challenging problems throughout a number of professional analytics and consulting roles. 

Today there are a larger number of professional science masters programs with curriculums similar to the programs I contemplated over 10 years ago. 

According to National Professional Science Master’s Association:

"Professional Science Master's (PSMs) are designed for students who are seeking a graduate degree in science or mathematics and understand the need for developing workplace skills valued by top employers. A perfect fit for professionals because it allows you to pursue advanced training and excel in science or math without a Ph.D., while simultaneously developing highly-valued business skills....PSM programs consist of two years of coursework along with a professional component that includes business, communications and/or regulatory affairs."

In 2012 there was an article in Science detailing these degrees and some data related to salaries which seemed attractive. According to the article the first program was officially offered in 1997, reaching 140 programs by 2009 with over 247 at the time of printing.

This commentary from the article corroborates how I feel about my experience:

“There is a tendency for students to buy into the line that if you don't get a Ph.D., you're not a serious professional, that you're wasting your mind,” she says. After spending a decade talking with PSM students and graduates, she is certain that’s not true. “There is so much potential for growth and satisfaction with a PSM degree. You can become a person you didn’t even know you wanted to be.”

Below are some programs that would look interesting to me that students interested in this option should check out.  (there is a program locator you can find here) . Many of these programs are a mash up of biology/biotech and applied economics and business degrees.

George Mason University- PSM Bioinformatics Management

University of Illinois - Agricultural Production

Cornell- MPS Agriculture and Life Sciences

Washington State University - PSM Molecular Biosciences

Middle Tennesee State University - PSM Biotechnology

California State - MS Biotechnology/MBA 

Johns Hopkins - MBA/MS Biotechnology

Rice - PSM Bioscience and Health Policy

North Carolina State University - MBA (Biosciences Mgt Concentration)

Purdue/Kelley - MS-MBA  (not a heavy science emphasis but a very cool degree regardles from great schools)

See also:
Analytical Translators
Why Study Agricultural/Applied Economics

Saturday, June 3, 2017

In Praise of The Citizen Data Scientist

There was actually a really good article I read over at Data Science Central titled "The Data Science Delusion." Here is an interesting slice:

"This democratization of algorithms and platforms, paradoxically, has a downside: the signaling properties of such skills have more or less been lost. Where earlier you needed to read and understand a technical paper or a book to implement a model, now you can just use an off-the-shelf model as a black-box. While this phenomenon affects many disciplines, the vague and multidisciplinary definition of data science certainly exacerbates the problem."

It is true there is some loss of signal. However, companies may need to look for new signals as technological change progresses and new forms of capital complements labor. Its this new labor complementing role of capital (in the form of open source statistical computing packages and computing power) that is creating demand for those that can leverage these tools competently, without knowing all  "the nitty-gritty mathematical academic formulas to everything about support vector machines or Kernels and stuff like that to apply it properly and get results."

Sure, as a result there are a lot of analytics programs popping up out there to take advantage of these advances, but its also the reason programs like applied economics are becoming so popular.  In fact, in promoting its program, Johns Hopkins University almost seems to echo some of the sentiment in the quotes above, but takes a positive spin:

"Economic analysis is no longer relegated to academicians and a small number of PhD-trained specialists. Instead, economics has become an increasingly ubiquitous as well as rapidly changing line of inquiry that requires people who are skilled in analyzing and interpreting economic data, and then using it to effect decisions about national and global markets and policy, involving everything from health care to fiscal policy, from foreign aid to the environment, and from financial risk to real risk." 

In fact, I admit for a while I was a little disappointed my alma mater did not embrace the data science/analytics degree trend, or offer more courses in applied programming or incorporate languages like R into more courses. However, now, while I think these things are great I realize the more important data science skills are related to the analytical thinking and firm theoretical, statistical, and quantitative foundations that programs in economics and finance already offer at the undergraduate and masters level. While formal data science training might be the way of the future, I would venture to say that the vast majority of today's 'data scientists' were academically trained in a quantitative discipline like the above and self trained (perhaps via coursera etc.) on the skills and tools most people think of when they think of data science.  As I have said before, sometimes you don't need someone with a PhD in computer science or an astrophysics. Sometimes you really just need a good MBA that understands regression and the basics of a left join.

The DSC article above concludes with a little jab at data science, that I tend to agree with wholeheartedly:

"Great data science work is being done in various places by people who go by other names (analyst, software engineer, product head, or just plain old scientist). It is not necessary to be a card-carrying data scientist to do good data science work. Blasphemy it may be to say so, but only time will tell whether the label itself has value, or is only helping create a delusion." 

See also:

What you really need to know to be a data scientist
Super Data Science podcast - credit scoring
How to think like a data scientist to become one
What makes a great data scientist
Are data scientists going extinct
More on data science from actual data scientists



Thursday, October 13, 2016

Why Data Science Needs Economics

Cleaning out my inbox recently, I ran across an article from 2015. Below is a link and excerpt:


http://dataconomy.com/4-predictions-for-big-data-in-2015-from-industry-leaders/


4. Data Science Will Belong to the Economists

"We will start to see data science (to the extent that it operates as a coherent entity) increasingly rely on the domain expertise of economists. The early days of data science were very math, statistics and programming oriented. Then there was the rise of the “computational social scientist,” which added sociology to the mix. Many trend setting data science places are finding that sociology, and similar disciplines, tend to be retrospective, while other fields, like economics, offer simulation and auction modeling and other techniques to get more proactive and predictive with data. Of course, most economists don’t have the programming chops to land most data science jobs, but I think we’ll see that start to change significantly."
 
I think coding is important...and it would be nice if students interested in a career in data science could get more exposure to coding (SQL/R/SAS/Python etc.) in the classroom as well as algorithmic approaches (decision trees, neural networks etc.). However, I think its more important to have the analytical thinking skills and grounding in statistical inference that they get from an economics program (both UG and GR). That's the skillset I think will differentiate the data scientists in the future from the very technical tools focused ones in demand today. 

Recently on EconTalk, Russ Roberts and Cathy O'Neil discuss her book Weapons of Math Destruction and they take on issues related to explaining vs predicting, causality vs fitting the data. (see also their previous episode with Susan Athey). The role of quasi-experimental methods and rigorous identification as well as theory was emphasized. And theory is something, having spent the better part of my career focusing on empirical methods (both causal inference and machine learning) that I have not given enough thought to until recently. But the more I think about it....the more I realize it is necessary. Can big data and algorithms deliver tighter, unbiased, and more truthful insights? This excerpt from the 10th edition of Heyne, Boettke, and Pryschitko's The Economic Way of Thinking leads me to think exactly the opposite:

"We can observe facts, but it takes a theory to explain the causes. It takes a theory to weed out the irrelevant facts from the relevant ones."

And they give an anecdote:

"although the facts clearly show that most pot smokers were former milk drinkers, milk drinking probably is not a relevant fact in explaining pot smoking; similarly, the Superbowl is likely irrelevant when explaining Wall Street Interactions"(even if the data does show that the Dow does well when an NFC team does well)."

And more about theory:


"Our observations of the world are in fact drenched with theory, which is why we can usually make sense out of the buzzing confusion that assaults our eyes and ears. Actually we observe only a small fraction of what we "know," a hint here and a suggestion there. The rest we fill in from the theories we hold: small and broad, vague and precise..."

Big data in many ways is buzzing confusion, and yes algorithmic approaches i.e machine learning can help us find patterns and relationships that can be useful. But relying totally on a data driven process devoid of theory is more often going to lead us down the wrong path depending on the questions we are trying to answer. Economics is a way of thinking and economic theory can help us make sense of what we find, it can help us ask better or important questions, and can help guide us to understand the answers to those questions. It is forward looking as the article above states.

Of course we can test theories using data, through some clever identification strategy or even employing methods from machine learning in conjunction with conventional econometric approaches. And this brings me back full circle to my previous post about what the most important skillsets for data scientists may be going forward, and how economics training, and in fact economic theory can help fill that niche.

See also:

Are Data Scientitsts Going Extinct
To Explain or Predict
Economists as Data Scientists
Why Study Agricultural and Applied Economics
Analytics vs Causal Inference
Culture War: Inferential Statistics vs Machine Learning
Big Data: Don't throw the baby out with the bathwater
Causal Inference and Quasi-Experimental Design Roundup 
Big Data: Causality and Local Expertise are Key in Agronomic Applications 
Data Scientists vs Algorithms vs Real Solutions to Real Problems 
Analytical Translators


Sunday, August 7, 2016

The State of Applied Econometrics-Imbens and Athey on Causality, Machine Learning, and Econometrics

I recently ran across:

The State of Applied Econometrics - Causality and Policy Evaluation
Susan Athey, Guido Imbens

https://arxiv.org/abs/1607.00699v1 

A nice read, although I skipped directly to the section on machine learning. A few interesting causality/machine learning comments.

They discussed some known issues related to estimating propensity scores using various machine learning algorithms in terms of the sensitivity of results, especially for propensity scores close to 0 or 1. They discuss trimming weights as one possible approach, which I have heard before in Angrist and Pischke and other work (see below). In fact, in a working paper where I employed gradient boosting to estimate propensity scores for IPTW regression, I trimmed weights. However, I did not trim them for the stratified matching estimator that I also used. I wish I still had the data because I would like to see the impact on my previous results.

Another interesting application discussed in this paper was a two (or 3?) stage LASSO estimation (they actually have a great overall discussion of penalized regression and regularization in machine learning) where they mention first running LASSO to select variables related to the outcome of interest, second running LASSO to select for variables related to selection, and finally running OLS to estimate a causal model that includes the selected variables from the previous LASSO methods.

The paper covers a range of other topics including decision trees, random forests, distinctions between traditional econometrics and machine learning, instrumental variables etc.

Some Additional Notes and References:

Multiple Algorithms (CART/Logistic Regression/Boosting/Random Forests) with PS weights and trimming:
http://econometricsense.blogspot.com/2013/04/propensity-score-weighting-logistic-vs.html

Following Angrist and Pischke I present results for regressions utilizing data that has been 'screened' by eliminating observations where ps > .90 or < .10 using the r 'matchit' package

http://econometricsense.blogspot.com/2015/03/using-r-matchit-package-for-propensity.html


Estimating the Causal Effect of Advising Contacts on Fall to Spring Retention Using Propensity Score Matching and Inverse Probability of Treatment Weighted Regression

Matt Bogard, Western Kentucky University

Abstract

In the fall of 2011 academic advising and residence life staff working for a southeastern university utilized a newly implemented advising software system to identify students based on attrition risk. Advising contacts, appointments, and support services were prioritized based on this new system and information regarding the characteristics of these interactions was captured in an automated format. It was the goal of this study to investigate the impact of this advising initiative on fall to spring retention rates. It is a challenge on college campuses to evaluate interventions that are often independent and decentralized across many university offices and organizations. In this study propensity score methods were utilized to address issues related to selection bias. The findings indicate that advising contacts associated with the utilization of the new software had statistically significant impacts on fall to spring retention for first year students on the order of a 3.26 point improvement over comparable students that were not contacted.

Suggested Citation

Matt Bogard. 2013. "Estimating the Causal Effect of Advising Contacts on Fall to Spring Retention Using Propensity Score Matching and Inverse Probability of Treatment Weighted Regression" The SelectedWorks of Matt Bogard
Available at: http://works.bepress.com/matt_bogard/25

Saturday, March 5, 2016

Machine Learning and Econometrics

Not long ago Tyler Cowen blogged at Marginal Revolution about a Quora post by Susan Athey discussing the impact of machine learning on econometrics, flavors of machine learning, and differences in the emphasis placed on tools and methodologies traditional in each field. The differences often hinge on whether one's intention is to explain or predict,  or if one is interested in causal inference vs analytics. I really liked the point about instrumental variables made in the snippet below:

"Yet, a cornerstone of introductory econometrics is that prediction is not causal inference, and indeed a classic economic example is that in many economic datasets, price and quantity are positively correlated.  Firms set prices higher in high-income cities where consumers buy more; they raise prices in anticipation of times of peak demand. A large body of econometric research seeks to REDUCE the goodness of fit of a model in order to estimate the causal effect of, say, changing prices. If prices and quantities are positively correlated in the data, any model that estimates the true causal effect (quantity goes down if you change price) will not do as good a job fitting the data….Techniques like instrumental variables seek to use only some of the information that is in the data – the “clean” or “exogenous” or “experiment-like” variation in price—sacrificing predictive accuracy in the current environment to learn about a more fundamental relationship that will help make decisions about changing price. This type of model has not received almost any attention in ML."

Tyler also points to a wealth of resources by Suan Athey here. And check out the mini-course she taught with Guido Imbens via NBER.

The differences and synergies between tools used in both econometrics and machine learning is something I have been interested in for a long time and have blogged about several times in the past. Kenneth Sanford and Hal Varian have also been writing about this as well. See related content below.

Related Content and Further Reading

Economists as Data Scientists http://econometricsense.blogspot.com/2012/10/economists-as-data-scientists.html

Econometrics, Math, and Machine Learning….what? http://econometricsense.blogspot.com/2015/09/econometrics-math-and-machine.html 

"Mathematical Themes in Economics, Machine Learning, and Bioinformatics" (2010)
Available at: http://works.bepress.com/matt_bogard/7/ 

Notes to 'Support' an Understanding of Support Vector Machines  http://econometricsense.blogspot.com/2012/05/notes-to-support-understanding-of.html

Culture War: Classical Statistics vs. Machine Learning http://econometricsense.blogspot.com/2011/01/classical-statistics-vs-machine.html

Analytics vs Causal Inference http://econometricsense.blogspot.com/2014/01/analytics-vs-causal-inference.html

Big Data: Don’t throw the baby out with the bath water http://econometricsense.blogspot.com/2014/05/big-data-dont-throw-baby-out-with.html

To Explain or Predict http://econometricsense.blogspot.com/2015/03/to-explain-or-predict.html 

Big Data: Causality and Local Expertise Are Key in Agronomic Applications http://econometricsense.blogspot.com/2014/05/big-data-think-global-act-local-when-it.html


Big Data:  New Tricks for Econometrics
Hal R. Varian
June 2013
Revised:  April 14, 2014
http://people.ischool.berkeley.edu/~hal/Papers/2013/ml.pdf

Is machine learning trending with economists? (Kenneth Sanford)  http://blogs.sas.com/content/subconsciousmusings/2015/06/05/is-machine-learning-trending-with-economists/

Wednesday, September 30, 2015

Big Data, IoT, Ag Finance, and Causal Inference

Over at my applied economics blog, I recently discussed an article from AgWeb; How the feds interest rate decision affects farmers. This actually got me questioning some of the ramifications of leveraging data analysis in the context of ag lending (from both a farmer and lender perspective), which ultimately lead to me thinking about some interesting questions that would be exciting to investigate:
  1.  Is there a causal relationship between producers that leverage IoT and Big Data analytics applications and farm output/performance/productivity
  2. How do we quantify the outcome-is it some measure of efficiency or some financial ratio?
  3. If we find improvements in this measure-is it simply a matter of selection? Are great producers likely to be productive anyway, with or without the technology?
  4. Among the best producers, is there still a marginal impact (i.e. treatment effect) for those that adopt a technology/analytics based strategy?
  5. Can we segment producers based on the kinds of data collected by IoT devices on equipment, aps, financial records, GPS etc.?  (maybe this is not that much different than the TrueHarvest benchmarking done at FarmLink) and are there differentials in outcomes, farming practices, product use patterns etc. by segment
See also:
Big Ag Meets Big Data (Part 1 & Part 2)
Big Data- Causality and Local Expertise are Key in Agronomic Applications
Big Ag and Big Data-Marc Bellemare
Other Big Data and Agricultural related Application Posts at EconometricSense
Causal Inference and Experimental Design Roundup

Saturday, September 5, 2015

Econometrics, Math, and Machine Learning...what?

In a recent Bloomberg View piece, Noah Smith wrote a piece titled "Economics has a Math Problem" that has caught a lot of attention lately.  There were three interesting arguments or subjects I found interesting in the piece.

#1 In economics, theory often takes a unique role in the determination of causality

"In most applied math disciplines -- computational biology, fluid dynamics, quantitative finance -- mathematical theories are always tied to the evidence. If a theory hasn’t been tested, it’s treated as pure conjecture....Not so in econ. Traditionally, economists have put the facts in a subordinate role and theory in the driver’s seat. "

This alone might seem controversial to some, but to many economists, causality is a theory driven phenomenon, and can never truly be determined by data. I won't expand on this any further. But the point is that often, economists, outside of a purely predictive or forecasting scenario, are interested in answering causal questions, and despite all the work since the credibility revolution in terms of quasi-experimental designs, theory still plays an important role in determine causality and the direction of effects.

 #2 In economics and econometrics, there is a huge emphasis on explaining causal relationships, both theoretically and empirically, but in machine learning the emphasis is prediction, classification, and pattern recognition devoid of theory or data generating processes

"Machine learning is a broad term for a collection of statistical data analysis techniques that identify key features of the data without committing to a theory. To use an old adage, machine learning “lets the data speak.”…machine learning techniques emphasized causality less than traditional economic statistical techniques, or what's usually known as econometrics. In other words, machine learning is more about forecasting than about understanding the effects of policy."

That really gets at what I have written before, about machine learning vs classical inference. (If Noah's article is interesting to you, then I highly recommend the Leo Brieman paper I reference in that post). Its true, at first it might seem that most economists interested in causal inference might sideline machine learning methods for their lack if emphasis on identification of causal effects or a data generating process. One of the biggest differences between the econometric theory most economists have been trained in and the new field of data science is in effect familiarity and use of methods from machine learning. But if they are interested strictly in predictive modeling and forecasting, these methods might be quite appealing. (I've argued before that economists are ripe for being data scientists). As we know, the methods and approaches we take to analyzing our data differ substantially depending on whether we are trying to explain vs. predict.

But then things start to get interesting:

#3 Recent work in econometrics has narrowed the gap between machine learning and econometrics

"But Athey and Imbens have also studied how machine learning techniques can be used to isolate causal effects, which would allow economists to draw policy implications."

I have not actually drilled into the references and details around this but it is interesting. Just thinking about it a little, I recalled that not long ago I worked on a project where I used gradient boosting (a machine learning algorithm) to estimate propensity scores to estimate treatment effects associated with a web ap.

Even one of the masters of metrics and causal inference, Josh Angrist is offering a course titled "Applied Econometrics:Mostly Harmless Big Data" via the MIT open course platform. And for a long time, economist Kenneth Sanford has been following this trend of emphasis on data science and machine learning in econometrics.

Overall, I think it will be interesting to see more examples of applications of machine learning in causal inference. But, when these applications involve big data and the internet of things, economists will really have to test their knowledge of a range of other big data tools that have little to do with building models or doing calculations.

See also:
Analytics vs Causal Inference
Big Data: Don't throw the baby out with the bath water
Propensity Score Weighting: Logistic vs CART vs Boosting vs Random Forests 
Data Cleaning
Got Data? Probably not like your econometrics textbook!
In God we trust, all others show me your code.
Data Science, 10% inspiration, 90% perspiration
Related:
 Big Ag Meets Big Data (Part 1 & Part 2)

Friday, July 3, 2015

Analytical Translators

A recent Deloitte Press article discusses a role that is becoming more and more important as a result of the explosion of big data and data science in industry. The article In praise of “light quants” and “analytical translators” discussed the important role of analtyical translators, which may be even harder to find than actual data scientists themselves.

“When we think about the types of people who make analytics and big data work, we typically think of highly quantitative or computational folks with hard knowledge and skills. You know the usual suspects: data scientists who can make Hadoop jump through hoops, statisticians who dream in SAS or R, data wizards who can extract two years of data from a medical device that normally dumps it after 20 minutes (a true request)....A “light quant” is someone who knows something about analytical and data management methods, and who also knows a lot about specific business problems. The value of the role comes, of course, from connecting the two."

Actually, in some of my previous ponderings and speculations about the coming convergence of big data, analytics, and genomics, I discussed the potential for such a role in the precision agriculture and data science space (See Big Data: Causality and Local Expertise Are Key in Agronomic Applications):

"as we think about all the data that can potentially be captured through the internet of things from seed choice, planting speed, depth, temperature, moisture, etc this could become especially important. This might call for a much more personal service including data savvy reps to help agronomists and growers get the most from these big data apps or the data that new devices and software tools can collect and aggregate.  Data savvy agronomists will need to know the assumptions and nature of any predictions or analysis, or data captured by these devices and apps to know if surrogate factors like Dan mentions have been appropriately considered. And agronomists, data savvy or not will be key in identifying these kinds of issues.  Is there an ap for that? I don't think there is an automated replacement for this kind of expertise, but as economist Tyler Cowen says, the ability to interface well with technology and use it to augment human expertise and judgement is the key to success in the new digital age of big data and automation. "

In fact, I recently discovered an actual position for a major player in this space with the title "BioAg Knowledge Transfer Agronomist" that seems to fit the bill. I think we will see more roles like this in the future.

Related:
Farm Link: The Rise of Data Science in Agriculture http://econometricsense.blogspot.com/2015/06/farmlink-and-rise-of-data-science-in.html

The Internet of Things, Big Data, and John Deere http://www.econometricsense.blogspot.com/2015/01/the-internet-of-things-big-data-and.html

The Use of Knowledge (in a) Big Data Society http://ageconomist.blogspot.com/2015/07/the-use-of-knowledge-in-big-data-society.html


Friday, June 19, 2015

Got Data? Probably not like your econometrics textbook!

Recently there has been a lot of discussion of the Angrist and Pischke piece entitled "Why Econometrics Teaching Needs an Overhaul." (read more...) and I have discussed before the large gap between theoretical and applied econometrics.

But here I plan to discuss another potential gap in teaching and application and this is a topic that often is not introduced at any point in a traditional undergraduate or graduate economics curriculum, and that is hacking skills. This becomes extremely important for economists that someday find themselves doing applied work in a corporate environment, or working in the area of data science. Drew Conway points out the there are three spheres of data science including hacking skills, math and statistics knowledge, and subject matter expertise. For many economists, the hacking sphere might be the weakest (read also Big Data Requires a New Kind of Expert: The Econinformatrician) while their quantitative training otherwise makes them ripe to become very good data scientists.

Drew Conway's Data Science Venn Diagram

In a recent whitepaper, I discuss this issue:

Students of econometrics might often spend their days learning proofs and theorems, and if they are lucky they will get their hands on some data and access to software to actually practice some applied work rather it be for a class project or part of a thesis or dissertation. I have written before about the large gap between theoretical and applied econometrics, but there is another gap to speak of, and it has nothing to do with theoretical properties of estimators or interpreting output from STATA, SAS or R. This has to do with raw coding, hacking, and data manipulation skills; the ability to tease out relevant observations and measures from both large structured transactional databases or unstructured log files or web data like tweet-streams. This gap becomes more of an issue as econometricians move from more academic environments to corporate environments and especially so for those economists that begin to take on roles as data scientists. In these environments, not only is it true that problems don’t fit the standard textbook solutions (see article ‘Applied Econometrics’), but the data doesn't look much like the simple data sets often used in textbooks either.  One cannot always expect their IT people to be able to just dump them a flat file with all the variables and formats that will work for your research project. In fact, the absolute best you might hope for in many environments is a SQL or Oracle data base with hundreds or thousands of tables and the tiny bits of information you need spread across a number of them. How do you bring all of this information together to do an analysis? This can be complicated, but for the uninitiated I will present some ‘toy’ examples to give a feel for executing basic database queries to bring together different pieces of information housed in separate tables in order to produce a ‘toy’ analytics ready data set.

I am certain that many schools actually do teach some of the basics related to joining and cleaning data sets, and if they don't then others might figure this out on the job or through one research project or another. I am not certain that this gap needs to be filled  necessarily as part of any econometrics course. However, it is something students need to be aware of and offering some sort of workshop, lab or formal course (maybe as part of a more comprehensive data science curriculum like this) would be very beneficial.

Read the whole paper here:

Matt Bogard. 2015. "Joining Tables with SQL: The most important econometrics lesson you may ever learn" The SelectedWorks of Matt Bogard
Available at: http://works.bepress.com/matt_bogard/29  

See also: Is Machine Learning Trending with Economists?

Wednesday, June 17, 2015

Farmlink and the Rise of Data Science in Agriculture


At a recent Global Ag Investing Conference Dave Gebhardt (Chief Strategy Officer for FarmLink ) spoke about the rise of data science in agriculture. You can read the story and find a link to the podcast here:

In the podcast he discusses the way data science is revolutionizing agriculture, and how we are at a "tipping point where advances in science, IT, technology, and computing power have put a whole new level of opportunities before us."

This sounds a lot like what I have previously discussed in relation to big data and the internet of things: 

Watch more about how FarmLink is leveraging IoT, big data, and advanced analytics:



Related:

 Big Ag Meets Big Data (Part 1 & Part 2)

Saturday, June 13, 2015

SAS vs R? The right answer to the wrong question?

For a long time I tracked a discussion on LinkedIn that consisted of various opinions about using SAS vs R. Some people can take this very personal.  Recently there was an interesting post at the DataCamp blog addressing this topic. They also provided an interesting infographic making some comparisons between SAS and R as well as SPSS.  Other popular debates also include python in the mix. (By the way, it is possible to integrate all three on the SAS platform and you can also run R via the open source integration node in SAS Enterprise Miner 13.1).

Aside: For older versions of SAS EM-can you drop in a code node and call R via PROC IML?

Anyway, getting back to the article, I tend to agree with this one point:

"While these debates are a good thing for the community and the programming language as a whole, they unfortunately also have a negative effect on those individuals that are just in the beginning of their data analytics career. Biased opinions on all sides of the table make it difficult for new data analysts to see the forest for the trees when choosing a statistical programming language."

While I agree with this notion, I want to reflect for a minute on the concept of a programming language. If you think of SAS as just a programming language, then perhaps these kinds of comparisons and discussions make sense, but for a data scientist, I think one's view of analtyics should transcend just a language. When we think of an overall analytical solution there is a lot to consider, from how the data is generated, how it is captured and warehoused, how it is extracted and cleaned and accessed by whatever programming tool(s), how it is visualized and analyzed, and ultimately, how do we operationalize the solution so that it can be consumed by business users.

So to me the relevant question is not, which programming language is preferred by data scientists, or which program is better for implementing specific machine learning algorithms; but perhaps what is the best analytical solutions platform for solving the problems at hand? 

Friday, May 8, 2015

Mendelian Instruments (Applied Econometrics meets Bioinformatics)

Recently I defended the use of quasi-experimental methods in wellness studies, and a while back I sort of speculated that genomic data might be useful in a quasi-experimental setting-but wasn’t sure how: 

If causality is the goal, then merge 'big data' from the gym app with biometrics and the SNP profiles and employ some quasi-expermental methodology to investigate causality.”

Then this morning at marginal revolution I ran across a link to a blog post that mentioned exploiting mendelian variation as instruments for a particular study related to alcohol consumption.

This piece gives a nice intro I think:

Stat Med. 2008 Apr 15;27(8):1133-63. Mendelian randomization: using genes as instruments for making causal inferences in epidemiology.

Lawlor DA1, Harbord RM, Sterne JA, Timpson N, Davey Smith G

Link: http://www.ncbi.nlm.nih.gov/pubmed/17886233

“Observational epidemiological studies suffer from many potential biases, from confounding and from reverse causation, and this limits their ability to robustly identify causal associations. Several high-profile situations exist in which randomized controlled trials of precisely the same intervention that has been examined in observational studies have produced markedly different findings. In other observational sciences, the use of instrumental variable (IV) approaches has been one approach to strengthening causal inferences in non-experimental situations. The use of germline genetic variants that proxy for environmentally modifiable exposures as instruments for these exposures is one form of IV analysis that can be implemented within observational epidemiological studies. The method has been referred to as 'Mendelian randomization', and can be considered as analogous to randomized controlled trials. This paper outlines Mendelian randomization, draws parallels with IV methods, provides examples of implementation of the approach and discusses limitations of the approach and some methods for dealing with these.”

Tuesday, April 28, 2015

Healthcare Analytics at SAS Global Forum 2015

I was not able to attend this year's SAS Global Forum, but have had a chance to browse the numerous session papers as well as enjoy some live content. The conference page has a searchable catalog and for each paper you will find links to similar sessions in the sidebar. There were around 5000 attendees at this year's conference and over 600 papers. For a 2 1/2 day conference that's more than 200 papers to cover per day. Below is a selection of some papers related to healthcare analytics. If we widen the search to include papers related to other fields, but applications in healthcare analtyics the selection would probably double. I'm sure I missed something, and would be glad to know if you had a favorite paper or presentation you'd like to share in the comments.


1329 - Causal Analytics: Testing, Targeting, and Tweaking to Improve Outcomes This session is an introduction to predictive analytics and causal analytics in the context of improving outcomes. The session covers the following topics: 1) Basic... View More 20 minutes Breakout Jason Pieratt

1340 - Using SAS® Macros to Flag Claims Based on Medical Codes Many epidemiological studies use medical claims to identify and describe a population. But finding out who was diagnosed, and who received treatment, isn't always simple.... View More 50 minutes Breakout Andy Karnopp

2382 - Reducing the Bias: Practical Application of Propensity Score Matching in Health-Care Program Evaluation To stay competitive in the marketplace, health-care programs must be capable of reporting the true savings to clients. This is a tall order, because most health-care programs... View More 20 minutes Breakout Amber Schmitz

2920 - Text Mining Kaiser Permanente Member Complaints with SAS® Enterprise Miner™ This presentation details the steps involved in using SAS® Enterprise Miner™ to text mine a sample of member complaints. Specifically, it describes how the Text Parsing, Text... View More 30 minutes E-Poster Amanda Pasch

3214 - How is Your Health? Using SAS® Macros, ODS Graphics, and GIS Mapping to Monitor Neighborhood and Small-Area Health Outcomes With the constant need to inform researchers about neighborhood health data, the Santa Clara County Health Department created socio-demographic and health profiles for 109... View More 20 minutes Breakout Roshni Shah

3254 - Predicting Readmission of Diabetic Patients Using the High-Performance Support Vector Machine Algorithm of SAS® Enterprise Miner™ 13.1 Diabetes is a chronic condition affecting people of all ages and is prevalent in around 25.8 million people in the U.S. The objective of this research is to predict the... View More 20 minutes Breakout Hephzibah Munnangi

3281 - Using SAS® to Create Episodes-of-Hospitalization for Health Services Research An essential part of health services research is describing the use and sequencing of a variety of health services. One of the most frequently examined health services is... View More 20 minutes Breakout Meriç Osman

3282 - A Case Study: Improve Classification of Rare Events with SAS® Enterprise Miner™ Imbalanced data are frequently seen in fraud detection, direct marketing, disease prediction, and many other areas. Rare events are sometimes of primary interest. Classifying... View More 20 minutes Breakout Ruizhe Wang

3411 - Identifying Factors Associated with High-Cost Patients Research has shown that the top five percent of patients can account for nearly fifty percent of the total healthcare expenditure in the United States. Using SAS® Enterprise... View More 30 minutes E-Poster Jialuo Cheng

3488 - Text Analytics on Electronic Medical Record Data This session describes our journey from data acquisition to text analytics on clinical, textual data. 50 minutes Breakout Mark Pitts

3560 - A SAS Macro to Calculate the PDC Adjustment of Inpatient Stays The Centers for Medicare & Medicaid Services (CMS) uses the Proportion of Days Covered (PDC) to measure medication adherence. There is also some PDC-related research based on... View More 20 minutes Breakout anping chang


3600 - When Two Are Better Than One: Fitting Two-Part Models Using SAS In many situations, an outcome of interest has a large number of zero outcomes and a group of nonzero outcomes that are discrete or highly skewed.  For example, in modeling... View More 20 minutes Breakout Laura Kapitula

3740 - Risk-Adjusting Provider Performance Utilization Metrics Pay-for-performance programs are putting increasing pressure on providers to better manage patient utilization through care coordination, with the philosophy that good... View More 50 minutes Breakout Tracy Lewis

3741 - The Spatio-Temporal Impact of Urgent Care Centers on Physician and ER Use The unsustainable trend in healthcare costs has led to efforts to shift some healthcare services to less expensive sites of care. In North Carolina, the expansion of urgent... View More 50 minutes Breakout Laurel Trantham

3760 - Methodological and Statistical Issues in Provider Performance Assessment With the move to value-based benefit and reimbursement models, it is essential toquantify the relative cost, quality, and outcome of a service. Accuratelymeasuring the cost... View More 50 minutes Breakout Daryl Wansink

SAS1855 - Using the PHREG Procedure to Analyze Competing-Risks Data Competing risks arise in studies in which individuals are subject to a number of potential failure events and the occurrence of one event might impede the occurrence of other... View More 20 minutes Breakout Ying So


SAS1900 - Establishing a Health Analytics Framework Medicaid programs are the second largest line item in each state’s budget. In 2012, they contributed $421.2 billion, or 15 percent of total national healthcare expenditures.... View More 20 minutes Breakout Krisa Tailor


SAS1951 - Using SAS® Text Analytics to Examine Labor and Delivery Sentiments on the Internet In today’s society, where seemingly unlimited information is just a mouse click away, many turn to social media, forums, and medical websites to research and understand how mothers feel about the birthing process. Mining the data in these resources helps provide an understanding of what mothers value and how they feel. This paper shows the use of SAS® Text Analytics to gather, explore, and analyze reports from mothers to determine their sentiment about labor and delivery topics. Results of this analysis could aid in the design and development of a labor and delivery survey and be used to understand what characteristics of the birthing process yield the highest levels of importance. These resources can then be used by labor and delivery professionals to engage with mothers regarding their labor and delivery preferences. View Less 20 minutes Breakout Michael Wallis


Thursday, March 26, 2015

To Explain or Predict

Some people may not make the important distinctions between prediction vs inference when it comes to modeling approaches/methodologies/data handling/assumptions.  I recently ran across a blog post by Rob J. Hyndman that pointed to the following article in the journal Statistical Science:

 Statist. Sci.
 Volume 25, Number 3 (2010), 289-310.

"Statistical modeling is a powerful tool for developing and testing theories by way of causal explanation, prediction, and description. In many disciplines there is near-exclusive use of statistical modeling for causal explanation and the assumption that models with high explanatory power are inherently of high predictive power. Conflation between explanation and prediction is common, yet the distinction must be understood for progressing scientific knowledge. While this distinction has been recognized in the philosophy of science, the statistical literature lacks a thorough discussion of the many differences that arise in the process of modeling for an explanatory versus a predictive goal. The purpose of this article is to clarify the distinction between explanatory and predictive modeling, to discuss its sources, and to reveal the practical implications of the distinction to each step in the modeling process."

This is a nice article which I think complements Leo Brieman's paper discussed here before regarding two cultures of predictive modeling. Rob gives a nice synopsis of some of the main points from the paper:

  1. The AIC is better suited to model selection for prediction as it is asymptotically equivalent to leave-​​one-​​out cross-​​validation in regression, or one-​​step-​​cross-​​validation in time series. On the other hand, it might be argued that the BIC is better suited to model selection for explanation, as it is consistent.
  2. P-​​values are associated with explanation, not prediction. It makes little sense to use p-​​values to determine the variables in a model that is being used for prediction. (There are problems in using p-​​values for variable selection in any context, but that is a different issue.)
  3. Multicollinearity has a very different impact if your goal is prediction from when your goal is estimation. When predicting, multicollinearity is not really a problem provided the values of your predictors lie within the hyper-​​region of the predictors used when estimating the model.
  4. An ARIMA model has no explanatory use, but is great at short-​​term prediction.
  5. How to handle missing values in regression is different in a predictive context compared to an explanatory context. For example, when building an explanatory model, we could just use all the data for which we have complete observations (assuming there is no systematic nature to the missingness). But when predicting, you need to be able to predict using whatever data you have. So you might have to build several models, with different numbers of predictors, to allow for different variables being missing.
  6. Many statistics and econometrics textbooks fail to observe these distinctions. In fact, a lot of statisticians and econometricians are trained only in the explanation paradigm, with prediction an afterthought. That is unfortunate as most applied work these days requires predictive modelling, rather than explanatory modelling.

Rob also links Galit Shmueli's web page, (the author of the article above) who apparently has done some extensive research related to these distinctions.  Lots of additional resources (blog) here in this regard. Galit states:

"My thesis is that statistical modeling, from the early stages of study design and data collection to data usage and reporting, takes a different path and leads to different results, depending on whether the goal is predictive or explanatory." 

I touched on these distinctions before, but did not realize the extent of the actual work being done in this area by Galit.

Analytics vs Causal Inference
Big Data: Don't throw the baby out with the bath water

See also: Paul Allison on multicollinearity

Friday, January 30, 2015

Data Science is 10% Inspiration and 90% Perspiration

“Success is 10 percent inspiration and 90 percent perspiration.” Thomas Alva Edison

Last fall there was a really good article in the New York Times:

For Big-Data Scientists, ‘Janitor Work’ Is Key Hurdle to Insights

"Yet far too much handcrafted work — what data scientists call “data wrangling,” “data munging” and “data janitor work” — is still required. Data scientists, according to interviews and expert estimates, spend from 50 percent to 80 percent of their time mired in this more mundane labor of collecting and preparing unruly digital data, before it can be explored for useful nuggets."

“Data wrangling is a huge — and surprisingly so — part of the job,” said Monica Rogati, vice president for data science at Jawbone, whose sensor-filled wristband and software track activity, sleep and food consumption, and suggest dietary and health tips based on the numbers. “It’s something that is not appreciated by data civilians. At times, it feels like everything we do.”

“It’s an absolute myth that you can send an algorithm over raw data and have insights pop up,” said Jeffrey Heer, a professor of computer science at the University of Washington and a co-founder of Trifacta, a start-up based in San Francisco.

This has always been true, even before 'Big Data' was a big deal. It was one of the first rude awakenings I had as a researcher straight out of graduate school (one of the other things was related to the large gap between econometric theory and applied econometrics). Thank goodness, I was lucky enough to work in a shop where this was much appreciated and I developed the necessary SAS and SQL (and later R) skills to deal with these issues. They just don't teach this stuff in school (I'm sure they might some places).

The article mentions some efforts to develop software to make these tasks simpler. I think there is a very fine line between the value gained from doing this grunt work vs. the savings of time an energy that we could yield if we flatten the cost curve when it comes to data prep. As the article says:

"Data scientists emphasize that there will always be some hands-on work in data preparation, and there should be. Data science, they say, is a step-by-step process of experimentation."

“You prepared your data for a certain purpose, but then you learn something new, and the purpose changes,” said Cathy O’Neil, a data scientist at the Columbia University Graduate School of Journalism, and co-author, with Rachel Schutt, of “Doing Data Science” (O’Reilly Media, 2013)."

See also:

In God We Trust, All Others Show Me Your Code


In God We Trust, All Others Show Me Your Code

There recently was a really interesting article at the Political Methodologist titled:

A Decade of Replications: Lessons from the Quarterly Journal of Political Science

They have high standards related to research documentation:

"Since its inception in 2005, the Quarterly Journal of Political Science (QJPS) has sought to encourage this type of transparency by requiring all submissions to be accompanied by a replication package, consisting of data and code for generating paper results. These packages are then made available with the paper on the QJPS website. In addition, all replication packages are subject to internal review by the QJPS prior to publication. This internal review includes ensuring the code executes smoothly, results from the paper can be easily located, and results generated by the replication package match those in the paper."

"Although the QJPS does not necessarily require the submitted code to access the data if the data are publicly available (e.g., data from the National Election Studies, or some other data repository), it does require that the dataset containing all of the original variables used in the analysis be included in the replication package. For the sake of transparency, the variables should be in their original, untransformed and unrecoded form, with code included that performs the transformations and recodings in the reported analyses. This allows replicators to assess the impact of transformations and recodings on the results."

From an efficiency standpoint, I don't know if this standard should be applied universally or not. We wouldn't want to bottleneck the body of peer reviewed literature contributing to society's pool of knowledge, but at the same time, some sort of filtration system might keep the murkiness out of the water so we can see more clearly the 'real' effects of policies and treatments.

I certainly know from personal (professional and academic) experience collaborating with others a lot of time and resources have been lost trying to reinvent the wheel, reconstruct the creation of some data set etc. because of lack of documentation around how data was pulled or cleaned. Better documentation and code sharing always seems better. Maybe everyone needs a Github account.

Tuesday, March 25, 2014

Institutional Research Presentations at SAS Global Forum

I'm not attending #SASGF14, but some of my colleagues in higher ed are. Here is what they are doing. If you are not attending global forum, or can't make their talks, I encourage you to check out their papers via the online proceedings once they are posted.

Tuesday March 25

Paper 1448 - From Providing Support to Driving Decisions: Improving the Value of Institutional Research For almost two decades, Western Kentucky University's Office of Institutional Research (WKU-IR) has used SAS® to help shape the future of the institution by providing faculty and administrators with information they can use to make a difference in the lives of their students. This presentation provides specific examples of how WKU-IR has shaped the policies and practices of our institution and discusses how WKU-IR moved from a support unit to a key strategic partner. In addition, the presentation covers the following topics: How the WKU Office of Institutional Research developed over time; Why WKU abandoned reactive reporting for a more accurate, convenient system using SAS® Enterprise Intelligence Suite for Education; How WKU shifted from investigating what happened to predicting outcomes using SAS® Enterprise Miner™ and SAS® Text Miner; How the office keeps the system relevant and utilized by key decision makers; What the office has accomplished and key plans for the future.


Paper 1638 - Institutional Research: Serving University Deans and Department Heads Administrators at Western Kentucky University rely on the Institutional Research department to perform detailed statistical analyses to deepen the understanding of issues associated with enrollment management, student and faculty performance, and overall program operations. This paper presents several instances of analyses performed for the university to help it identify and recruit suitable candidates, uncover root causes in grade and enrollment trends, evaluate faculty effectiveness, and assess the impact of student characteristics, programs, or student activities on retention and graduation rates. The paper briefly discusses the data infrastructure created and used by Institutional Research. For each analysis performed, it reviews the SAS® program and key components of the SAS code involved. The studies presented include the use of SAS® Enterprise Miner™ to create a retention model incorporating dozens of student background variables. It shows an examination of grade trends in the same courses taught by different faculty and subsequent student behavior and success, providing insights into the nuances and subtleties of evaluating faculty performance. Another analysis uncovers the possible influence of fraternities and sororities in freshmen algebra courses. Two investigations explore the impact of programs on student retention and graduation rates. Each example and its findings illustrate how Institutional Research can support the administration of university operations. The target audience is any SAS professional interested in learning more about Institutional Research in higher education and how SAS software is used by an Institutional Research department to serve its organization.


Monday March 24

Paper 1689 - Simple ODS Tips to Get RWI (Really Wonderful Information) SAS® continues to expand and improve its reporting capability. With new SAS® 9.4 enhancements in ODS (Output Delivery System), the opportunity to create stunning reports has expanded even further. If you are charged with creating relevant, informative, easy-to-read reports for clients or administrators, then the ODS Report Writing Interface, ODS LAYOUT enhancements, and the new ODSTEXT procedure are important tools to use. These tools allow you to create reports in a smart, eye-catching format that can be turned around quite quickly and programmed to provide optimum flexibility. How many times have you worked hours to tweak and fine-tune a report directly in Microsoft Excel, Microsoft Word, Microsoft Power Point or some other similar software only to be asked for a “quick update”, which would then take hours to recreate because you are manually transferring data? Do you ever dread receiving the compliment, “This is really wonderful information!!!!” because you know it will be followed by “Can you run this for EVERY region?” Well, dread no more, because when you harness the power of SAS® ODS, you can create first-rate, flexible, fabulous reports! Join me as I share with you two real-world examples of ODS capabilities using (1) a marketing piece I designed to help the president of our university spotlight county- and region-specific data as he recruited across the state and (2) our academic program review form, a multi-page report that outputs to Word so that program coordinators can add personalized commentary to support their program’s effectiveness.