Showing posts with label econometrics. Show all posts
Showing posts with label econometrics. Show all posts

Sunday, May 31, 2026

Second-Handedness, AI and the Production and Use of Knowledge in Society

Almost 3 years ago I wrote a post entitled "If Applied Econometrics Were Easy LLMs Could Do it." Since then the technical capabilities of AI have progressed, but the advancements only reinforce some of the main takeaways from that post:

"There are risks when these tools are used like Dunning-Kruger-as-a-Service (DKaaS), where the critical thinking and actual learning begins and ends with prompt engineering and a response. We have to be very careful to recognize as Philip Tetlock describes in his book "Superforecasters" that there is a difference between mimicking and reflecting meaning vs. originating meaning.  To recognize that it’s not just what you know that matters, but how you know what you know. The second-handed tendency to believe that we can or should be outsourcing our thinking to AI… is philosophically and epistemically disturbing."

The Main Takeaway or BLUF: 

The overall implication of this post is that how we use AI can impact what we learn and how we learn. At a certain point the how starts to matter more than the what, undermining our long term growth and capabilities as individuals, profitability of businesses, and eventually in society overall.  

In this post I want to expand on this epistemically disturbing theme from my prior post given how rapidly AI capabilities are advancing. This is a long post - some may want to skip to the summary and conclusions at the end of the post and then come back to sections of interest. Or use AI to summarize the main points :) 

So What's New Besides Even More Advanced AI?

Since my last post, recent publications in this space have expanded on the consequences of use of more advanced AI in society. Specifically I will be drawing from a recent NBER working paper: AI, Human Cognition and Knowledge Collapse as well as other related work.  In this paper authors consider how generative AI, and in particular agentic AI, shapes human learning incentives and the long-run evolution of society’s information ecosystem.  In this paper they build a dynamic model of learning and decision-making and discuss the implications. They discuss how there is dynamic tension in that AI can improve decision quality today, but erode learning incentives that sustain long term collective knowledge, potentially even leading to a total knowledge collapse where "in the long-run equilibrium all human knowledge is destroyed." 

In this post I am not setting out to prove anything, or empirically defend any specific hypotheses (I'll leave that to the AI researchers and academic economists). My goal is to only to draw parallels between this recent work and build on my prior thoughts on the implications of the use of AI and knowledge in society. 

AI, Human Cognition, and Knowledge Collapse - Summary

The paper discusses the role of substitution effects, complements, economies of scope, and externalities in the production and use of knowledge and decision making as it relates to AI. When people put forth the effort to learn without AI, there is a private benefit in that what is learned helps make better decisions. This private knowledge is also complemented by the existing stock of public knowledge. AI can leverage public knowledge and produce context specific (local) knowledge and recommendations to individuals.  This also supports better decision making, but at a lower cost because AI substitutes for individual learning effort. It is important to note that without AI, individual learning often contributes a marginal amount of new knowledge to society's stock of general knowledge. This joint production of individually useful specific and public general knowledge represents economies of scope in the production of knowledge. We know that new knowledge plays an important role in human progress and sustainable economic growth over time as pointed out by Arrow and his work related to learning and doing and the role of knowledge (1962) and more recently Romer's growth models with endogenous technological change. At the same time, individuals don't necessarily directly benefit from their contribution of new knowledge to the public stock of knowledge (also discussed in Arrow). So private production of public  knowledge comes at an uncompensated cost resulting in a positive externality to society. The private benefit and lower cost of learning that AI delivers reduces individual effort in knowledge production given the uncompensated positive externality. I'll stop there and return to the paper's treatment of macro level impacts later. First I want to discuss the individual and firm level implications of the model.

AI and Applied Econometrics

My prior post also gets into lots of other things like AI and causality and working with AI mostly in the context of on doing applied econometrics. If you want to get a flavor of just how much AI may be influencing the way econometrics gets done, check out some of Scott Cunningham's work or Claude Blattman by Chris Blattman.  

Tyler Cowen at Marginal Revolution has had several posts discussing how AI is impacting economic research like this post - Will AI Kill the Research Paper?  In her post AI, Price Theory, and the Future of Economics Research, Lynn Kiesling offers a perspective focusing on the impact of AI on workflows and what skills will become differentiators for economists of the future, with a Hayekian take of course.  Brian Albrecht chimes in on this too. Both Kiesling and Albrecht discuss how AI can change workflows and reduce the costs of execution, but this will actually make economic reasoning more important. 

Albrecht states: "The question I would focus on is...whether the world still needs people who can hear a claim about the economy and ask whether it makes sense." 

His post makes the answer an obvious yes: "Those are not questions that more data answers. They require economic reasoning about what’s generating the patterns in the first place...automating the technique doesn’t automate the reasoning about whether the technique’s output makes sense. It increases the volume of output that needs reasoning applied to it."

I think the question behind the question Albrecht asks above, and a key theme of this post is, whether the use of AI will eventually erode our ability to provide that kind of mainline economic reasoning? Or human reasoning in general for that matter?

Individual Level Impacts

In my last post I called out a few examples of how we might use AI at the individual level and where things can go wrong. One example is attempting to use AI as a research assistant:

[What this leaves out is] how much you get out of putting your hands on a paper or book and going through it and wrestling with the ideas, the paths leading from from hypotheses to the conclusions, and how the cited references let you retrace the steps of the authors to understand why, either slowly nudging your priors in new directions or reinforcing your existing perspective, and synthesizing these ideas with your own. Then summarizing and applying and communicating this synthesis with others. ChatGPT might give the impression that is what it is doing in a fraction of the time you could do it (literally seconds vs. hours or days)....There is a big difference between the learning that takes place when you go through this process of integrative complex thinking vs. just reading a summary delivered on a silver platter from chatGPT.  I’m skeptical what I’m describing can be outsourced to AI without losing something important....How much knowledge and important nuance is lost with every [updated query to AI]? What is missed? Thinking! [and learning]

This parallels much of what is discussed in the NBER working paper. They state these sorts of issues more formally in their modeling assumptions as they relate to the substitution and crowding out effects of AI and knowledge generation. 

There is also other evidence related to negative individual effects of using AI called out within the NBER paper. 

In Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task, Kosmyna et. al discuss the impact of using AI for writing tasks:

While LLMs offer immediate convenience, our findings highlight potential cognitive costs. Over four months, LLM users consistently underperformed at neural, linguistic, and behavioral levels. These results raise concerns about the long-term educational implications of LLM reliance and underscore the need for deeper inquiry into AI's role in learning.

In AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking, Gerlich finds a significant negative correlation between frequent AI tool usage and critical thinking abilities with a worse effect among younger learners compared to older subjects. Other researchers Budiyono et al. (2025) have made similar findings: 

"Reliance on AI writing tools significantly reduced cognitive effort and creativity, overshadowed personal writing styles, and led to a decline in confidence and skill retention. These results suggest that, while AI tools enhance efficiency and technical accuracy, over-reliance on them may hinder the development of critical thinking, creativity, and independent writing skills" 

This surfaces the importance of critical thinking in the face of increased reliance on AI tools and the need to mitigate the negative effects of AI on those thinking skills. As Kiesling notes in her blog post, "the profession will have to rethink how it cultivates judgment when many traditional apprenticeship tasks have been automated." That likely goes for all professions and something businesses need to think about when it comes to developing talent in our future workforce. 

Impact at the Business Level

In my pior post I noted: 

AI is not capable of doing these things [actual thinking tasks], and believing and even attempting or pretending that we can get these things on a second-handed basis from an AI tool will ultimately erode the real human skills and capabilities essential to real productivity and growth over the long run.  If we fail to accept this we will hear a giant sucking sound that is the ROI we thought we were going to get from AI in the short run by attempting to automate what can't be automated. That is the false promise of a tools and technology mindset.

These seem to be related to the implications of substitution effects and crowding out in the paper, but impacting the firm level.

I also discussed a point made by Cassie Kozrykov in a video where she discussed these issues: 

"that may be the biggest problem, that management has not learned how to manage thinking...vs. what you can measure easily....thinking is something you can't force, you can only get in the way of it."

She elaborates a bit more about this in her LinkedIn post"A misguided view of productivity could mean lost jobs for workers without whom organizations won't be able to thrive in the long run - what a painful mistake for everyone."

In her blog post mentioned earlier, Kiesling makes an important observation related to this line of thinking: "If AI cheapens formalized information processing, then tacit knowledge, local knowledge, judgment, and institutional understanding may rise in relative value." But recent research (Ide, 2025) indicates that one of the downsides of reliance on AI is that "tools that automate entry-level tasks"  are "likely to disrupt the diffusion of tacit knowledge" especially to novice workers. 

This could eventually lead to less productive firms as overtime the work force becomes less knowledgeable than the least knowledgeable pre-AI solvers. In other words, AI will put a premium on expertise, but if the level of available expertise erodes over time with the use of AI, it could ultimately erode productivity and firm value. (see also Ide and Talamas, 2025 for more implications). Next I will turn to the potential aggregate impact of these forces on society overall. 

Impacts on Society

The paper discusses how more accurate AI benefits the individual and reduces their required effort to learn (direct effect). However this reduction in private effort comes at a potential cost to society - the loss or crowding out of of any marginal production of new knowledge (indirect effect). 

The substitution and crowding out effects of AI can lead to long term reductions in the stock of public knowledge, and in certain situations the author's model shows this can lead to a knowledge collapse. 

Specifically this is tied to the level of AI accuracy. When AI recommendations exceed an accuracy threshold, the economy can tip into a knowledge collapse steady state in which general knowledge vanishes ultimately despite high quality personalized advice. That implies the better and more accurate the AI the worse things can get. 

I'll have more to discuss about societal impacts below. 

Conclusions

So what is my perspective on the takeaways for individuals, firms, and society? First an important distinction. In my earlier post I talked about the distinctions Cassie Kozrykov made between thinking and what she called thunking:

"[thunking includes] things that consume our time and resources but don't require thinking. Having done your homework, the kind of summary information you get from an LLM can help reinforce your thinking and learnings and save time in terms of manually googling or looking up a lot of things you once knew but have forgotten."

So when AI is used for thunking, for example scanning a form to make sure it is complete or checking for errors, or extracting key topics from a chat or call transcript, etc.,  the substitution and crowding out effects in the paper would be minimal and so would not have the detrimental impacts on learning and society's stock of knowledge. The negative effects of AI arise when AI is used for thinking tasks. That distinction is really the lens for all of the above and the implications below. 

Impacts on Early Career Individuals

For those early in their career the substitution and crowding out effects mentioned above may be challenging and require making strategic tradeoffs. They need to think carefully about how they use AI. Fully embracing AI for thunking makes sense, but they should be cautious about using it for thinking tasks where they may miss out on learning, personal growth and development opportunities.  A bigger challenge may be that opportunities for learning and development could be eliminated through automation (as discussed in Ide, 2025). 

Impacts on Seasoned Professionals and Expertise

I might speculate that for those that have already gained lots of training and experience before AI, the substitution and crowding out effects would be minimal. They might use AI for thinking tasks, and going forward not notice overtime any depreciation in their skills.  In other words, they already have sufficient human capital to draw on and could complement that with AI giving them a competitive advantage (capital complements labor). However I think their personal contribution to society's stock of knowledge would be ambiguous. 

Will AI reduce the demand for expertise? In the short term we might be deluded by AI to think that it is mimicking expertise. So at first returns to expertise may drop as companies try to cut costs in the short run. However, as noted in Ide (2025) and if we think about the points made by Albrecht and Kiesling above, as AI shifts margins, certain kinds of expertise become important, the kinds of knowledge and judgment that isn't going to be in any training data for AI to access and learn from. 

"In a world where production becomes abundant, discernment becomes relatively scarce and thus relatively more valuable. What matters more is the ability to decide....what assumptions are plausible, which results travel across contexts, and what pattern in the evidence actually matters." 

This really gets to the heart of Hayek's knowledge problem, know how and know what are still going to be dispersed across many minds, and unavailable to any centralized decision maker with or without AI no matter how powerful AI becomes. The problem of managing dispersed knowledge remains. And one of the key points to this whole discussion is accepting the fact that how you know what you know matters as much as what you know when it comes to making better decisions. So expertise focused on managing and solving these problems ('solvers' as denoted in Ide, 2025) and the need for advice will still command a premium in a world with AI - especially if AI leads to an erosion of the general stock of knowledge and expertise in the future according to the model discussed above. As Ide and Talamas (2025) note - more knowledgeable workers will likely benefit disproportionately from AI. This is emphasized more in the discussion about implications for businesses below.

Impacts on Businesses

The use of AI for thunking tasks will be areas where there is obvious business value from AI. But a big challenge for business firms will be how do you take adavantage of productivity gains of AI and remain competitive while cultivating knowledge and judgment among employees if you are also automating away opportunities to learn? How do you avoid eroding the stock of knowledge at the firm level?

We know, taking a knowledge based theory of the firm, that the value of the firm is the sum of its decisions, and better decisions require knowledge. A firm's portfolio of knowledge assets becomes a source of value and competitive advantage (Grant, 2010). If the model in the NBER article is realistic, and there are substitution, externalities, and crowding out effects from AI, how do firms manage this portfolio in an age of AI without cannibalizing their most precious assets?

Comments from Kiesling are worth repeating: "If AI cheapens formalized information processing, then tacit knowledge, local knowledge, judgment, and institutional understanding may rise in relative value."

Ide (2025) emphasizes the importance of "expanding novices’ access to high-quality mentorship" from experts that have likely accumulated tacit knowledge and expertise over their careers prior to relying on AI. This will put a premium on expertise and experience, while at the same time require investing in the professional development of novices whose learning opportunities are being automated away. How do firms encourage workers to invest in learning and producing knowledge essential to growth and competitive advantage? Without the right incentive structures and professional development strategies, opportunities will likely be automated away and/or workers will take advantage of the substitution effects of AI. If this is the direction AI takes us, the next generation of workers will lack expertise, and they won't contribute to the growing stock of knowledge necessary for sustained competitive advantage at the firm level. 

Impacts on Society

If we think about growth models in economics we have to wonder if AI will enhance economic growth through technological change, or will the use of AI actually lead to knowledge collapse (as in the NBER paper) and stagnation? As Robert Lucas once said regarding economic growth and development: "The consequences for human welfare involved in questions like these are simply staggering: Once one starts to think about them, it is hard to think about anything else."

I think this is hard to really know which margins will really change, and which forces will dominate, and what will actually be the long term (or even short term) impacts of AI. The dominant narrative I hear often is that AI will help us solve problems we could never solve before and transform the fields of science, medicine, business and warfare. Many of us can already point to use cases that have benefited us personally. There is also a darker narrative about unemployment and loss of purpose. 

While the dynamics discussed in the NBER paper seem plausible, and correspond a lot with my prior thoughts on AI, I can't say for sure if I think knowledge collapse is inevitable or not. Regardless, the authors of the NBER paper propose some policy ideas to prevent knowledge collapse. They propose a two phased approach that starts with fully suppressing AI in order to rebuild the stock of general knowledge, followed by a phase of capping or 'garbling' the precision or accuracy of AI to maintain the general stock of knowledge. Similar to calls for moratoriums on data centers, these seem like blunt tools if not impractical. 

Going back to my original post - I think again Cassie Kozrykov makes an important point: 

"when you are not the one making the decision and it looks like the machine is doing it, there is someone who is actually making that decision for you...and I think that we have been complacent and we have allowed our technology to be faceless....how will we hold them accountable....for wisdom...thinking is our responsibility"

As I said in that post - thinking is a moral responsibility. Outsourcing our thinking and fooling ourselves into believing that we can get knowledge and wisdom and judgment second-handed from a summary written by an AI tool, believing that is the same thing and provides the same value as what we could produce as thinking humans, is a dangerous illusion.  

In 2020 former President Barak Obama emphasized the importance of thinking in a democracy: 

"if we do not have the capacity to distinguish what's true from what's false, then by definition the marketplace of ideas doesn't work. And by definition our democracy doesn't work. We are entering into an epistemological crisis." 

Thinking is the means by which the human race and civil society thrives and survives. That may not be a solution that can easily be turned into a business strategy or law, but it is the answer. 

Disclaimer: AI was not used in any direct way to write this post. For the specific purposes of this post, any related Google searches were appended with '-ai' to avoid inadvertent influence of default AI summaries generated by a search.


Afterward: Some Connections in Literature, Philosophy, and Religion

In this section I want to discuss some loose but related connections I have made from literature, philosophy and religion. 

  • In many thoughts and discussions about AI, I can't help but think about this quote from Dune, by Frank Herbert: “Once men turned their thinking over to machines in the hope that this would set them free. But that only permitted other men with machines to enslave them.” 

  • In The Fountainhead, Ayn Rand emphasizes the importance of individuality and thinking for oneself vs. relying on others to think for you - an act she refers to as second-handedness: 
“That, precisely, is the deadliness of second-handers… Not to judge, but to repeat. Not to do, but to give the impression of doing….What would happen to the world without those who do, think, work, produce?…You don't think through another's brain and you don't work through another's hands. When you suspend your faculty of independent judgment, you suspend consciousness. To stop consciousness is to stop life.”

  • Can AI actually think? A lot of the discussion in all of the above is in a sense about the tensions between using AI for thinking vs. thunking.  Again, Cassie Kozrykov has a position on this: "AI does not automate thinking. It doesn't! There is a lot of strange rumblings about this that sound very odd to me who has been in this space for 2 decades." That may be a good reason why she advocates for using AI for what she calls thunking tasks but against thinking tasks. 
  • From a purely metaphysical perspective, there may be good reason to believe that no matter what advancements are made in neuroscience or computer science, machines will never truly be able to think as humans do. In his book Immortal Souls, philosopher Ed Feser makes this case. 
  • In his critique of AI, Feser states: "The contemporary obsession with computers as a model for the human mind is a wild goose chase."  If I were to crudely summarize some of his arguments I would start by considering what does it mean to think? Thoughts are required to think. What do thoughts require? Thoughts require things like abstract concepts and universals all of which are immaterial - they have no matter and take up no space. It follows that formal thought processes cannot be material. Ergo machines, which are wholly material cannot have thought processes and cannot think. 
  • Another way of thinking about this is in terms of immanent vs. transuent causation which Feser discusses in more detail in his book Aristotle's Revenge: The Metaphysical Foundations of Physical and Biological Science. Feser describes an immanent causal process as one that originates within an agent on its own. It is a teleological process that points to or aims toward the realization of ends. It is basically having an intention and acting on it - which is what we think of minds being able to do. Transuent causal processes are imposed on objects and terminate outside an agent. This would be like a boulder rolling down a hill or gears in clocks keeping time,..physical processes like computers executing code. Thinking he argues, requires immanent causal processes. 
  • But with advances in computer science and our understanding of neuroscience, could machines actually think if we make them complex enough? Could thinking be an emergent property of physical processes? Feser Argues that increasing complexity is simply a matter of increasing the complexity of transuent causal processes. He states: "you can add to a transuent causal process all the further transuent causal processes you like but you will never get immanent causation out of it. The most you will get is something that might look like immanent causation, just as a polygon with sufficiently many sides might look like a circle...thinking is an activity that cannot be coherently analyzed in terms of transuent causation alone." As philosopher J.P Morland states: "pointing to emergence is simply to slap a label on a problem rather than solve it." 
  • One might attempt to bypass Feser's arguments by denying the distinction between transuent and immanent causal processes and simply eliminate immanent causal processes from our picture of reality. But this is hard to do coherently. As Feser argues: "the eliminativist has to carry out immanent causal activity in the very act of denying that there is such a thing as immanent causal activity. His position is incoherent." As M.R. Bennett and P.M.S. Hacker have noted "the eliminativist saws off the branch on which he is seated."

  • In his article "Idols of the Valley", Yuval Levin writes about Pope Leo XIV's encyclical about AI, Magnifica Humanitas. He discusses, in a sense, the moral and religious implications of the substitution effects (or shortcuts) of AI, as a form of idolatry: 

    "...the danger to which Pope Leo is pointing...is the danger of turning our tools into idols, and thereby of becoming little more than tools ourselves. It is a danger that afflicts those who make these idols, and also threatens those who put their trust in them. The appeal of idols has always been that they offer shortcuts. The God of the Bible demands that you live in a way that forms your mind and heart and soul toward your fullest human potential. This requires hard work but it yields a kind of person both capable and worthy of a flourishing life. The idol offers the material benefits of such a life without that formative work...This plainly rhymes with some of the deepest moral challenges posed to us by artificial intelligence. AI, at least used a certain way, offers us shortcuts around formative work, matching outputs with inputs without the need for the interceding effort of mind, heart, and soul. If all you care about are the outputs, not the form of your mind, heart, and soul, then the offer is awfully hard to resist....various idolatries offer us shortcuts that promise the benefit without the work: Just turn yourself into a tool and you will be more productive without more effort. This is of course just what Magnifica Humanitas warns of. It is what AI at its most idolatrous and dangerous can offer. That doesn’t have to be what AI is in our experience — not at all. But it can be if we aren’t careful."

    Related Posts

    If Applied Econometrics Were Easy, LLMs Could Do It https://econometricsense.blogspot.com/2023/07/if-applied-econometrics-were-easy-llms.html

    Statistics is a Way of Thinking Not a Just a Box of Tools. https://econometricsense.blogspot.com/2020/04/statistics-is-way-of-thinking-not-just.html 

    Will There Be a Credibility Revolution in Data Science and AI? https://econometricsense.blogspot.com/2018/03/will-there-be-credibility-revolution-in.html 

    R.A. Fisher, Big Data, and Pretended Knowledge. https://econometricsense.blogspot.com/2021/07/ra-fisher-big-data-and-thinking-like.html 

    Experimentation and Causal Inference Meet the Knowledge Problem. https://econometricsense.blogspot.com/2020/04/the-value-of-business-experiments-and.html

    References

    Kenneth J Arrow. The economic implications of learning by doing. The review of economic studies, 29(3):155–173, 1962a.

    Philosophical Foundations of Neuroscience. 1st Ed. M. R. Bennett, P. M. S. Hacker. Blackwell. 2003

    Herman Budiyono, M Pudjaningsih, B Prastio, and A Maulidina. Exploring the long-term impact of ai writing tools on independent writing skills: a case study of indonesian language education students. International Journal of Information and Education Technology, 15(5):1003–1013, 2025.

    Aristotle's Revenge: The Metaphysical Foundations of Physical and Biological Science. Edward Feser. 2019.

    Immortal Souls: A Treatise on Human Nature. Edward Feser. 2024.

    Michael Gerlich. Ai tools in society: Impacts on cognitive offloading and the future of critical
    thinking. Societies, 15(1):6, 2025.

    Grant, Robert M. Contemporary Strategy Analysis. 7th Edition. John Wiley and Sons. U.K. (2010).

    The Use of Knowledge in Society. F. A. Hayek. The American Economic Review, Vol. 35, No. 4. (Sep., 1945), pp. 519-530

    Enrique Ide. Automation, ai, and the intergenerational transmission of knowledge. arXiv preprint arXiv:2507.16078, 2025. Journal of Political Economy, 133(12):3762–3800, 2025.

    Enrique Ide and Eduard Talam`as. Artificial intelligence in the knowledge economy. Journal of Political Economy, 133(12):3762–3800, 2025.

    Nataliya Kosmyna, Eugene Hauptmann, Ye Tong Yuan, Jessica Situ, Xian-Hao Liao, Ashly Vivian Beresnitzky, Iris Braunstein, and Pattie Maes. Your brain on chatgpt: Accumulation of cognitive debt when using an ai assistant for essay writing task. arXiv preprint arXiv:2506.08872, 2025. 

    Thunking vs Thinking: Whose Job Does AI Automate? Which tasks are on AI’s chopping block? Cassie Kozrykov. https://kozyrkov.medium.com/thunking-vs-thinking-whose-job-does-ai-automate-959e3585877b


    Sunday, April 6, 2025

    Agricultural Economics as a Poster Child of Applied Economics

     Abstract

    Agricultural economists have embodied the notions of applied economics for a long time. They have used economic principles to address real-world problems, integrating economics and scientific knowledge. Applied economics tends to be multidisciplinary and develop applied concepts, theories, and tools. Some, like human capital, diffusion of innovation, contingent valuation, and numerous numerical and econometric techniques have spread throughout economics. Agricultural economic research has been data intensive, and improved information technologies strengthen this tendency. Yet data without theory is of limited use and coevolution of theory and data are essential. Empirical analysis should incorporate quantitative information as well as narratives. We are challenged to understand the coevolution of business, supply chains, and technology, and how they are affected by policies and affect markets. Research should integrate agriculture, energy, and the environment and develop tools to analyze and regulate the emerging bio-economy integrating biotech and infotech.

    Zilberman, D. (2019), Agricultural Economics as a Poster Child of Applied Economics: Big Data & Big Issues1. American Journal of Agricultural Economics, 101: 353-364. https://doi.org/10.1093/ajae/aay101

    Monday, December 16, 2019

    Some Recommended Podcasts and Episodes on AI and Machine Learning

    Something I have been interested in for some time now is both is the convergence of big data and genomics and the convergence of causal inference and machine learning. 

    I am a big fan of the Talking Biotech Podcast which allows me to keep up with some of the latest issues and research in biotechnology and medicine. A recent episode related to AI and machine learning covered a lot of topics that resonated with me. 

    There was excellent discussion on the human element involved in this work, and the importance of data data prep/feature engineering (the 80% of work that has to happen before the ML/AI can do its job) and the challenges of non-standard 'omics' data.  Also the potential biases that researchers and developers can inadvertently introduce in this process. Much more including applications of machine learning and AI in this space and best ways to stay up to speed on fast changing technologies without having to be a heads down programmer. 

    I've been in a data science role since 2008 and have transitioned from SAS to R to python. I've been able to keep up within the domain of causal inference to the extent possible, but I keep up with broader trends I am interested in via podcasts like Talking Biotech. Below is a curated list of my favorites related to data science with a few of my favorite episodes highlighted.


    1) Casual Inference - This is my new favorite podcast by two biostatisticians covering epidemiology/biostatistics/causal inference - and keeping it casual.

    Fairness in Machine Learning with Sherri Rose | Episode 03 - http://casualinfer.libsyn.com/fairness-in-machine-learning-with-sherri-rose-episode-03

    This episode was the inspiration for my post: When Wicked Problems Meet Biased Data.





    #093 Evolutionary Programming - 


    #266 - Can we trust scientific discoveries made using machine learning



    How social science research can inform the design of AI systems https://www.oreilly.com/radar/podcast/how-social-science-research-can-inform-the-design-of-ai-systems/ 



    #37 Causality and potential outcomes with Irineo Cabreros - https://bioinformatics.chat/potential-outcomes  


    Andrew Gelman - Social Science, Small Samples, and the Garden of Forking Paths https://www.econtalk.org/andrew-gelman-on-social-science-small-samples-and-the-garden-of-the-forking-paths/ 
    James Heckman - Facts, Evidence, and the State of Econometrics https://www.econtalk.org/james-heckman-on-facts-evidence-and-the-state-of-econometrics/


    Thursday, January 24, 2019

    Modeling Healthcare Claims as a Dependent Variable

    Healthcare claims present challenges to the applied econometrician. Claims costs typically exhibit a large number of zero values (high zero mass), extreme skewness, and heteroskedasticity. Below is a histogram depicting the distributional properties typical of claims data.




    The literature (see references below) addresses a number of approaches (i.e. log models, GLM, and two part models) often used for modeling claims data. However, without proper context the literature can leave one with a lot of unanswered questions, or several seemingly plausible answers to the same question.

    The department of Veteran's Affairs runs a series of healthcare econometrics cyberseminars covering these topics. Particularly, they have two video lectures devoted to modeling healthcare costs as a dependent variable.

    https://www.hsrd.research.va.gov/cyberseminars/series.cfm#hec3

    Principles discussed include:

    1) Despite what is taught in a lot of statistics classes about skewed data, in claims analysis we usually DO want to look at MEANS not MEDIANS.

    2) Why logging claims and then running analysis on the logged data to deal with skewness is probably not the best practice in this context.

    3) How adding a small constant number to zero values prior to logging can lead to estimates that are very sensitive to the choice of constant value.

    4) Why in many cases it could be a bad idea to exclude ‘high cost claimants’ from an analysis without good reasons. This probably should not be an arbitrary routine practice.

    5)When and why you may or may not prefer ‘2-part models’

    Note: Utilization data like ER visits, primary care visits and hospital admissions are also typically non-negative and skewed with high mass points.  Utilization can be modeled as counts using poisson, negative binomial, or zero-inflated poisson and zero inflated negative binomial models in a GLM framework although not discussed here.

    References:

    Mullahy, John. "Much Ado Abut Two: Reconsidering Retransformation And The Two-Part Model In Health Econometrics," Journal of Health Economics, 1998, v17(3,Jun), 247-281.

    Liu L, Cowen ME, Strawderman RL, Shih Y-CT. A Flexible Two-Part Random Effects Model for Correlated Medical Costs. Journal of health economics. 2010;29(1):110-123. doi:10.1016/j.jhealeco.2009.11.010.

    Too much ado about two-part models
    and transformation? Comparing methods of modeling Medicare expenditures
    Melinda Beeuwkes Buntin a,∗, Alan M. Zaslavsky
    Journal of Health Economics 23 (2004) 525–542

    REVIEW OF STATISTICAL METHODS FOR ANALYSING HEALTHCARE RESOURCES AND COSTS
    BORISLAVA MIHAYLOVAa,, ANDREW BRIGGSb, ANTHONY O’HAGANcand SIMON G. THOMPSON
    Health Econ. 20: 897–916 (2011)

    Generalized modeling approaches to risk adjustment of skewed outcomes data.
    J Health Econ. 2005 May;24(3):465-88.
    Manning WG1, Basu A, Mullahy J.

    Econometric Modeling of Health Care Costs and Expenditures: A Survey of Analytical Issues and Related Policy Considerations . John Mullahy. Medical Care. Vol. 47, No. 7, Supplement 1: Health Care Costing: Data, Methods, Future Directions (Jul., 2009), pp. S104-S108

    Analyzing Health Care Costs: A Comparison of
    Statistical Methods Motivated by Medicare Colorectal Cancer Charges. MICHAEL GRISWOLD, GIOVANNI PARMIGIANI,ARNIE POTOSKY,JOSEPH LIPSCOMB. Biostatistics (2004), 1, 1, pp. 1–23

    Estimating log models: to transform or not to transform? Willard G. Manning and John Mullahy. Journal of Health Economics 20 (2001) 461–494

    Angrist, J.D. Estimation of Limited Dependent Variable Models With Dummy Endogenous Regressors: Simple Strategies for Empirical Practice. Journal of Business & Economic Statistics January 2001, Vol. 19, No. 1.

    P Dier, D Yanez, A Ash, M Hornbrook, DY Lin. Methods for analyzing health care utilization and costs Ann Rev Public Health (1999) 20:125-144
    Lachenbruch P. A. 2001. “Comparisons of two-part models with competitors” Statistics in Medicine, 20:1215–1234.

    Lachenbruch P.A. 2001. “Power and sample size requirements for two-part models” Statistics in Medicine, 20:1235–1238.

     Diehr,P. ,Yanez,D. Ash, A. Hornbrook, M. & Lin, D. Y. 1999 “Methods for analyzing health care utilization and costs.” Annu. Rev. Public Health, 20:125–44.

    Saturday, October 20, 2018

    Power and Sample Size Analysis in Applied Econometrics

    Recently I was thinking about a conversation from an episode of the EconTalk podcast with Russ Roberts and John Ioannidis where the topic of power came up. Russ Roberts says:

    “though I was trained as a Ph.D., got a Ph.D. in economics at the U. of Chicago, I never heard that phrase, 'power,' applied to a statistical analysis. What we did--and I think what most economists, many economists, still do, is: we had a data set; we had something we wanted to discover and test or examine or explore, depending on the nature of the problem.”

    This made me think, but how can I know something about this but well known trained PhD economists who went to better schools than me weren't taught this? So I went back and looked at all of my copies of econometrics textbooks. These are well known and have been commonly used by masters and PhD graduate students in economics. Econometric Analysis by Greene, Econometric Analysis of Cross Section and Panel Data by Wooldridge,  A Course in Econometrics by Goldberger, A Guide to Econometrics by Kennedy, Using Econometrics by Studenmund. I even threw in Mastering 'Metrics and Mostly Harmless Econometrics by Angrist and Pischke.

    While Wooldridge did discuss clustering and stratified sampling, most of the emphasis was placed on getting the correct standard errors and appropriate weighting. From my previous years of referencing these texts, as well as a cursory review again of the index and chapters of each one I could not find any treatment of power or sample size calculations.

    So I thought, maybe this is something covered in prerequisite courses. Going back to the undergraduate level in economics I recall very little about this. Checking a popular text, Statistics for Business and Economics by Anderson, Sweeney, Williams, Camm, and Cochran I did find a basic example in relation to power and sample sizes for a t-test.  What about a graduate level pre-requisite for econometrics? In my first year of graduate school I took a graduate level course in mathematical statistics (this was a course doing business under a research methods title) that used Degroot's text Probability and Statistics. Definitely a lot about the concept of power in theory, but no emphasis on various calculations for sample size. One textbook I own with treatment of this is Principles and Procedures of Statistics, A Biometrical Approach by Steel, Torrie, and Dickey. But that does not count because that was the text used in my experimental design course in graduate school in the department of agriculture. Not part of a standard econometrics curriculum.

    I've come to the conclusion that power and sample size analysis may not be widely emphasized in graduate econometrics training across the board in all programs. It's not something missed in a lecture a decade ago. Similar to advanced specialized topics like spatial econometrics, details related to power and sample size analysis, survey design, stratified random sampling etc. are likely covered depending on one's specialty in the field and the program.

    However,  it is evident that some economists do this kind of work.

    For instance, here is an example from a paper with food economist Jayson Lusk:

    "However, there are many economic problems where sample size directly affects a benefit or loss function. In these cases, sample size is an endogenous variable that should be considered jointly with other choice variables in an optimization problem. In this article we introduce an economic approach to sample size determination utilizing a Bayesian decision theoretic framework."

    As well as healthcare economist Austin Frakt. 

    So why do we care about power and sample size and what is 'power'?

    Jim Manzi, Author of Uncontrolled: The Surprising Payoff of Trial-and-Error for Business, Politics, and Society offers the following analogy in an Econ Talk podcast:

    “Well, the power in a statistical experiment, and I often use this analogy, is sort of like the magnification power on the microscope you probably used in high school biology. It has on the side, 4x, 8x, 16x, which is how many times it can increase the apparent size of a physical object. And the metaphor I'd use is, if I try and use a child's microscope to carefully observe a section of a leaf looking for an insect that's a little smaller than an ant, and I don't observe the ant, I can reliably say: I don't see the insect, and therefore there is no bug there. If I use that exact same microscope to try and find on that exact same piece of leaf, not a bug but a tiny microbe that's, you know, smaller than a speck of dust, I'll look at it and I'll say: it's all kind of fuzzy, I see a lot of squiggly things; I think that little squiggle might be something or it might not. I don't see the microbe, but I can't reliably say that therefore there is no microbe there, because trying to zoom in closer and closer to look for something that small, all I see is a bunch of fuzz. So my failure to see the microbe is a statement about the precision of my instrument, not about whether there's really a microbe on the leaf.”

    So, if we have a sample that is ‘not sufficiently powered’ it is possible that we could fail to find a relationship between treatment and outcome, even if one actually exists. Equivalently, our estimated regression coefficient may not be statistically significant when a relationship actually does exist. Increasing sample size is one primary way to increase power in an experiment. So the question becomes how large does ‘n’ have to be to have a sample sufficiently powered to detect the effect of a treatment on an outcome (at some stated level of significance)?

    So how do you do these calculations? If you can't find examples in your econometrics textbook (if you do find one let me know!) there are plenty of texts in the biostatistics genre that probably cover this. Principles and Procedures of Statistics, A Biometrical Approach by Steel, Torrie, and Dickey is one example that I started with. Cochran, W (1977). Sampling. Techniques, 3rd ed. is another often cited source.

    It is important to emphasize, power analysis is much more than simply finding and using the right formula to perform a calculation. It is a structured way to think about how to connect a business problem or research question to an experimental design, analysis, or methodology and requires critical thinking about our data and the question we are trying to ask.

    See also: Andrew Gelman on Econtalk discussing "what does not kill my statistical significance makes it stronger"

    Thursday, May 24, 2018

    Statistical Inference vs. Causal Inference vs. Machine Learning: A motivating example

    In his well known paper, Leo Breiman discusses the 'cultural' differences between algorithmic (machine learning) approaches and traditional methods related to inferential statistics. Recently, I discussed how important understanding these kinds of distinctions are when it comes to understanding how current automated machine learning tools can be leveraged in the data science space.

    In his paper Leo Breiman states:

    "Approaching problems by looking for a data model imposes an apriori straight jacket that restricts the ability of statisticians to deal with a wide range of statistical problems."

    On the other hand, Susan Athey's work highlights the fact that no one has developed the asymptotic theory necessary to adequately address causal questions using methods from machine learning (i.e. how does a given machine learning algorithm fit into the context of the Rubin Causal Model/potential outcomes framework?)

    Dr. Athey is working to bridge some of this gap, but it's very complicated. I think there is a lot that can also be done, just understanding and communicating about the differences between inferential and causal questions vs. machine learning/predictive modeling questions. When should each be used for a given business problem? What methods does this entail?

    In an MIT Data Made to Matter podcast, economist Joseph Doyle discusses his paper investigating the relationship between more aggressive (and expensive) treatments by hospitals and improved outcomes for medicare patients. Using this as an example, I hope to broadly illustrate some of these differences looking at this problem through all three lenses.

    Statistical Inference

    Suppose we just want to know if there is a significant relationship between aggressive treatments 'A' and health outcomes (mortality) 'M.' We might estimate a regression equation (similar to one of the models in the paper) such as:

    M = b0 + b1*A + b2*X + e where X is a vector of relevant controls.

    We would be very careful about the nature of our data, correct functional form, and getting our standard errors correct to make valid inferences about our estimate 'b1' of the relationship between aggressive treatments A and mortality M. A lot of this is traditionally taught in econometrics, biostatistics, and epidemiology (things like heteroskedasticity, multicollinearity, distributional assumptions related to the error terms etc.)

    Causal Inference

    Suppose we wanted to know if the estimate b1 in the equation above is causal. In Doyle's paper they discuss some of the challenges:

    "A major issue that arises when comparing hospitals is that they may treat different types of patients. For example, greater treatment levels may be chosen for populations in worse health. At the individual level, higher spending is strongly associated with higher mortality rates, even after risk adjustment, which is consistent with more care provided to patients in (unobservably) worse health. At the hospital level, long-term investments in capital and labor may reflect the underlying health of the population as well. Differences in unobservable characteristics may therefore bias results toward finding no effect of greater spending."

    One of the points he is making is that even if we control for everything we typically measure in these studies (captured by X above) there are unobservable characteristics related to patients that weaken our estimate of b1. Recall that methods like regression and matching (which are two flavors of identification strategies based on selection on observables) achieve identification by assuming that conditional on observed characteristics (X), selection bias disappears.  We want to make conditional on X comparisons of Y (or M in the model above) that mimic as much as possible the experimental benchmark of random assignment (see more on matching estimators here.)

    However, if there are important characteristics related to selection that we don't observe and can't include in X, then in order to make valid causal statements about our results, we need a method that identifies treatment effects within a selection on 'un'-observables framework. (examples include difference-in-differences, fixed effects, and instrumental variables).

    In Doyle's paper, they used ambulance service as an instrument for hospital choice to make causal statements about A.

    Machine Learning/Predictive Modeling

    Suppose we just want to predict mortality by hospital to support some policy or operational objective where the primary need is accurate predictions. A number of algorithmic methods might be exploited including logistic regression, decision trees, random forests, neural networks etc. Based on the mixed findings in the literature, a machine learning algorithm may not exploit 'A' at all even though Doyle finds a significant causal effect based on his instrumental variables estimator. The point is, in many cases a black box algorithm that includes or excludes treatment intensity as a predictor doesn't really care about the significance of this relationship or its causal mechanism, as long as at the end of the day the algorithm predicts well out of sample and maintains reliability and usefulness in application over time.

    Discussion

    If we wanted to know if the relationship between intensity of care 'A' was statistically significant or causal, we would not rely on machine learning methods. At least nothing available on the shelf today pending further work by researchers like Susan Athey. We would develop the appropriate causal or inferential model designed to answer the particular question at hand. In fact, as Susan Athey points out in a past Quora commentary, models used for causal inference could possibly give worse predictions:

    "Techniques like instrumental variables seek to use only some of the information that is in the data – the “clean” or “exogenous” or “experiment-like” variation in price—sacrificing predictive accuracy in the current environment to learn about a more fundamental relationship that will help make decisions...This type of model has not received almost any attention in ML."

    The point is, for the data scientist caught in the middle of so much disruption related to tools like automated machine learning, as well as technologies producing and leveraging large amounts of data, it is important to focus on business understanding and map the appropriate method to address what is trying to be achieved. The ability to understand the differences in tools and methodologies related to statistical inference, causal inference, and machine learning and explaining those differences to stakeholders will be important to prevent 'straight jacket' thinking about solutions to complex problems.

    References:

    Doyle, Joseph et al. “Measuring Returns to Hospital Care: Evidence from Ambulance Referral Patterns.” The journal of political economy 123.1 (2015): 170–214. PMC. Web. 11 July 2017.
    https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4351552/

    Matt Bogard. "A Guide to Quasi-Experimental Designs" (2013)
    Available at: http://works.bepress.com/matt_bogard/24/

    Sunday, March 18, 2018

    Will there be a credibility revolution in data science and AI?

    Summary: Understanding where AI and automation are going to be the most disruptive to data scientists in the near term relates to understanding methodological differences between explaining and predicting. It will require the ability to ask a different kind of question than machine learning algorithms are capable of answering off of the shelf today. At the end of the day we have to think about how the world adjusts along multiple margins. This may put a premium on soft skills. At the heart of causal inference, it's not about data and algorithms. As Judea Pearl says, it's about something that is not in the data to begin with. A credibility revolution in AI that embraces causal inference will only get us as far as good theory can take us.

    There is a lot of enthusiasm about the disruptive role of automation and AI in data science. Products like H20ai and DataRobot offer tools to automate or fast track many aspects of the data science work stream. Combined with other AI based applications and machine learning, large language models (LLMs) will likely help us free up resources from routine tasks, gather information, and have more time to focus and ask better questions. If this trajectory continues, what will the work of the future data scientist look like?

    Many have already pointed out the very difficult task of automating the soft skills possessed by data scientists. In a previous LinkedIn post I discussed this in the trading space where automation and AI could create substantial disruptions for both data scientists and traders. Here I quoted Matthew Hoyle:

    "Strategies have a short shelf life-what is valuable is the ability and energy to look at new and interesting things and put it all together with a sense of business development and desire to explore"

    Understanding disruption from Ai and automation will also depend largely on making a distinction between explaining and predicting. 

    Once armed with predictions, businesses will start to ask questions about 'why'. This will transcend prediction or any of the visualizations of the patterns and relationships coming out of black box prediction algorithms or simply summarizing and parroting pre-existing information. Decision makers will need to know what decisions or factors are moving the needle on revenue or customer satisfaction and engagement or improved efficiencies. Essentially they will want to ask questions related to causality [i.e. compared to what? what makes a difference?]  There is a significant difference between understanding what drivers correlate with or 'predict' the outcome of interest and what actually makes a difference in the outcome. What they will be asking for is a different paradigm or a credibility revolution in AI and data science.

    What do we mean by a credibility revolution?

    Economist Jayson Lusk puts it well:

    "Fortunately economics (at least applied microeconomics) has undergone a bit of credibility revolution.  If you attend a research seminar in virtually any economi(cs) department these days, you're almost certain to hear questions like, "what is your identification strategy?" or "how did you deal with endogeneity or selection?"  In short, the question is: how do we know the effects you're reporting are causal effects and not just correlations."

    Healthcare Economist Austin Frakt has a similar take:

    "A “research design” is a characterization of the logic that connects the data to the causal inferences the researcher asserts they support. It is essentially an argument as to why someone ought to believe the results. It addresses all reasonable concerns pertaining to such issues as selection bias, reverse causation, and omitted variables bias. In the case of a randomized controlled trial with no significant contamination of or attrition from treatment or control group there is little room for doubt about the causal effects of treatment so there’s hardly any argument necessary. But in the case of a natural experiment or an observational study causal inferences must be supported with substantial justification of how they are identified. Essentially one must explain how a random experiment effectively exists where no one explicitly created one."

    How are these questions and differences unlike your typical machine learning application? Susan Athey does a great job explaining in a Quora response about how causal inference is different from off the shelf machine learning methods (the kind being automated today):

    "Sendhil Mullainathan (Harvard) and Jon Kleinberg with a number of coauthors have argued that there is a set of problems where off-the-shelf ML methods for prediction are the key part of important policy and decision problems.  They use examples like deciding whether to do a hip replacement operation for an elderly patient; if you can predict based on their individual characteristics that they will die within a year, then you should not do the operation...Despite these fascinating examples, in general ML prediction models are built on a premise that is fundamentally at odds with a lot of social science work on causal inference. The foundation of supervised ML methods is that model selection (cross-validation) is carried out to optimize goodness of fit on a test sample. A model is good if and only if it predicts well. Yet, a cornerstone of introductory econometrics is that prediction is not causal inference.....Techniques like instrumental variables seek to use only some of the information that is in the data – the “clean” or “exogenous” or “experiment-like” variation in price—sacrificing predictive accuracy in the current environment to learn about a more fundamental relationship that will help make decisions...This type of model has not received almost any attention in ML."

    Developing an identification strategy, as Jayson Lusk discussed above, and all that goes along with that (finding natural experiments or valid instruments, or navigating the garden of forking paths related to propensity score matching or a number of other quasi-experimental methods) involves careful considerations and decisions to be made and defended in ways that would be very challenging to automate. Even when human's do this there is rarely a single best approach to these problems. They are far from routine. Just ask anyone that has been through peer review or given a talk at an economics seminar or conference. 

    The kinds of skills that will be useful in this space would be similar to those of the econometrician or epidemiologist or any quantitative decision maker comfortable with the norms and practices that have evolved out of the credibility revolution.. as data science thought leader Eugene Dubossarsky puts it:

    “the most elite skills…the things that I find in the most elite data scientists are the sorts of things econometricians these days have…bayesian statistics…inferring causality” 

    Noone has a crystal ball.  It is not to say that the current advances in automation are falling short on creating value [just look again at the growing interest in LLMs]. They should no doubt create value like any other form of capital complementing the labor and soft skills of the data scientist. And as mentioned above they could free up more resources to focus on more causal questions that previously may not have been answered. I discussed this type of synergy previously in a related post before LLMs when the hype was mostly focused on 'big data':

     "correlations or 'flags' from big data might not 'identify' causal effects, but they are useful for prediction and might point us in directions where we can more rigorously investigate causal relationships if interested" 

    If automation of causal inference is possible, it will require a different approach than what we have seen so far. We might look to the pioneering work that Susan Athey is doing converging machine learning and causal inference:

    "I’m also working on developing statistical theory for some of the most widely used and successful estimators, like random forests, and adapting them so that they can be used to predict an individual’s treatment effects as a function of their characteristics. For example, I can tell you for a particular individual, given their characteristics, how they would respond to a price change, using a method adapted from regression trees or random forests. This will come with a confidence interval as well." 

    Or - ongoing research related to causal structure discovery (CSD) as discussed recently in Nature.

    At the end of the day however, when thinking about the disruption of AI and automation we have to think about how the world adjusts along multiple margins. This may put a premium on soft skills. At the heart of causal inference, its not about data and algorithms, instead as Judea Pearl says, its about something that is not in the data to begin with. It is a way of thinking. As noted before from the 10th edition of Heyne, Boettke, and Pryschitko's The Economic Way of Thinking:

    "We can observe facts, but it takes a theory to explain the causes. It takes a theory to weed out the irrelevant facts from the relevant ones....Our observations of the world are in fact drenched with theory, which is why we can usually make sense out of the buzzing confusion that assaults our eyes and ears. Actually we observe only a small fraction of what we "know," a hint here and a suggestion there. The rest we fill in from the theories we hold: small and broad, vague and precise..."

    A credibility revolution in AI that embraces causality will only get us as far as good theory can take us. 


    [This post was updated on June 9, 2023]

    Additional References:

    Shen, X., Ma, S., Vemuri, P. et al. Challenges and Opportunities with Causal Discovery Algorithms: Application to Alzheimer’s Pathophysiology. Sci Rep 10, 2975 (2020). https://doi.org/10.1038/s41598-020-59669-x

    From 'What If?' To 'What Next?' : Causal Inference and Machine Learning for Intelligent Decision Making https://sites.google.com/view/causalnips2017

    Susan Athey on Machine Learning, Big Data, and Causation http://www.econtalk.org/archives/2016/09/susan_athey_on.html 

    Machine Learning and Econometrics (Susan Athey, Guido Imbens) https://www.aeaweb.org/conference/cont-ed/2018-webcasts 

    Related Posts:

    Why Data Science Needs Economics
    http://econometricsense.blogspot.com/2016/10/why-data-science-needs-economics.html

    To Explain or Predict
    http://econometricsense.blogspot.com/2015/03/to-explain-or-predict.html

    Culture War: Classical Statistics vs. Machine Learning: http://econometricsense.blogspot.com/2011/01/classical-statistics-vs-machine.html 

    HARK! - flawed studies in nutrition call for credibility revolution -or- HARKing in nutrition research  http://econometricsense.blogspot.com/2017/12/hark-flawed-studies-in-nutrition-call.html

    Econometrics, Math, and Machine Learning
    http://econometricsense.blogspot.com/2015/09/econometrics-math-and-machine.html

    Big Data: Don't Throw the Baby Out with the Bathwater
    http://econometricsense.blogspot.com/2014/05/big-data-dont-throw-baby-out-with.html

    Big Data: Causality and Local Expertise Are Key in Agronomic Applications
    http://econometricsense.blogspot.com/2014/05/big-data-think-global-act-local-when-it.html

    The Use of Knowledge in a Big Data Society II: Thick Data
    https://www.linkedin.com/pulse/use-knowledge-big-data-society-ii-thick-matt-bogard/ 

    The Use of Knowledge in a Big Data Society
    https://www.linkedin.com/pulse/use-knowledge-big-data-society-matt-bogard/ 

    Big Data, Deep Learning, and SQL
    https://www.linkedin.com/pulse/deep-learning-regressionand-sql-matt-bogard/

    Economists as Data Scientists
    http://econometricsense.blogspot.com/2012/10/economists-as-data-scientists.html 

    Monday, August 7, 2017

    Confidence Intervals: Fad or Fashion

    Confidence intervals seem to be the fad among some in pop stats/data science/analytics. Whenever there is mention of p-hacking, or the ills of publication standards, or the pitfalls of null hypothesis significance testing, CIs almost always seem to be the popular solution.

    There are some attractive features of CIs. This paper provides some alternative views of CIs, discusses some strengths and weaknesses, and ultimately proposes that they are on balance superior to p-values and hypothesis testing. CIs can bring more information to the table in terms of effect sizes for a given sample however some of the statements made in this article need to be read with caution. I just wonder how much the fascination with CIs is largely the result of confusing a Bayesian interpretation with a frequentist application or just sloppy misinterpretation. I completely disagree that they are more straight forward to students (compared to interpreting hypothesis tests and p-values as the article claims).

    Dave Giles gives a very good review starting with the very basics of what is a parameter vs. an estimator vs. an estimate, sampling distributions etc. After reviewing the concepts key to understanding CIs he points out two very common interpretations of CIs that are clearly wrong:

    1) There's a 95% probability that the true value of the regression coefficient lies in the interval [a,b].
    2) This interval includes the true value of the regression coefficient 95% of the time.

    "we really should talk about the (random) intervals "covering" the (fixed) value of the parameter. If, as some people do, we talk about the parameter "falling in the interval", it sounds as if it's the parameter that's random and the interval that's fixed. Not so!"

    In Robust misinterpretation of confidence intervals, the authors take on the idea that confidence intervals offer a panacea for interpretation issues related to null hypothesis significance testing (NHST):

    "Confidence intervals (CIs) have frequently been proposed as a more useful alternative to NHST, and their use is strongly encouraged in the APA Manual...Our findings suggest that many researchers do not know the correct interpretation of a CI....As is the case with p-values, CIs do not allow one to make probability statements about parameters or hypotheses."

    The authors present evidence about this misunderstanding by presenting subjects with a number of false statements regarding confidence intervals (including the two above pointed out by Dave Giles) and noting the frequency of incorrect affirmations about their truth.

    In Mastering 'Metrics, Angrist and Pishcke give a great interpretation of confidence intervals that doesn't lend itself in my opinion as easily to abusive probability interpretations:

    "By describing a set of parameter values consistent with our data, confidence intervals provide a compact summary of the information these data contain about the population from which they were sampled"

    Both hypothesis testing and confidence intervals are statements about the compatibility of our observable sample data with population characteristics of interest. The ASAreleased a set of clarifications on statements on p-values. Number 2 states that "P-values do not measure the probability that the studied hypothesis is true." Nor does a confidence interval (again see Ranstan, 2014).

    Venturing into the risky practice of making imperfect analogies, take this loosely from the perspective of criminal investigations. We might think of confidence intervals as narrowing the range of suspects based on observed evidence, without providing specific probabilities related to the guilt or innocence of any particular suspect. Better evidence narrows the list, just as better evidence in our sample data (less noise) will narrow the confidence interval.

    I see no harm in CIs and more good if they draw more attention to practical/clinical significance of effect sizes. But I think the temptation to incorrectly represent CIs can be just as strong as the temptation to speak boldly of 'significant' findings following an exercise in p-hacking or in the face of meaningless effect sizes. Maybe some sins are greater than others and proponents feel more comfortable with misinterpretations/overinterpretations of CIs than they do with misinterpretations/overinterpretaions of p-values.

    Or as Briggs concludes about this issue:

    "Since no frequentist can interpret a confidence interval in any but in a logical probability or Bayesian way, it would be best to admit it and abandon frequentism"


    Methods of Psychological Research Online 1999, Vol.4, No.2 © 1999 PABST SCIENCE PUBLISHERS Confidence Intervals as an Alternative to Significance Testing Eduard Brandstätter1 Johannes Kepler Universität Linz

    J. Ranstam, Why the -value culture is bad and confidence intervals a better alternative, Osteoarthritis and Cartilage, Volume 20, Issue 8, 2012, Pages 805-808, ISSN 1063-4584, http://dx.doi.org/10.1016/j.joca.2012.04.001 (http://www.sciencedirect.com/science/article/pii/S1063458412007789)

    Robust misinterpretation of confidence intervals
    Rink Hoekstra & Richard D. Morey & Jeffrey N. Rouder &
    Eric-Jan Wagenmakers Psychon Bull Rev
    DOI 10.3758/s13423-013-0572-3 2014

    Friday, July 21, 2017

    Regression as a variance based weighted average treatment effect

    In Mostly Harmless Econometrics Angrist and Pischke discuss regression in the context of matching. Specifically they show that regression provides variance based weighted average of covariate specific differences in outcomes between treatment and control groups. Matching gives us a weighted average difference in treatment and control outcomes weighted by the empirical distribution of covariates. (see more here). I wanted to roughly sketch this logic out below.

    Matching

     δATE = E[y1i | Xi,Di=1] - E[y0i | Xi,Di=0] = ATE

    This gives us the average difference in mean outcomes for treatment and control  (y1i,y0i ⊥ Di) i.e. in a randomized controlled experiment potential outcomes are independent from treatment status

    We represent the matching estimator empirically by:

     Σ δx P(Xi,=x) where δx is the difference in mean outcome values between treatment and control units at a particular value of X, or  difference in outcome for a particular combination of covariates (y1,y0 ⊥ Di|xi) i.e. conditional independence assumed- hence identification is achieved through a selection on observables framework.

    
Average differences δx are weighted by  the distribution of covariates via the term P(Xi,=x).

    Regression

    We can represent a regression parameter using the basic formula taught to most undergraduates:

    Single Variable: β = cov(y,D)/v(D)
    Multivariable:  βk = cov(y,D*)/v(D*)

    where  D* = residual from regression of D on all other covariates and 
E(X’X)-1E(X’y) is a vector with the kth element cov(y,x*)/v(x*) where x* is the residual from regression of that particular ‘x’ on all other covariates.

    We can then represent the estimated treatment effect from regression as:

     δR = cov(y,D*)/v(D*) = E[(Di-E[Di|Xi])E[yiIDiXi] / E[(Di-E[Di|Xi])^2]  assuming (y1,y0 ⊥ Di|xi)

    Again regression and matching rely on similar identification strategies based on selection on observables/conditional independence.

    Let E[yi | DiXi] = E[yi | Di =0,Xi] + δx Di

    Then with more algebra we get: δR = cov(y,D*)/v(D*) = E[σ^2D(Xi) δx]/ E[σ^2D(Xi)]

    where σ^2D(Xi) is the conditional variance of treatment D given X or  E{E[(Di –E[Di|Xi])^2|Xi]}.

    While the algebra is cumbersome and notation heavy, we can see that the way most people are familiar with viewing a regression estimate cov(y,D*)/v(D*)  is equivalent to the term (using expectations)  E[σ2D(Xi) δx]/ E[σ2D(Xi)] , and we can see that this term contains the product of the conditional variance of D and our covariate specific differences in treatment and controls δx.

    Hence, regression gives us a variance based weighted average treatment effect, whereas matching provides a distribution weighted average treatment effect.

    So what does this mean in practical terms? Angrist and Piscke explain that regression puts more weight on covariate cells where the conditional variance of treatment status is the greatest, or where there are an equal number of treated and control units. They state that differences matter little when the variation of δx is minimal across covariate combinations.

    In his post The cardinal sin of matching, Chris Blattman puts it this way:

    "For causal inference, the most important difference between regression and matching is what observations count the most. A regression tries to minimize the squared errors, so observations on the margins get a lot of weight. Matching puts the emphasis on observations that have similar X’s, and so those observations on the margin might get no weight at all....Matching might make sense if there are observations in your data that have no business being compared to one another, and in that way produce a better estimate" 

    Below is a very simple contrived example. Suppose our data looks like this:
    We can see that those in the treatment group tend to have higher outcome values so a straight comparison between treatment and controls will overestimate treatment effects due to selection bias:

     E[Y­­­i|di=1] - E[Y­­­i|di=0] =E[Y1i-Y0i]  +{E[Y0i|di=1] - E[Y0i|di=0]}

     However, if we estimate differences based on an exact matching scheme, we get a much smaller estimate of .67. If we run a regression using all of the data we get .75. If we consider 3.78 to be biased upward then both matching and regression have significantly reduced it, and depending on the application the difference between .67 and .75 may not be of great consequence. Of course if we run the regression including only matched variables, we get exactly the same results. (see R code below). This is not so different than the method of trimming based on propensity scores suggested in Angrist and Pischke.


    Both methods rely on the same assumptions for identification, so noone can argue superiority of one method over the other with regard to identification of causal effects.

    Matching has the advantage of having a nonparametric, alleviating concerns with functional form. However, there are lots of considerations to work through in matching (i.e. 1:1, 1:many, optimal caliper width, variance/bias tradeoff and kernel selection etc.). While all of these possibilities might lead to better estimates, I wonder if they don't sometimes lead to a garden of forking paths.

    See also: 

    For a neater set of notes related to this post, see:

    Matt Bogard. "Regression and Matching (3).pdf" Econometrics, Statistics, Financial Data Modeling (2017). Available at: http://works.bepress.com/matt_bogard/37/

    Using R MatchIt for Propensity Score Matching

    R Code:

    # generate demo data
    x <- c(4,5,6,7,8,9,10,11,12,1,2,3,4,5,6,7,8,9)
    d <- c(1,1,1,1,1,1,1,1,1,0,0,0,0,0,0,0,0,0)
    y <- c(6,7,8,8,9,11,12,13,14,2,3,4,5,6,7,8,9,10)

    summary(lm(y~x+d)) # regression controlling for x