Showing posts with label agricultural economics. Show all posts
Showing posts with label agricultural economics. Show all posts

Sunday, April 6, 2025

Agricultural Economics as a Poster Child of Applied Economics

 Abstract

Agricultural economists have embodied the notions of applied economics for a long time. They have used economic principles to address real-world problems, integrating economics and scientific knowledge. Applied economics tends to be multidisciplinary and develop applied concepts, theories, and tools. Some, like human capital, diffusion of innovation, contingent valuation, and numerous numerical and econometric techniques have spread throughout economics. Agricultural economic research has been data intensive, and improved information technologies strengthen this tendency. Yet data without theory is of limited use and coevolution of theory and data are essential. Empirical analysis should incorporate quantitative information as well as narratives. We are challenged to understand the coevolution of business, supply chains, and technology, and how they are affected by policies and affect markets. Research should integrate agriculture, energy, and the environment and develop tools to analyze and regulate the emerging bio-economy integrating biotech and infotech.

Zilberman, D. (2019), Agricultural Economics as a Poster Child of Applied Economics: Big Data & Big Issues1. American Journal of Agricultural Economics, 101: 353-364. https://doi.org/10.1093/ajae/aay101

Saturday, March 13, 2021

Why Study Economics/Applied Economics?

Applied Economics is a broad field with many applications.

Applied Economics is a broad field of study covering many topics. Recognizing the wide range of applications has led departments of Agricultural Economics across numerous universities to change their degree program names to Applied Economics.  In 2008, the American Agricultural Economics Association changed its name to the Agricultural and Applied Economics Association (AAEA).

This trend is noted in research published in the journal Applied Economic Perspectives and Policy:

"Increased work in areas such as agribusiness, rural development, and environmental economics is making it more difficult to maintain one umbrella organization or to use the title “agricultural economist” ... the number of departments named" Agricultural Economics” has fallen from 36 in 1956 to 9 in 2007."

This brief podcast from the University of Minnesota's Department of Applied Economics is an example of this trend: 


It discusses the breadth of questions and problems applied economists address in their work including obesity and food systems, environmental and water resource economics, development, growth, trade, and technological change; public sector economics, health policy and management, human resources and industrial relations. Applied research in this area is often interdisciplinary including biology, engineering, health and animal sciences, and nutrition as an example. 

Why study applied economics?  A few inspiring quotes from Southern Illinois University introduction to their programs in Agribusiness Economics:

If you want to prove sustainable resource use saves money and protects the land…
If you understand that the wheat crop here can make a difference for a hungry child across the ocean…  

Applied Economics emphasizes quantitative and analytics skills ideal for careers in data science

Many applied economics master's degrees are designed to serve as a very attractive terminal degree for professionals. 

To quote, from Johns Hopkins University’s Applied Economics program home page:

“Economic analysis is no longer relegated to academicians and a small number of PhD-trained specialists. Instead, economics has become an increasingly ubiquitous as well as rapidly changing line of inquiry that requires people who are skilled in analyzing and interpreting economic data, and then using it to effect decisions ………Advances in computing and the greater availability of timely data through the Internet have created an arena which demands skilled statistical analysis, guided by economic reasoning and modeling.”

Many applied economics programs are STEM designated programs reflecting the emphasis that applied economics places on quantitative and analytics skills. The University of Pittsburg has designed their STEM designated M.S. in Quantitative Economics specifically with data science roles in mind. Virginia Tech offers an online Master of Ag and Applied Economics, the first I have seen in an Agricultural and Applied Economics department specifically designed to incorporate economics with data science and programming.


The focus on causality differentiates economics from other fields.

Once armed with predictions from machine learning an AI, businesses will start to ask questions about  what decisions or factors are moving the needle on revenue or customer satisfaction and engagement or improved efficiencies. Essentially they will want to ask questions related to causality, which requires a completely different paradigm for data analysis.

In a KDnuggets interview, Economist Scott Nicholson (Chief Data Scientist at Accretive Health and formerly at LinkedIn) comments on the differences between economists and data scientists: 

 "In terms of applied work, economists are primarily concerned with establishing causation. This is key to understanding what influences individual decision-making, how certain economic and public policies impact the world, and tells a much clearer story of the effects of incentives. With this in mind, economists care much less about the accuracy of the predictions from their econometric models than they do about properly estimating the coefficients, which gets them closer to understanding causal effects. At Strata NYC 2011, I summed this up by saying: If you care about prediction, think like a computer scientist, if you care about causality, think like an economist."

As data science thought leader Eugene Dubossarsky puts it in a SuperDataScience podcast:

“the most elite skills…the things that I find in the most elite data scientists are the sorts of things econometricians these days have…bayesian statistics…inferring causality” 

Nobel Prize Laureate Joshua Angrist discussed the new opportunities for students graduating with economics and quantitative skills that are available at firms like Amazon because of their interest in causal questions and running experiments:

   

In another interview Angrist emphasizes opportunities for Economics bachelor's degree holders:

"There's a very strong private sector market for economics undergrad especially economics undergrads who have good training in econometrics...like Amazon and Google and Facebook and Trip Adviser they are looking for people that can do some statistics but a lot of the questions that they are interested in are causal questions. What will be the consequences of changing prices for example or changing marketing strategies and these companies have discovered that the best training for that is undergrad work in economics or econometrics. We really specialize in causality in a way regular data science does not.....someone who trains in data science might learn a lot about machine learning but won't necessarily learn about for example instrumental variables or regression discontinuity methods and those turn out to be very useful for the tech sector."

A post at the Uber Engineering blog explains how they find these skills to be valuable in a business setting: 

"One of the most exciting areas we’ve been working on is causal inference, a category of statistical methods that is commonly used in behavioral science research to understand the causes behind the results we see from experiments or observations...causal inference helps us provide a better user experience for customers on the Uber platform. The insights from causal inference can help identify customer pain points, inform product development, and provide a more personalized experience...At a higher level, causal inference provides information that is critical to both improving the user experience and making business decisions through better understanding the impact of key initiatives."

Economics provides a foundation with long lasting value and offers a bright future.

Economics combines mathematically precise theories (like microeconomics) and empirically sound methods (like econometrics) to study people's choices and how they are made compatible. As a social and behavioral science and a quantitative and technical field, learning to think like an economist and applying those skills will never go out of fashion. There are a number of both undergraduate and graduate degree programs in economics and applied economics across the country and I would encourage you to check them out. I've listed a few more examples of applied economics programs below.

***This post is an update to an original post made in September 2010 found here.
 
Related Posts: 


Economists as Data Scientists http://econometricsense.blogspot.com/2012/10/economists-as-data-scientists.html   

References:

'What is the Future of Agricultural Economics Departments and the Agricultural and Applied Economics Association?' By Gregory M. Perry. Applied Economic Perspectives and Policy (2010) volume 32, number 1, pp. 117–134.

Additional Graduate Programs in Applied Economics and Related Fields

Western Kentucky University - M.A. in Applied Economics (Also UG and GR options in Agriculture and Food Science
Murray State University - M.S. Agriculture/Agribusiness Economics 
Virginia Tech - M.S. Ag and Applied Economics
University of Cincinnati - M.S. Applied Economics
Clemson University - M.S. Applied Economics and Statistics
Montana State University - M.S. Applied Economics
Cornell University - M.S. & M.P.S. in Applied Economics and Management 
Oklahoma State University - M.S. Agricultural Economics  and MAg in Agribusiness
Texas A&M - M.S. in Agricultural Economics
North Dakota State University - M.S. Agribusiness and Applied Economics
University of Illinois - M.S. Agricultural and Applied Economics
University of Missouri - Agricultural and Applied Economics
Auburn University - M.S. Agricultural Economics and Rural Sociology (various programs)
AAEA  - Directory of additional programs at the graduate and undergraduate levels


Tuesday, June 6, 2017

Professional Science Master's Degree Programs in Biotechnology and Management

As an undergraduate I always had an interest in biotechnology and molecular genetics. However, lab work did not particularly appeal to me. I also recognized early on that science does not occur in a vacuum- its subject to social, political, economic, and financial forces. This drew me to the field of economics, specifically public choice theory.

When it came time for graduate school I was still torn. I really wasn't interested in an MBA and didn't really have the background to work in a lab or do field work in genetic research. I really liked economics. The combination of mathematically precise theories (microeconomics/game theory) and empirically sound methods (econometrics) provided a powerful framework for applied problem solving.

I had two advisers make recommendations that got me thinking outside the box. One suggested ultimately I would find a niche that combined both economics and genetics. The other suggested I look at programs like the Bioscience Management program that was being offered at the time at George Mason University (now Bioinformatics Management). While there were not a lot of programs like that being offered at the time, the Agriculture Department at Western Kentucky University provided enough flexibility in their masters program to include courses in biostatistics, genetics, and applied economics. I was able to work on research projects analyzing consumer perceptions of biotechnology and biotech trait resistance management using tools from econometrics, game theory, and population genetics.  Additionally I took courses in applied economics and finance from both the Department of Agriculture and College of Business where I was exposed to tools related to investment analysis, options pricing, and analysis and valuation of biotech companies as well as the impacts of technological change and biotechnology on food and economic development.

With this combination of quantitative training and applied work I have been able to leverage SAS, R, and Python to solve a number of challenging problems throughout a number of professional analytics and consulting roles. 

Today there are a larger number of professional science masters programs with curriculums similar to the programs I contemplated over 10 years ago. 

According to National Professional Science Master’s Association:

"Professional Science Master's (PSMs) are designed for students who are seeking a graduate degree in science or mathematics and understand the need for developing workplace skills valued by top employers. A perfect fit for professionals because it allows you to pursue advanced training and excel in science or math without a Ph.D., while simultaneously developing highly-valued business skills....PSM programs consist of two years of coursework along with a professional component that includes business, communications and/or regulatory affairs."

In 2012 there was an article in Science detailing these degrees and some data related to salaries which seemed attractive. According to the article the first program was officially offered in 1997, reaching 140 programs by 2009 with over 247 at the time of printing.

This commentary from the article corroborates how I feel about my experience:

“There is a tendency for students to buy into the line that if you don't get a Ph.D., you're not a serious professional, that you're wasting your mind,” she says. After spending a decade talking with PSM students and graduates, she is certain that’s not true. “There is so much potential for growth and satisfaction with a PSM degree. You can become a person you didn’t even know you wanted to be.”

Below are some programs that would look interesting to me that students interested in this option should check out.  (there is a program locator you can find here) . Many of these programs are a mash up of biology/biotech and applied economics and business degrees.

George Mason University- PSM Bioinformatics Management

University of Illinois - Agricultural Production

Cornell- MPS Agriculture and Life Sciences

Washington State University - PSM Molecular Biosciences

Middle Tennesee State University - PSM Biotechnology

California State - MS Biotechnology/MBA 

Johns Hopkins - MBA/MS Biotechnology

Rice - PSM Bioscience and Health Policy

North Carolina State University - MBA (Biosciences Mgt Concentration)

Purdue/Kelley - MS-MBA  (not a heavy science emphasis but a very cool degree regardles from great schools)

See also:
Analytical Translators
Why Study Agricultural/Applied Economics

Sunday, February 12, 2017

Molecular Genetics and Economics

A really interesting article in JEP:

A slice:

"In fact, the costs of comprehensively genotyping human subjects have fallen to the point where major funding bodies, even in the social sciences, are beginning to incorporate genetic and biological markers into major social surveys. The National Longitudinal Study of Adolescent Health, the Wisconsin Longitudinal Study, and the Health and Retirement Survey have launched, or are in the process of launching, datasets with comprehensively genotyped subjects…These samples contain, or will soon contain, data on hundreds of thousands of genetic markers for each individual in the sample as well as, in most cases, basic economic variables. How, if at all, should economists use and combine molecular genetic and economic data? What challenges arise when analyzing genetically informative data?"


Link:

https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3306008/


Reference:
Beauchamp JP, Cesarini D, Johannesson M, et al. Molecular Genetics and Economics. The journal of economic perspectives : a journal of the American Economic Association. 2011;25(4):57-82.

Wednesday, November 4, 2015

'Big' Data vs. 'Clean' Data

I've previously written about the importance of data cleaning, and recently I was reading a post- Data Science Can Transform Agriculture, If We Get It Right on FarmLink's blog and I was impressed by the following:

"We believe it's a transformational time for this industry – call it Ag 3.0 – when the combination of human know-how and insight, coupled with robust data science and analytics will change the productivity, profitability and sustainability of agriculture."

This reminds me, as I have discussed before in relation to big data in agriculture, of Economist Tyler Cohen's comments on an EconTalk podcast, " the ability to interface well with technology and use it to augment human expertise and judgement is the key to success in the new digital age of big data and automation."

 But in relation to data cleaning, I thought this was really impressive:

"...we disqualified more than two-thirds of the data collected during our first year. Now, we inspect each combine before use following a 50 point check list to identify any problem that could affect accuracy of collection, have developed a world class Quality Assurance process to test the data, and created IP addressable access to our combines to be able to identify and compensate for operator error. As a result, last year over 95% of collected data met our standard for being actionable. Admittedly, our first year data was “big.” But we chose to view it as largely worthless to our customers, just as much of the data being collected through farmer exchanges, open APIs, or memory sticks for example will be. It simply lacks the rigor to justify use in such important undertakings."

It takes patience and discipline to sometimes to make the necessary sacrifices and put the necessary resources into data quality, and it looks like this company gets it. Data cleaning isn't just academic. It's serious. Maybe it's time to replace #BigData with #CleanData.
 
Related:
Big Data
Data Cleaning
Got Data? Probably not like your econometrics textbook!
Big Ag Meets Big Data (Part 1 & Part 2)
Big Data- Causality and Local Expertise are Key in Agronomic Applications
Big Ag and Big Data-Marc Bellemare
Big Data, IoT, Ag Finance, and Causal Inference
In God we trust, all others show me your code.
Data Science, 10% inspiration, 90% perspiration

Wednesday, October 7, 2015

Metrics Monday with Marc Bellemare

 I have been following Marc Bellemare for a while now on twitter (@mfbellemare) and really became interested in his blog because of the proliferate amount of very good posts related to applied econometrics. He also writes about a number of interesting topics related to applied economics and in areas related to his own research. In the last few weeks (months?), he as been running a series of posts titled 'Metrics Monday where he addresses lots of issues related to applied econometrics that aren't always addressed in typical theory based courses. I think every advanced undergraduate, graduate student, or any 'metrics or analytics practitioner should read all of his econometrics related posts. Below are some links to selected posts. I'll probably add to this list as he posts more, and as I discover older related posts I have not yet read.

Some of my favorite 'Metrics Monday posts: 

Friends *do* let friends do IV

When is heteroskedasticity (not) a problem

Hypothesis Testing in Theory and Practice

Data Cleaning

Outliers

Rookie mistakes in empirical analysis

What to do with missing data

Other Applied Econometrics Posts by Marc Bellemare:

Love it or logic, Or: people really care about binary dependent variables

A rant on estimation with binary dependent variables

In defense of the cookbook approach to econometrics

Econometrics teaching needs an overhaul

Do Both

Wednesday, September 30, 2015

Big Data, IoT, Ag Finance, and Causal Inference

Over at my applied economics blog, I recently discussed an article from AgWeb; How the feds interest rate decision affects farmers. This actually got me questioning some of the ramifications of leveraging data analysis in the context of ag lending (from both a farmer and lender perspective), which ultimately lead to me thinking about some interesting questions that would be exciting to investigate:
  1.  Is there a causal relationship between producers that leverage IoT and Big Data analytics applications and farm output/performance/productivity
  2. How do we quantify the outcome-is it some measure of efficiency or some financial ratio?
  3. If we find improvements in this measure-is it simply a matter of selection? Are great producers likely to be productive anyway, with or without the technology?
  4. Among the best producers, is there still a marginal impact (i.e. treatment effect) for those that adopt a technology/analytics based strategy?
  5. Can we segment producers based on the kinds of data collected by IoT devices on equipment, aps, financial records, GPS etc.?  (maybe this is not that much different than the TrueHarvest benchmarking done at FarmLink) and are there differentials in outcomes, farming practices, product use patterns etc. by segment
See also:
Big Ag Meets Big Data (Part 1 & Part 2)
Big Data- Causality and Local Expertise are Key in Agronomic Applications
Big Ag and Big Data-Marc Bellemare
Other Big Data and Agricultural related Application Posts at EconometricSense
Causal Inference and Experimental Design Roundup

Friday, September 11, 2015

Mastering Metrics....and the Grain Markets

I recently just finished two great books, Mastering 'Metrics, and Mastering the Grain Markets.

Mastering the Grain Markets

While I have a background in agricultural and applied economics, my interest was always related to the public choice and the environmental implications of biotechnology, as well as econometrics (hence this blog). So, I didn't really have much formal background related to commodity markets, other than a little exposure to options through a couple of finance classes.  I have certainly read some really good extension publications related to futures, options, and hedging but Mastering the Grain markets by Elaine Kub really brings these issues to life. She brought me back to my crop scouting days in her many discussions of corn production and the agronomics of our major commodities. She also tackles some major issues and controversies associated with modern agriculture, everything from speculation, to biotech to sustainability issues, gluten fad diets and more. Prepare for a trip from gate to plate in this book that teaches like a textbook but reads like a novel!

Even if you think all you are interested in are the specifics around how futures and options work, you'll end up being convinced that the holistic approach is essential. To borrow one quote:

"..any participation in the grain markets is a form of participation in agriculture, and it should be regarded as one piece of a beautiful, challenging, miraculous whole."
 
A couple areas that struck me as particularly interesting were her discussions of counterparty risk and over the counter contracts. I'll probably have a separate post on this blog or my ag econ blog regarding counterparty risk.

So, why share a review about a grain markets book on an applied econometrics blog? Well, all the discussion about OTCs and risk management rekindled my interest in copulas, which I have blogged about before, and also made me a little more curious about index based crop insurance. Risk modeling in commodities go hand in hand with econometrics. Oh, and she even hits on precision agriculture and alludes to big data in agriculture:

"At the end of the growing season, he has every data point he could possibly need (seed population, seed depth, input rates, final yield, soil moisture, etc.) to fine tune his production practices on each GPS mapped square foot of his farm."

Mastering Metrics

Before reading MM, I had previously read Angrist and Pischke's Mostly Harmless Econometrics. It was my first rigorous introduction to the potential outcomes framework and causal inference. It took me a while to work through and I still reference it often. Even though Mastering 'Metrics was supposed to be a 'lite' version or maybe an undergraduate version of MHE, reading in 'reverse' order worked out well. What I really liked was their intro to regression, and the presentation of regression as a matching estimator becomes even more crystal clear to me than it did in MHE. To borrow a quote:

"Specifically, regression estimates are weighted averages of multiple matched comparisons"

I really think a lot of people I encounter have a hard time thinking about that. I also got better insight and clarification on a number of issues related to instrumental variables, regression discontinuity, and difference-in-differences.  Within the IV discussion, I really like the causal-chain of effects presentation and discussion of 'intent to treat', and better understand all the things related to compliers and noncompliers etc. They also really got me up to speed with regard to the differences in parametric vs. non-parametric RD and an important distinction between fuzzy and sharp RD:

"....with fuzzy, applicants who cross a threshold are exposed to a more intense treatment, while with a sharp design, treatment switches cleanly on or off at the cutoff."

Another thing that stood out with me, in their DD chapter they made some clarifications about weighted regression and clustered standard errors that seemed very helpful. Other things in general, I really liked their treatment of the regression anatomy formula and understood it much better in this reading.  Their basic review and treatment of inference, standard errors, and t-statistics is really great and a good way to segway an undergraduate student from an introductory statistics class into the more advanced topics they present later in the text. I could also see certain graduate programs, even outside of economics making use of this text.

Both Mastering the Grain Markets and Mastering Metrics end with a final chapter tying everything together.

I highly recommend both books.

More thoughts....

So above I mentioned that risk modeling and econometrics go hand in hand, but have been thinking, were any of the techniques covered in MM useful for work related to the commodities markets? In terms of informing marketing and risk management strategies, I'm not sure. Maybe some readers have some idea. But, in terms of policy analysis as it relates to commodity markets, perhaps. There are some that advocate that we should restrict speculation in commodity markets. Scott Irwin looked at the impact of index funds on commodity markets, using granger causality (although granger causality was not discussed in MHE or MM).  Other work has relied on panel methods. A quick google search reveals some related work using instrumental variables discussed in MM. For now I'll just say to be continued.....

References:

Irwin, S. H. and D. R. Sanders (2010), “The Impact of
Index and Swap Funds on Commodity Futures Markets:
Preliminary Results”, OECD Food, Agriculture and
Fisheries Working Papers, No. 27, OECD Publishing.
doi: 10.1787/5kmd40wl1t5f-en


Wednesday, August 12, 2015

Index Based Crop Insurance and Big Data

 There is some interesting work going on currently in relation to risk management in the agriculture space as it relates to 'big data.'

"Agriculture risk management is about having access to ‘big data’ since growth conditions, risk types, climate and insurance terms vary largely in space. Solid crop models are based on large databases including simulated weather patterns with tempo-spatial correlations, crop planting areas, soil types, irrigation application, fertiliser use, crop rotation and planting calendars. … Similarly, livestock data need to include livestock densities which drive diseases, disease spread vectors and government contingency plans to address outbreaks of highly contagious diseases…Ultimately, big data initiatives will support selling agriculture insurance policies via smart phones based on highly sophisticated indices and will make agriculture insurance and risk management rapidly scalable and accessible to a large majority of those mostly affected – the farmers." - from The Actuary, March 2015

Similarly, in another issue of The Actuary there is more discussion related to this:

"In more recent years IBI (Indemnity Based Insurance) has received a renewed interest, largely drivenby advances in infrastructure (i.e., weather stations), technology (i.e., remote sensing
and satellites), as well as computing power, which has enabled the development of new statistical and mathematical models. With an IBI contract, indemnities are paid based on some index level, which is highly correlated to actual losses. Possible indices include rainfall, yields, or vegetation levels measured by satellites. When an index exceeds a certain predetermined threshold, farmers receive a fast, efficient payout, in some cases delivered via mobile phones. "


The article notes several benefits related to IBI products, including decreased moral hazard and adverse selection as well as the ability to transfer risk. However some challenges were noted related to 'basis' risk, where the index used to determine payments may not be directly linked to actual losses. In such cases, a farmer may recieve a payment when no loss is realized, or may actually experience loss but the index values don't trigger a payment. The farmer is left feeling like they have paid for something without benefit in the latter case. The article discusses three types of basis risk; variable, spatial, and temporal. Variable risk occurs when other unmeasured factors impact a peril not captured by the index. Maybe its wind speed during pollination or some undocumented pest damage or something vs measured items like temperature or humidity. An example of spatial risk might be related to cases where index data may be data generated from meteorological stations too far from the field location to accurately trigger payments for perils related to rain or temperature.  Temporal risk is really interesting to me in terms of the potential for big data:

"The temporal component of the basis risk is related to the fact that the sensitivity of yield to the insured peril often varies over the crops’ stages of growth. Factors such as changes in planting dates, where planting decisions are made based on the onset of rains, for example, can have a substantial impact on correlation as they can shift critical growth stages, which then do not align with the critical periods of risk assumed when the crop insurance product was designed." 

It would seem to me that the kinds of data elements being capture by services offered by companies like Climate Corp, Farmlink, John Deere etc. in combination of other aps (drones/smartphones/other modes of censoring/data collection) might be informative to creating and monitoring the performance of better indexes to help mitigate the basis risk associated with IBI related products.

References:

New frontiers in agricultural insurance . The Actuary. March 2015. DR AUGUSTE BOISSONNADE
http://www.theactuary.com/features/2015/03/new-frontiers-in-agriculture/

AGRICULTURAL INSURANCE— MORE ROOM TO GROW? The Actuary- May 2015.
Lysa Porth and Ken Seng Tan

See also:

Copula Based Agricultural Risk Models
 Big Ag Meets Big Data (Part 1 & Part 2)

Copula Based Agricultural Risk Models

I have written previously about copulas, with some very elementary examples (see here and here). Below are some papers I have added to my reading list with applications in agricultural risk management. I'll likely followup in the future with some annotation/review.

Zimmer, D. M. (2015), Crop price comovements during extreme market downturns. Australian Journal of Agricultural and Resource Economics. doi: 10.1111/1467-8489.12119

Energy prices and agricultural commodity prices: Testing correlation using copulas method
Krishna H Koirala, Ashok K Mishra, Jeremy M D 'antoni, Joey E Mehlhorn
Energy 01/2015; DOI:10.1016/j.energy.2014.12.055 · 

Xiaoguang Feng, Dermot J. Hayes
Diversifying Systemic Risk in Agriculture: A Copula-based Approach

Mixed-Copula Based Extreme Dependence Analysis: A Case Study of Food and Energy Price Comovements
Feng Qiu and Jieyuan Zhao
Selected Paper prepared for presentation at the Agricultural & Applied Economics Association’s 2014 AAEA Annual Meeting, Minneapolis, MN, July 27-29, 2014.

Price asymmetry between different pork cuts in the USA: a copula approach
Panagiotou and Stavrakoudis Agricultural and Food Economics (2015) 3:6
DOI 10.1186/s40100-015-0029-2

Copula-Based Models of Systemic Risk in U.S. Agriculture: Implications for Crop Insurance and
Reinsurance Contracts Barry K. Goodwin
October 22, 2012

Friday, June 19, 2015

Got Data? Probably not like your econometrics textbook!

Recently there has been a lot of discussion of the Angrist and Pischke piece entitled "Why Econometrics Teaching Needs an Overhaul." (read more...) and I have discussed before the large gap between theoretical and applied econometrics.

But here I plan to discuss another potential gap in teaching and application and this is a topic that often is not introduced at any point in a traditional undergraduate or graduate economics curriculum, and that is hacking skills. This becomes extremely important for economists that someday find themselves doing applied work in a corporate environment, or working in the area of data science. Drew Conway points out the there are three spheres of data science including hacking skills, math and statistics knowledge, and subject matter expertise. For many economists, the hacking sphere might be the weakest (read also Big Data Requires a New Kind of Expert: The Econinformatrician) while their quantitative training otherwise makes them ripe to become very good data scientists.

Drew Conway's Data Science Venn Diagram

In a recent whitepaper, I discuss this issue:

Students of econometrics might often spend their days learning proofs and theorems, and if they are lucky they will get their hands on some data and access to software to actually practice some applied work rather it be for a class project or part of a thesis or dissertation. I have written before about the large gap between theoretical and applied econometrics, but there is another gap to speak of, and it has nothing to do with theoretical properties of estimators or interpreting output from STATA, SAS or R. This has to do with raw coding, hacking, and data manipulation skills; the ability to tease out relevant observations and measures from both large structured transactional databases or unstructured log files or web data like tweet-streams. This gap becomes more of an issue as econometricians move from more academic environments to corporate environments and especially so for those economists that begin to take on roles as data scientists. In these environments, not only is it true that problems don’t fit the standard textbook solutions (see article ‘Applied Econometrics’), but the data doesn't look much like the simple data sets often used in textbooks either.  One cannot always expect their IT people to be able to just dump them a flat file with all the variables and formats that will work for your research project. In fact, the absolute best you might hope for in many environments is a SQL or Oracle data base with hundreds or thousands of tables and the tiny bits of information you need spread across a number of them. How do you bring all of this information together to do an analysis? This can be complicated, but for the uninitiated I will present some ‘toy’ examples to give a feel for executing basic database queries to bring together different pieces of information housed in separate tables in order to produce a ‘toy’ analytics ready data set.

I am certain that many schools actually do teach some of the basics related to joining and cleaning data sets, and if they don't then others might figure this out on the job or through one research project or another. I am not certain that this gap needs to be filled  necessarily as part of any econometrics course. However, it is something students need to be aware of and offering some sort of workshop, lab or formal course (maybe as part of a more comprehensive data science curriculum like this) would be very beneficial.

Read the whole paper here:

Matt Bogard. 2015. "Joining Tables with SQL: The most important econometrics lesson you may ever learn" The SelectedWorks of Matt Bogard
Available at: http://works.bepress.com/matt_bogard/29  

See also: Is Machine Learning Trending with Economists?

Wednesday, June 17, 2015

Farmlink and the Rise of Data Science in Agriculture


At a recent Global Ag Investing Conference Dave Gebhardt (Chief Strategy Officer for FarmLink ) spoke about the rise of data science in agriculture. You can read the story and find a link to the podcast here:

In the podcast he discusses the way data science is revolutionizing agriculture, and how we are at a "tipping point where advances in science, IT, technology, and computing power have put a whole new level of opportunities before us."

This sounds a lot like what I have previously discussed in relation to big data and the internet of things: 

Watch more about how FarmLink is leveraging IoT, big data, and advanced analytics:



Related:

 Big Ag Meets Big Data (Part 1 & Part 2)

Sunday, March 8, 2015

Are observational studies off the table or can you have your eggs and eat them too?


From the New York Times:

"But the primary problem is that nutrition policy has long relied on a very weak kind of science: epidemiological, or “observational,” studies in which researchers follow large groups of people over many years. But even the most rigorous epidemiological studies suffer from a fundamental limitation. At best they can show only association, not causation. Epidemiological data can be used to suggest hypotheses but not to prove them."


I remember being in a discussion once about the safety of GMO foods and someone made the remark that epidemiological studies were unreliable. I really was perplexed that someone could throw an entire field like epidemiology under the bus. They were most likely thinking about some of the claims made in an article I have referred to before here recently:

Deming, data and observational studies: A process out of control and needing fixing. 2011 Royal Statistical Society.  (link)

“Any claim coming from an observational study is most likely to be wrong.” Startling, but true. Coffee causes pancreatic cancer. Type A personality causes heart attacks. Trans-fat is a killer. Women who eat breakfast cereal give birth to more boys. All these claims come from observational studies; yet when the studies are carefully examined, the claimed links appear to be incorrect. What is going wrong? Some have suggested that the scientific method is failing, that nature itself is playing tricks on us. But it is our way of studying nature that is broken and that urgently needs mending, say S. Stanley Young and Alan Karr; and they propose a strategy to fix it."

This criticism of observational studies is interesting. I have discussed this before in relation to wellness program evaluation studies and generally concluded that those criticisms sort look like straw man arguments trying to hold observational studies against the golden standard of the randomized clinical trial. Obviously you would expect results from a RCT to be more reliable and replicable. Does this mean that we refuse to ask interesting questions or analyze important policy decisions just because a RCT is too expensive or impractical, when there are quasi-experimental approaches that could be used to advance our knowledge and understanding of important issues? Or in these cases are we just better off basing our understanding of the world on circumstantial and anecdotal evidence alone? 

Jayson Lusk has blogged about this recently: 

"What about the health impacts of meat consumption?  It is true that many observational, epidemiological studies show a correlation between red meat eating and adverse health outcomes (interestingly there is a fair amount of overlap on the authors of the dietary studies and the environmental studies on meat eating).  But, this is a pretty weak form of evidence, and much of this work reminds of the kinds of regression analyses done in the 1980s and 90s in economics before the so-called "credibility revolution."

So, perhaps some bad practices associated with p-hacking and lack of identification has given epidemiological research, and observational studies in general, a worse reputation than they deserve. And if Jayson is correct, a lot of these poor study designs lacking identification may have misled us about our diets and healthy food choices.

Friday, January 30, 2015

Considerations in Propensity Score Matching

A while back I stumbled across a paper by Liu & Lynch: Do Agricultural Land Preservation Programs Reduce Farmland Loss?

Link:http://www.ncsu.edu/cenrep/research/documents/LiuandLynchAgpreservationLandeconomics_RR_final.pdf 

They use a really long panel and propensity score matching and highlight some important considerations in propensity score matching applications:

1) They used unrestricted matching (which basically ignored the time component, allowed an individual to actually be matched to themselves if propensity scores across time periods matched up) and restricted matching which required matches to be made only between treatments and controls within a given time period/census period.

2) They provide a  interesting discussion of the variance/bias tradeoff associated with bandwidth and kernel selection:

 “Bandwidth and kernel type selection is an important issue in choosing matching method. Generally speaking, a large bandwidth leads to a larger bias but smaller variance of the estimated average treatment effect of the PDR programs; a small bandwidth leads to a smaller bias but a larger variance. The differences among kernel types are embedded in the weights they assign to non-PDR county observations whose estimated propensity score are farther away from that of their matched PDR county observations.”

3) They discussed the  use of a leave one out cross-validation mechanism  to choose the ‘best matching’ method (combination of matching method i.e. nearest neighbor, kernel, local linear & combination with 5 possible kernel types and 6 bandwidths) optimized based on MSE criteria.  They site a references for this:

Racine, J. S. and Q. Li. 2004. “Nonparametric estimation of regression functions with both categorical and continuous data.” Journal of Econometrics 119 (1): 99-130.

Black, D., and J. Smith. 2004. “How Robust is the Evidence on the Effects of College Quality? Evidence from Matching.” Journal of Econometrics, 121(1-2): 99-124.

4) They state: “Matching with replacement performs as well or better than matching without replacement”– based on:

Dehejia, R., and S. Wahba. 2002. “Propensity score matching methods for non-experimental causal studies.” The Review of Economics and Statistics 84: 151-161.

Rosenbaum, P. 2002. Observational Studies (2nd edition). New York: Springer Verlag.

5) They also write that  the “selection of matching methods depends on the distribution of the estimated propensity score”- i.e. 

Kernel Matching works well with asymmetric distributions, excludes bad matches

The Local Linear Estimator may be more efficient than standard kernel matching when there is a large concentration of observations with propensity scores near 1 or 0:

Also from: McMillen, D. P., and J.F. McDonald. 2002. “Land values in a newly zoned city.” Review of Economics and Statistics 84(1): 62–72.

They also discuss the fact that nearest neighbor matching is more biased if propensity score  distributions are not very compatible.

6) Balancing Tests – A lot of practitioners implement more subjective evaluations of balance based on data visualization, but in this paper formal tests are discussed.

“After matching, we check again whether the two matched groups are the same on their observed characteristics. If unbalanced, the estimated ATT may not be solely the impact of PDR programs. Instead, it may be a combination of the impacts of PDR programs and the unbalanced variables. We rely on two of the balancing tests that exist in the empirical literature: the standardized difference test and a regression-based test. The first method is a t-test for the equality of the means for each covariate in the matched PDR and non-PDR counties. The regression test estimates coefficients for each covariate on polynomials of the estimated propensity scores…. and the interaction of these polynomials with the treatment binary variable, ….If the estimated coefficients on the interacted terms are jointly equal to zero according to an F-test, the balancing condition is satisfied.”

7) They also offer some really nice explanations of how kernel matching works:

“Kernel matching and local linear matching techniques match each PDR county with all non-PDR counties whose estimated propensity scores fall within a specified bandwidth (Heckman, Ichimura and Todd, 1997). The bandwidth is centered on the estimated propensity score for the PDR county. The matched non-PDR counties are weighted according to the density function of the kernel types.11 The closer a non-PDR county’s estimated propensity score is to the matched PDR county’s propensity score, the more similar the non-PDR county is to the matched PDR county and therefore it is assigned a larger weight calculated from a kernel functions defined in each method. More non-PDR counties are utilized under the kernel and local linear matching as compared to nearest neighbor matching.”


Saturday, January 24, 2015

The Internet of Things, Big Data, and John Deere

What is the internet of things? It's essentially the proliferation of smart products with connectivity and data sharing capabilities that are changing the way we interact with them and the rest of the world. We've all experienced IOT via smartphones, but the next generation of smart products will include our cars, homes, and appliances. A recent Harvard a Business Review article discusses what this means for the future of industry and the economy: 

How Smart, Connected products are Transforming the Competition by Porter and Heppelmann (the same porter of Porter's 5 competitive forces) : 

"IT is becoming an integral part of the product itself. Embedded sensors, processors, software, and connectivity in products (in effect, computers are being put inside products), coupled with a product cloud in which product data is stored and analyzed and some applications are run, are driving dramatic improvements in product functionality and performance....As products continue to communicate and collaborate in networks, which are expanding both in number and diversity, many companies will have to reexamine their core mission and value proposition"

In the article, they view how IOT is reshaping the competitive framework through the lense of Porter's five competitive forces. 

HBR continually uses John Deere as a case study example of a firm that is leading the industry leveraging big data and analytics as a successful model for an IOT strategy, particularly the integration of IOT applications through connected tractors and implements. So, the competitive space expands from a singular focus on a line of equipment, to optimization of performance and interoperability within a connected system of systems: 

"The function of one product is optimized with other related products...The manufacturer can now offer a package of connected equipment and related services that optimize overall results. Thus in the farm example, the industry expands from tractor manufacturing to farm equipment optimization."

A 'system of systems' approach means not only smart connected tractors and implements, but layers of connectivity and data related to weather, crop prices, and agronomics. That changes value proposition not just for the farm implement business, but also the seed, input, sales, and crop consulting as well. When you think about the IOT, suddenly Monsanto's purchase of Climate Corporation make sense. Suddenly the competitive landscape has changed in agriculture. Both Monsanto and John Deere are offering advanced analytic and agronomics consulting services based on the big data generated from the IOT. John Deere has found value creation in these expanded economies of scope given the date generated from their equipment. Seed companies see added value in optimizing and developing customized genetics as farmers begin to farm (collectin data) literally inch by inch (vs field by field). Both new alliances and rivalrys are being drawn between what was once rather distinct lines of business. 

And, as the HBR article notes, there is a lot of opportunity in the new era of Big Data and the IOT.  The convergence of biotechnology, genomics, and big data has major implications for economic development as well as environmental sustainability. 

Additional Reading:

This article does a great job breaking down insight from the HBR paper: How John Deere is Using APIs to Grow the World's Food Supply  
From the Motley Fool: The Internet of Things is Changing How Your Food is Grown

Thursday, January 22, 2015

Fat Tails, Kurtosis, and Risk

When I think of 'fat tails' or 'heavy tails' I typically think of situations that can be described by probability distributions with heavy mass in the tails. This might imply that the tails are 'fat' or 'thicker' than other situations with less mass in the tails. (for instance, a normal distribution might be said to have thin tails while a distribution with more mass in the tails than the normal might be considered a 'fat tailed' distribution.)

Basically when I think of tails I think of 'extreme events', some occurrence with an excessive departure from what's expected on average. So, if there is more mass in a the tail of a probability distribution (it is a fat or thick tailed distribution) that implies that extreme events will occur with greater probability.  So in application, if I am trying to model or assess the probability of some extreme event (like a huge loss on an investment) then I better get the distribution correct in terms of tail thickness.

See here for a very  nice explanation of fat tailed events from the Models and Agents blog: http://modelsagents.blogspot.com/2008/02/hit-by-fat-tail.html

Also here is a great podcast from the CME group related to tail hedging strategies: (May 29 2014) http://www.cmegroup.com/podcasts/#investorsAndConsultants

And here's a twist, what if I am trying to model the occurrence of multiple events simultaneously? (a huge loss in commodities and equities simultaneously, or losses on real estate in Colorado and Florida simultaneously). I would want to model a multivariate process that captures the correlation or dependence between multiple extreme events (in other words between the tails of their distributions). Copulas offer an approach to modeling tail dependence, and again, getting the distributions correct matters.

When I think of how do we assess or measure tail thickness in a data set, I think of kurtosis. Rick Wicklin recently had a nice piece discussing the interpretation of kurtosis and relating kurtosis to tail thickness.

"A data distribution with negative kurtosis is often broader, flatter, and has thinner tails than the normal distribution."

" A data distribution with positive kurtosis is often narrower at its peak and has fatter tails than the normal distribution."

The connection between kurtosis can be tricky and kurtosis  cannot be interpreted this way universally in all situations. Rick gives some good examples if you want to read more.  But this definition of kurtosis from his article seems keep us honest:

"kurtosis can be defined as "the location- and scale-free movement of probability mass from the shoulders of a distribution into its center and tails. " with the caveat - "the peaks and tails of a distribution contribute to the value of the kurtosis, but so do other features."

But Rick also had an earlier post on fat and long tailed distributions where he puts all of this into perspective in terms of the connection to modeling extreme events as well as a more rigorous discussion and definition of tails and what 'fat' or 'heavy' tailed means:

"Probability distribution functions that decay faster than an exponential are called thin-tailed distributions. The canonical example of a thin-tailed distribution is the normal distribution, whose PDF decreases like exp(-x2/2) for large values of |x|. A thin-tailed distribution does not have much mass in the tail, so it serves as a model for situations in which extreme events are unlikely to occur.

Probability distribution functions that decay slower than an exponential are called heavy-tailed distributions. The canonical example of a heavy-tailed distribution is the t distribution. The tails of many heavy-tailed distributions follow a power law (like |x|–α) for large values of |x|. A heavy-tailed distribution has substantial mass in the tail, so it serves as a model for situations in which extreme events occur somewhat frequently."

Also, somewhat related, Nassim Taleb, in his paper on the precautionary principle and GMOs (genetically modified organisms) discusses such concepts as ruin, harm,  fat tails and fragility, tail sensitivity to uncertaintly etc.  He uses very rigorous definitions of these terms and determines that there are certain things like GMOs that would require a non-naive application of the precautionary principle while other things like nuclear energy would not. (also tune into his discussion of this with Russ Roberts on Econtalk- more discussion on this actual application at my applied economics blog Economic Sense).

Friday, June 27, 2014

Is distance a proxy for pesticide exposure and is it related to ASD? Some thoughts...


Recently a paper has made some headlines, and the message getting out seems to be that living near a farm field where there has been pesticide applications has been found to increase the risk of Autism spectrum disorder. A few things about the paper. First, one of the things I admire about econometric work is the attempt to make use of some data set, some variable, or some measurement to estimate the effect of some intervention or policy, in a world where we can’t always get our hands on the thing we are really trying to measure. The book Freakonomics comes to mind, or quasi-experimental designs and the use of instrumental variables.

Second, I’m not an epidemiologist, entomologist, or have a background in toxicology,  but my expertise is more focused on statistical methods so I will comment on the article from that perspective. While the authors could not  (or simply did not) actually measure pesticide exposure in any medical or biological sense, they attempted to infer that distance from an agricultural field might correlate well enough to proxy for exposure. That is a large assumption and perhaps one of the greatest challenges of the study. It is not a study on actual exposure. So I’ll  try to only refer to exposure from this point in quotes.  But the authors did make clever use of some interesting data sources. They matched up required reported pesticide applications and report dates with zipcodes of the study respondents and reported pregnancy stages to determine distance from application and at what point of their pregnancy they were exposed.  They reported distance in three bands  or buffer zones of 1.25, 1.5, & 1.75 km   This was actually nice work, if distance could be equated to some known level of exposure. Unfortunately, while they cited some other work attempting to tie exposure to ASD, I did not see a citation in the body of the text where any work had been done justifying the use of distance as a proxy, or those particular bands. More on this later. They also attempted to control for a number of confounders, applied survey weighting to ‘weight up’ the effects to reflect the parent population, and in addition, at least based on my reading, may have even tried to control for some level of selection bias by using IPTW regression with SAS.

Discussion of Results

There were at least four major findings in the paper:

(1) Proximity to organophosphates at some point during gestation was associated with a 60% increased risk for ASD

(2) higher for 3rd trimester exposures [OR = 2.0, 95% confidence interval (CI) = (1.1, 3.6)],

(3) and 2nd trimester chlorpyrifos applications: OR = 3.3 [95% CI = (1.5, 7.4)].

(4)Children of mothers residing near pyrethroid insecticide applications just prior to conception or during 3rd trimester were at greater risk for both ASD and DD, with OR's ranging from 1.7 to 2.3.

So where do we go with these results? First off all of these findings are based on odds ratios. The reported odds ratio in the first finding above was 1.60 which implies a [1.6-1.0]*100 = 60% increase in odds of ASD for ‘exposed’ vs ‘non-exposed’ children. This is an increase in odds, and does not have the exact same interpretation as an increase in probability. (see more about logistic regression and odds ratios here). Some might read the headline and walk away with the wrong idea that living within proximity of farm fields with organophosphate applications constitutes  ‘exposure’ to organophosphates  and is associated with a 60% increased probability of ASD, but that is stacking one large assumption on top of another misinterpretation.

However, these findings are but a slice of the full results reported in the paper. Table 3 reports a number of findings across the distance bands, types of pesticide, and pregnancy stage. One thing about odds ratios, an odds ratio of ‘1’ implies no effect. The vast majority of these findings were associated with odds ratios with 95% confidence intervals containing 1, or very very close to 1. For those that like to interpret p-values, a 95% CI for an odds ratio that contains 1 implies that the estimated regression coefficient in the model has a p-value > .05, i.e. non-significant results.

Another interesting thing about the table, is that there doesn’t seem to be any pattern of distance/pregnancy stage/chemistry associated with the estimated effects or odds ratios. A point made well in a recent blog post regarding this study at scienceblogs.com here.

Sensitivity

From the paper: “In additional analyses, we evaluated the sensitivity of the estimates to the choice of buffer size, using 4 additional sizes between 1 and 2km: results and interpretation remained stable (data not shown).”

That’s unfortunate too. Given the previous discussion of odds ratios, lack of empirical support or literature related to using distance as a proxy for exposure, you would think more sensitivity analysis would be merited to show robustness to all of these assumptions even if and especially if there is no previous precedent in the literature related to distance.  This in combination with the previous discussion regarding the large number of insignificant odds ratios and select reporting of the marginally significant results is probably what fueled accusations of data drudging.

Omitted Controls

From the Paper: “Primarily, our exposure estimation approach does not encompass all potential sources of exposure to each of these compounds: among them external non-agricultural sources (e.g. institutional use, such as around schools); residential indoor use; professional pesticide application in or around the home for gardening, landscaping or other pest control; as well as dietary sources (Morgan 2012).”

So, there are a number of important routes of exposure that were not controlled for, or perhaps a good deal of omitted variable bias and unobserved heterogeneity.  The point of my post is not to pick apart a study linking pesticides to ASD. There are no perfect data sets and no perfect experimental designs. All studies have weaknesses, and my interpretation of this study certainly has flaws. The point is, while this study has made some headlines with some media outlets, and seems scary; it is not one that should be used to draw sharp conclusions or to run to your legislator for new regulations.
This reminds me of a quote I have shared here recently:
"Social scientists and policymakers alike seem driven to draw sharp conclusions, even when these can be generated only by imposing much stronger assumptions than can be defended. We need to develop a greater tolerance for ambiguity. We must face up to the fact that we cannot answer all of the questions that we ask." (Manski, 1995)

References:
Manski, C.F. 1995. Identification Problems in the Social Sciences. Cambridge: Harvard University Press.

Neurodevelopmental Disorders and Prenatal Residential Proximity to Agricultural Pesticides: The CHARGE Study
Janie F. Shelton, Estella M. Geraghty, Daniel J. Tancredi, Lora D. Delwiche, Rebecca J. Schmidt, Beate Ritz, Robin L. Hansen, and Irva Hertz-Picciotto
Environmental Health Perspectives.   June 23, 2014