Showing posts with label hypothesis testing. Show all posts
Showing posts with label hypothesis testing. Show all posts

Friday, December 21, 2018

Thinking About Confidence Intervals: Horseshoes and Hand Grenades

In a previous post, Confidence Intervals: Fad or Fashion I wrote about Dave Giles' post on interpreting confidence intervals. A primary focus of these discussions was how confidence intervals are often mis-interpreted. For instance the two statements below are common mischaracterizations of CIs:

1) There's a 95% probability that the true value of the regression coefficient lies in the interval [a,b].
2) This interval includes the true value of the regression coefficient 95% of the time.

You can read the previous post or Dave's post for more details. But in re-reading Dave's post myself recently one statement had me thinking:

"So, the first interpretation I gave for the confidence interval in the opening paragraph above is clearly wrong. The correct probability there is not 95% - it's either zero or 100%! The second interpretation is also wrong. "This interval" doesn't include the true value 95% of the time. Instead, 95% of such intervals will cover the true value."

I like the way he put that...'95% of such intervals' distinguishing this from a particular observed/calculated confidence interval. I think someone trained to think about CIs in the incorrect probabilistic way may have trouble getting at this. So how might we think about this in a way that captures CIs in a way that is still useful, but doesn't get us tripped up with incorrect probability statements?

My favorite statistics text is Degroot's Probability and Statistics. In the 4th edition they are very careful about explaining confidence intervals:

"Once we compute the observed values of a and b, the observed interval (a,b) is not so easy to interpret....Before observing the data we can be 95% confident that the random interval (A,B) will contain mu, but after observing the data, the safest interpretation is that (a,b) is simply the observed value of the random interval (A,B)"

While Degroot is careful, it still may not be very intuitive. However, in Principles and Procedures of Statistics: A Biometrical Approach (Steel, Torie, and Dickey) they present a more intuitive explanation.

"since mu will either be or not be in the interval, that is P=0 or 1, the probability will actually be a measure of confidence we placed in the procedure that led to the statement. This is like throwing a ring at a fixed post; the ring doesn't land in the same position or even catch on the post every time. However we are able to say that we can circle the post 9 times out of 10, or whatever the value should be for the measure of our confidence in our proficiency."

The ring tossing analogy seems to work pretty well. I'll customize it by using horseshoes instead. Yes 95 out of 100 times you might throw a ringer (in the game of horseshoes that is when the horse shoe circles the peg or stake when you toss it). You know this before you toss it. And to use Dave Giles language, *before* calculating a confidence interval we know that 95% of such intervals will cover the population parameter of interest. And, after we toss the shoe, it either circles the peg or not, that is a 1 or a 0 in terms of probability. Similarly, *after* computing a confidence interval, the true mean or population parameter of interest is covered or not with a probability of 0 or 100%.

This isn't perfect, but thinking of confidence intervals this way at least keeps us honest about making probability statements.

Going back to my previous post, I still like the description of confidence intervals Angrist and Pishke provide in Mastering 'Metrics, that is 'describing a set of parameter values consistent with our data.' 

For instance if we run the regression:

y = b0 + b1X + e  to estimate y = B0 + B1 + e

and get our parameter estimate b with a 95% confidence interval like (1.2,1.8), we can say that our sample data is consistent with any population that has a B taking a value that falls in the interval. That implies there are a number of populations that our data would be consistent with. Narrower intervals imply very similar populations, very similar values of B, and speaks to more precision in our estimate of B.

I really can't make an analogy for hand grenades. It just gave me a title with a ring to it.

See also:
Interpreting Confidence Intervals
Bayesian Statistics Confidence Intervals and Regularization
Overconfident Confidence Intervals

Sunday, December 31, 2017

HARK! - flawed studies in nutrition call for credibility revolution -or- HARKing in nutrition research

There was a nice piece over at the Genetic Literacy Project I read just recently: Why so many scientific studies are flawed and poorly understood. (link). They gave a fairly intuitive example of false positives in research using coin flips. I like this because I used the specific example of flipping a coin 5 times in a row to demonstrate basic probability concepts in some of the stats classes I used to teach. Their example might make a nice extension:

"In Table 1 we present ten 61-toss sequences. The sequences were computer generated using a fair 50:50 coin. We have marked where there are runs of five or more heads one after the other. In all but three of the sequences, there is a run of at least five heads. Thus, a sequence of five heads has a probability of 0.55=0.03125 (i.e., less than 0.05) of occurring. Note that there are 57 opportunities in a sequence of 61 tosses for five consecutive heads to occur. We can conclude that although a sequence of five consecutive heads is relatively rare taken alone, it is not rare to see at least one sequence of five heads in 61 tosses of a coin."

In other words, a 5 head run in a sequence of 61 tosses (as evidence against a null hypothesis of p(head) = .5 i.e. a fair coin) is their analogy for a false positive in research. Particularly they relate this to nutrition research where it is popular to use large survey questionnaires that consist of a large number of questions:

"asking lots of questions and doing weak statistical testing is part of what is wrong with the self-reinforcing publish/grants business model. Just ask a lot of questions, get false-positives, and make a plausible story for the food causing a health effect with a p-value less than 0.05"

It is their 'hypothesis' that this approach in conjunction with a questionable practice referred to as 'HARKing' (hypothesizing after the results are known) is one reason we see so many conflicting headlines about what we should and should not eat or benefits or harms of certain foods and diets. There is some damage done in terms of peoples' trust in science as a result.  They conclude:

"Curiously, editors and peer-reviewers of research articles have not recognized and ended this statistical malpractice, so it will fall to government funding agencies to cut off support for studies with flawed design, and to universities to stop rewarding the publication of bad research. We are not optimistic."

More on HARKing.....

A good article related to HARKing is a paper written by Norbert L. Kerr.  By HARKing he specifically discusses it as the practice of proposing one hypothesis (or set of hypotheses) but later changing the research question *after* the data is examined. Then presenting the results *as if* the new hypothesis were the original.  He does distinguish this from a more intentional exercise in scientific induction, inferring some relation or principle post hoc from a pattern of data. This is more like exploratory data analysis.

I discussed exploratory studies and issues related to multiple testing in a previous post:  Econometrics, Multiple Testing, and Researcher Degrees of Freedom. 

To borrow a quote from this post- "At the same time, we do not want demands of statistical purity to strait-jacket our science. The most valuable statistical analyses often arise only after an iterative process involving the data" (see, e.g., Tukey, 1980, and Box, 1997).

To say the least, careful consideration of tradeoffs should be made in the way research is conducted, and as the post discusses in more detail, the garden of forking paths involved.

I am not sure to what extent the credibility revolution has impacted nutrition studies, but the lessons apply here.

References:

HARKing: Hypothesizing After the Results are Known
Norbert L. Kerr
Personality and Social Psychology Review
Vol 2, Issue 3, pp. 196 - 217
First Published August 1, 1998

Monday, March 6, 2017

Interpreting Confidence Intervals

From: Handbook of Biological Statistics http://www.biostathandbook.com/confidence.html
"There is a myth that when two means have confidence intervals that overlap, the means are not significantly different (at the P<0.05 level)… (Schenker and Gentleman 2001, Payton et al. 2003); it is easy for two sets of numbers to have overlapping confidence intervals, yet still be significantly different by a two-sample t–test; conversely… Don't try compare two means by visually comparing their confidence intervals, just use the correct statistical test."
A really cogent note related to this from the Cornell Statistical Consulting unit can be found here.
"Generally, when comparing two parameter estimates, it is always true that if the confidence intervals do not overlap, then the statistics will be statistically significantly different. However, the converse is not true. That is, it is erroneous to determine the statistical significance of the difference between two statistics based on overlapping confidence intervals."

More details using basic math here.

A 2005 article in Psychological Methods indicates a large number of researchers don't interpret them correctly. In an interesting American Psychologist article (2005) researchers determined that under a number of broadly applicable conditions 95% confidence intervals can overlap as much as 25% for groups that are actually significantly different at the 5% level and that with zero overlap statistical significance is actually at the 1% level (p~.01).


Dave Giles also brings up an Insect Science paper by Payton et al in discussion related to using CI to determine statistical significance that relates to this:
http://davegiles.blogspot.com/2017/01/hypothesis-testing-using-non.html?_sm_au_=iVVFv75rHPD1TSHs
"Here's a well-known result that bears on this use of the confidence intervals. Recall that we're effectively testing H0: μ1 = μ2, against HA: μ1 ≠ μ2. If we construct the two 95% confidence intervals, and they fail to overlap, then this does not imply rejection of H0 at the 5% significance level. In fact the correct significance is roughly one-tenth of that.  Yes, 0.5%!
If you want to learn why, there are plenty of references to help you. For instance, check out McGill et al. (1978), Andrews et al. (1980), Schenker and Gentleman (2001), Masson and Loftus (2003), and Payton et al. (2003) - to name a few. The last of these papers also demonstrates that a rough rule-of-thumb would be to use 84% confidence intervals if you want to achieve an effective 5% significance level when you "try" to test H0 by looking at the overlap/non-overlap of the intervals."
Actually according to the last paper mentioned (Payton,2003) the 84% CI is adjusted depending on the ratio of standard errors from the two populations you are comparing:
But all of the work above is predicated on a comparison of two populations. Considerations of multiple comparisons complicate things further. (see Rick Wicklin's post on doing this in SAS). Perhaps if a visual presentation is what we want we plot the CIs (as much as we may not like dynamite plots) but denote which groups are significantly different based on the properly specified tests (per the note from the handbook above).  Something like below:
http://freakonomics.com/2008/07/30/how-big-is-your-halo-a-guest-post/

References:
Belia, S, Fidler, F, Williams, J, Cumming, G (2005). Researchers misunderstand confidence intervals and standard error bars Psychological Methods, 10 (4), 389-396
Am Psychol. 2005 Feb-Mar;60(2):170-80. Inference by eye: confidence intervals and how to read pictures of data. Cumming G(1), Finch S
Payton, M. E., M. H. Greenstone, and N. Schenker, 2003. Overlapping confidence intervals or standard error intervals: What do they mean in terms of statistical significance? Journal of Insect Science, 3, 1–6.


Saturday, March 5, 2011

Student's t, Normality, & the Slutsky Theorems

 Often in an analysis we are using s2 instead of σ2 (as the true population variance is often unknown). Yet we want to construct confidence intervals, or equivalently conduct hypothesis tests. The t distribution is the ratio of a normally distributed variable and chi-square distributed variable ( DeGroot, 2002).  If our data is distributed exactly normal, we can rely on using the t-table for constructing confidence intervals. These are exact results, as the t-ratio is exactly distributed t given that the underlying data is distributed exactly normal. 

But what if we don’t know the distribution of the data we are working with, or don’t feel comfortable making assumptions of normality. Usually we have to estimate σ2 with s2. Of course if n is large enough, reliance on the t-table or the standard normal table will give similar results. But, Goldberger offers the following anecdote:

“ There is no good reason to rely routinely on a t-table rather than a normal table unless Y itself is normally distributed” ( Goldberger, 1991).

So, how do you justify this?   In this case there are some powerful theorems regarding asymptotic properties of sample statistics known as the Slutsky Theorems. These are outlined in Goldberger, 1991. The following sequence of steps using these theorems is based on Goldberger and my lecture notes from ECO 603 Research Methods for Economics, which was actually a mathematical statistics course taught by Dr. Christopher Bollinger at the University of Kentucky. Any errors or mistakes are completely my own.

GivenΘ^ is an estimator for the population parameter Θp  implies convergence in probability, and d implies convergence in distribution:

S1: If Θ^ p Θ  then for any continuous function h (Θ^)→ p h(Θ).

S2: If   Θ1^ and Θ2^  converge in probability to (Θ1, Θ2), then
 h (Θ1^, Θ2^) p  h(Θ1, Θ2).


S3: If Θ^ p Θ  and Zn  d N(0,1)  then (Θ^  +    Znd  N( Θ, 1 )

S4: If Θ^ p Θ  and Zn  d N(0,1)  then Θ^  Zn   d  N( 0, Θ2 )

S5: If n1/2 (Θ^ - Θ ) / s1/2 ~A N(0,1) then for continuous functions of Θ,

        n1/2 (h(Θ^) – h( Θ )) / s1/2 ~A N(0, h’(Θ)2Σ2)

( Goldberger, 1991).

So now if we want to use s2 to estimate  σ2 we form the statistic

Z^ = ( Xbar - μ)2 / (s2 / n)1/2  = n1/2  ( Xbar - μ)/ s    (1)

This looks like the t- statistic, but if we can’t make the assumption of normality, the exact results of the t-distribution do not apply. In this case we rely on results of both the CLT and Slutsky theorems.
Given the traditional  standardized normal variable formulation: 

 Z = n1/2  ( Xbar - μ) σ 

Algebraic manipulation shows that

( σ/ s)  n1/2  ( Xbar - μ)/ σ  = n1/2  ( Xbar - μ)/ s

then   ( σ/s)  Z = Z^  where Z^ is defined in (1) above

By the CLT,  Z d N( 0,1)

It can be shown that  s2 p  σ2

If we define Θ^ as  s then we can view  ( σ/s) as a function h (Θ^)

Then by S1  h (Θ^)→ p h(Θ)  which implies that ( σ/s) p ( σ/ σ) = 1

Given S1S4 gives the following result:  Θ^ Zn   d  N( 0, Θ2 ) which implies that
 ( σ/s)Z d  N( 0, ( σ/ σ)2 ) = N( 0,1)

And therefore Z^ d N(0,1).


Therefore, by the Central Limit and Slutsky theorems (S1 and S4) one can use the asymptotic properties of the statistic Z^ = n1/2  ( Xbar - μ)/s to form confidence intervals based on the standard normal distribution without making any assumptions about the distribution of the sample data and using s2 to estimate   σ2.

How large does n have to be before asymptotic properties apply?  From Kennedy, A Guide to Econometrics 5th Edition:

 "How large does the sample size have to be for estimators to display their asymptotic properties? The answer to this crucial question depends on the characteristics of the problem at hand. Goldfeld and Quandt (1972, p.277) report an example in which a sample size of 30 is sufficiently large and an example in which a sample of 200 is required."

An important note to remember, it is often the case that people say 'as n becomes large the normal distribution approximates the t-distribution', but in fact, as shown above, as n-becomes large the formulation above (Z^) actually approximates the normal distribution (again based on the CLT and the Slutsky theorems).

References:

A Guide to Econometrics, Kennedy 2003.
A Course in Econometrics, Goldberger 1991.
Economics 603 Research Methods and Procedures in Economics Course Notes. University of Kentucky. Taught by Dr. Christopher Bollinger (2002).