----- Original Message ----- From: "Teaching in the Psychological Sciences digest" <[EMAIL PROTECTED]> To: "tips digest recipients" <[EMAIL PROTECTED]> Sent: Thursday, November 11, 2004 12:00 AM Subject: tips digest: November 10, 2004
>> ---------------------------------------------------------------------- > > Subject: "Marginal Significance" > From: Todd Nelson <[EMAIL PROTECTED]> > Date: Wed, 10 Nov 2004 11:40:28 -0800 > X-Message-Number: 2 I realize that I enter this thread late and that I'll re-iterate some of the points others have made but perhaps I'll make one or two novel ones. > Dear TIPSters, > > Throughout my ph.d. Training in social psychology at Michigan State > University in the early 90's, it was common parlance to refer to statistical > findings with P-values of between .05 and .10 as "marginally significant." > It was a very common term in all the major social psychological journals > (and I dare say I even remember this term used in the more general journals, > such as Psychological Review, Psych Bulletin, Psych Science, etc.) and among > my professors. The use of such a phrase is dependent upon "local customs" with respect to talking about statistics and statistical analyses. Part of it is based on social convention, part of it is based on one's knowledge, experience, orientation towards statistics. Another example of this is the use of the word "reliable" in referring to a statistically significant result. During my training I had rarely experienced this usage. It wasn't until I met some people who made a common practice of using it that I realized that I didn't quite understand what they meant by it. In my head, I made the equation of "reliable = statistically significant" but it became obvious that people meant something more than this, perhaps implicitly making the "argument" that the result was somehow more real than the statistics indicated. Given that any statistical test is subject to two types of errors (i.e., Type I - false rejection of a true null hypothesis, and Type II - failure to reject a false null hypothesis), neither of which is zero (indeed, the Type II error rate is usually much greater than the Type I error rate, with Type II error rate = 1- Power), I was and am hardpressed to understand how a single result, without replication, is "reliable" in any sense of the word. It seems to me to be more of a statement of "faith" than of statistics. > The other day, in a thesis defense I was chairing, I was dumfounded when the > other thesis committee members strongly objected to the term "marginally > significant" in the student's results section (or even in the discussion), > both saying they had NEVER heard of the term (!) Again, this is a reflection of differences in convention and experience. The real question is what is their attitude towards dealing with results where ".05< p < .10"? If they take a strict, old-fashioned Fisherian approach to statistics, then they are probably following the heuristic of "If p<.05, then reject null hypothesis else fail to reject". If they are more modern, they may be more concerned with evaluating the effect size that is the focus of the statistical test and whether one has sufficient statistical power to detect it.. They would probably follow a rule like "if .05< p < .10, then evaluate the possibility of an underpowered statistically significant result." In this context, the concern is whether one had misspecified the effect size used in determining what the sample size should be for the study (under the assumption that a power analysis was conducted prior to beginning the study) Without a power analysis, even a post hoc one, it is difficult to reach any firm conculsion about what a "marginal" or non-significant result means. However, as I'll point out below, the situation may actually be somewhat more complicated. > Wanting to make sure I was still on planet Earth, I consulted several > statistics and research textbooks in my office, and found a few that > referred to "marginal significance" and several articles by noted > statisticians who make the case for discussing results in the .05-.10 range > (e.g., Hunter, Cohen, Rosenthal). I use the term when talking about exploratory research where the Power of the statistical test is known and not very high (in general, less that 95% - a Type II error rate greater than 5%). However, as I'll point out below, I think this is a misplaced emphasis. What's important is not the statistical significance of a result, the important question is what is the effect size that one has detected in their study. > I am wondering: have you heard of this term (marginally significant)? Do you > use it and teach it? If so, why? If not, what is your objection? > > I am very interested to hear everyone's views.... I'm giving readers fair warning here: I'm going to go into an extended discussion here and what I have to say will probably be of interest only to statistics and methods geeks. With that out of the way. I believe that there are two issues that are raised here: (1) Does a single experimental result provide a definitive conclusion, and (2) How do we deal with negative reults or marginally significant results, especially when we realize that in the context of meta-analysis individual studies that have "non-significant/marginal" results can still produce statistically significant result when the results are pooled. With respect to point (1) above, it seems to me that many psychologists have adopted a type of "rugged individualism" when it comes to the conduct of research which I believe helps to promote an overly optimistic view of what the results of a experiment might mean. If one has 99.99% or better power AND one is dealing with huge effect sizes AND one has outstanding control over source of error and threats to internal validity, then using the Fisherian rule stated above makes sense (recall that Fisher didn't believe in statistical power). However, this is rarely the case. Power is often much less than 99.99%, if it is determined at all. Effect sizes may not be known even if there is prior research providing estimates. And it is rare that all threats to internal validity and sources of error have been controlled. Given all this uncertainty, it is hardly reasonable for one to claim that a single specific study has produced a definitive result, though we may want to act that way. Even if we have a statistically significant result, we have a 5% chance of making an error (that is, if one set the alpha/Type I error rate to .05). This is why independent replication by others is so important: as the number of similar replicated results increases, the probability of having made a Type I errors gets smaller and smaller. However, with respect to point (2) above, the rise of meta-analysis has made it clear that no single experiment is sufficient in and of itself, all it does is provide a single set of estimates of what the values are for certain population parameters and an effect size (Note: there is issue of whether one should conceive of an effect as a "fixed effect" or a "random effect", but I'll leave that can of worms for another day). In the context of meta-analysis, even non-significant results are important if they come from under-powered studies because pooling the results across studies can overcome the lack of power of individual studies. A worked out example of this is provided in the following: Ingelfinger, J.A., Mosteller, F., Thibodeau, L.A., & Ware, J.H. (1994). Biostatistics in Clinical Medicine. New York: McGraw-Hill. (See Chapter 14). A more general framework for "research synthesis" is presented in a variety of sources, making the result of any specific study just a data point in the context of related studies. One source that I would suggest to people who are interested in this is the following which follows the Cochrane Collaboration's framework: Egger, M., Smith, G.D., & Altman, D.G. (2001). Systematic Reviews in Health Care: Meta-analysis in Context (2nd Ed.). Cornwall, England: BMJ press. I've taken a long route to make what may seem to be rather simple points: whether one uses the terms "marginally significant" for the ".05< p <.10" situation depends upon how one understands and conceptualizes what the statistical test is expressing and how this result relates to other studies. The question of assumptions, unexamined or otherwise is pertinent. Mike Palij New York University [EMAIL PROTECTED] s --- You are currently subscribed to tips as: [EMAIL PROTECTED] To unsubscribe send a blank email to [EMAIL PROTECTED]
