----- Original Message ----- 
From: "Teaching in the Psychological Sciences digest"
<[EMAIL PROTECTED]>
To: "tips digest recipients" <[EMAIL PROTECTED]>
Sent: Thursday, November 11, 2004 12:00 AM
Subject: tips digest: November 10, 2004

>> ----------------------------------------------------------------------
>
> Subject: "Marginal Significance"
> From: Todd Nelson <[EMAIL PROTECTED]>
> Date: Wed, 10 Nov 2004 11:40:28 -0800
> X-Message-Number: 2

I realize that I enter this thread late and that I'll re-iterate
some of the points others have made but perhaps I'll make
one or two novel ones.

> Dear TIPSters,
>
> Throughout my ph.d. Training in social psychology at Michigan State
> University in the early 90's, it was common parlance to refer to
statistical
> findings with P-values of between .05 and .10 as "marginally significant."
> It was a very common term in all the major social psychological journals
> (and I dare say I even remember this term used in the more general
journals,
> such as Psychological Review, Psych Bulletin, Psych Science, etc.) and
among
> my professors.

The use of such a phrase is dependent upon "local customs" with respect
to talking about statistics and statistical analyses.  Part of it is based
on
social convention, part of it is based on one's knowledge, experience,
orientation towards statistics.  Another example of this is the use of
the word "reliable" in referring to a statistically significant result.
During
my training I had rarely experienced this usage. It wasn't until I met
some people who made a common practice of using it that I realized
that I didn't quite understand what they meant by it.  In my head, I
made the equation of "reliable = statistically significant" but it became
obvious that people meant something more than this, perhaps implicitly
making the "argument" that the result was somehow more real than the
statistics indicated.  Given that any statistical test is subject to two
types
of errors (i.e., Type I - false rejection of a true null hypothesis, and
Type II - failure to reject a false null hypothesis), neither of which is
zero (indeed, the Type II error rate is usually much greater than the
Type I error rate, with Type II error rate = 1- Power), I was and am
hardpressed to understand how a single result, without replication, is
"reliable" in any sense of the word.  It seems to me to be more of a
statement of "faith" than of statistics.

> The other day, in a thesis defense I was chairing, I was dumfounded when
the
> other thesis committee members strongly objected to the term "marginally
> significant" in the student's results section (or even in the discussion),
> both saying they had NEVER heard of the term (!)

Again, this is a reflection of differences in convention and experience.
The real question is what is their attitude towards dealing with results
where ".05< p < .10"?  If they take a strict, old-fashioned Fisherian
approach to statistics, then they are probably following the
heuristic of  "If p<.05, then reject null hypothesis else fail
to reject".  If they are more modern, they may be more concerned
with evaluating the effect size that is the focus of the statistical
test and whether one has sufficient statistical power to detect
it..  They would probably follow a rule like "if .05< p < .10, then
evaluate the possibility of an underpowered statistically significant
result."  In this context, the concern is whether one had
misspecified the effect size used in determining what the
sample size should be for the study (under the assumption that
a power analysis was conducted prior to beginning the study)
Without a power analysis, even a post hoc one,  it is difficult
to reach any firm conculsion about what a "marginal" or
non-significant result means.  However, as I'll point out
below, the situation may actually be somewhat more complicated.

> Wanting to make sure I was still on planet Earth, I consulted several
> statistics and research textbooks in my office, and found a few that
> referred to "marginal significance" and several articles by noted
> statisticians who make the case for discussing results in the .05-.10
range
> (e.g., Hunter, Cohen, Rosenthal).

I use the term when talking about exploratory research where the
Power of the statistical test is known and not very high (in general,
less that 95% - a Type II error rate greater than 5%).  However, as
I'll point out below, I think this is a misplaced emphasis.  What's
important is not the statistical significance of a result, the important
question is what is the effect size that one has detected in their study.

> I am wondering: have you heard of this term (marginally significant)? Do
you
> use it and teach it? If so, why? If not, what is your objection?
>
> I am very interested to hear everyone's views....

I'm giving readers fair warning here:  I'm going to go into
an extended discussion here and what I have to say will probably
be of interest only to statistics and methods geeks.  With that
out of the way.

I believe that there are two issues that are raised here:

(1)  Does a single experimental result provide a definitive
conclusion,

and

(2)  How do we deal with negative reults or marginally
significant results, especially when we realize that in
the context of meta-analysis individual studies that have
"non-significant/marginal" results can still produce
statistically significant result when the results are pooled.

With respect to point (1) above, it seems to me that
many psychologists have adopted a type of "rugged
individualism" when it comes to the conduct of research
which I believe helps to promote an overly optimistic
view of what the results of a experiment might mean.
If one has 99.99% or better power AND one is dealing
with huge effect sizes AND one has outstanding control
over source of error and threats to internal validity,
then using the Fisherian rule stated above makes sense
(recall that Fisher didn't believe in statistical power).
However, this is rarely the case.  Power is often much
less than 99.99%, if it is determined at all.  Effect
sizes may not be known even if there is prior research
providing estimates.  And it is rare that all threats to
internal validity and sources of error have been controlled.
Given all this uncertainty, it is hardly reasonable for
one to claim that a single specific study has produced
a definitive result, though we may want to act that
way.  Even if we have a statistically significant result,
we have a 5% chance of making an error (that is, if
one set the alpha/Type I error rate to .05).  This is
why independent replication by others is so important:
as the number of similar replicated results increases,
the probability of having made a Type I errors gets
smaller and smaller.

However, with respect to point (2) above, the rise of
meta-analysis has made it clear that no single experiment
is sufficient in and of itself, all it does is provide a single
set of estimates of what the values are for certain population
parameters and an effect size (Note:  there is issue of
whether one should conceive of an effect as a "fixed effect"
or a "random effect", but I'll leave that can of worms
for another day).  In the context of meta-analysis,
even non-significant results are important if they
come from under-powered studies because pooling
the results across studies can overcome the lack of
power of individual studies.  A worked out example
of this is provided in the following:

Ingelfinger, J.A., Mosteller, F., Thibodeau, L.A., &
Ware, J.H. (1994).  Biostatistics in Clinical Medicine.
New York:  McGraw-Hill. (See Chapter 14).

A more general framework for "research synthesis" is
presented in a variety of sources, making the result of
any specific study just a data point in the context of
related studies.  One source that I would suggest to
people who are interested in this is the following which
follows the Cochrane Collaboration's framework:

Egger, M., Smith, G.D., & Altman, D.G. (2001).
Systematic Reviews in Health Care:  Meta-analysis
in Context (2nd Ed.).  Cornwall, England: BMJ press.

I've taken a long route to make what may seem to be
rather simple points:  whether one uses the terms
"marginally significant" for the ".05< p <.10" situation
depends upon how one understands and conceptualizes
what the statistical test is expressing and how this result
relates to other studies.  The question of assumptions,
unexamined or otherwise is pertinent.

Mike Palij
New York University
[EMAIL PROTECTED]

s



---
You are currently subscribed to tips as: [EMAIL PROTECTED]
To unsubscribe send a blank email to [EMAIL PROTECTED]

Reply via email to